A method for identifying land space interference elements in a cascaded security bottom line scene classification
By constructing a convolutional self-attention network and an interference element instance segmentation network through a cascaded safety baseline scene classification method, the problem of insufficient accuracy in interference element identification in complex geographical environments by traditional remote sensing technology is solved, and high-precision interference element identification and rapid technology application are achieved.
Patent Information
- Application Number
- CN202510993943.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-07-18
AI Technical Summary
Traditional remote sensing technology struggles to effectively identify interference elements in complex geographical environments, especially in areas where multiple baselines intersect, leading to unclear identification of interference elements and impacting the implementation of land space planning. Furthermore, existing depth change detection networks lack sufficient accuracy in safety baseline scenarios.
A cascaded safety baseline scenario classification method is adopted. By constructing a convolutional self-attention safety baseline scenario detection network and a land space interference element instance segmentation network, and combining it with a reinforced sample set to train the model, the safety baseline scenarios are identified and interference elements are constrained. Self-attention features and geospatial association algorithms are used to improve the recognition accuracy.
It improves the accuracy of interference element identification and the ability to resist background interference, reduces the false alarm rate, quickly locates interference elements, achieves high-precision interference element target identification and type refinement, and assists in field investigation.
Smart Images

Figure CN120808167B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of remote sensing land space interference element target detection technology, and in particular to a land space interference element identification method based on cascaded security baseline scene classification. Background Technology
[0002] With the successful development and completion of the construction of the national spatial planning system and the first round of national spatial planning compilation and approval, the focus of planning management has gradually shifted to implementation supervision. The most important aspect of effective supervision of the implementation of national spatial planning is to quickly and accurately monitor various factors (national spatial interference factors) that affect the implementation of national spatial planning, and to efficiently identify various human activities, as well as disturbances and changes in the types, intensity, and patterns of national spatial use caused by changes in the Earth's ecological environment and natural conditions. This is of great significance for promoting high-quality development of national spatial planning and improving the effectiveness of its implementation.
[0003] With the rapid development of remote sensing technology, its advantages of being non-contact, low-cost, and wide-ranging have rapidly promoted the efficient identification and monitoring of interference elements in national land space. However, traditional identification methods mainly rely on GIS overlay analysis and static threshold determination, which are difficult to adapt to the needs of identifying dynamic interference elements in complex geographical environments. Domestic and international studies have shown that approximately 67% of national land space conflicts stem from the unclear identification of interference elements in areas with multiple baseline intersections. This not only fails to fully leverage the advantages of remote sensing applications but also seriously affects the construction process of my country's national land space planning implementation network.
[0004] The rapid development of artificial intelligence technology has also rapidly promoted revolutionary reforms in remote sensing and land space interference element identification. Based on the construction of a deep change detection network model, it is possible to quickly identify multi-temporal land space interference elements. However, due to the relatively dispersed remote sensing target features of interference elements, the features affecting the interference category cannot be well learned and trained. As a result, in relatively complex security baseline scenarios, it can only basically achieve general interference element detection. At the same time, due to interference from complex natural backgrounds, its accuracy has not reached the requirements of engineering applications.
[0005] Therefore, to address the above issues, simply using a conventional deep change detection network framework to directly detect interfering elements ignores the constraints and control of the safety baseline scenario. It is difficult to make good use of the feature enhancement capabilities under the coupling of scenario knowledge. It is necessary to further cascade the spatial constraint capabilities of land interference elements and integrate them into the land spatial interference element target state recognition process. This will overcome complex backgrounds, focus on feature state learning, and further improve the boundary accuracy and type refinement of interference elements. Summary of the Invention
[0006] The purpose of this invention is to provide a method for identifying interference elements in territorial space based on cascaded safety baseline scenario classification, thereby solving the aforementioned problems existing in the prior art.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0008] A method for identifying interference elements in national land space based on cascaded safety baseline scenario classification includes the following steps:
[0009] S1. Enhanced Semantic Recognition Sample Set for Safety Bottom Line Scenarios: Based on Previous High-Resolution Remote Sensing Images I Pre High-resolution remote sensing images after time phase I Post Acquiring Safety Bottom Line Mixed Image Data Source I 安全底线 Using a safety baseline-based hybrid image data source I 安全底线 Acquired image data source I′ that accurately identifies the safety baseline 安全底线 Safety bottom line scenario marker data VEC scene Creating a safety baseline scene recognition sample set S base_label The sample set S for identifying safety baseline scenarios base_label Enhancement processing is performed to generate an enhanced semantic recognition sample set S′ for the safety baseline scenario. base_label ;
[0010] S2. Construction of the safety bottom-line scene detection model: Construct a convolutional self-attention safety bottom-line scene detection network BaseNet, and utilize the enhanced safety bottom-line scene semantic recognition sample set S′ base_label Perform iterative training to obtain the safety bottom line scenario detection model M. Base And reasoning is used to obtain the target vector result set R of the safety baseline within the demonstration area. Base ;
[0011] S3. Enquiry of Enhanced Interference Element Instance Segmentation Sample Set: Utilizing Post-Temporal High-Resolution Remote Sensing Imagery I Post And interference element target vector label data VEC Inf Create an instance segmentation sample set S of interference elements ele_label ; Segment the sample set S for interference element instances ele_label Perform enhancement processing to generate an enhanced sample set S′ of interference element instances. ele_label ;
[0012] S4. Construction of the Land Spatial Disturbance Element Instance Segmentation Network Model: An InterfNet-based land spatial disturbance element instance segmentation network is constructed using a deep vision-based model. The enhanced disturbance element instance segmentation sample set S′ is then utilized. ele_label Iterative training is performed to obtain the land space interference element instance segmentation network model M. Inf ;
[0013] S5. Acquisition of the target set of land space interference elements under the control of the safety bottom line: Based on the vector result set of the safety bottom line target in the demonstration area R Base High-resolution remote sensing images after time phase I Post Using the instance segmentation network model M of land space interference elements Inf The result set R of the safety baseline target vector is obtained through reasoning and annotation. Base The semantic element set Rset of interference elements in constrained scenarios Inf And based on the semantic element set Rset of interference elements Inf Obtain the target set Rset′ of land space interference elements under the final safety baseline control Inf .
[0014] Preferably, step S1 specifically includes the following:
[0015] S11, The previous high-resolution remote sensing images I within the demonstration area Pre High-resolution remote sensing images after time phase I Post Perform downsampling processing based on low-frequency information preservation, and obtain the results after 4-fold downsampling as I. Pre4 and I Post4 To jointly build a safe baseline for hybrid image data sources I 安全底线 ;
[0016] S12, Mix the security baseline image data source I 安全底线 By fusing and analyzing the safety baseline vector auxiliary data, false safety baseline image content is eliminated, resulting in a usable image data source I′ for accurately identifying safety baselines. 安全底线 ;
[0017] S13. The available image data source I′ accurately identifies the safety baseline. 安全底线 Safety bottom line scenario marker data VEC scene As data input, the image and labeled vector data are traversed within a certain spatial range and step size to generate a safety baseline scene recognition sample set S. base_label ;
[0018] S14. Sample set S for identifying safety baseline scenarios base_label Perform enhancement processing to generate an enhanced semantic recognition sample set S′ for the safety bottom line scenario. base_label .
[0019] Preferably, the safety baseline vector auxiliary data is a polygonal data range, and the content range it points to is the activity range of frequent human activities, such as urban boundaries, rural boundaries, and factory sites.
[0020] Step S12 specifically involves generating the safety baseline vector auxiliary data and the safety baseline hybrid image data source I. 安全底线 Auxiliary weights T of the same size 辅助权重 By mixing image data sources at the safety baseline 安全底线 With auxiliary weight T 辅助权重 By superimposing and multiplying, we obtain an image data source I′ that can accurately identify the safety baseline. 安全底线 ;
[0021]
[0022] I' 安全底线 =I 安全底线 °T 辅助权重
[0023] Wherein, when the safety baseline vector auxiliary data is located within the auxiliary polygon, the auxiliary weight T 辅助权重 The value is 0.85; otherwise, the auxiliary weight T 辅助权重 The value is 0.15.
[0024] Preferably, the enhancement process includes single-sample random enhancement and multi-sample mosaic semantic enhancement;
[0025] The single-sample random augmentation includes geometric transformation and radiometric transformation; geometric transformation includes random rotation around the sample center and symmetric transformation along the vertical or horizontal axis; radiometric transformation includes blurring, brightness, contrast, sharpening, and adding noise;
[0026] Obtain the enhanced safety baseline scene recognition sample set S′ base_label The geometric transformation involved in the single-sample random augmentation is as follows: the sample image is randomly rotated counterclockwise by a certain angle θ starting from its own center point, and the sample image is rotated and augmented at 30-degree intervals while retaining its original size. Thus, the 360-degree angle is divided into 12 parts, and each augmentation will randomly select one of the angles θ for augmentation.
[0027]
[0028] θ∈[0,30,60,90,120,150,180]
[0029] Obtain the enhanced interference element instance segmentation sample set S′ ele_label The geometric transformation involved in the single-sample random augmentation is as follows: the sample image is randomly rotated counterclockwise by a certain angle α starting from its own center point, and the sample image is rotated and augmented at 15-degree intervals while retaining its original size. This divides the 360-degree area into 24 parts, and each augmentation will randomly select one of the angles α for augmentation.
[0030]
[0031] α∈[0,15,30,45,60,75,90,105,120,135,150,165,180]
[0032] Among them, S img S' represents the matrix of sample images to be enhanced; θ and α are the rotation angles; S' img This is a matrix of sample images after rotation enhancement;
[0033] The multi-sample mosaic semantic enhancement includes three enhancement modes: 1*3, 2*2, and 3*3. The three fused enhanced samples are arranged randomly, while the parts without real semantics are filled with background 0.
[0034] Preferably, the backbone network of the convolutional self-attention safety bottom-line scene detection network BaseNet is jointly constructed by ResNet18 and VIT-Base networks; its classification head network adopts a 4-class semantic segmentation network to form a specific semantic segmentation model suitable for safety bottom-line scene detection.
[0035] In the BaseNet convolutional self-attention safety bottom-line scene detection network, the input data is extracted by 8 layers of ResNet18, and then the feature map is transformed into an Embedding feature block with the same input dimension as the VIT network through 2D convolution. Then it is connected to the Transformer Encoder module, and finally passed through the MLP network to output the final classification scene semantic single channel 4-class scene value feature probability map, completing the forward calculation process of sample data in the network.
[0036] The safety bottom line scene detection model M is obtained by iteratively training the BaseNet convolutional self-attention safety bottom line scene detection network. Base During the process, the total number of training iterations was 7 epochs, the initial learning rate was 0.001, and the learning rate was optimized and adjusted with linear warm-up and cosine annealing as the training progressed.
[0037] Preferably, the InterfNet instance segmentation network for land space interference elements uses Swin-Large as the feature backbone network to extract features from sample images. It includes three head networks for detecting the spatial location, category, and semantic mask information of interference elements.
[0038] The instance segmentation target category is set to 5, and the semantic segmentation map within each localization box is a binary segmentation feature map to extract the target semantic information of the interfering elements.
[0039] Preferably, the safety baseline scene recognition sample set S base_label and safety baseline target vector result set R BaseAll of them have the safety baseline scenario category attribute. Mark 1 represents the safety baseline for cultivated land, 2 represents the safety baseline for ecological areas, 3 represents the safety baseline for flood risk, and 0 represents the background that is not marked.
[0040] The interference element instance segmentation sample set S ele_label It has the category attribute of interference elements. Marker 0 represents roads, Marker 1 represents large-scale construction land, Marker 2 represents regular artificial lakes, Marker 3 represents large-scale bare land construction, and Marker 4 represents other typical structures.
[0041] Preferably, step S5 specifically includes the following:
[0042] S51, Within the demonstration area, the target vector result set R is based on the safety baseline. base As a constraint, with the later-phase high-resolution image I Post Jointly using the instance segmentation network model M of territorial spatial interference elements Inf Inference and annotation are performed to obtain the target vector result set R of the safety baseline. base Semantic feature set Rset of interference elements in constrained scenarios Inf ;
[0043] S52, Add the semantic element set Rset for interfering elements. Inf The safety bottom line type field is obtained by analyzing the semantic feature set Rset for each interfering feature. Inf Target bounding box location information and safety baseline target vector result set R base Spatial correlation is performed to obtain the target types of the safety baseline, which are then populated into the type field. Combined with its own target types, the final target set Rset′ of land space interference elements under the safety baseline control is obtained. Inf .
[0044] Preferably, during the reasoning and annotation process in step S51, the target vector result set R of the safety baseline is... base High-resolution post-temporal images under spatial mask I Post As the image to be inferred, the semantic geographic information of the interfering elements obtained through block-based multi-threaded parallel inference is spatially vectorized to obtain the semantic element set Rset of the interfering elements. Inf Spatial range: Based on the semantic labels of interfering elements, their category information is indexed and filled in the category attribute field of the interfering information, thereby obtaining the safety baseline target vector result set R. base Semantic feature set Rset of interference elements in constrained scenarios Inf .
[0045] Preferably, step S52 specifically involves first processing the semantic element set Rset of interfering elements. InfA safety baseline type field is added. Next, the target bounding box location information for each element is defined as the minimum bounding rectangle, and the four-dimensional range of the specific interfering element is obtained. Then, for each interfering element, the corresponding safety baseline target vector result from the previous time phase is correlated with the four-dimensional spatial range to obtain the corresponding safety baseline target type, which is then filled into the current element's safety baseline target type field. Finally, by combining its own target type, the final target set Rset′ of the territorial spatial interference elements under safety baseline control is obtained. Inf .
[0046] The beneficial effects of this invention are as follows: 1. This invention employs a deep concatenated network that couples safety baseline scene discrimination with interference element instance segmentation for automatic identification of interference elements in national land space. Compared with conventional single change detection network models, it has higher recognition accuracy and better resistance to safety baseline background interference. 2. This invention first uses auxiliary data and a safety baseline scene recognition framework to lock the actual safety baseline scene range in the preceding and following time phases. Then, it extracts specific interference element information within the real scene. The model inference takes into account the spatial constraints of the safety baseline scene and reduces the image range for inference. This not only reduces the false alarm rate of interference elements caused by erroneous safety baseline targets but also accelerates the inference speed of interference element targets. 3. The composite convolutional neural network with phase-wise downsampling and a self-attention feature-based safety baseline scene semantic recognition network used in this invention further enhances the safety baseline target recognition capability while maintaining good model convergence with only a small number of iterative training iterations. 4. This invention uses the four spatial ranges of interference elements and geospatial correlation algorithms to reverse track the safety baseline scenario types of interference elements, reducing the problem of incorrect source identification of interference elements in multiple safety baseline scenarios, improving the accuracy of identifying the safety baseline subordination relationship of interference elements, and realizing high-precision identification and source basis of interference elements, which can quickly locate and better assist in field investigations. Attached Figure Description
[0047] Figure 1 This is a flowchart of the identification method in an embodiment of the present invention;
[0048] Figure 2 This is a schematic diagram of three modes in multi-sample mosaic semantic enhancement in an embodiment of the present invention;
[0049] Figure 3 This is a schematic diagram of the BaseNet network structure for detecting safety baseline scenarios in an embodiment of the present invention;
[0050] Figure 4 The images used in this embodiment of the invention are the original high-resolution remote sensing image (A) and the low-resolution remote sensing image (B) after ordinary 4x downsampling. C is a remote sensing image image that retains low-frequency information as a safety baseline and then undergoes 4x downsampling.
[0051] Figure 5 This is a remote sensing dataset that can be used for training, obtained by combining auxiliary data to remove false safety baseline image targets from the images in this embodiment of the invention; A is the later time-phase data, and B is the earlier time-phase data;
[0052] Figure 6 The training sample set (C) is generated in this embodiment of the invention by combining the safety baseline vector data (A) and the safety baseline remote sensing dataset (B).
[0053] Figure 7 The results are the enhanced remote sensing safety baseline sample set in the embodiments of the present invention; A is the original sample, B is the sample result after radiometric enhancement, C is the sample result after random counterclockwise rotation of 45° around the image center, and D is the 1*3 mosaic semantic enhancement sample result.
[0054] Figure 8 The result of inference from the post-temporal remote sensing image in this embodiment of the invention is the result of the safety bottom line training model. A is the original remote sensing image, and the dark line in B is the range of the inferred arable land red line.
[0055] Figure 9 These are samples of interference elements and their enhancement effects in embodiments of the present invention; A, top and bottom, are sample images of the later time phase and instance sample markers of interference elements located therein; B, top and bottom, are the results after randomly rotating 30° counterclockwise around the center of the sample image of the later time phase and instance sample markers of interference elements located therein after the same rotation; C, top and bottom, are image samples of interference elements enhanced with 2*2 mosaic samples and corresponding instance marker sample results.
[0056] Figure 10 This is the result of inference of interference elements in the remote sensing image of the later time phase under the safety baseline constraint in the embodiment of the present invention (B), where A is the case of the result being applied to the earlier time phase;
[0057] Figure 11 This is the final identification result of interference elements in the post-temporal remote sensing image in this embodiment of the invention; A is the geographical correlation between the spatial range of the interference elements and the safety baseline constraints, and B is the content of the attribute table of the final interference elements. Detailed Implementation
[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0059] Example 1
[0060] To overcome the limitations of single-change detection interference element network identification, this embodiment provides a cascaded safety baseline scene classification method for identifying land space interference elements. This method targets the identification of land space interference elements in complex backgrounds of remote sensing imagery. First, it utilizes low-resolution, large-scale remote sensing images from previous time periods within the demonstration area, combined with auxiliary data such as known cultivated land and ecological red lines, to construct a realistic scene recognition model and obtain the true and effective spatial range of the safety baseline within the monitoring area. Second, within this baseline scene range, it uses high-resolution remote sensing images from the demonstration area to create small-scale interference element status instance segmentation samples, train and infer target detection networks, and obtain the target location and semantic information of interference elements. Finally, it spatially erases and associates the spatial locations of interference element targets with the safety baseline range, forming a monitoring process for changes in interference element targets within the demonstration area. The overall algorithm, by coupling spatial knowledge of the safety baseline with previous time-phase attributes, further cascades and precisely identifies interference elements at fine scales, forming a land space interference element target identification algorithm with multi-task scene cascading and geographic attribute spatial association. Figure 1 As shown, the method specifically includes the following parts:
[0061] I. Acquisition of the Enhanced Semantic Recognition Sample Set for Security Bottom Line Scenarios
[0062] Based on previous high-resolution remote sensing images I Pre High-resolution remote sensing images after time phase I Post Acquiring Safety Bottom Line Mixed Image Data Source I 安全底线 Using a safety baseline-based hybrid image data source I 安全底线 Acquired image data source I′ that accurately identifies the safety baseline 安全底线 Safety bottom line scenario marker data VEC scene Creating a safety baseline scene recognition sample set S base_label The sample set S for identifying safety baseline scenarios base_label Enhancement processing is performed to generate an enhanced semantic recognition sample set S′ for the safety baseline scenario. base_label . Specifically,
[0063] 1.1 The previous high-resolution remote sensing images of the demonstration area I Pre High-resolution remote sensing images after time phase I Post Perform downsampling processing based on low-frequency information preservation, and obtain the results after 4-fold downsampling as I. Pre4 and I Post4 To jointly build a safe baseline for hybrid image data sources I 安全底线 .
[0064] I 安全底线 This includes high-resolution remote sensing images from previous and subsequent time periods, filtered by a 3x3 filter F.l After completing the retention of low-frequency information at the safety baseline, I is obtained by downsampling by 4 times. Pre4 and I Post4 The mixed data composed of F l The expression is,
[0065]
[0066] The 4x downsampling method used is the average of max pooling downsampling with a kernel size of 2 and a step size of 2, and average pooling downsampling. This method is based on previous high-resolution remote sensing images I. Pre Taking 4x downsampling as an example, the result I obtained is Pre4 The expression is,
[0067] I Pre4 =(Maxpooling(2,2)(I Pre )+Avepooling(2,2)(I Pre )) / 2.
[0068] 1.2. Put I 安全底线 By fusing and analyzing the safety baseline vector auxiliary data, false safety baseline image content is eliminated, resulting in a usable image data source I′ for accurately identifying safety baselines. 安全底线 .
[0069] Among them, the safety baseline vector auxiliary data is usually a polygonal data range, and the content range it points to is the activity range of frequent human activities such as urban boundaries, rural boundaries and factory sites.
[0070] The fusion analysis in this step, which removes false security baseline image content, is mainly accomplished through an algorithm that locks high-weight parameters. The specific execution process is as follows:
[0071] Generate safety baseline vector auxiliary data and I 安全底线 Auxiliary weights T of the same size 辅助权重 , mix image data sources with safety baseline I 安全底线 With auxiliary weight T 辅助权重 By superimposing and multiplying (using matrix element-wise multiplication here), a usable and accurate image data source I′ for identifying the safety baseline is obtained. 安全底线 The specific method for setting the auxiliary weights is as follows:
[0072]
[0073] I' 安全底线 =I 安全底线 °T 辅助权重
[0074] Wherein, when the safety baseline vector auxiliary data is located within the auxiliary polygon, the auxiliary weight T 辅助权重 The value is 0.85; otherwise, the auxiliary weight T 辅助权重 The value is 0.15.
[0075] 1.3. Change I′ 安全底线 Safety bottom line scenario marker data VEC scene As data input, the image and labeled vector data are traversed within a certain spatial range and step size to generate a safety baseline scene recognition sample set S. base_label .
[0076] Safety bottom line scenario recognition sample set S base_label The dimensions are 512px * 512px, and the safety baseline scenarios corresponding to the type tags are shown in Table 1.
[0077] Table 1 shows the safety baseline scenarios corresponding to the type labels.
[0078] Serial Number Tag value Safety bottom line scenario name 1 0 background 2 1 Safety baseline for arable land 3 2 Ecological safety bottom line 4 3 Flood risk safety bottom line
[0079] 1.4 Sample set S for identifying safety baseline scenarios base_label Perform enhancement processing to generate an enhanced semantic recognition sample set S′ for the safety bottom line scenario. base_label .
[0080] By S base_label Generate an enhanced safety baseline scene recognition sample set S′ base_label The enhancement algorithms involved are divided into two main aspects: single-sample randomness enhancement and multi-sample mosaic semantic enhancement.
[0081] (1) The randomness enhancement of a single sample includes geometric transformation (random rotation around the sample center, symmetrical transformation of the vertical or horizontal axis) and radiation transformation (blurring, brightness, contrast, sharpening, adding noise, etc.). The principle of rotation transformation in geometric space is to randomly rotate the sample image counterclockwise by a certain angle θ from the center point of the sample image itself, and perform rotation processing on the sample while retaining the original size at 30-degree intervals. Thus, the 360 degrees are divided into 12 parts, and each enhancement will randomly select one angle for enhancement processing. The calculation formula is as follows.
[0082]
[0083] θ∈[0,30,60,90,120,150,180]
[0084] Among them, S img S′ represents the matrix of sample images to be enhanced, where θ is the rotation angle. imgThis is the matrix of sample images after rotation enhancement. Since no translation is involved, the third column of the rotation matrix is all 0. Also, the rotation angle has angular symmetry properties, so only the range of 0-180° needs to be considered.
[0085] (2) Multi-sample mosaic semantic enhancement, including three enhancement modes: 1*3, 2*2, and 3*3, such as... Figure 2 As shown. Unlike the traditional 1*3 pattern, the three fused augmented samples are not arranged horizontally as in the traditional way, but are arranged randomly, and the parts without real semantics are filled with background 0.
[0086] II. Construction of Safety Bottom Line Scenario Detection Model
[0087] Construct a convolutional self-attention safety bottom-line scene detection network BaseNet, and utilize the enhanced safety bottom-line scene semantic recognition sample set S′ base_label Perform iterative training to obtain the safety bottom line scenario detection model M. Base And reasoning is used to obtain the target vector result set R of the safety baseline within the demonstration area. Base .
[0088] The backbone of the safety bottom line scenario detection network BaseNet is jointly constructed by ResNet18 and VIT-Base networks. The classification head network adopts a 4-class semantic segmentation network to form a specific semantic segmentation model suitable for safety bottom line scenario detection. The model can quickly converge based on a large range of small-scale training.
[0089] The specific BaseNet network structure is as follows: Figure 3 As shown, the input data is processed through 8 layers of ResNet18 feature extraction, and then the feature map is converted into an Embedding feature block with the same input dimension as the Vit network through 2D convolution. Then it is connected to the Transformer Encoder module, and finally passed through the MLP network to output the final single-channel 4-class scene value feature probability map of the classification scene semantics, thus completing the forward calculation process of the sample data in the network.
[0090] The safety bottom line scenario detection model M is obtained by iteratively training the BaseNet safety bottom line scenario detection network. Base The total number of training iterations (epochs) was set to 7, and the initial learning rate was set to 0.001. The learning rate was optimized and adjusted using linear warm-up and cosine annealing as training progressed.
[0091] Using the safety bottom line scenario detection model M Base The safety baseline scene detection result R was obtained by reasoning and identifying remote sensing images of the demonstration area. baseIt is a vector file format and has a safety baseline scenario category attribute (see Table 1). 1 represents the farmland safety baseline, 2 represents the ecological safety baseline, and 3 represents the flood risk safety baseline, while the background 0 is not marked.
[0092] III. Obtaining the Enhanced Interference Element Instance Segmentation Sample Set
[0093] Using post-temporal high-resolution remote sensing image I Post And interference element target vector label data VEC Inf Create an instance segmentation sample set S of interference elements ele_label ; Segment the sample set S for interference element instances ele_label Perform enhancement processing to generate an enhanced sample set S′ of interference element instances. ele_label .
[0094] Interference element instance segmentation sample set S ele_label The tag files are in txt format, and there are 5 categories, as shown in Table 2.
[0095] Table 2 Types of Interference Elements
[0096] Serial Number Tag value Interference Element Types 1 0 the way 2 1 Large-scale construction land 3 2 Regular artificial lake 4 3 Large-scale construction on bare land 5 4 Other typical structures
[0097] By S ele_label Generate an enhanced segmentation sample set S′ of the interference element instances. ele_label The enhancement process involved is the same as that in the first part. The difference is that when performing rotation enhancement on samples under geometric transformation, considering that the spatial distribution of interfering elements is small and easily affected by the background environment, the rotation enhancement is performed on the samples at 15-degree intervals with the center of the sample image as the reference. Thus, the 360 degrees are divided into 24 parts, and each enhancement will randomly select one of the angles α for enhancement.
[0098]
[0099] Due to the principle of point symmetry of angles, the angle in the formula only needs to be controlled within the range of 0-180 degrees, that is, α∈[0,15,30,45,60,75,90,105,120,135,150,165,180].
[0100] IV. Construction of a Network Model for Segmenting Instances of Disturbance Elements in Territorial Space
[0101] An InterfNet-based spatial disturbance element instance segmentation network based on a deep vision model is constructed, and the enhanced disturbance element instance segmentation sample set S′ is utilized. ele_label Iterative training is performed to obtain the land space interference element instance segmentation network model M. Inf .
[0102] The InterfNet spatial interference element instance segmentation network uses Swin-Large as the feature backbone network to extract features from sample images, and three head networks are used to detect the spatial location, category and semantic mask information of interference elements.
[0103] In this context, the target category for instance segmentation is set to 5, and the semantic segmentation map within each localization box is a binary segmentation feature map to extract the target semantic information of interfering elements.
[0104] V. Acquisition of Target Set of Disturbance Elements in Territorial Space under Safety Bottom Line Control
[0105] Based on the safety baseline target vector result set R in the demonstration area Base High-resolution remote sensing images after time phase I Post Using the instance segmentation network model M of land space interference elements Inf The result set R of the safety baseline target vector is obtained through reasoning and annotation. Base The semantic element set Rset of interference elements in constrained scenarios Inf And based on the semantic element set Rset of interference elements Inf Obtain the target set Rset′ of land space interference elements under the final safety baseline control Inf Specifically,
[0106] 5.1 Within the demonstration area, the target vector result set R is based on the safety baseline. base As a constraint, with the later-phase high-resolution image I Post Jointly using the instance segmentation network model M of territorial spatial interference elements Inf Inference and annotation are performed to obtain the target vector result set R of the safety baseline. base Semantic feature set Rset of interference elements in constrained scenarios Inf .
[0107] When performing inference and annotation, the target vector result set R of the safety baseline is... base High-resolution post-temporal images under spatial mask I Post As the image to be inferred, block-based multi-threaded parallel inference is used (where the number of parallel threads is consistent with the number of GPUs). The obtained semantic geographic information of the interfering elements is spatially vectorized to obtain the semantic feature set Rset of the interfering elements. Inf Spatial range: Based on the semantic labels of interfering elements, their category information is indexed and filled in the category attribute field of the interfering information, thereby obtaining the safety baseline target vector result set R. base Semantic feature set Rset of interference elements in constrained scenarios Inf .
[0108] 5.2 Add a set of semantic elements (Rset) to the set of interfering elements. Inf The safety bottom line type field is obtained by analyzing the semantic feature set Rset for each interfering feature. Inf Target bounding box location information and safety baseline target vector result set R base Spatial correlation is performed to obtain the target types of the safety baseline, which are then populated into the type field. Combined with its own target types, the final target set Rset′ of land space interference elements under the safety baseline control is obtained. Inf .
[0109] Specifically, firstly, the semantic element set Rset of the interfering elements... Inf A safety baseline type field is added. Next, the target bounding box location information for each element is defined as the minimum bounding rectangle, and the four-dimensional range of the specific interfering element is obtained. Then, for each interfering element, the corresponding safety baseline target vector result from the previous time phase is correlated with the four-dimensional spatial range to obtain the corresponding safety baseline target type, which is then filled into the current element's safety baseline target type field. Finally, by combining its own target type, the final target set Rset′ of the territorial spatial interference elements under safety baseline control is obtained. Inf .
[0110] Example 2
[0111] In this embodiment, remote sensing image data from the ZY-3 satellite with a spatial resolution of 2 meters is used to conduct an experiment on the extraction of interference elements in national land space using the method of the present invention, so as to better illustrate the execution process and advantages of the method of the present invention.
[0112] I. Acquisition of the Enhanced Semantic Recognition Sample Set for Security Bottom Line Scenarios
[0113] Within the demonstration area, select front- and back-temporal true-color satellite remote sensing images with a spatial resolution of 2 meters and a size of 4137px*3254px. Pre and I Post After undergoing a 4x downsampling processing algorithm based on low-frequency information preservation, I is obtained. Pre4 and I Post4 The image size becomes 1110px*873px, and the spatial resolution becomes 8 meters, together forming I 安全底线 .like Figure 4 As shown.
[0114] to I 安全底线 The data is processed using an algorithm that locks high-weight parameters based on urban vector auxiliary data, and I... 安全底线 The urban building area is subjected to background removal to weaken the false safety baseline image portion that affects the identification of the safety baseline range, resulting in I′. 安全底线 At this time, I′ 安全底线The dimensions remain 1110px * 873px, but the spatial resolution becomes 8 meters. For example... Figure 5 As shown.
[0115] Using I′ 安全底线 and real-world safety baseline scenario marker vector data VEC scene Create and generate S-level safety baseline scenario samples base_label Since this demonstration area only includes the arable land safety baseline, the sample label value is 1, and the sample size is 512*512px. After increasing the step size and expanding replication, a total of 418 sample pairs were generated. Figure 6 As shown.
[0116] The obtained 418 pairs of safety bottom line sample sets S base_label After performing single-sample randomness enhancement (including geometric transformations: random rotation around the sample center, symmetrical transformation along the vertical or horizontal axis; radiometric changes such as blurring, brightness, contrast, sharpening, and adding noise) and multi-sample mosaic semantic enhancement, a reinforced safety baseline scene recognition sample set S′ is generated. base_label The augmented sample size was increased tenfold to 4180 pairs, and then divided into training and validation sets in a 9:1 ratio, resulting in 3762 pairs and 418 pairs respectively. The sample size remained 512px * 512px. Figure 7 As shown.
[0117] II. Construction of Safety Bottom Line Scenario Detection Model
[0118] Using the obtained S′ base_label The self-attention safety bottom-line scene detection network BaseNet was trained with an initial learning rate of 0.001 for 7 epochs. After training, it quickly converged to a validation accuracy of 0.976, yielding the safety bottom-line scene detection model file M. Base The model file size is 175MB and it is saved to disk.
[0119] Using the obtained safety baseline scenario detection model file M Base By inferring the later-phase remote sensing images within the demonstration and verification area, the target vector result set R of the safety baseline within the demonstration area is obtained. base Since this demonstration zone only contains arable land, the safety baseline identification category is 1. For example... Figure 8 As shown.
[0120] III. Obtaining the Enhanced Interference Element Instance Segmentation Sample Set
[0121] Using post-temporal high-resolution remote sensing image I Post And interference element target vector label data VEC Inf Create an instance segmentation sample set S of interference elements ele_labelA total of 2080 pairs of interference element instance samples were generated, with a sample size of 256px*256px. Through single-sample randomness enhancement (including geometric transformations: random rotation around the sample center, symmetrical transformation along the vertical or horizontal axis; radiation changes such as blurring, brightness, contrast, sharpening, and noise addition) and multi-sample mosaic instance segmentation enhancement, an enhanced interference element instance segmentation sample set S′ was generated. ele_label At this point, the sample size expands to 18,720 pairs, allocated according to an 8:2 training-to-validation ratio, resulting in 14,976 and 3,744 pairs of interference element training and validation samples, respectively. For example... Figure 9 As shown.
[0122] IV. Construction of a Network Model for Segmenting Instances of Disturbance Elements in Territorial Space
[0123] Using the obtained S′ ele_label The constructed InterfNet network for segmenting instances of spatial interference elements was trained with an initial learning rate of 0.01, a batch size of 8, and 120 epochs. The final result was a spatial interference element instance segmentation network model M with a validation set mAP of 0.85. Inf The model file is 853MB.
[0124] V. Acquisition of Target Set of Disturbance Elements in Territorial Space under Safety Bottom Line Control
[0125] Using the obtained safety baseline target vector result set R within the demonstration area base and interference element instance segmentation network model file M Inf Using R base Constrained post-temporal image block reasoning yields the semantic element set Rset of post-temporal interference elements. Inf In this example, five interference elements were identified, all of which were large-scale construction sites and had a label value of 1. Figure 10 As shown.
[0126] The semantic feature set Rset that traverses the interfering features Inf For each interference element, spatially correlate the four corners of its instance detection target with the farmland safety baseline and fill the established "safety baseline" field. This yields an interference element map for this demonstration area, with the annotation "Large-scale building land interference located at the farmland safety baseline." The annotations for the other four interference elements are "Not within any safety baseline field; discard this building land interference" (e.g., ...). Figure 11 As shown in B), the most accurate interference element set Rset′ is obtained through safety baseline scenario constraint analysis. Inf .like Figure 11 As shown in Table 3, the constraint analysis of the interference factors is presented.
[0127] Table 3 Constraint Analysis of Interference Factor Results
[0128]
[0129] This invention addresses the identification of spatial interference elements in complex backgrounds of remote sensing imagery. First, it utilizes low-resolution, large-scale remote sensing imagery from earlier time periods, combined with auxiliary data from known urban and rural activity zones, to intelligently identify the true and effective spatial extent of the safety baseline within the monitoring area using a safety baseline scene model. Second, within this baseline scene, it uses high-resolution remote sensing imagery from later time periods to create small-scale instance segmentation samples of current status of interference elements, train and infer target detection networks, and obtain the target location and semantic information of the interference elements. Finally, it spatially correlates the spatial attributes of the interference element targets at corresponding temporal locations with the safety baseline range, forming a monitoring process of attribute transitions and changes in interference element targets within the demonstration area. This invention's multi-task scene cascading and geographic attribute spatial correlation algorithm model for identifying spatial interference elements in land use further refines the accurate identification of interference elements at small scales by coupling spatial knowledge of the safety baseline with earlier time-phase attributes. It overcomes the difficulties of traditional methods, such as weak spatial attribute correlation within the safety baseline scene and inability to judge attribute changes in interference elements, providing strong data support and technical assurance for the construction of a monitoring network for land use planning implementation and ensuring high-quality development of land use.
[0130] By adopting the above-disclosed technical solution of this invention, the following beneficial effects are obtained:
[0131] This invention provides a method for identifying interference elements in national land space based on cascaded safety baseline scene classification. This invention employs a deep cascaded network that couples safety baseline scene discrimination with interference element instance segmentation for automatic identification of interference elements in national land space. Compared with conventional single change detection network models, it has higher recognition accuracy and better resistance to background interference from the safety baseline. This invention first uses auxiliary data and a safety baseline scene recognition framework to lock the actual safety baseline scene range in the preceding and following time phases. Then, it extracts specific interference element information within the real scene. The model inference takes into account the spatial constraints of the safety baseline scene and reduces the image range for inference, which not only reduces the false alarm rate of interference elements caused by erroneous safety baseline targets but also accelerates the inference speed of interference element targets. The composite convolutional neural network with phase-wise downsampling and a self-attention feature-based semantic recognition network for safety baseline scenes further enhances the ability to identify safety baseline targets while maintaining good model convergence with only a small number of iterations. This invention uses the four spatial ranges of interference elements and geospatial correlation algorithms to reverse track the safety baseline scenario types of interference elements, reducing the problem of incorrect source identification of interference elements in multiple safety baseline scenarios, improving the accuracy of identifying the safety baseline subordination relationship of interference elements, and realizing high-precision identification and source basis of interference elements, which can quickly locate and better assist in field investigations.
[0132] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for identifying a land space interference factor of a cascaded security baseline scene classification, characterized in that: Comprising the following steps, S1. Enhanced Semantic Recognition Sample Set for Safety Bottom Line Scenarios: Based on Previous High-Resolution Remote Sensing Images I Pre High-resolution remote sensing image I Post Acquiring Safety Bottom Line Mixed Image Data Source I 安全底线 Using a safety baseline-based hybrid image data source I 安全底线 Acquired image data source I′ that accurately identifies the safety baseline 安全底线 Safety bottom line scenario marker data VEC scene Creating a safety baseline scene recognition sample set S base_label For the safety baseline scenario identification sample set S base_label Enhancement processing is performed to generate an enhanced semantic recognition sample set S′ for the safety baseline scenario. base_label ; S2, safety baseline scene detection model construction: construct a convolutional self-attention safety baseline scene detection network BaseNet, and use the reinforced safety baseline scene semantic recognition sample set S' base_label Iterative training is performed to obtain a safety baseline scene detection model M Base and reasoning to obtain a safety baseline target vector result set R Base in the demonstration area. S3, obtaining the enhanced interference factor instance segmentation sample set: using the post-phase high-resolution remote sensing image I Post and the interference factor target vector marking data VEC Inf making the interference factor instance segmentation sample set S ele_label ; to the interference factor instance segmentation sample set S ele_label performing enhancement processing to generate the enhanced interference factor instance segmentation sample set S' ele_label ; S4, land space interference factor instance segmentation network model construction: construct a land space interference factor instance segmentation network InterfNet based on a deep vision basic model, and use the reinforced interference factor instance segmentation sample set S' ele_label Iterative training is performed to obtain a land space interference factor instance segmentation network model M Inf ; S5, obtaining a target set of land space interference elements under the safety bottom line management: based on the safety bottom line target vector result set R in the demonstration area Base and the post-phase high-resolution remote sensing image I Post using the land space interference element instance segmentation network model M Inf reasoning and labeling to obtain the safety bottom line target vector result set R Base the interference element semantic element set Rset under the constraint scene Inf and based on the interference element semantic element set Rset Inf obtaining the final target set of land space interference elements under the safety bottom line management Rset' Inf .
2. The method of claim 1, wherein the method further comprises: Step S1 specifically includes the following, S11, the high-resolution remote sensing image I in the demonstration area before the phase Pre and the high-resolution remote sensing image I after the phase Post , the low-frequency information is reserved, and the down-sampling processing is carried out, and the results after 4 times down-sampling are I Pre4 and I Post4 , to jointly construct a safety bottom line mixed image data source I 安全底线 ; S12, mixing the security baseline image data source I 安全底线 With the security baseline vector auxiliary data fusion analysis, the false security baseline image content is removed, and the usable accurate security baseline recognition image data source I' is formed 安全底线 ; S13, available accurate safety line recognition image data source I' 安全底线 and safety line scene marking data VEC scene As data input, iterate through the image and marking vector data with certain spatial range and step length, generate safety line scene recognition sample set S base_label ; S14, a security baseline scene recognition sample set S base_label performing an enhancement process to generate an enhanced security baseline scene semantic recognition sample set S' base_label .
3. The method of claim 2, wherein the method further comprises: determining a land use type of the land parcel; and determining a land use type of the land parcel based on the land use type of the land parcel and the land use type of the land parcel. The security baseline vector auxiliary data is a planar polygon data range, and the content range pointed by the security baseline vector auxiliary data is an activity range of human activities such as city boundaries, village boundaries and factory sites; Step S12 is specifically to generate the security baseline vector auxiliary data and the security baseline mixed image data source I 安全底线 The auxiliary weight T 辅助权重 has the same size 安全底线 . By superimposing and multiplying the security baseline mixed image data source I 辅助权重 and the auxiliary weight T 安全底线 , the available image data source I' for accurately identifying the security baseline is obtained. wherein the value of the auxiliary weight T 辅助权重 is 0.85 when the security floor vector auxiliary data is located within the auxiliary polygon; otherwise, the value of the auxiliary weight T 辅助权重 is 0.
15.
4. The method of claim 1, wherein the method further comprises: The enhancement processing includes single-sample random enhancement and multi-sample mosaic semantic enhancement; The single-sample random enhancement includes geometric transformation and radiation transformation; the geometric transformation includes random rotation with a sample center, longitudinal axis or transverse axis symmetry transformation; the radiation transformation includes blurring, brightness, contrast, sharpening and adding noise; Obtaining the reinforced security bottom line scene recognition sample set S' base_label The geometric transformation in the single-sample random enhancement involves randomly rotating a certain angle θ counterclockwise from the center point of the sample image itself, and performing a reserved original size rotation enhancement processing on the sample image at an interval of 30 degrees. Then, 360 degrees are divided into 12 parts, and one angle θ is randomly selected for enhancement processing each time. θ∈[0,30,60,90,120,150,180] obtaining the enhanced interference element instance segmentation sample set s′ ele_label The geometric transformation in the single-sample random enhancement involves randomly rotating a certain angle a counterclockwise from the center point of the sample image itself, and performing a reserved original size rotation enhancement processing on the sample image at intervals of 15 degrees. 360 degrees are divided into 24 parts, and each enhancement randomly selects one angle a for enhancement processing. α∈[0,15,30,45,60,75,90,105,120,135,150,165,180] where S img is the sample picture matrix to be enhanced; θ and α are the rotation angles; S′ img is the sample picture matrix after rotation enhancement; The multi-sample mosaic semantic enhancement includes three modes of enhancement: 1*3, 2*2 and 3*3; the three fused enhanced samples are arranged in a random manner, and the part without real semantics is filled with background 0.
5. The method of claim 1, wherein the method further comprises: The backbone network part of the convolutional self-attention security baseline scene detection network BaseNet is jointly constructed by the networks of ResNet18 and VIT-Base; the classification head network adopts a 4-class semantic segmentation network, forming a specific semantic segmentation model suitable for security baseline scene detection; In the convolutional self-attention security baseline scene detection network BaseNet, after the input data is subjected to 8-layer feature extraction by ResNet18, the feature map is converted into an Embedding feature block with consistent input dimension of the VIT network through 2-dimensional convolution, and then is connected to the Transformer Encoder module, and finally is subjected to the MLP network to output the final classification scene semantic single-channel 4-class scene value feature probability map, completing the forward calculation process of the sample data in the network; A safety bottom line scene detection model M is obtained through iterative training of a convolution self-attention safety bottom line scene detection network BaseNet Base In the process, the total training iteration round number epoch is 7, the initial learning rate is 0.001, and the iteration learning rate is optimized and adjusted in a linear preheating and cosine annealing manner as the training proceeds.
6. The method of claim 1, wherein the method further comprises: The land space interference element instance segmentation network InterfNet adopts Swin-Large as the feature backbone network to extract sample image features, which includes three head networks for detecting the spatial position, category and semantic mask information of the interference elements; The instance segmentation target category is set to 5, and the semantic segmentation graph in each positioning box is a binary segmentation feature graph, which extracts the target semantic information of the interference elements.
7. The method of claim 1, wherein the method further comprises: The security bottom line scene recognition sample set S base_label And the security bottom line target vector result set R Base All have security bottom line scene category attributes, and label 1 represents the security bottom line of the farmland range, 2 represents the security bottom line of the ecological range, 3 represents the security bottom line of the flood risk, and 0 represents the background not being marked; The interference element instance segmentation sample set S ele_label The interference element class attribute is provided, and the label 0 represents a road, the label 1 represents a large-scale building site, the label 2 represents a regular artificial lake, the label 3 represents a large-range bare land construction, and the label 4 represents other typical structures.
8. The method of claim 1, wherein the method further comprises: determining a land use type of the land parcel; and determining a land use type of the land parcel. Step S5 specifically includes the following, S51, a safety bottom line target vector result set R in the demonstration area base As a constraint condition, a high-resolution image I of a later time phase Post Jointly, using a land space interference factor instance segmentation network model M Inf Reasoning and labeling are performed, and then a safety bottom line target vector result set R is obtained base Interference factor semantic factor set Rset in the constraint scene Inf ; S52, increase the interference element semantic element set Rest Inf security bottom line type field, by the target frame positioning information of each interference element semantic element set Rset Inf spatial correlation with the security bottom line target vector result set R base obtain the security bottom line target type, and fill it into the type field, and obtain the final land space interference element target set Rset' under the security bottom line management and control Inf .
9. The method of claim 8, wherein the method further comprises: The reasoning and labeling in step S51 is performed on the safety bottom line target vector result set R base The post-phase high-resolution image I under the spatial mask Post As the image to be reasoned, the interference factor semantic geographic information obtained by using block multi-thread parallel reasoning is spatially vectorized to obtain the interference factor semantic element set Rset Inf The spatial range is indexed according to the category information of the interference factor semantic label, and the interference information category attribute field is filled in, and then the safety bottom line target vector result set R is obtained base The interference factor semantic element set Rset in the constraint scene Inf .
10. The method of claim 8, wherein the method further comprises: Step S52 is specifically, first, the interference element semantic element set Rset Inf Add the safety baseline type field, secondly, define the target frame positioning information of each element as the minimum bounding rectangle, and obtain the four-position range of the specific interference element, thirdly, for each interference element, obtain the corresponding safety baseline target type by associating the safety baseline target vector result of the corresponding previous phase through the four-position space range, and fill it into the current element safety baseline target type field, and then combine the target type itself to obtain the final land space interference element target set Rset' under the safety baseline management Inf .
Citation Information
Patent Citations
Remote sensing image semantic segmentation method combined with super-resolution technology
CN118014844A
Remote sensing image building change detection method and system based on morphological constraint
CN118736425A