Building collapse scene level sparse target detection method based on small sample training
By using a small sample training method, multi-source satellite remote sensing images and deep neural network models, the data scarcity and complex scene modeling problems of sparse target detection in building collapse scenes are solved, and high-precision building collapse area detection is achieved.
Patent Information
- Application Number
- CN202510743331.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-26
AI Technical Summary
Target detection in building collapse scenarios faces the problems of data scarcity, poor sparse target detection performance, and insufficient complex scene modeling capabilities. Existing detection algorithms have difficulties in data acquisition, limited receptive fields, and insufficient multimodal data fusion capabilities, resulting in insufficient model generalization capabilities and low detection accuracy.
A small sample training-based method is adopted to generate high-resolution multispectral images through preprocessing of multi-source satellite remote sensing images. A deep neural network model is constructed, including feature extraction, scene-level sparse feature attention module and Transformer classification network. A pixel-level and scene-level label dual-constrained loss function is designed, and target domain alignment is performed to improve the accuracy of the model in sparse target detection.
It achieves high-precision detection of sparse targets in building collapse scenarios, improves the model's generalization ability and ability to understand complex scenarios, and can generate accurate detection results under small sample conditions.
Smart Images

Figure CN120708078A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of remote sensing monitoring, and in particular relates to a building collapse scene-level sparse target detection method based on small sample training. Background Art
[0002] Building collapse is a major disaster that threatens human life and urban resilience. Its suddenness and widespread destruction are compounded by the sparse distribution and diverse morphology of objects (such as survivors, structural debris, and cracks) in post-disaster scenarios, posing significant challenges to traditional detection methods. With the advancement of artificial intelligence and remote sensing technologies, deep learning-based visual inspection technology has gradually become a core tool for post-disaster emergency response. However, its application in building collapse scenarios still faces the following bottlenecks:
[0003] (1) Data scarcity and high labeling costs. Building collapse scenarios are highly random and have low repeatability, making it difficult to collect large amounts of labeled data. Traditional deep learning models (such as Faster R-CNN and YOLO) rely on a large number of labeled samples, but data acquisition in actual disaster scenarios is difficult, resulting in insufficient model generalization capabilities.
[0004] (2) Poor performance in sparse target detection. Key targets in building collapse scenarios (such as small cracks and partially exposed steel bars) are small, sparsely distributed, and often obscured by debris and smoke. Existing detection algorithms are prone to missed detections and false detections due to limited receptive fields or insufficient feature extraction.
[0005] (3) Insufficient modeling capabilities for complex scenes. Geographic scenes have variable lighting and cluttered backgrounds (e.g., twisted steel bars and cracked concrete), and scene-level analysis requires integration of 3D structural information (e.g., the stacking state of debris). Existing methods often rely on 2D images and lack the ability to fusion multimodal data and model global context. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a building collapse scene-level sparse target detection method based on small sample training.
[0007] First, a building collapse scene-level sparse target detection method based on small sample training is provided, including:
[0008] Step 1: Obtain multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets;
[0009] Step 2: Enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples;
[0010] Step 3: Build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module;
[0011] Step 4: Design a loss function with dual constraints of scene-level and pixel-level labels, and train the deep neural network model using the enhanced training data sample set;
[0012] Step 5: Perform target domain alignment on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
[0013] Preferably, in step 1, the multi-source satellite remote sensing images include panchromatic images and multispectral images; and the preprocessing includes radiometric calibration, FLASH atmospheric correction, orthorectification, geometric correction, registration, fusion and cropping.
[0014] Preferably, in step 2, enhancing the positive samples in the first training data sample set includes:
[0015] Before training, the number of positive and negative samples is counted, and random inversion and upsampling operations are performed on the positive samples to expand the number of positive samples to the same number of negative samples; the formula is as follows:
[0016] S total =S neg ∪S aug (S pos )
[0017] Among them, S total The training data sample set after enhancement consists of negative samples S neg and the enhanced positive sample S aug (S pos ) composition, S neg With S aug (S pos )The sample size is consistent.
[0018] Preferably, in step 3, the feature extraction module is used to extract deep semantic features of the image, the scene-level sparse feature attention module focuses on building collapse features through weight distribution, the Transformer-based classification network module performs global semantic modeling on the features, and the information reconstruction module outputs pixel-level and scene-level detection results.
[0019] Preferably, in step 4, the loss function includes pixel-level supervision loss and scene-level supervision loss, and balances the effects of the two through a weight coefficient. The formula is expressed as follows:
[0020] L total =λL pixel +(1-λ)L scene
[0021] Among them, L pixel is the pixel-level supervision loss, calculated by the binary cross entropy loss, L scene is the scene-level supervision loss, and λ is the loss weight.
[0022] Preferably, in step 5, performing target domain alignment on the test data sample set includes:
[0023] Principal component analysis is used to extract principal component basis vectors of the enhanced training data sample set and the test data sample set, an orthogonal domain transformation matrix is constructed to eliminate domain differences, and the source domain is projected onto the feature subspace of the target domain; the source domain and target domain are the test data sample set.
[0024] Preferably, in step 5, the accuracy index includes OA and Kappa index.
[0025] In a second aspect, a building collapse scene-level sparse target detection device based on small sample training is provided, which is used to execute any of the methods described in the first aspect, including:
[0026] The acquisition unit is used to acquire multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets;
[0027] The enhancement unit is used to enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples;
[0028] The construction unit is used to build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module;
[0029] A design unit, configured to design a loss function with dual constraints of scene-level and pixel-level labels, and train the deep neural network model using the enhanced training data sample set;
[0030] The input unit is used to perform target domain alignment processing on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
[0031] According to a third aspect, a computer storage medium is provided, wherein a computer program is stored in the computer storage medium; when the computer program is executed on a computer, the computer executes any one of the methods described in the first aspect.
[0032] In a fourth aspect, an electronic device is provided, including:
[0033] Memory, used to store computer programs;
[0034] A processor is used to execute the computer program to implement any method as described in the first aspect.
[0035] The beneficial effects of the present invention are as follows: the present invention adopts a small sample learning framework based on the attention mechanism. In the training stage, the present invention adjusts the training strategy to address the problem of only a few samples available for training due to the sparsity of targets in large-scale scenes. First, a variety of sample enhancement methods such as oversampling and random inversion are used to adjust the ratio of positive and negative samples in the input network to achieve a balanced state, thereby optimizing the model's learning effect on the data. The input data is first fed into the ResNet18 feature extraction network model, which can deeply mine the intrinsic semantic information of the sample and extract highly representative deep semantic features. Then, in response to the problem of sparse building collapse targets in the scene, these deep semantic features will be input into the scene-level sparse feature attention module. This module enhances the original deep semantic features of the building collapse target and reconstructs representative high-level semantic vector features by accurately calculating the attention weights of key features in the scene, further highlighting the feature information that has an important impact on the model's decision-making. Next, the reconstructed deep features and the original semantic features are jointly input into the Transformer-based cross attention module, combined with contextual information, to learn global information and obtain a richer and more accurate feature representation. Subsequently, these fused features are fed into the convolution-based Information Reconstruction Module (IRM). The IRM performs fine processing and reconstruction of the features through a series of convolution operations, and ultimately generates the model's detection results. To ensure that the model can output accurate results, the model adopts a dual constraint mechanism of pixel-level labels and scene-level labels. It supervises the network's learning process from both micro and macro levels, guiding the model to pay more attention to sparse building collapse targets, effectively improving the network's understanding and prediction capabilities for complex scenes. During the testing phase, in view of the differences in the characteristics of training data and test data, a domain alignment method based on principal component analysis is adopted to narrow the gap between domains and improve the model's detection effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 Flowchart of the building collapse scene-level sparse target detection method based on small sample training provided by the present invention;
[0037] Figure 2A schematic diagram of the structure of the scene-level sparse feature attention module provided by the present invention;
[0038] Figure 3 Schematic diagram of the Transformer-based classification deep network model provided by the present invention;
[0039] Figure 4 Visualization results for the building collapse detection example application; (a) image to be tested; (b) label; (c) detection result. DETAILED DESCRIPTION
[0040] The present invention will be further described below with reference to the following examples. The following examples are provided only to facilitate understanding of the present invention. It should be noted that, without departing from the principles of the present invention, it is possible for a person skilled in the art to make various modifications to the present invention, and such improvements and modifications fall within the scope of the claims of the present invention.
[0041] Example 1:
[0042] To solve the problems of the prior art, Example 1 of the present application provides a building collapse scene-level sparse target detection method based on small sample training, including:
[0043] Step 1: Obtain multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets.
[0044] Specifically, a batch of multispectral and panchromatic remote sensing images from domestically produced satellites were acquired. These images were then processed using radiometric calibration, FLASH atmospheric correction, orthorectification, and geometric correction. Low signal-to-noise ratio bands in the multispectral images were removed. A pan-sharpening algorithm was used to generate high-spatial-resolution multispectral remote sensing images. The processed images were then annotated at the scene level to construct and divide the initial training and test data sample sets.
[0045] The multispectral and panchromatic data come from different satellites, and the sources of multi-source data sets further increase the difficulty of cross-scene detection. The data source is screened by screening earthquake events in recent years. It is worth noting that the scene-level sparse target detection method for building collapse based on small sample training proposed in the present invention is a scene-level sparse target detection algorithm for building collapse based on single-phase remote sensing images, that is, the multi-source panchromatic image and multispectral image are finally fused to obtain a multispectral remote sensing image with high spatial resolution.
[0046] Step 2: Enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples.
[0047] Due to the serious imbalance in scene categories and the extremely sparse samples of building collapse scenes, this paper adopts a positive sample enhancement training strategy. Before training, the number of positive and negative samples is counted, and random inversion and upsampling are performed on positive samples to expand the number of positive samples to the same number of negative samples. The formula is as follows:
[0048] S total =S neg ∪S aug (S pos ) (1)
[0049] Among them, S total The training data sample set after enhancement consists of negative samples S neg and the enhanced positive sample S aug (S pos ) composition, S neg With S aug (S pos )The sample size is consistent.
[0050] Step 3: Build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module.
[0051] In step 3, the feature extraction module is used to extract deep semantic features of the image, the scene-level sparse feature attention module focuses on building collapse features through weight distribution, the Transformer-based classification network module performs global semantic modeling on the features, and the information reconstruction module outputs pixel-level and scene-level detection results.
[0052] Specifically, the feature extraction module uses an improved ResNet18 convolutional neural network. Compared with the original ResNet18 convolutional neural network, this application adds a final layer of 1×1 convolution and adjusts the output channel to 64. This model can deeply mine the intrinsic semantic information of the sample and extract highly representative deep semantic features. The calculation formula of the feature extraction module can be expressed as:
[0053] Y=ResNet18(X) (2)
[0054] Among them, Y represents the deep semantic features output by the ResNet18 convolutional neural network, is the input of the network, H and W represent the number of rows and columns of the input image respectively.
[0055] In addition, the scene sparse feature attention module extracts features and assigns weights to the input image, so that the model focuses on the important features of the building collapse target in the image and reduces attention to irrelevant features, thereby improving the performance and accuracy of the model. The specific process is as follows: Figure 2 As shown. The model first uses the ResNet18 deep network model to extract deep semantic features Transform the input Y into Then, point-wise convolution is used to obtain L semantic groups, each representing a semantic concept. We then use the Softmax function operating on the H / 4×W / 4 dimensions of each semantic group to calculate the spatial attention map. This step further refines the input features and homogenizes them into L regions. The attention weights of the building collapse features in each region are learned through point-wise convolution. Subsequently, the attention weights are matrix multiplied with the input deep features, and the attention map is used to calculate the weighted average sum of the pixels, enhancing the features of the building collapse area and providing support for subsequent Transformer-based classification tasks. The process can be expressed as follows:
[0056] Y z =(A) T Y′=(σ(φ(Y′;W))) T Y′ (3)
[0057] In the above formula, is the enhanced feature of sparse building collapse features, is the input deep semantic feature, φ is the point-wise convolution, and σ is the Softmax function. In this invention, L is set to 16.
[0058] Furthermore, the Transformer-based classification deep network model consists of a Transformer encoder and a decoder. The Transformer encoder consists of a multi-head self-attention (MSA) and a feedforward network, such as Figure 3 As shown. Unlike the original Transformer that uses post-norm residual units, this model follows the VisionTransformer feature and uses pre-norm residual units (batch normalization is performed before the Attention operation), that is, layer normalization is performed before MSA. Pre-norm has been proven to be more stable and superior than post-norm. The context between tokens of these building collapse enhancement features is modeled using the Transformer encoder. The starting point is that the Transformer can make full use of the global semantic relationship based on tokens to generate contextualized tokens representations for each small scene collapse feature. As shown Figure 4 As shown in (a), we first enhance the sparse collapsed feature Y zInput to the Transformer decoder, and finally obtain the feature Y that establishes the contextual relationship E The calculation process can be expressed as:
[0059] Q=Y z W q (4)
[0060] K=Y z W k (5)
[0061] V=Y z W v (6)
[0062]
[0063] MSA(Y z )=Concat(head1,...,head r )W o (8)
[0064] head i =Att(Y z W i q ,Y z W i k ,Y z W i v ) (9)
[0065] Y E =norm(FFN(norm(MSA(Y z )))) (10)
[0066] Where MSA stands for multi-head self-attention, r is the number of attention heads (r is set to 8 in this paper), FFN is a feedforward network, and norm is layer normalization. The Transformer decoder consists of a multi-head cross-attention (MA) and a feedforward network. Its function is to match the semantic enhancement of building collapse with the original deep-level features through cross-attention, obtaining the final global combined features for subsequent detection and reconstruction. Its calculation process can be expressed as:
[0067] MA(Y,Y z )=Concat(head1,...,head r )W o (11)
[0068] head i =Att(YWi q ,Y z W i k ,Y z W i v ) (12)
[0069] Y D =FC(norm(norm(FFN(norm(MA(Y,Y z )))))) (13)
[0070] Among them, Y D is the final decoding feature, FC is the fully connected layer, and MA is the multi-head cross attention.
[0071] In addition, the information reconstruction module is the biggest difference from other reconstruction modules in that it can obtain pixel-level detection results and scene-level detection results at the same time as reconstruction and decoding. It is specifically composed of 2 layers of 3×3 convolution and 2 layers of upsampling layers, which sequentially convert Y D Reconstruct the two-channel pixel detection results of the same size as the input image, then use the Argmax function to obtain the pixel-level detection results, and finally map them to the scene-level detection results through the fully connected layer and the Sigmoid layer. The calculation process can be expressed as:
[0072] P pixel =argmax(IRM(Y D )) (14)
[0073] P scene =sigmoid(FC(P pixel )) (15)
[0074] Among them, IRM is the information reconstruction module.
[0075] Step 4: Design a loss function with dual constraints of scene-level and pixel-level labels, and use the enhanced training data sample set to train the deep neural network model.
[0076] Step 5: Perform target domain alignment on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
[0077] Example 2:
[0078] Based on Example 1, Example 2 of the present application provides a more specific building collapse scene-level sparse target detection method based on small sample training, including:
[0079] Step 1: Obtain multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets.
[0080] In step 1, the multi-source satellite remote sensing image includes a panchromatic image and a multispectral image; the preprocessing includes radiometric calibration, FLASH atmospheric correction, orthorectification, geometric correction, registration, fusion and cropping.
[0081] Step 2: Enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples.
[0082] Step 3: Build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module.
[0083] Step 4: Design a loss function with dual constraints of scene-level and pixel-level labels, and use the enhanced training data sample set to train the deep neural network model.
[0084] In step 4, the loss function includes pixel-level supervision loss and scene-level supervision loss, which enables the model to supervise and guide at both the scene and pixel levels, making the model pay more attention to the sparse building collapse features, and balance the effects of the two through the weight coefficient, further solving the sample imbalance problem. The formula is expressed as follows:
[0085] L total =λL pixel +(1-λ)L scene (16)
[0086] Among them, L pixel is the pixel-level supervision loss, calculated by the binary cross entropy loss, L scene is the scene-level supervision loss, and λ is the loss weight. pixel and L scene The formula is defined as follows:
[0087]
[0088] Where N is the total number of pixels, is the predicted probability of the i-th pixel, is the true label of the i-th pixel, M is the total number of scenes, is the predicted value of the jth scene, is the true label of the jth scene.
[0089] Step 5: Perform target domain alignment on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
[0090] Specifically, in step 5, the cross-scene detection method based on principal component analysis and combined building range can eliminate the domain difference between the test data sample set and the training data sample set. The specific method is as follows: during the test, the test data sample set is first changed according to the data characteristics of the training set, so that the characteristics of the test data sample set are close to the training data sample set. Figure 1 As shown in the test phase, the method first targets the training data sample set and test data sample set The principal components are extracted by PCA to obtain the basis vectors of the target domain (training data sample set) and the source domain (test data sample set). Where d is the number of bands of the training set samples and the test sample, n t is the number of training set samples, n s is the number of test set samples, and the process is first decentralized:
[0091]
[0092] Next, we calculate the covariance matrix and perform eigendecomposition to obtain the basis vectors of the target and source domains:
[0093]
[0094] Among them, C t and C s are the covariance matrices of the training data sample set and the test data sample set, U t and U t are the corresponding eigenvector matrices, Λ t and Λ s are the corresponding eigenvalue diagonal matrices, B t and B s are the basis vectors of the target domain and the source domain respectively. Then, through domain transformation matrix learning, the difference between the source domain and the target domain basis is minimized to generate the orthogonal domain transformation matrix A. The process can be expressed as:
[0095]
[0096] Finally, the source domain (test data sample set) is projected onto the target domain (training data sample set) subspace through the domain transformation matrix. The process can be expressed as:
[0097]
[0098] Finally, the test data X′ after domain alignment s Input into the trained model for detection to achieve cross-scene detection.
[0099] In addition, in step 5, the OA and kappa indicators are used to evaluate the accuracy of the model.
[0100] For example, the effect of the present invention is further analyzed through specific experimental results. This application selects the image of the 4.5 magnitude earthquake in Yinchuan on November 20, 2022, obtained by Gaofen-7 for example detection application, with a size of 43240×53360. First, we preprocess the hyperspectral image, align and transform the domain, and finally cut and input it into the trained model. The final detection results are shown in Table 1:
[0101] Table 1 Figure 4 Quantitative accuracy evaluation of sparse target change event detection in building collapse images
[0102]
[0103] Figure 4 Table 1 shows the experimental visualization results and quantitative accuracy evaluation of the selected areas in the large-scale building collapse sparse target change event detection dataset. Figure 4 It can be seen that the proposed model can obtain better detection results (overall accuracy (OA) and Kappa coefficient) for sparse target detection at the building collapse scene level, and can achieve relatively accurate building collapse scene detection when the target is so sparse. In summary, the present invention provides a building collapse scene-level sparse target detection method based on small sample training, which can achieve high-precision building collapse scene-level sparse target detection in large-scene remote sensing images and has important application value. It should be understood that the parts not elaborated in detail in this specification belong to the prior art.
[0104] It should be noted that the parts in this embodiment that are the same or similar to those in Example 1 can be referenced to each other and will not be described in detail in this application.
[0105] Example 3:
[0106] Based on Example 2, Example 3 of the present application provides a building collapse scene-level sparse target detection device based on small sample training, including:
[0107] The acquisition unit is used to acquire multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets;
[0108] The enhancement unit is used to enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples;
[0109] The construction unit is used to build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module;
[0110] A design unit, configured to design a loss function with dual constraints of scene-level and pixel-level labels, and train the deep neural network model using the enhanced training data sample set;
[0111] The input unit is used to perform target domain alignment processing on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
[0112] It should be noted that the device provided in this embodiment is a device corresponding to the method provided in Example 2. Therefore, the parts in this embodiment that are the same or similar to those in Example 2 can be referenced to each other and will not be repeated in this application.
[0113] In summary, the present invention's scene-level sparse building collapse object detection method based on small-sample training focuses on: 1) scene-level building collapse target event detection technology; 2) training strategy adjustments to address small sample and class imbalance issues; 3) algorithmic models for handling the sparsity of building collapse target events; and 4) research on cross-scene detection issues. The research goal is to develop a method for detecting and identifying sparsely distributed building collapse targets in single-temporal remote sensing imagery.
Claims
1. A building collapse scene-level sparse target detection method based on small sample training, characterized by: include: Step 1: Obtain multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets; Step 2: Enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples; Step 3: Build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module; Step 4: Design a loss function with dual constraints of scene-level and pixel-level labels, and train the deep neural network model using the enhanced training data sample set; Step 5: Perform target domain alignment on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
2. The building collapse scene-level sparse target detection method based on small sample training according to claim 1 is characterized in that: In step 1, the multi-source satellite remote sensing image includes a panchromatic image and a multispectral image; the preprocessing includes radiometric calibration, FLASH atmospheric correction, orthorectification, geometric correction, registration, fusion and cropping.
3. The building collapse scene-level sparse target detection method based on small sample training according to claim 2 is characterized in that: In step 2, the enhancing of positive samples in the first training data sample set includes: Before training, the number of positive and negative samples is counted, and random inversion and upsampling operations are performed on the positive samples to expand the number of positive samples to the same number of negative samples; the formula is as follows: S total =S neg ∪S aug (S pos ) Among them, S total The training data sample set after enhancement consists of negative samples S neg and the enhanced positive sample S aug (S pos ) composition, S neg With S aug (S pos )The sample size is consistent.
4. The building collapse scene-level sparse target detection method based on small sample training according to claim 3 is characterized in that: In step 3, the feature extraction module is used to extract deep semantic features of the image, the scene-level sparse feature attention module focuses on building collapse features through weight distribution, the Transformer-based classification network module performs global semantic modeling on the features, and the information reconstruction module outputs pixel-level and scene-level detection results.
5. The building collapse scene-level sparse target detection method based on small sample training according to claim 4 is characterized in that: In step 4, the loss function includes pixel-level supervision loss and scene-level supervision loss, and balances the effects of the two through weight coefficients. The formula is expressed as follows: THE total =λL pixel +(1-λ)L scene Among them, L pixel is the pixel-level supervision loss, calculated by the binary cross entropy loss, L scene is the scene-level supervision loss, and λ is the loss weight.
6. The building collapse scene-level sparse target detection method based on small sample training according to claim 5 is characterized in that: In step 5, the target domain alignment processing is performed on the test data sample set, including: Principal component analysis is used to extract principal component basis vectors of the enhanced training data sample set and the test data sample set, an orthogonal domain transformation matrix is constructed to eliminate domain differences, and the source domain is projected onto the feature subspace of the target domain; the source domain and target domain are the test data sample set.
7. The building collapse scene-level sparse target detection method based on small sample training according to claim 6 is characterized in that: In step 5, the accuracy indicators include OA and Kappa indicators.
8. Building collapse scene-level sparse target detection device based on small sample training, characterized by: Used to perform the method according to any one of claims 1 to 7, comprising: The acquisition unit is used to acquire multi-source satellite remote sensing images, preprocess them to generate high-resolution multispectral images, perform scene-level annotation, and construct initial training data sample sets and test data sample sets; The enhancement unit is used to enhance the positive samples in the initial training data sample set, balance the ratio of positive and negative samples, and obtain an enhanced training data sample set with a balanced ratio of positive and negative samples; The construction unit is used to build a deep neural network model, including a feature extraction module, a scene-level sparse feature attention module, a Transformer-based classification network module, and an information reconstruction module; A design unit, configured to design a loss function with dual constraints of scene-level and pixel-level labels, and train the deep neural network model using the enhanced training data sample set; The input unit is used to perform target domain alignment processing on the test data sample set, input the processed test data set into the trained model for detection, obtain the detection results of the building collapse area, and obtain the vector coordinates of the building collapse area, and then evaluate the model based on the accuracy index.
9. A computer storage medium, characterized in that The computer storage medium stores a computer program; when the computer program is run on a computer, the computer executes the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Collapsed building detection method and system based on cross domain teacher-student mutual training
CN116778335A
Three-dimensional reconstruction method, apparatus and device, and computer readable storage medium
CN116977548A
Remote sensing image change detection method and system
CN117745673A
Building roof automatic extraction method and system suitable for high-scene satellite
CN118691984A
Semantic scene migration-based post-earthquake damaged building identification method and system
CN118968281A