High-resolution remote sensing image construction site and state accurate extraction method
Through multi-agent learning and full convolutional neural network, combined with bare soil and board room spectral extraction index, the problems of low efficiency and insufficient accuracy of construction sites and state extraction in the existing technology are solved, and efficient and accurate construction sites and state recognition and extraction are achieved.
Patent Information
- Application Number
- CN202510005220.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is inefficient, subjective and costly in the extraction of high-resolution remote sensing images in construction sites and states, and cannot meet the requirements of fast and accurate data extraction.
By analyzing the element characteristics of the construction site, the bare soil spectrum extraction index KVL and the board room spectrum extraction index CKL are constructed, and combined with the pixel-level classification of full convolutional neural multi-agents, the efficient extraction of the construction site and state is achieved.
It realizes high efficiency and high-precision identification and extraction of construction sites and states, reduces the dependence of manual interpretation, and improves the accuracy and efficiency of data extraction.
Smart Images

Figure CN120088666A_ABST
Abstract
Description
Technical Field
[0001] This application relates to a method for extracting construction sites from remote sensing, in particular to a method for accurately extracting construction sites and their states from high-resolution remote sensing images, belonging to the technical field of remote sensing image recognition. Background Art
[0002] With the accelerating urbanization process, the scale of infrastructure construction has been increasing year by year, which will inevitably bring hazards such as environmental pollution, noise pollution, and surface damage, having many adverse effects on people's daily lives. Rapidly extracting construction sites and their states provides a basis for pipeline safety inspections, land use safety law enforcement, and construction progress control, and has an important impact on the planning and development of engineering construction.
[0003] In today's era of rapid development of data technology, how to quickly and accurately obtain ground spatio-temporal data has become a key research issue in the field of geospatial data, and the requirements for its currency, real-time nature, and accuracy are also getting higher and higher. UAV low-altitude remote sensing technology can fly at low altitude under any complex terrain conditions and under clouds to obtain high-resolution images. It also has the advantages of low cost, simple and convenient operation, and can take off at any time according to the task schedule. Therefore, this application combines high-resolution remote sensing images and UAV images as the data source for extracting construction sites and their states to achieve the extraction of construction sites and their states.
[0004] There has been little research on the extraction direction of construction sites and their states in the prior art. Most of them still rely mainly on visual interpretation of high-resolution images, manually interpreting the construction sites and their states in the images, and then completing tasks such as construction site work progress and construction site safety monitoring based on the interpretation results. The visual interpretation process has relatively high requirements for the professional background knowledge and strict logical thinking ability of the interpreters. When detecting construction sites over long distances and large areas, the high-resolution images used are massive. The method of manual interpretation is inefficient, subjective, and costly, and cannot meet the requirements of data extraction tasks. Therefore, this application starts from analyzing the extraction elements in construction sites and their image features, and develops a method for extracting construction sites and their states from massive high-resolution remote sensing images using the feature characteristics of construction site extraction elements.
[0005] In image extraction and classification, compared with traditional mode extraction, the biggest difference in multi-agent learning is that it can automatically learn features from big data. By performing layer-by-layer feature transformation on the initial signal, the sample features are transformed from the original feature space to a new feature space, and hierarchical feature representations are automatically learned, which is more conducive to image classification and extraction detection. By constructing appropriate multi-agents and optimization strategies, the extraction of target data can be achieved.
[0006] Problems to be solved in the extraction of construction sites and their status from remote sensing images of the prior art and key technical difficulties of this application include:
[0007] (1) In the prior art, there is little research on the extraction direction of construction sites and their status. Most of them still rely on visual interpretation of high-resolution images. Manual interpretation is used to judge the construction sites and their status in the images, and then tasks such as the work progress of the construction site and the safety monitoring of the construction site are completed based on the judgment results. The visual interpretation process has high requirements for the professional background knowledge and strict logical thinking ability of the interpreters. When detecting construction sites over long distances and large areas, a large amount of high-resolution images are used. The method of manual interpretation is inefficient, subjective, and costly, and cannot meet the requirements of data extraction tasks. Therefore, starting from analyzing the extraction elements and their image features in the construction site, this application develops a method for extracting construction sites and their status from a large amount of high-resolution remote sensing images by using the feature of the extraction elements of the construction site.
[0008] (2) The prior art lacks an extraction model structure corresponding to the extraction elements of the construction site, lacks a training network for high-resolution remote sensing image data, and lacks a set of high-quality algorithms for extracting construction sites and their status from high-resolution images, and cannot achieve the rapid extraction of construction sites from high-resolution remote sensing images. There is a lack of extraction and analysis of the extraction elements and their image features of the construction site. For the extraction of the construction site, there is a lack of analysis of the interpretation signs such as the color, shape, and texture of the ground object data in the construction site, and the feature analysis of the extraction elements of the construction site is carried out from three aspects: bare soil, construction machinery, and construction board houses. It is impossible to extract different indices by using the features of the extraction elements on the image, and it is impossible to increase learning features for the training of construction site extraction, and the extraction accuracy is poor. The prior art does not select appropriate cost functions and loss functions, and updating parameters by adjusting the learning rate can greatly improve the learning rate.
[0009] (3) There are many machine learning models in the prior art. Each network structure is different, and there are many parameters involved. Different structures also have different extractions for a certain target. Therefore, research on selecting a suitable network structure for the extraction of construction sites and their status is required. The prior art lacks the analysis of the characteristics of each typical structure of machine learning for the extraction elements of the construction site, and does not compare and analyze their adaptability. The prior art lacks an effective machine learning model for extracting construction sites and their status from remote sensing images, lacks a suitable output excitation model and overfitting prevention strategy, and cannot use the combination of high-resolution remote sensing images and UAV images as the data source for extracting construction sites and their status, and cannot achieve high-efficiency and high-precision recognition and extraction of construction sites and their status. Summary of the Invention
[0010] This application adaptively learns the target features in remote sensing big data, and based on the analysis and calculation of the extraction elements of the construction site, establishes a method for extracting the construction site and its status from high-resolution images based on multi-agent learning. First, analyze the element features of the construction site on high-resolution remote sensing images, determine the extraction elements, and based on the performance characteristics of each extraction element on the image, establish the features for construction site extraction, and construct the bare soil spectral extraction index KVL; considering the characteristics that the construction site is an irregular extraction target, use the fully convolutional neural multi-agent for pixel-level classification; analyze the input data and output data features in the fully convolutional multi-agent, crop and screen the multi-source high-resolution images to obtain k 500*500 pixel image data, select m images as the experimental dataset and perform preprocessing, construct the fully convolutional neural multi-agent for construction site and status extraction, establish the construction site and status extraction process method, and use the combination of high-resolution remote sensing images and UAV images as the data source for construction site and status extraction, realizing the high-efficiency, high-precision identification and extraction of the construction site and its status.
[0011] To achieve the above technical effects, the technical solutions adopted in this application are as follows:
[0012] A method for accurately extracting the construction site and its status from high-resolution remote sensing images, which adaptively learns the target features in remote sensing big data, and based on the analysis and calculation of the extraction elements of the construction site, establishes a method for extracting the construction site and its status from high-resolution images based on multi-agent learning: First, analyze the element features of the construction site on high-resolution remote sensing images, determine the extraction elements, and based on the performance characteristics of each extraction element on the image, establish the features for construction site extraction, and construct the bare soil spectral extraction index KVL; considering the characteristics that the construction site is an irregular extraction target, use the fully convolutional neural multi-agent for pixel-level classification; analyze the input data and output data features in the fully convolutional multi-agent, crop the multi-source high-resolution images into k 500*500 pixel image data, select m images as the experimental dataset and perform preprocessing, construct the fully convolutional neural multi-agent for construction site and status extraction, and establish the construction site and status extraction process method;
[0013] (1) Construct the feature factors for construction site status extraction: Analyze and calculate the extraction elements of the construction site, determine three major extraction elements: bare soil, construction board houses, and construction machinery, calculate and analyze the spectral features, texture features, and image occupancy ratios of each extraction element in the image, compare and fuse the extraction algorithms for the features of each extraction element, establish the construction site feature extraction algorithm, calculate the reflection spectral curves of different types of soil based on the board house spectral extraction index CKL, and establish the bare soil spectral extraction index KVL;
[0014] (2) Constructing a fully convolutional multi-agent for construction site and state extraction: A fully convolutional neural multi-agent for pixel-level classification is used. Based on the irregular contour shape of the construction site, the initial image spatial structure data is retained for pixel-level classification. The fully connected layers ZN6 and ZN7 after the 5th pooling structure are converted into convolutional layers conv6 and conv7. The output is converted into 4096 feature maps, where ZN6 and ZN7 are fully connected layers containing 4096 neurons, and conv6 and conv7 are convolutional layers containing 4096 convolution kernels with a size of l×1. Finally, a convolutional layer with 2 convolution kernels is used to regress the 4096 vectors of each node into a 2-dimensional target category. The category confidence of each node is predicted and normalized. The size of the feature map obtained is proportional to the input construction site sample image.
[0015] (3) Construction site and state extraction based on fully convolutional multi-agent: The obtained high-resolution images are cropped and screened as experimental data sets and preprocessed accordingly. By outputting the incentive model, preventing overfitting strategy, and agent updating strategy, an optimization strategy is established for the multi-agent. A fully convolutional neural multi-agent is constructed for construction site and state extraction, and a construction site and state extraction process is established.
[0016] Preferably, multi-agent learning feature extraction is used for learning and extraction of construction machinery. For construction board houses and bare soil areas, the board house spectrum extraction index CKL and soil extraction index KVL are used for image feature enhancement. The processed image feature map is used as the input of the multi-agent to enhance the feature learning ability of the multi-agent.
[0017] Constructing the bare soil spectrum extraction index KVL: Extracting the homogeneous characteristics of the soil reflectance spectrum in the visible light band, the reflectivity increases with the increase of wavelength, and the increase is significant. The red band is amplified to establish the bare soil spectrum extraction index KVL, which is expressed as formula 1:
[0018] KVL=(RB)+(RG) Formula 1
[0019] Among them, R, G, and B refer to the reflectivity values of the red, green, and blue bands of visible light, respectively.
[0020] Preferably, the spectral extraction index CKL of the board room is constructed: the spectral reflectance of the construction board room is low in the blue, green and red bands of visible light, and the spectral extraction index CKL of the board room is constructed accordingly:
[0021] CKL=(BG)+(BR) Formula 2
[0022] Among them, R, G, and B refer to the reflectance values of the red, green, and blue bands of visible light, respectively. The CKL index is aimed at the special color data of the construction board room. It normalizes the color bands and amplifies the characteristics of the target object to achieve differentiated extraction.
[0023] Preferably, construct a full convolution multi-agent for construction site and state extraction: establish three full convolution multi-agents from coarse to fine, and name the convolution agents according to the receptive field of a single pixel in the output image, namely: ZNT-32s, ZNT-16s, ZNT-8s. The numbers in the names represent the receptive field range of a single pixel in the output image. The smaller the range, the higher the accuracy of the network for contour extraction. The receptive field size corresponding to a single pixel in the output image of ZNT-32s is 32×32 (that is, the output pixel unit corresponds to an image block of 32×32 in the initial image). It is composed of 16 convolution layers, 15 ReLU activation layers, 5 pooling structures, 2 dropout layers, 1 deconvolution layer and 1 cropping layer, a total of 40 layers;
[0024] The output layer of the ZNT-32s network is a convolution layer. The value range of each pixel value in the output image is between 0 and 1, representing the probability that it belongs to the construction site. There are 5 pooling structures in the whole network, so that a single pixel in the output image corresponds to 32×32 pixels in the initial image. The size of the output image is 1 / 32 of the original image. The network output is upsampled by the deconvolution layer using the bilinear interpolation method to obtain the output result with the same size as the input image;
[0025] On the basis of ZNT-32s, the network is further improved to improve the resolution. A single pixel in the output result of the pooling structure pool4 of ZNT-16s corresponds to 16×16 pixels in the initial image. First, the conv7 convolution layer is upsampled by 2 times, and then fused with the pool4 layer to obtain the network output result, which not only retains the detailed data in the pooling structure pool4, but also reaches the image classification accuracy of the ZNT-32s network.
[0026] Preferably, remote sensing data set collection and processing: for UAV image data, according to the camera parameters (camera principal point, focal length, distortion parameters, etc.) obtained by camera calibration, correct the principal point offset and distortion of the obtained UAV images, and then perform color homogenization processing on the UAV images to ensure the accuracy of image stitching. Secondly, use image matching software for feature-based image matching, including establishing a scale space, feature point positioning, calculation of the main angle of feature points, calculation of feature descriptors, and then perform aerial triangulation. Use the camera parameter file and image point coordinate file to solve the exterior orientation elements of each image through bundle adjustment in the region. Finally, use DEM generation software to calculate the corresponding ground discrete point coordinates according to the matched homologous points and the solved image orientation elements, generate DEM from the ground discrete points, and resample to generate an orthophoto image;
[0027] Perform pixel-level calibration. The cropped image size is 500×500 pixels. Manually calibrate the samples of the cropped image. The specific calibration standard for the sample dataset is to construct a label map with the same size as the original image, where each pixel corresponds one-to-one with the pixel at the corresponding position in the original image. If the pixel area in the original image belongs to the construction site, the corresponding label image is marked white; otherwise, it is marked black, and the marked map corresponding to the original image is obtained.
[0028] Preferably, output excitation model: In the extraction of construction site and status, a simple and fast non-linear output excitation model is used to achieve non-linear transformation:
[0029]
[0030] x and y are the input and output of the model respectively, reducing the gradient dispersion generated during network training and enabling the network to introduce sparsity by itself.
[0031] Preferably, overfitting prevention strategy: Modify the agent itself. In the training of a multi-agent, randomly update the agent in each iteration, and use the introduced randomness to increase the generalization ability of the network. During the network training stage, temporarily discard the neurons of the network from the network with a certain probability, keep the input and output layers unchanged, and update the weights in the multi-agent according to the error backpropagation algorithm to prevent the mutual adaptation between nodes.
[0032] Preferably, assume that the entire network has n parameters, and the selection of each parameter is a fixed probability, then there are 2 n kinds of selection possibilities, corresponding to 2 n sub-networks. When n is very large, the iteratively updated sub-networks do not repeat, avoiding the situation that a certain network is overfitted to the training set;
[0033] Multiply the output during testing by the fixed probability p used during training. Through one test, all the parameters of the original 2 n sub-networks are considered. Assume that x and y are the input and output respectively, and w is the weight parameter of the previous layer. w| p is the subset of weight parameters obtained by sampling with a fixed probability p, expressed as Equation 4:
[0034]
[0035] When writing the program, use y = w·p*x = w′*x during the testing stage, and modify the formula during training to Equation 5:
[0036]
[0037] The weight parameters obtained through training are w'.
[0038] Preferably, the agent update strategy: for a set of training samples data = x, label = y, define a prediction strategy such that F θ (x) is as close as possible to y, that is, the loss function L[y, F θ (x)] is as small as possible. The loss function is the second norm, expressed as Equation 6:
[0039] L[y, F θ (x)] = (y - F θ (x)) 2 Equation 6
[0040] During the training process, continuously modify the parameter θ in F to make L continuously decrease:
[0041] ΔL ≈ dL Equation 7
[0042]
[0043] Get:
[0044]
[0045] To make △L < 0, only need:
[0046]
[0047] Add a learning factor α (α > 0) to make the learning of θ smoother, get:
[0048]
[0049] Set the learning factor using empirical values, and then correct the learning factor in the test experiment;
[0050] Place the fully convolutional multi-agent in the Pascal VOC data sets natural image data set for pre-training to obtain the representation ability of the target, and then use the trained model weight parameters as the initial parameters of the construction site and state extraction multi-agent, and perform learning and training on the collected high-resolution remote sensing image data.
[0051] Preferably, the training of the fully convolutional multi-agent: First, place the fully convolutional multi-agent on the Pascal VOC datasets natural image dataset for pre-training on the image segmentation task. The Pascal VOC data sets dataset does not contain the construction site target to be extracted. Then, use the collected high-resolution remote sensing image as the dataset and send it into the multi-agent obtained by pre-training for learning the construction site. During the training process, use the parameters of the multi-agent obtained by pre-training as the initial parameters of the multi-agent for training the construction site:
[0052] Step 1: Preprocess the sample data set by cropping the picture frames. Crop the high-resolution remote sensing images into a sample set of 500×500 pixels so as to feed them into the network for training and learning;
[0053] Step 2: Manually calibrate the construction sites on the collected high-resolution remote sensing images;
[0054] Construct a preliminary data set through the above two steps. Select images with high quality and suitable scenes as the training data set. Each high-resolution remote sensing image in the data set corresponds to a binary image containing 0 and 1 as the label;
[0055] Organize the data set in the format of Pascal VOC data sets. The data set is distributed in three folders: JPEGImages, which stores the initial image data in JPEG format; SegmentationClass, which stores the label data corresponding to the initial images. The corresponding file names are the same, but stored in PNG format; ImageSets\Segmentation, which stores the data files of the picture names of each image, stored by text application files. This part contains two text files, train.txt and val.txt, corresponding to the file names of the training images and test images respectively. Input the classified data set into the constructed fully convolutional multi-agent for agent learning;
[0056] Adopt a strategy of progressive training from coarse to fine. First, train the ZNT-32s network. Use the stochastic gradient descent method to realize the iterative update of the network weights. The training samples used are k images. Perform k*1, k*2, …, k*30 iterations on the training samples. In addition, m images are added for testing. The accuracy is the best when the number of iterations of the experimental data is 3 times the amount of experimental data. Set the number of iterations in the experiment to 3*k times;
[0057] After training the ZNT-32s network, assign the weights of the ZNT-32s network to the ZNT-16s, and let the ZNT-16s perform iterative training based on the ZNT-32s. Similarly, perform 3*k iterations. Similarly, the ZNT-8s is iteratively updated based on the training results of the ZNT-16s. Finally, three trained networks are obtained, realizing the training of multi-agents with three levels of fineness from coarse to fine. Then select the network with the best experimental effect, add the construction shed spectral extraction index CKL and the bare soil spectral extraction index KVL, and perform iterative update based on the training results of this network to train the multi-agent for extracting construction sites and their states proposed in this application.
[0058] Compared with the prior art, the innovation points and advantages of this application are:
[0059] (1) This application adaptively learns the target features in remote sensing big data, and based on the analysis and calculation of the extraction elements of the construction site, establishes a method for extracting the construction site and its status from high-resolution images based on multi-agent learning. First, analyze the element features of the construction site on high-resolution remote sensing images, determine the extraction elements, and based on the performance characteristics of each extraction element on the image, establish the features for construction site extraction, and construct the bare soil spectral extraction index KVL; considering the characteristic that the construction site belongs to an irregular extraction target, use the fully convolutional neural multi-agent for pixel-level classification; analyze the input data and output data features in the fully convolutional multi-agent, crop the multi-source high-resolution images into k 500*500 pixel image data, screen out m images as the experimental data set and perform preprocessing, construct the fully convolutional neural multi-agent for construction site and status extraction, establish the extraction process method for the construction site and its status, and use the combination of high-resolution remote sensing images and UAV images as the data source for construction site and status extraction, realizing the high-efficiency and high-precision identification and extraction of the construction site and its status.
[0060] (2) This application constructs the feature factors for construction site status extraction, analyzes and calculates the extraction elements of the construction site, and determines three major extraction elements: bare soil, construction board houses, and construction machinery. Calculate and analyze the spectral features, texture features, and image occupancy ratios of each extraction element in the image, compare and fuse the extraction algorithms for the features of each extraction element, establish the construction site feature extraction algorithm, calculate the reflection spectral curves of different types of soil based on the board house spectral extraction index CKL, establish the bare soil spectral extraction index KVL, and verify the experimental data with this index. The result obtained by feeding the construction board house spectral extraction index CKL and the bare soil spectral extraction index KVL into the fully convolutional multi-agent on the basis of the fully convolutional multi-agent is better than the extraction effect obtained by directly feeding the original image into the fully convolutional multi-agent.
[0061] (3) This application designs a multi-agent for construction site and status extraction, and establishes a complete technical process to apply the fully convolutional multi-agent for pixel-level classification to construction site and status extraction. Designs the number of network layers, the number of nodes, etc. of the fully convolutional neural multi-agent, and selects the appropriate output excitation model, agent update strategy, etc. for multi-agent optimization through the research of optimization strategy methods. Use the board house spectral extraction index CKL and the bare soil spectral extraction index KVL as the input of the multi-agent to improve the construction site extraction accuracy, construct a complete technical process for high-resolution image construction site and status extraction based on multi-agent learning, and realize the accurate and rapid extraction of the construction site and its status using high-resolution remote sensing images.
[0062] (4) This application constructs characteristic factors for extracting the state of the construction site. Based on the spectral extraction index CKL of the plank house, the reflected spectral curves of different types of soil are calculated, and the bare soil spectral extraction index KVL is established. A full convolution multi-agent for extracting the construction site and its state is constructed. For the extraction of the construction site and its state based on the full convolution multi-agent, through the output excitation model, overfitting prevention strategy, and agent update strategy, an optimization strategy for the multi-agent is established. A full convolution neural multi-agent for extracting the construction site and its state is constructed, and a flow scheme for extracting the construction site and its state is established. The experiment and analysis of the construction site and its state extraction are completed, proving that the algorithm proposed in this application has achieved the best results in terms of accuracy, recall rate, and comprehensive evaluation index, and can well realize the extraction of the construction site and its state. Description of the Drawings
[0063] Figure 1 are the reflected spectral curve diagrams of several different types of soil.
[0064] Figure 2 is an experimental example diagram of the soil recognition index KVL.
[0065] Figure 3 is a schematic diagram of the multi-agent network construction structure of this application.
[0066] Figure 4 is a schematic diagram of the remote sensing data set acquisition, processing, and calibration process.
[0067] Figure 5 is a comparison diagram before and after the remote sensing data set acquisition, processing, and calibration.
[0068] Figure 6 is a schematic diagram of the construction site extraction results of ZNT-8s, ZNT-16s, and ZNT-32s.
[0069] Figure 7 is a schematic diagram of the construction site extraction results of the algorithm of this application and ZNT-8s.
[0070] Figure 8 is a precision comparison diagram of the construction site extraction results of the algorithm of this application and ZNT-8s, ZNT-16s, and ZNT-32s. Detailed Embodiment
[0071] Next, in combination with the drawings, the technical solutions of the method for accurately extracting the construction site and its state from high-resolution remote sensing images provided by this application will be further described, so that those skilled in the art can better understand this application and be able to implement it.
[0072] This application adaptively learns the target features in remote sensing big data, and establishes a method for extracting construction sites and their states from high-resolution images based on multi-agent learning on the basis of parsing and calculating the extraction elements of construction sites: First, analyze the element features of construction sites on high-resolution remote sensing images, determine the extraction elements, and establish features for construction site extraction based on the performance features of each extraction element on the image, and construct the bare soil spectral extraction index KVL; considering the characteristic that the construction site is an irregular extraction target, use the fully convolutional neural multi-agent for pixel-level classification; analyze the input data and output data features in the fully convolutional multi-agent, crop the multi-source high-resolution images into k 500*500 pixel image data, screen out m images as the experimental dataset and perform preprocessing, construct a fully convolutional neural multi-agent for extracting construction sites and their states, and establish a method for extracting construction sites and their states.
[0073] (1) Construct the feature factors for extracting the construction site state: Parse and calculate the extraction elements of the construction site, determine the three major extraction elements: bare soil, construction board houses, and construction machinery, calculate and analyze the spectral features, texture features, and image occupancy ratios of each extraction element in the image, compare and fuse the extraction algorithms for the features of each extraction element, establish the construction site feature extraction algorithm, calculate the reflection spectral curves of different types of soils based on the board house spectral extraction index CKL, and establish the bare soil spectral extraction index KVL;
[0074] (2) Construct the fully convolutional multi-agent for extracting construction sites and their states: Use the fully convolutional neural multi-agent for pixel-level classification. Based on the irregular contour shape of the construction site, retain the initial image spatial structure data for pixel-level classification, convert the fully connected layers ZN6 and ZN7 after the 5th pooling structure into convolutional layers conv6 and conv7, and the output is also converted from the original one-dimensional vector with a length of 4096 to 4096 feature maps. Among them, ZN6 and ZN7 are fully connected layers containing 4096 neurons, and conv6 and conv7 are convolutional layers containing 4096 convolutional kernels with a size of l×1. Finally, through a convolutional layer containing 2 convolutional kernels, the 4096 vectors of each node are regressed into a 2-dimensional target category, the class confidence of each node is predicted and normalized, and the size of the obtained feature map is proportional to the input construction site sample image.
[0075] (3) Extraction of construction sites and their states based on the fully convolutional multi-agent: Crop the obtained high-resolution images into k image data, screen out m images as the experimental dataset and perform corresponding preprocessing. Through the output excitation model, overfitting prevention strategy, and agent update strategy, an optimization strategy for the multi-agent is established, a fully convolutional neural multi-agent for extracting construction sites and their states is constructed, and a method for extracting construction sites and their states is established.
[0076] (4) Construction Site and State Extraction Experiment and Analysis
[0077] A construction site and state extraction process plan is established. Multiple full convolution multi-agent algorithms are trained and tested on the selected sample data set, with 645 images used for training and 633 images used for testing. By analyzing the test results and accuracy evaluation results of multiple experiments, the multi-agent parameters are adjusted and optimized accordingly to determine the construction site and state extraction algorithm model of this application.
[0078] I. Construction of Feature Factors for Extracting Construction Site State
[0079] The construction of feature factors for extracting construction site state includes: bare soil, construction board houses, and construction machinery. The KVL, a spectral extraction index for bare soil, is constructed to extract the bare soil part in the construction site. Based on the low proportion of construction machinery in the image data but its distinguishable external contour and irregular geometric shape, multi-agent learning feature extraction is used for the learning and extraction of construction machinery. For the construction board houses and bare soil areas, the CKL, a spectral extraction index for board houses, and the KVL, a soil extraction index, are used to enhance the image features. The processed image feature map is used as the input of the multi-agent to enhance the feature learning ability of the multi-agent.
[0080] (I) Construction of the KVL Spectral Extraction Index for Bare Soil
[0081] Extract the homogeneous characteristics of the soil reflection spectrum in the visible light band. The reflectance increases with the increase of wavelength and the increase is significant. The reflection spectral curve is as Figure 1 shown.
[0082] The red band is amplified to establish the KVL spectral extraction index for bare soil, expressed as Equation 1:
[0083] KVL = (R - B) + (R - G) Equation 1
[0084] Where R, G, and B respectively refer to the reflectance values of the visible light red, green, and blue bands.
[0085] To verify the effectiveness of this index, this application calculates the KVL index for 350 sample data and obtains the feature map results as Figure 2 shown. The 350 initial image sample data containing bare soil in the construction site are calculated with the image feature map after being processed by the KVL index. The results show that the accuracy rate of using the KVL index for extracting the construction board house area can reach 94%, which can be used as an effective method for extracting bare soil in the construction site.
[0086] (II) Construction of the CKL Spectral Extraction Index for Board Houses
[0087] The spectral reflectance of construction prefabricated houses is low in the visible light blue, green, and red bands. Based on this, the spectral extraction index CKL of the prefabricated houses is constructed:
[0088] CKL = (B - G) + (B - R) Equation 2
[0089] Among them, R, G, and B respectively refer to the reflectance values in the visible light red, green, and blue bands. The CKL index is for the special color data of construction prefabricated houses, and performs a normalization operation on the color bands to amplify the characteristics of the target object to achieve differential extraction.
[0090] II. Constructing a full-convolutional multi-agent for construction site and status extraction
[0091] Based on the irregular contour shape of the construction site, the initial image spatial structure data is retained for pixel-level classification. The fully connected layers ZN6 and ZN7 after the 5th pooling structure are converted into convolutional layers conv6 and conv7, and the output is also converted from a one-dimensional vector with a length of 4096 to 4096 feature maps. Among them, ZN6 and ZN7 are fully connected layers containing 4096 neurons, and conv6 and conv7 are convolutional layers containing 4096 convolutional kernels with a size of l×1. Finally, through a convolutional layer containing 2 convolutional kernels, the 4096 vectors of each node are regressed into a 2D target category, the class confidence of each node is predicted and normalized, and the size of the obtained feature map is proportional to the input construction site sample image.
[0092] The full-convolutional multi-agent has two obvious advantages: one is that it can input images of any size, that is, the training and test images do not have to have the same size. The second is to avoid the storage duplication problem caused by the convolutional operation in pixel block calculation, making the training process more efficient. The feature of the full-convolutional multi-agent is that the output result is a two-dimensional image and has spatial symmetry with the input image, so that a pixel on the output feature map corresponds to the pixel at the corresponding position of the input image, and the spatial feature is well retained.
[0093] Three full-convolutional multi-agents from coarse to fine are established. The convolutional agents are named according to the receptive field of a single pixel of the output image, namely: ZNT-32s, ZNT-16s, ZNT-8s. The numbers in the name represent the receptive field range of a single pixel of the output image. The smaller the range, the higher the accuracy of the network for contour extraction. The receptive field size corresponding to a single pixel of the output image of ZNT-32s is 32×32 (that is, the output pixel unit corresponds to an image block of 32×32 in the initial image), and it is composed of a total of 40 layers including 16 convolutional layers, 15 ReLU activation layers, 5 pooling structures, 2 dropout layers, 1 deconvolutional layer, and 1 cropping layer.
[0094] Figure 3The output layer of the ZNT-32s network is a convolutional layer. The value range of each pixel of the output image is between 0 and 1, representing the probability that it belongs to the construction site. There are 5 pooling structures in the whole network, making a single pixel of the output image correspond to 32×32 pixels in the original image. The size of the output image is 1 / 32 of the original image. The network output is upsampled using a deconvolutional layer by bilinear interpolation to obtain an output result with the same size as the input image.
[0095] Since construction machinery and prefabricated houses in the construction site are refined extraction elements, the above resolution cannot meet the requirements. Therefore, the network is further improved based on ZNT-32s to improve the resolution. As Figure 3 shown, a single pixel in the output result of the pool4 of the ZNT-16s pooling structure corresponds to 16×16 pixels in the original image. First, the conv7 convolutional layer is upsampled by 2 times, and then fused with the pool4 layer to obtain the network output result, which not only retains the detailed data in the pool4 of the pooling structure but also achieves the image classification accuracy of the ZNT-32s network.
[0096] ZNT-8s further improves the resolution based on ZNT-16s. The fusion result in the ZNT-16s network is upsampled by 2 times, and then fused with the pool3 of the pooling structure to obtain a network output result with a resolution of 8 pixels.
[0097] III. Extraction of Construction Site and Its State Based on Fully Convolutional Multi-Agent
[0098] (I) Collection and Processing of Remote Sensing Datasets
[0099] For UAV image data, according to the camera parameters (camera principal point, focal length, distortion parameters, etc.) obtained from camera calibration, the obtained UAV images are corrected for principal point offset and distortion, and then the UAV images are color-equalized to ensure the accuracy of image stitching. Secondly, an image matching software is used for feature-based image matching, including establishing a scale space, feature point localization, calculation of the main angle of feature points, calculation of feature descriptors. Then, aerial triangulation is carried out. The exterior orientation elements of each image are solved by bundle adjustment using the camera parameter file and the image point coordinate file. Finally, a DEM generation software is used to calculate the corresponding ground discrete point coordinates according to the matched homologous points and the solved image orientation elements. A DEM is generated from the ground discrete points and resampled to generate an orthoimage.
[0100] Perform radiometric correction and geometric correction on high-resolution remote sensing satellite images. Radiometric correction includes radiometric calibration and atmospheric correction. The minimum value removal method and histogram matching method are used to complete the DN value conversion to eliminate the influence of atmospheric illumination factors on ground object reflection. After radiometric correction, correct the geometric distortion in the initial image, including orthorectification, geometric registration, image mosaicking, and image fusion. Secondly, use geometric registration to overlay and match images with different time phases, different sensor types, or different imaging conditions, so that the images in the same area or the overlapping areas of adjacent area images overlap. Then, piece together several images with overlapping or adjacent areas into a complete image to complete the image mosaicking process. Perform image fusion on high-resolution panchromatic data and multispectral data to form a fused image with both high spatial resolution and multispectral characteristics. Finally, crop the processed image to complete the preprocessing process of the image.
[0101] Perform pixel-level calibration. The cropped image sizes are all 500×500 pixels. Manually calibrate the cropped images as samples. The specific calibration standard for the sample dataset is to construct a label map with the same size as the original image, where each pixel corresponds one-to-one with the pixel at the corresponding position in the initial image. If the pixel area in the initial image belongs to the construction site, the corresponding label image is marked white, otherwise it is marked black to obtain the marked map corresponding to the initial image. The calibration process is as Figure 4 shown. Figure 5 Show the comparison diagrams before and after calibration. The initial image and the calibrated samples are used in subsequent training.
[0102] (2) Multi-agent optimization method
[0103] 1. Output excitation model
[0104] In the extraction of construction sites and their states, use a simple and fast non-linear output excitation model to achieve non-linear transformation:
[0105]
[0106] Reduce the gradient dispersion generated during network training and enable the network to introduce sparsity by itself.
[0107] 2. Strategies to prevent overfitting
[0108] In the multi-agent training for the extraction of construction sites and their states, the multi-agent hierarchy is relatively deep and involves more parameters, while the collected and calibrated training samples are relatively few, which will lead to overfitting, that is, the phenomenon that the gap between the training error and the validation data error curves is large when they are stable. The reason for its occurrence is that the multi-agent learning model has a strong expressive ability. However, with less training data, the network has a good memory of the training data and can have a good effect on the training data, but the model does not have good generalization for non-training data, thus resulting in a large error.
[0109] Modify the agent itself. During the training of a multi-agent system, randomly update the agent in each iteration. Introduce randomness to enhance the generalization ability of the network. During the network training phase, temporarily discard the neurons of the network from the network with a certain probability, keeping the input and output layers unchanged. Update the weights in the multi-agent according to the error backpropagation algorithm to prevent the mutual adaptation between nodes.
[0110] Suppose the entire network has n parameters, and the selection of each parameter is a fixed probability, then there are 2 n selection possibilities, corresponding to 2 n sub-networks. When n is very large, the iteratively updated sub-networks do not repeat, avoiding the situation where a certain network is overfitted to the training set.
[0111] Multiply the output during testing by the fixed probability p used during training. Through one test, all the parameters of the original 2 n sub-networks are taken into account. Suppose x and y are the input and output respectively, and w is the weight parameter of the upper layer. w| p is the subset of weight parameters obtained by sampling with a fixed probability p, expressed as Equation 4:
[0112]
[0113] When writing the program, use y = w * p * x = w' * x in the testing phase, and modify the formula during training to Equation 5:
[0114]
[0115] The weight parameter obtained by training is w'.
[0116] 3. Agent Update Strategy
[0117] For a set of training samples data = x, label = y, define a prediction strategy such that F θ (x) is as close as possible to y, that is, the loss function L[y, F θ (x)] is as small as possible. The loss function is the two-norm, expressed as Equation 6:
[0118] L[y, F θ (x)] = (y - F θ (x)) 2 Equation 6
[0119] During the training process, continuously modify the parameter θ in F to make L continuously decrease:
[0120] ΔL ≈ dL Equation 7
[0121]
[0122] Obtained:
[0123]
[0124] To make △L < 0, it is only necessary that:
[0125]
[0126] Add a learning factor α (α > 0) to make the learning of θ smoother, and obtain:
[0127]
[0128] Use empirical values to set the learning factor, and then correct the learning factor in the test experiment.
[0129] Place the fully convolutional multi-agent in the Pascal VOC data sets natural image data set (excluding construction site image data) for pre-training to obtain the representation ability of the target, and then use the trained model weight parameters as the initial parameters of the construction site and state extraction multi-agent, and perform learning and training on the collected high-resolution remote sensing image data.
[0130] (III) Training of the fully convolutional multi-agent
[0131] First, place the fully convolutional multi-agent in the Pascal VOC data sets natural image data set for pre-training on the image segmentation task. The Pascal VOC data sets data set does not contain the construction site target to be extracted. Then, use the collected high-resolution remote sensing images (including high-resolution remote sensing satellite images and UAV images) as the data set and send them into the multi-agent obtained by pre-training for learning about the construction site. During the training process, use the parameters of the multi-agent obtained by pre-training as the initial parameter values of the multi-agent for training the construction site. This strategy can ensure that a relatively ideal result can be learned in the case of insufficient sample data sets. The construction of the training data set in this application is described in detail below:
[0132] The first step: Perform frame cropping preprocessing on the sample data set, and crop the high-resolution remote sensing images into a sample set of 500×500 pixels for feeding into the network for training and learning;
[0133] The second step: Manually calibrate the sample data set. Since there is no publicly available detection data set for the extracted ground object target - the construction site, it is necessary to manually calibrate the construction site on the collected high-resolution remote sensing images.
[0134] Construct a preliminary data set through the above two steps, select images with high quality and suitable scenes as the training data set, and each high-resolution remote sensing image in the data set corresponds to a binary image containing 0 and 1 as the label;
[0135] Organize the data set in the format of Pascal VOC data sets. The data set is distributed in three folders: JPEGImages, which stores the initial image data in JPEG format; SegmentationClass, which stores the label data corresponding to the initial images, with the same file name but in PNG format; ImageSets\Segmentation, which stores the data files of the picture names (without suffix) of each image, stored by the text application. This part contains two text files, train.txt and val.txt, corresponding to the file names of the training images and test images respectively. Input the classified data set into the constructed fully convolutional multi-agent for agent learning;
[0136] Adopt a strategy of progressive training from coarse to fine. First, train the ZNT-32s network, and use the stochastic gradient descent method to realize the iterative update of the network weights. The training samples used are k images, and the training samples are iterated k*1, k*2, …, k*30 times. In addition, m images are added for testing. The accuracy is the best when the number of iterations of the experimental data is 3 times the amount of experimental data. Therefore, set the number of iterations in the experiment to 3*k times;
[0137] After training the ZNT-32s network, assign the weights of the ZNT-32s network to ZNT-16s, and let ZNT-16s perform iterative training on the basis of ZNT-32s, also iterating 3*k times. Similarly, ZNT-8s is iteratively updated on the basis of the training results of ZNT-16s. Finally, three trained networks are obtained, realizing the training of multi-agents with three levels of fineness from coarse to fine. Then select the network with the best experimental effect, add the construction board spectral extraction index CKL and the bare soil spectral extraction index KVL, and perform iterative update on the basis of the training results of this network to train the multi-agent for construction site and status extraction proposed in this application.
[0138] (IV) Construction Site and Status Extraction Process
[0139] After the fully convolutional agent learns the construction site and status in the training samples, a multi-agent for construction site and status extraction is formed. Send the cropped test data and the corresponding calibration samples into the four groups of trained fully convolutional multi-agents for testing, and finally calculate and evaluate the test results.
[0140] IV. Construction Site and Status Extraction Experiment
[0141] Software environment: The operating system of the workstation is Ubuntu 14.04 based on Linux. The third-party open-source convolutional multi-agent framework Caffe. The third-party open-source multi-agent learning toolkit DeepLearnToolboxo
[0142] (I) Comparative experiment
[0143] In this application, the semantic segmentation algorithm (ZNT) of the fully convolutional multi-agent based on natural images is selected for comparison and directly used for the construction site detection of high-resolution remote sensing images. ZNT contains three different multi-agents, which are the 32s network from coarse to fine, with a classification fineness of 32 pixels; the 16s network, with a classification fineness of 16 pixels; and the 8s network, with a classification fineness of 8 pixels. First, the three multi-agents of ZNT are used for the learning and training of the construction site and its status in high-resolution remote sensing images, and then the test results are compared to select the multi-agent with the strongest learning ability. This model is compared with the multi-agent proposed in this application that includes the construction board spectral extraction index CKL and the bare soil spectral extraction index KVL algorithm in a comparative experiment.
[0144] (II) Experimental results and analysis
[0145] Three different multi-agents ZNT-32, ZNT-16, and ZNT-8 in the fully convolutional multi-agent are introduced into the extraction of the construction site and its status in high-resolution remote sensing images. 645 images out of the 1278 collected image samples are used as the input values of the multi-agent to learn the corresponding multi-agent parameters, and then the remaining 633 images are used as the sample set for multi-agent testing and input into the trained multi-agent for model testing.
[0146] Figure 6 Four representative detection results are shown. For easy viewing, the output images are all subjected to image binarization processing. Figure 6 From left to right are: the initial input image, the Ground truth marked image, and the experimental results of ZNT-8s, ZNT-16s, and ZNT-32s.
[0147] This group of experiments proves that the fully convolutional multi-agent is applied to the extraction of the construction site and its status, and among them, the ZNT-8s multi-agent has the best effect on the extraction of the construction site and its status. Then, the construction board spectral extraction index CKL and the bare soil spectral extraction index KVL are used to replace the blue and red bands in the ZNT-8s network, and training and learning are carried out again. The detection data is sent into the trained model and compared with the detection results of ZNT-8s.
[0148] Figure 7From left to right are: the initial input image, the Ground truth labeled image, the output result image of the algorithm of this application, and the output result image of the ZNT-8s algorithm.
[0149] As can be seen from the output result image, on the basis of the full convolution multi-agent, the result obtained by feeding the construction site spectral extraction index CKL and the bare soil spectral extraction index KVL into the full convolution multi-agent is better than the extraction effect obtained by directly feeding the initial image into the full convolution multi-agent. Next, the experimental results will be analyzed from the aspect of accuracy evaluation.
[0150] This application selects to compare with three different construction site extraction methods based on the full convolution multi-agent, and calculates the extraction accuracy. The results are as Figure 8 shown. The experiment proves that among the three multi-agents of the full convolution multi-agent, ZNT-8s and ZNT-32s can better realize the extraction of the construction site and its state. The algorithm proposed in this application has achieved the best results in terms of accuracy, recall rate and comprehensive evaluation index, and can well realize the extraction of the construction site and its state.
[0151] Due to the high similarity between the construction site and the bare soil, it is difficult for general algorithms to distinguish them. The biggest difference is that there are usually blue prefabricated houses and construction machinery in the construction site. Such semantic data helps to distinguish the construction site from general bare soil. The remote sensing methods of the existing technology can only extract the color data or texture data of pixels, and it is difficult to mine the semantic data of images. Multi-agent learning can effectively mine the deep features of the target and propose abstract semantic data to improve the extraction and classification accuracy. Combining the experimental results with the evaluation accuracy shows that this application introduces the prefabricated house index CKL and the soil enhancement KVL index into the multi-agent learning model, and achieves a good effect on the extraction of the construction site. The experimental results prove that adding the influence features of the construction site can effectively improve the recall ability of the full convolution multi-agent for the target, has good extraction ability on this data set, and has strong robustness.
Claims
1. A method for accurately extracting construction sites and states from high-resolution remote sensing images, characterized in that: Adaptive learning of target features in remote sensing big data, on the basis of analyzing and calculating the extraction elements of construction sites, a method based on multi-agent learning for construction site and state extraction from high-resolution images is established: firstly, the element features of the construction site on the high-resolution remote sensing image are analyzed, the extraction elements are determined, and based on the performance characteristics of each extraction element on the image, the features for construction site extraction are established, and the bare soil spectral extraction index KVL is constructed; considering the characteristics of the construction site as an irregular extraction target, a full convolution neural multi-agent for pixel-level classification is adopted; the input data and output data features in the full convolution neural multi-agent are analyzed, the multi-source high-resolution images are cropped into k 500*500 pixel image data, m images are selected as the experimental data set and preprocessed, a full convolution neural multi-agent for construction site and state extraction is constructed, and a construction site and state extraction process method is established; (1) Constructing construction site state extraction feature factors: Analyze and calculate the construction site extraction elements, determine the three major extraction elements: bare soil, construction board house, and construction machinery, calculate and analyze the spectral characteristics, texture characteristics, and image proportion of each extraction element in the image, compare and integrate the extraction algorithms of each extraction element feature, establish the construction site feature extraction algorithm, calculate the reflectance spectrum curves of different types of soil based on the board house spectrum extraction index CKL, and establish the bare soil spectrum extraction index KVL; (2) Constructing a fully convolutional multi-agent for construction site and state extraction: A fully convolutional neural multi-agent for pixel-level classification is used. Based on the irregular contour shape of the construction site, the initial image spatial structure data is retained for pixel-level classification. The fully connected layers ZN6 and ZN7 after the 5th pooling structure are converted into convolutional layers conv6 and conv7. The output is converted into 4096 feature maps, where ZN6 and ZN7 are fully connected layers containing 4096 neurons, and conv6 and conv7 are convolutional layers containing 4096 convolution kernels with a size of l×1. Finally, a convolutional layer with 2 convolution kernels is used to regress the 4096 vectors of each node into a 2-dimensional target category. The category confidence of each node is predicted and normalized. The size of the feature map obtained is proportional to the input construction site sample image. (3) Construction site and state extraction based on fully convolutional multi-agent: The obtained high-resolution images are cropped and screened as experimental data sets and preprocessed accordingly. By outputting the incentive model, preventing overfitting strategy, and agent updating strategy, an optimization strategy is established for the multi-agent. A fully convolutional neural multi-agent is constructed for construction site and state extraction, and a construction site and state extraction process is established.
2. According to the method for accurately extracting construction sites and states from high-resolution remote sensing images of claim 1, it is characterized by: Multi-agent learning feature extraction is used to learn and extract construction machinery. For construction board houses and bare soil areas, the board house spectrum extraction index CKL and soil extraction index KVL are used to enhance image features. The processed image feature map is used as the input of the multi-agent to enhance the feature learning ability of the multi-agent. Constructing the bare soil spectrum extraction index KVL: Extracting the homogeneous characteristics of the soil reflectance spectrum in the visible light band, the reflectivity increases with the increase of wavelength, and the increase is significant. The red band is amplified to establish the bare soil spectrum extraction index KVL, which is expressed as formula 1: KVL = (RB) + (RG) Formula 1 Among them, R, G, and B refer to the reflectivity values of the red, green, and blue bands of visible light, respectively.
3. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1 is characterized by: Construct the spectral extraction index CKL of the board room: The spectral reflectance of the construction board room is low in the blue, green and red bands of visible light, and the spectral extraction index CKL of the board room is constructed accordingly: CKL=(BG)+(BR) Formula 2 Among them, R, G, and B refer to the reflectance values of the red, green, and blue bands of visible light, respectively. The CKL index targets the special color data of construction panels, normalizes the color bands, and amplifies the features of the target object to achieve differentiated extraction.
4. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1 is characterized in that: Constructing a fully convolutional multi-agent for construction site and status extraction: Establish three fully convolutional multi-agents from coarse to fine, and name the convolutional agents according to the receptive field of a single pixel of the output image: ZNT-32s, ZNT-16s, and ZNT-8s. The numbers in the names represent the receptive field range of a single pixel of the output image. The smaller the range, the higher the accuracy of the network for contour extraction. The receptive field size corresponding to a single pixel of the output image of ZNT-32s is 32×32 (that is, the output pixel unit corresponds to an image block of 32×32 in the initial image). It consists of 40 layers, including 16 convolutional layers, 15 ReLU activation layers, 5 pooling structures, 2 discard layers, 1 deconvolution layer, and 1 cropping layer; The output layer of the ZNT-32s network is a convolution layer. The value of each pixel in the output image ranges from 0 to 1, representing the probability that it belongs to the construction site. The entire network contains 5 pooling structures, so that a single pixel of the output image corresponds to 32×32 pixels in the initial image. The size of the output image is 1 / 32 of the original image. The network output is upsampled using a deconvolution layer through quadratic linear interpolation to obtain an output result of the same size as the input image. Based on ZNT-32s, the network is further improved to improve the resolution. In the output result of the ZNT-16s pooling structure pool4, a single pixel corresponds to 16×16 pixels in the initial image. The conv7 convolutional layer is first upsampled by 2 times, and then fused with the pool4 layer to obtain the network output result, which not only retains the detailed data in the pooling structure pool4, but also achieves the image classification accuracy of the ZNT-32s network.
5. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1 is characterized in that: Remote sensing data set acquisition and processing: For drone image data, the principal point offset and distortion correction of the acquired drone images are performed according to the camera parameters (camera principal point, focal length, distortion parameters, etc.) obtained by camera calibration, and then the drone images are processed for uniform color to ensure the accuracy of image stitching. Next, image matching software is used to perform feature-based image matching, including establishing scale space, feature point positioning, feature point principal angle calculation, and feature descriptor calculation. Then, aerial triangulation is performed, and the camera parameter file and image point coordinate file are used to solve the exterior orientation elements of each image through bundle method regional adjustment. Finally, DEM generation software is used to calculate the corresponding ground discrete point coordinates according to the matched same-name points and the solved image orientation elements, and DEM is generated from the ground discrete points, and resampling is used to generate orthophotos; Pixel-level calibration is performed, and the size of the cropped images is 500×500 pixels. Manual sample calibration is performed on the cropped images. The specific calibration standard of the sample data set is to construct a label map of the same size as the original image, in which each pixel corresponds one-to-one to the pixel at the corresponding position in the initial image. If the pixel area in the initial image belongs to the construction site, the corresponding label image is marked in white, otherwise it is marked in black, and the label map corresponding to the initial image is obtained.
6. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1, characterized in that: Output excitation model: In the construction site and state extraction, a simple and fast nonlinear output excitation model is used to achieve nonlinear transformation: x and y are the input and output of the model respectively, which reduce the gradient dispersion generated during network training and make the network introduce sparsity by itself.
7. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1, characterized in that: Strategy to prevent overfitting: modify the agent itself. When training a multi-agent, randomly update the agent in each iteration, use the introduced randomness to increase the generalization ability of the network, and temporarily discard the neurons of the network from the network with a certain probability during the network training phase, keep the input and output layers unchanged, and update the weights in the multi-agent according to the error back propagation algorithm to prevent mutual adaptation between nodes.
8. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 7, characterized in that: Assuming that the entire network has n parameters, and each parameter is selected with a fixed probability, there are 2 n There are 2 possible constituencies, corresponding to n When n is large, the iteratively updated subnetworks are not repeated to avoid a certain network being overfitted to the training set. Multiply the test output by the fixed probability p used during training, and pass a test to convert the original 2 n All the parameters of the sub-network are taken into account. Assume that x and y are the input and output respectively, w is the weight parameter of the previous layer, w| p is the subset of weight parameters obtained by sampling with a fixed probability p, expressed as formula 4: When writing a program, use y=w·p*x=w′*x in the test phase and correct the training formula to Formula 5: The weight parameter obtained through training is w'.
9. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1, characterized in that: Agent update strategy: For a set of training samples data = x, label = y, define a prediction strategy so that F θ (x) is as close to y as possible, that is, the loss function L[y, F θ (x)] is as small as possible, and the loss function is the two-norm, expressed as Equation 6: L[y, F θ (x)] = (y - F θ (x)) 2 Equation 6 During the training process, the parameter θ in F is constantly modified so that L is continuously reduced: ΔL≈dL Equation 7 get: To make ΔL < 0, just: Adding a learning factor α (α>0) to make the learning of θ smoother, we get: Use empirical values to set learning factors, and then make corrections to the learning factors during test experiments; The fully convolutional multi-agent is placed in the Pascal VOC data sets for pre-training to obtain the target representation capability. Then the trained model weight parameters are used as the initial parameters of the construction site and state extraction multi-agent, and learning and training are performed on the collected high-resolution remote sensing image data.
10. The method for accurately extracting construction sites and states from high-resolution remote sensing images according to claim 1, characterized in that: Fully convolutional multi-agent training: First, the fully convolutional multi-agent is placed on the Pascal VOC data sets natural image dataset for pre-training of image segmentation tasks. The Pascal VOC data sets do not contain the construction site targets that need to be extracted. Then, the collected high-resolution remote sensing images are sent as datasets to the pre-trained multi-agent for learning about the construction site. During the training process, the parameters of the pre-trained multi-agent are used as the initial parameters of the multi-agent for training the construction site: Step 1: Perform frame cropping preprocessing on the sample data set, cropping the high-resolution remote sensing images into sample sets of 500×500 pixels so that they can be sent to the network for training and learning; Step 2: Manually calibrate the construction site using the collected high-resolution remote sensing images; Through the above two steps, a preliminary data set is constructed, and high-quality and scene-appropriate images are selected as the training data set. Each high-resolution remote sensing image in the data set corresponds to a binary image containing 0 and 1 as a label; The dataset is organized in the format of Pascal VOC data sets. The dataset is distributed in three folders: JPEGImages, which stores the initial image data in JPEG format; SegmentationClass, which stores the label data corresponding to the initial image. The corresponding file names are the same, but stored in PNG format; ImageSets\Segmentation, this folder contains the image name data files of each image, which are stored by text application files. This part contains two text files, train.txt and val.txt, which correspond to the file names of training images and test images respectively. The classified data sets are input into the constructed full convolution multi-agent for agent learning; The strategy of progressive training from coarse to fine is adopted. The ZNT-32s network is trained first. The random gradient descent method is used to iteratively update the network weights. The training samples used are k images. The training samples are iterated k*1, k*2, ..., k*30 times. In addition, m images are added for testing. The accuracy is best when the experimental data is iterated three times the amount of experimental data. The number of iterations in the experiment is set to 3*k times. After training the ZNT-32s network, assign the weight of the ZNT-32s network to ZNT-16s, and let ZNT-16s perform iterative training based on ZNT-32s, and iterate 3*k times. Similarly, ZNT-8s is iteratively updated based on the training results of ZNT-16s, and finally three trained networks are obtained to realize the training of multi-agents with three levels of precision from coarse to fine. Then, the network with the best experimental effect is selected, and the construction board room spectral extraction index CKL and the bare soil spectral extraction index KVL are added. Iterative updates are performed based on the training results of the network to train the construction site and status extraction multi-agent proposed in this application.