Retrieval and evaluation method for building damaged by unmanned aerial vehicle remote sensing disaster image
By applying drone remote sensing technology and deep learning models (such as GoogleLeNet and AlexNet) to perform target detection of damaged buildings in disaster images, the problem of traditional methods of detection accuracy and efficiency in disaster images is solved, and more efficient identification and positioning of damaged buildings is achieved.
Patent Information
- Application Number
- CN202510127068.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-28
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art has problems of low accuracy and low efficiency in the target detection of damaged buildings in disaster images, especially in disaster images with large data volume, large information and wide scale span. The traditional image feature extraction method is time-consuming and has poor results.
The feature extraction method of damaged building targets based on drone remote sensing is used, and the feature extraction is performed using GoogleLeNet and AlexNet models, and the damage zone is generated through regional prediction, and the regression and fusion of the detection results is performed in combination with the focus cross-border ratio and improved non-maximum suppression method.
It improves the accuracy and efficiency of target detection of damaged buildings in disaster images, can accurately identify and locate images after obtaining disaster areas, and assists the rescue department to formulate disaster relief plans.
Smart Images

Figure CN120219984A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a method for retrieving damaged buildings from remote sensing images, in particular to a method for retrieving and evaluating damaged buildings from UAV remote sensing disaster images, and belongs to the technical field of remote sensing damaged building retrieval and evaluation. Background Art
[0002] With the rapid development of remote sensing technology, the ways to obtain disaster images are more diverse, and the imaging quality of the images is getting higher and higher. Various feature information in the images can be used to perform multi-angle and multi-level object detection on the ground objects in the disaster images.
[0003] Traditional technical means artificially design the features of images from the aspects of color, shape, etc. in the remote sensing images. Since remote sensing images, especially disaster images, themselves contain a lot of background information, the artificially designed features cannot maintain the generalization characteristics in this case, and at the same time, a lot of energy is required for feature adjustment, and often an ideal object detection effect cannot be obtained.
[0004] Traditional object detection methods design highly robust and representative image features that are different from each other based on various characteristics such as the color space, edge features, and texture features of the object itself, and then design a classifier for image classification according to the designed image features. Such methods usually require a lot of time for feature design. Because the features of the image itself are uncertain, the original image operators cannot be directly applied to relatively complex detection scenarios, so a certain prior basis and a lot of time are required for various attempts of feature design. The characteristics of the image features themselves have a greater impact on the final detection result. And disaster images themselves have characteristics such as large data volume, much information contained, and wide scale span. Applying traditional image feature extraction methods will consume more time in the feature design stage, and at the same time, more energy is required for feature effect testing, and the final object detection result is not very ideal. Therefore, there is an urgent need for deep image features with performance superior to traditional image processing methods, and use hardware acceleration to solve the problem that feature design is too time-consuming.
[0005] With the development of computer graphics, object recognition has shown great promise. If it can be applied to the detection of damaged buildings in disaster images, and at the same time use GPU parallel computing to improve the computing efficiency, it will greatly improve the accuracy of object detection, and can also have a good improvement in the efficiency of target feature generation. Using this detection system, it is possible to accurately identify and locate the disaster area after obtaining the disaster area images, which is convenient for assisting the rescue department to formulate corresponding disaster relief plans in a timely manner.
[0006] The problems that need to be solved by the complex image key area detection method of the prior art and the key technical difficulties of the present application include:
[0007] (1) Existing technologies design features of remote sensing images manually from perspectives such as color and shape in the remote sensing images. Since remote sensing images, especially disaster images, themselves contain a lot of background information, manually designed features cannot maintain generalization in such cases. At the same time, a large amount of effort is required for feature adjustment, and often an ideal target detection effect cannot be obtained. The target detection methods of existing technologies design highly robust and representative image features that are distinguishable from each other based on various characteristics of the target itself, such as color space, edge features, and texture features. Then, a classifier for image classification is designed based on the designed image features. This type of method requires a large amount of time for feature design because the features of the image itself are uncertain, and the original image operators cannot be directly applied to relatively complex detection scenarios. Therefore, a certain prior basis and a large amount of time are required for various attempts at feature design. The characteristics of the image features themselves have a great impact on the final detection result. Disaster images themselves have characteristics such as a large amount of data, a lot of information contained, and a wide scale span. Applying traditional image feature extraction methods will consume more time in the feature design stage, and at the same time, more effort is required for feature effect testing, and the final target detection result is not very ideal. Therefore, there is an urgent need for deep image features with performance superior to traditional image processing methods and the use of hardware acceleration to solve the problem that feature design is too time-consuming. Existing technologies lack GPU parallel computing, the accuracy of target detection is low, and the efficiency of generating target features is also low, and it is impossible to accurately identify and locate the disaster area after obtaining disaster area images.
[0008] (2) The key problems in identifying damaged buildings in disaster images currently are as follows: First, the ground object information in disaster images is complex. How to select positive and negative samples so that the model can more accurately distinguish damaged building targets and other background ground objects. At the same time, the multi-pose problems of different scales and ground object categories will also have a significant impact on the feature learning process. Second, image data often contains a large number of irrelevant pixels, and the damaged buildings that really need to be detected may only be a small part of the area. Therefore, before detection, an efficient method is needed to pre-screen the image for regions to obtain regions that may have damaged building targets. The regions generated by this method need to ensure comprehensiveness and at the same time have a certain degree of accuracy. Third, there are multiple detection result bounding boxes on the target in the region detection result. How to use an algorithm to eliminate or fuse the duplicate and less effective results among them is another difficult point that needs to be solved.
[0009] (3) The overall level of intelligent processing of geological disaster remote sensing image data is currently very limited, and a large amount of valuable data itself cannot be fully utilized and cannot play its role. The existing technology lacks feature learning of damaged buildings in disaster images, and lacks the two models based on GoogLeNet and AlexNet to detect damaged buildings. First, there is a lack of damaged building target feature extraction based on drone remote sensing, and it is impossible to obtain more generalized damaged building target features; second, there is a lack of damaged area generation based on regional prediction, lack of regional prediction methods, and it is impossible to generate a large number of prediction areas, and it is impossible to ensure the diversity and multi-scale characteristics of the prediction areas; third, there is a lack of precise positioning of target detection results based on regional prediction, and a lack of regional prediction methods for the regression fusion of the focus frame of the detected damaged building targets. The retrieval accuracy and efficiency of damaged buildings are relatively low. Summary of the invention
[0010] This application conducts multi-angle and multi-factor analysis of disaster images, manually extracts basic land object types, and accurately completes the integration and processing of a certain scale of data sets, laying the foundation for analysis and detection. Different training strategies are formulated for different models, and the detection effects of the same model under different training modes and the detection effects of different models under the same training mode are compared and analyzed. Multi-indicator detection of damaged residential buildings is carried out in multiple regions, and the impact of the data set size selection strategy on the accuracy of similar land object detection in different regions is calculated from the perspective of model generalization. For specific detection targets, a specific target to be tested is generated using a selective search method, and the overlapping detection results in the detection results are regressed and fused using the focused intersection-union ratio and improved non-maximum suppression to improve the detection performance of the GoogLeNet and AlexNet fine-tuning models; advanced tools and means in the field of artificial intelligence are used in the field of disaster image processing, and the accuracy level that can be put into practical application is achieved, which has a certain positive effect on subsequent target detection, pattern recognition, complex scene recognition and other fields.
[0011] In order to achieve the above technical effects, the technical solutions adopted in this application are as follows:
[0012] The damaged building retrieval and evaluation method of UAV remote sensing disaster images is used to learn the features of damaged buildings in disaster images. The two training modes of weight full initialization and fine-tuning training are used to detect damaged buildings based on the GoogLeNet and AlexNet models. The deeper fine-tuning training network model has better detection performance under large-scale data sets.
[0013] 1) Damaged building target feature extraction based on UAV remote sensing: GoogLeNet and AlexNet are used to extract damaged building target features to obtain more generalized damaged building target features;
[0014] 2) Generation of damaged areas based on region prediction: By using the method of region prediction, a large number of prediction regions are generated. When using UAV remote sensing for target detection, only the prediction regions are detected and classified, ensuring the diversity and multi-scale characteristics of the prediction regions and improving the retrieval efficiency;
[0015] 3) Precise positioning of object detection results based on region prediction: A method based on region prediction is used to perform regression fusion of the focus boxes for the detected damaged building targets. The JBJ focus intersection over union and the improved BKE method are used to fuse and eliminate the detection results with high overlap on the same target in UAV remote sensing object detection.
[0016] Preferably, for the generation of damaged areas based on region prediction: Detect the damaged areas on the image to be measured, including two points: one is the possible orientation of the target to be measured, and predicting this improves the detection efficiency; the other is the true position of the target to be measured, that is, judging multiple adjacent damaged areas to weaken the situation of edge bounding and improve the bounding accuracy; in the case where the high-level image features during detection, that is, the damaged area targets, are fixed, the whole of the image to be measured or the self-extracted part is used as the input element of the recognition model, and the category corresponding to each input damaged area image and the probability of being determined as this category are obtained.
[0017] Preferably, for the generation of damaged areas based on geological disaster images: First, based on graph representation image segmentation, the image to be measured is split to obtain a certain number of specific regions with special properties, and then the damaged area targets are extracted from them. In image segmentation, the segmentation boundaries between different regions are determined, and the boundaries are established based on the determination between specific regions under graph representation. The greedy algorithm is used for screening and then segmenting the image, selectively ignoring the details of the regions with relatively large detail changes while retaining the details of the regions with small detail changes, so that the regions with rich diversity are aggregated, and visually making the specific regions more consistent; in terms of time, the efficiency of splitting into multiple regions is linearly related to the number of pixels of the original image;
[0018] The graph representation expresses the image in a broad sense as a graph in graph theory, which is composed of several given points and the connections between points. Suppose there is a graph G=(V, E), and each pixel in the image is taken as v i ∈V, and each two adjacent pixel points form an edge e i =(v i , v j )∈E, and the difference in pixel color values constitutes the weight w(e i ), and w is used as the value between v i and v jThe core index for measuring the difference, the weight is positively correlated with the pixel points. Divide G=(V, E) into a series of non-overlapping C∈M, where M is the set of subgraphs G'. All elements in M are independent of each other and satisfy non-overlapping, that is, similar v i are aggregated;
[0019] Assume that G has been simplified to the minimum spanning tree MST(C, E). The internal difference Int(C) of the divided C is the maximum weight of E contained in the current C, as shown in Equation 1:
[0020]
[0021] The difference between C1 and C2 is calculated by the weight of the minimum edge connecting their vertices. When there is no connecting edge between C1 and C2, define Dif(C1, C2)=∞, as shown in Equation 2:
[0022]
[0023] The determination of the boundary between two divided regions is to compare the sizes of Dif and the minimum division internal difference MInt, including Equation 3 and Equation 4:
[0024]
[0025] MInt(C1, C2)=min(lnt(C1)+τ(C1), lnt(C2)+τ(C2)) Equation 4
[0026] If the boundary assumption is correct, then C1 and C2 are divided, otherwise they are merged, where the threshold function τ controls the difference between the two. The threshold function is associated with the region size to better control the definition of the divided region boundary, as shown in Equation 5:
[0027] τ(C)=k / |C| Equation 5
[0028] |C| refers to the number of pixel points, and k is a parameter adjusted according to the image size. The larger k is, the larger the obtained elements are. In the implementation process, Gaussian smoothing is performed on the image, and some noises are removed on the basis of minimizing the change of the image, and the overall effect of the image remains unchanged;
[0029] Preferably, the steps of the overall algorithm are as follows:
[0030] The first step: For E∈G, sort in ascending order according to W;
[0031] The second step: S[0] is the initialization of the division, that is, each pixel v is an independent division region C;
[0032] Step 3: Let q = 1, 2, 3, …, m (where m is the number of edges). According to the segmentation result of the previous S[q - 1], select an o[q](v i ,v j ). If v i and v j belong to two non - overlapping segmented regions C respectively, and if o[q](v i ,v j ) < MInt, merge these two regions. If the determination is negative, continue to perform the repetitive operation;
[0033] Step 4: All the finally obtained C i ∈ S = S[m] is the segmentation set that is neither too fine nor too rough and is the one we seek;
[0034] After obtaining the preliminary segmentation set, since the set still has a large cardinality, calculate the approximation degree of the original segmented regions in the set. Two multi - entropy strategies are adopted to enhance the generation of multi - entropy in the damaged area of the image, including color - space multi - entropy and approximation - degree calculation multi - entropy. The former is applied to the generation of the original region set in image segmentation, while the latter is applied to the merging of the original regions;
[0035] Preferably, in the color - space multi - entropy, multiple color spaces including RGB, grayscale, Lab, normalized RG channel plus grayscale, HSV, normalized RGB, H, S, V are adopted to ensure a high degree of diversity in the generation of the original damaged area, and multiple color spaces are considered for the image scene under lighting conditions;
[0036] In the approximation - degree calculation multi - entropy, the specific calculation process is as follows:
[0037] (1) Color approximation: S color (r i ,r j ) is used to calculate the color approximation of r i and r j in the regions to be merged. Each such region obtains a color - distribution histogram When the number of color channels is 1, n = 25; if there are 3 color channels, then n = 75; Next, normalize the obtained color histogram, as shown in Equation 6:
[0038]
[0039] After calculating the color approximation, merge r i and r j to obtain r t . The calculation method of the color - distribution histogram of the new region r t is shown in Equation 7:
[0040] Ct =(size(r i )×C i +size(r j )×C j ) / size(r t ) Equation 7
[0041] The denominator in the equation is the new size of r t , that is, size(r t ) = size(r i ) + size(r j );
[0042] (2) Texture approximation: S texture (r i , r j ) is used to calculate the texture approximation between regions. For each region, take 8 directions for each color channel, Gaussian smoothing with a variance of 1, obtain a histogram of 10 bins for each color in each channel, normalize with 3 color channels, and obtain a vector of n = 8 * 3 * 10 = 240 dimensions The calculation method of the texture feature of the new region is similar to that of the color approximation. See Equation 8:
[0043]
[0044] Similar to color, texture is transferred;
[0045] (3) Size approximation: S size (r i , r j ) is used to calculate the size approximation between regions. The size is the number of pixel points contained in the region. By calculating the ratio of the sum of the sizes of two regions to the total number of pixels, it is judged whether it is a small region and whether it should be merged as early as possible. See Equation 9 for details:
[0046]
[0047] where size(im) represents the total number of pixels in the original image;
[0048] (4) Filling approximation: S fill (r i , r j ) is used to calculate the filling approximation between regions, measure whether there is an intersection or inclusion relationship between two regions. If so, merge them first to avoid blank regions in the output result. Calculate the minimum enclosing focused region BB i of r j and r ij , size(BB ij)The smaller it is, the higher the filling approximation degree, that is, the greater the possibility that the two regions intersect or contain each other. Subtract the number of pixels in the smallest enclosing region by r i and r j The sum of the number of points of and r i and r j Whether they are in a relationship of mutual inclusion or intersection can be judged by comparing the obtained value with the total number of pixels. See Equation 10 for details:
[0049]
[0050] Finally, all four calculation methods are combined together, and the structure is as shown in Equation 11:
[0051] S(r i , r j ) = a1S color (r i , r j ) + a2S texture (r i , r j ) + a3S size (r i , r j ) + α4S fill (r i , r j ) Equation 11
[0052] Among them, a i ∈{0, 1} indicates that the four approximation principles of color, texture, size, and filling are all optional;
[0053] Through image segmentation, the original damaged area and the fusion of the damaged area approximation degree are generated. Finally, a set of focus frames that may contain objects are obtained. The average highest overlap rate ABO obtained by comparing the focus frames generated by the algorithm with the area where the real situation is located is used to measure. For each determined category C, each real situation is expressed as G is the target area where the object is located, and 1 is a sub - element of the set L of calculated position hypotheses. Then the expression of ABO is Equation 12:
[0054]
[0055] The calculation method of the overlap rate is Equation 13:
[0056]
[0057] Calculate the respective ABO for each category in turn, and then use the average value MBAO of the ABO of all categories for evaluation;
[0058] Preferably, at the early stage of the algorithm, the similarity calculation method and the types of feature calculation methods are set in advance to limit the number of focus boxes obtained finally. The size critical values of the damaged area are initialized to 250 and 400, and finally about 200 focus boxes are obtained.
[0059] Object detection takes into account the multi-level and diversity of the prediction region at the same time. First, the original segmentation region is obtained through image segmentation based on graph representation, and the regions with large differences are segmented. Then, classification and merging are performed according to the approximation calculation method of multiple indicators, and the regions with high approximation are merged. Through multiple data iterations, the final set of focus boxes is obtained.
[0060] Preferably, for the real position positioning of the damaged area: the intersection over union JBJ is used to calculate the overlap rate between the detected focus box and the real calibration data to measure the accuracy of object detection. Set the focus box G1{(minx1, miny1)(maxx1, maxy1)} to be detected and the real position G2{(minx2, miny2)(maxx2, maxy2)}. Assume that G1 and G2 have an intersection G u {(minx, miny)(maxx, maxy)}, where the specific values are shown in Equations 14 to 16:
[0061] minx = max(minx1, minx2)
[0062] miny = max(miny1, miny2)
[0063] maxx = min(maxx1, maxx2)
[0064] maxy = min(maxy1, maxy2) Equation 14
[0065] S(G u ) = (maxx - minx)(maxy - miny) Equation 15
[0066]
[0067] Where S(G u ), S(G1), and S(G2) are the areas of the intersection G u , G1, and G2 respectively. G u = G1 ∩ G2. If G1 and G2 do not intersect, there is minx > maxx or miny > maxy;
[0068] At the same time, the critical value of JBJ is set through prior knowledge and the relative relationship between the image size and the target size. In this application, the critical value is set to 0.6.
[0069] Preferably, for the overlap of the focus areas and the region merging, a large number of candidate boxes obtained by processing selective search are divided into three steps:
[0070] Step 1: Generate a large number of candidate windows through a sliding window or a candidate target generation scheme for selective search;
[0071] Step 2: Use the classifier generated by the support vector machine to classify the candidate windows generated previously, and give each window a score as a subsequent judgment criterion;
[0072] Step 3: Fuse the detection results to obtain the final detection result;
[0073] The obtained data includes the position information and scores of the focus boxes screened by the classifier. Fuse the focus boxes on the same target. Suppose there are 8 focus boxes that frame the object to be measured, namely A, B, C, D, E, F, G, a total of 8 focus boxes. Set the intersection-over-union critical value as a. The specific steps are as follows:
[0074] Step 1: Sort all the boxes in ascending order according to the scores given by the classifier;
[0075] Step 2: Starting from the highest-scoring box A, calculate its intersection-over-union with the other remaining boxes. If the ratio of the intersection-over-union to the area of A is greater than the given critical value a, then put this box into group G A , and the members within the group all have a large overlap with A; if there is no such box, then A forms a group by itself, and the remaining boxes B - G are grouped in the same way to obtain other groups (if any, denoted as G B -G G );
[0076] Step 3: For each set obtained after grouping, calculate the optimal regression box for each group, denoted as T, T = {T1, T2,..., T n}, and t n is the optimal regression box of the nth group.
[0077] Step 4: When grouping and calculating, the boxes with a high intersection-over-union are gathered towards a specific target, and there may still be a situation where there are multiple boxes on a single target. To solve this problem, use a gradient intersection-over-union critical value, which decreases step by step with the number of iterations, and then repeat Steps 1, 2, and 3: regroup T according to the boxes in T and calculate the optimal solution, and finally completely separate the detection results to obtain a one-to-one corresponding result.
[0078] Preferably, the training process for damaged buildings:
[0079] 1) Convert the image dataset from the image format to a high-performance database to facilitate the efficient reading of data by the Caffe framework, and at the same time generate file labels;
[0080] 2) Calculate the mean file using the Caffe toolset;
[0081] 3) Select a network model and modify the model definition file;
[0082] 4) Develop corresponding training strategies;
[0083] 5) Select a training mode, randomly generate a weight initialization model, further train on the initialized model, and fine-tune the training of the model.
[0084] During the training process, the propagation weights are updated through error backpropagation. At the same time, the momentum factor is used to smooth the weights to accelerate convergence. Dropout is used to randomly make the weights of some hidden layer nodes in the network not work. The non-working part of the nodes is temporarily considered not to be part of the network structure, but the weights are retained and not updated this time, and will be updated when the next wave of sample data is input.
[0085] Preferably, the damaged building detection process: After the training process is completed, a detection model based on damaged residential buildings is obtained. The detection process uses this model to perform target detection on the image to be tested, specifically including the following steps:
[0086] 1) Use the selective search algorithm to generate a large number of focus boxes on the image of the area to be tested, record and retain the positions of the focus boxes;
[0087] 2) The trained detection model performs class determination on the image data inside each focus box. The retained data are the features of the damaged residential building category. If it is determined to be other feature types such as landslides, factory buildings, and intact residential buildings, they are directly excluded from the results first;
[0088] 3) The content inside the remaining focus boxes is determined to be damaged residential buildings, but there is a large amount of overlap on the same target. The cross-use and focus intersection over union and improved non-maximum suppression methods are used to group and fuse all the remaining focus boxes, and the redundant boxes are eliminated by considering the area and the model score respectively. Finally, ensure that there is only one focus box for each target;
[0089] 4) Finally, calculate the recall rate and precision rate of the target detection, and calculate them separately for different models and different focus box production methods.
[0090] Compared with the prior art, the innovation points and advantages of this application are:
[0091] (1) This application proposes a method for extracting the characteristics of damaged buildings based on UAV remote sensing. By improving the UAV remote sensing recognition model network, through the overall processing process and data flow of the data in the network model, two training modes of combining GoogLeNet and AlexNet with weight initialization and model fine-tuning are adopted to extract the characteristics of damaged buildings. Second, it proposes a method for generating damaged area regions based on region prediction: using a data-driven focus box generation method to obtain damaged area regions, replacing the traditional damaged area generation method, effectively improving the retrieval efficiency, and using an improved non-maximum suppression method to regress the measured damaged area regions. Third, through the detection experiment process and results of damaged buildings by UAV remote sensing, the necessary ground object characteristics for training are extracted, and a parallel processing training framework is used to perform feature learning on the UAV remote sensing network. By comparing the performance of different models under different learning strategies, the optimal method for the retrieval accuracy and efficiency of damaged buildings is obtained.
[0092] (2) This application conducts feature learning on damaged buildings in disaster images. Two training modes of completely initializing weights and fine-tuning training are formulated, and object detection of damaged buildings is carried out based on two models, GoogLeNet and AlexNet. The deeper fine-tuning training network model has better detection performance under large-scale datasets. For the extraction of damaged building target features based on UAV remote sensing, GoogLeNet and AlexNet are used to extract damaged building target features to obtain more generalized damaged building target features; for the generation of damaged areas based on region prediction, the method of region prediction is adopted to generate a large number of prediction regions. When using UAV remote sensing for object detection, only the prediction regions are detected and classified, ensuring the diversity and multi-scale characteristics of the prediction regions and improving the retrieval efficiency; for the precise positioning of the object detection results based on region prediction, the method of region prediction is used to perform focus box regression fusion on the detected damaged building targets. The JBJ focus intersection over union and the improved BKE method are used to fuse and eliminate the detection results with high overlap on the same target in UAV remote sensing object detection, effectively improving the detection effect and also having a certain improvement in the object detection positioning accuracy. The results show that this retrieval method based on damaged buildings has good performance and application prospects in object detection of disaster images.
[0093] (3) This application conducts multi-angle and multi-factor analysis on disaster images, manually extracts basic land object types, and accurately completes the integration and processing of a certain scale of data sets, laying the foundation for analysis and detection. Different training strategies are formulated for different models, and the detection effects of the same model under different training modes and the detection effects of different models under the same training mode are compared and analyzed. Multi-index detection of damaged residential buildings is carried out in multiple regions, and the impact of the data set size selection strategy on the accuracy of detection of similar land objects in different regions is calculated from the perspective of model generalization. For specific detection targets, a specific target to be tested is generated using a selective search method, and the overlapping detection results in the detection results are regressed and fused using focused intersection-over-union and improved non-maximum suppression to improve the detection performance of the GoogLeNet and AlexNet fine-tuning models; advanced tools and means in the field of artificial intelligence are used in the field of disaster image processing, and the accuracy level that can be put into practical application is achieved, which has a certain positive effect on subsequent target detection, pattern recognition, complex scene recognition and other fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] Figure 1 It is a technical roadmap for retrieval and assessment of damaged buildings using UAV remote sensing disaster images.
[0095] Figure 2 It is a schematic diagram of the focused intersection-between ratio (JBJ) calculation.
[0096] Figure 3 It is a flowchart for locating the actual location of the damaged area.
[0097] Figure 4 It is a schematic diagram of the damaged building retrieval process.
[0098] Figure 5 It is the histogram of area A and area B in the fine-tuning training of AlexNet and GoogLeNet models.
[0099] Figure 6 This is a comparison chart of GoogLeNet's weight initialization training and model fine-tuning training results.
[0100] Figure 7 It is the histogram of the training results of GoogLeNet model weight initialization and model fine-tuning in area A and area B.
[0101] Figure 8 This is a graph of the training results of different sample sizes under fine-tuning of the GoogLeNet model in region A.
[0102] Figure 9 This is a comparison chart of the detection accuracy of the color histogram model, GoogLeNet, and AlexNet. DETAILED DESCRIPTION
[0103] The technical solution of the method for retrieving and evaluating damaged buildings in disaster images provided by this application will be further described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand this application and be able to implement it.
[0104] Currently, emerging technologies such as artificial intelligence and unmanned aerial vehicles (UAVs) are developing rapidly, and the acquisition of remote sensing images is more convenient, and the quality of the data itself is also increasing steadily. However, the overall development level of the intelligent processing of this type of remote sensing image data is very limited, and a large amount of valuable data itself cannot be fully utilized and cannot play its role.
[0105] This application conducts feature learning on damaged buildings in disaster images. From two training modes of fully initializing weights and fine-tuning training, based on two models, GoogLeNet and AlexNet, target detection of damaged buildings is carried out. The deeper fine-tuning training network model has better detection performance under large-scale data sets. The research roadmap of this application is as Figure 1 shown.
[0106] 1) Feature extraction of damaged building targets based on UAV remote sensing: Use GoogLeNet and AlexNet to extract the target features of damaged buildings and obtain more generalized target features of damaged buildings;
[0107] 2) Generation of damaged areas based on region prediction: Adopt the method of region prediction to generate a large number of prediction regions. When using UAV remote sensing for target detection, only detect and classify the prediction regions to ensure the diversity and multi-scale features of the prediction regions and improve the retrieval efficiency;
[0108] 3) Precise positioning of the target detection results based on region prediction: Use a method based on region prediction to perform regression fusion of the focus boxes on the detected damaged building targets. The JBJ focus intersection over union and the improved BKE method are used to fuse and eliminate the detection results with high overlap on the same target in UAV remote sensing target detection. It effectively improves the detection effect and also has a certain improvement in the target detection and positioning accuracy.
[0109] The results show that this retrieval method based on damaged buildings has good performance and good application prospects in the target detection of disaster images.
[0110] I. Generation of damaged areas based on region prediction
[0111] The damaged areas are detected on the image to be measured, including two points: one is the possible orientation of the target to be measured, and predicting this can improve the detection efficiency; the other is the true position of the target to be measured, that is, judging multiple adjacent damaged areas to weaken the situation of edge bounding and improve the bounding accuracy; when the high-level image features in the detection, that is, the damaged area targets are fixed, the whole of the image to be measured or the self-extracted part is used as the input element of the recognition model, and the category corresponding to each input damaged area image and the probability of being judged as this category are obtained. The selection of damaged areas greatly affects the final result of target detection and the running efficiency of the detection model.
[0112] (1) Generation of damaged areas based on geological disaster images
[0113] First, based on graph representation image segmentation, the image to be measured is split into a certain number of specific areas with special properties, and then the damaged area targets are extracted from them. In image segmentation, the segmentation boundaries between different areas are determined. The boundary establishment is based on the determination between specific areas under graph representation. The greedy algorithm is used for screening and then segmenting the image. The details of the areas with relatively large changes are selectively ignored, while the details of the areas with small changes are retained, so that the areas with rich diversity are aggregated, making the specific areas more consistent visually; in terms of time, the efficiency of splitting into multiple areas is linearly related to the number of pixels in the original image.
[0114] The graph representation expresses the image in a broad sense as a graph in graph theory, which is composed of several given points and the connections between points. Suppose there is a graph G=(V, E), and each pixel in the image is regarded as v i ∈V, and each two adjacent pixel points form an edge e i =(v i , v j )∈E. The difference in pixel color values constitutes the weight w(e i ) of E. w is the core index for measuring the difference between v i and v j . The weight is positively correlated with the pixel points. G=(V, E) is segmented into a series of non-intersecting C∈M, where M is the set of subgraphs G'. All elements in M are independent of each other and satisfy non-intersection, that is, similar v i are aggregated;
[0115] Assume that G has been simplified to the minimum spanning tree MST(C, E). The internal difference Int(C) of the segmented C is the maximum weight of E contained in the current C, as shown in Equation 1:
[0116]
[0117] The difference between C1 and C2 is calculated by the weight value of the minimum edge connecting their vertices. When there is no connecting edge between C1 and C2, Dif(C1, C2) = ∞ is defined, as shown in Equation 2:
[0118]
[0119] The determination of the boundary between two segmented regions is to compare the magnitudes of Dif and the minimum internal difference of the segmentation MInt, including Equations 3 and 4:
[0120]
[0121] MInt(C1, C2) = min(lnt(C1) + τ(C1), lnt(C2) + τ(C2)) Equation 4
[0122] If the boundary assumption is correct, then C1 and C2 are segmented; otherwise, they are merged. Among them, the threshold function τ controls the difference between the two. The threshold function is associated with the region size to better control the definition of the segmented region boundary, as shown in Equation 5:
[0123] τ(C) = k / |C| Equation 5
[0124] |C| refers to the number of pixel points, and k is a parameter adjusted according to the image size. The larger k is, the larger the obtained elements. During the implementation process, Gaussian smoothing is performed on the image to remove some noise with the least change to the image, and the overall effect of the image remains unchanged;
[0125] The steps of the overall algorithm are as follows:
[0126] Step 1: For E ∈ G, sort in ascending order according to W;
[0127] Step 2: S[0] is the initialization of the segmentation, that is, each pixel v is an independent segmented region C;
[0128] Step 3: q = 1, 2, 3,..., m (m is the number of edges). According to the segmentation result of the previous S[q - 1], select an o[q](v i ,v j ). If v i and v j belong to two non - intersecting segmented regions C respectively, and if o[q](v i ,v j ) < MInt, merge these two regions. If the determination is no, continue to perform the repeated operation;
[0129] Step 4: All the finally obtained C i ∈ S = S[m] is the set of segmentations that are neither too fine nor too rough obtained;
[0130] After obtaining the preliminary segmentation set, the set still has a large cardinality. Calculate the approximation degree of the original segmentation regions in the set, and adopt two multivariate entropy strategies to enhance the generation of multivariate entropy for the damaged areas of the image, including the multivariate entropy of the color space and the multivariate entropy of the approximation degree calculation. The former is applied to the generation of the original region set in image segmentation, while the latter is applied to the merging of the original regions;
[0131] In the multivariate entropy of the color space, multiple color spaces including RGB, grayscale, Lab, normalized RG channels plus grayscale, HSV, normalized RGB, H, S, and V are adopted to ensure a high degree of diversity in the generation of the original damaged areas, and various color spaces are used to consider the image scenes of the lighting conditions;
[0132] In the multivariate entropy of the approximation degree calculation, the specific calculation process is as follows:
[0133] (1) Color approximation degree: S color (r i , r j ) is used to calculate the color approximation degree of r i and r j in the regions to be merged, and each such region obtains a color distribution histogram When the number of color channels is 1, n = 25; if there are 3 color channels, then n = 75; next, normalize the obtained color histogram, as shown in Equation 6:
[0134]
[0135] After calculating the color approximation degree, merge r i and r j to obtain r t , and the calculation method of the color distribution histogram of the new region r t is shown in Equation 7:
[0136] C t = (size(r i ) × C i + size(r j ) × C j ) / size(r t ) Equation 7
[0137] The denominator in the equation is the new size of r t , that is, size(r t ) = size(r i ) + size(r j );
[0138] (2) Texture approximation degree: S texture (r i,r j ) It is used to calculate the texture approximation between regions. For each region, Gaussian smoothing with a variance of 1 is taken in 8 directions for each color channel, and a histogram with 10 bins is obtained for each color in each channel. After normalization with 3 color channels, a vector with n = 8 * 3 * 10 = 240 dimensions is obtained. The calculation method of the texture feature of the new region is similar to that of the color approximation, as shown in Equation 8:
[0139]
[0140] Similar to color, texture is transferred;
[0141] (3) Size approximation: S size (r i ,r j ) It is used to calculate the size approximation between regions. The size is the number of pixel points contained in the region. By calculating the ratio of the sum of the sizes of two regions to the total number of pixel points, it is judged whether the region is small and whether it should be merged as early as possible. Specifically, see Equation 9:
[0142]
[0143] Where size(im) represents the total number of pixels in the original image;
[0144] (4) Filling approximation: S fill (r i ,r j ) It is used to calculate the filling approximation between regions, and measures whether there is an intersection or inclusion relationship between two regions. If so, they are merged first to avoid blank regions in the output result. Calculate the minimum enclosing focused region BB i of r j and r ij (focus box). The smaller the size(BB ij ), the higher the filling approximation, that is, the greater the possibility of intersection or inclusion between the two regions. Subtract the sum of the number of points of r i and r j from the number of pixel points of the minimum enclosing region, and compare the obtained value with the total number of pixel points to judge whether r i and r j are in a relationship of mutual inclusion or intersection. Specifically, see Equation 10:
[0145]
[0146] Finally, all four calculation methods are combined together, and the structure is as shown in Equation 11:
[0147] S(r i ,r j ) = a1Scolor (r i ,r j )+a2S texture (r i ,r j )+a3S size (r i ,r j )+a4S fill (r i ,r j ) Equation 11
[0148] where a i ∈{0, 1} indicates that the four approximation principles of color, texture, size, and filling are all optional;
[0149] Through image segmentation, the original damaged area and the approximation fusion of the damaged area are generated. Finally, a set of focus boxes that may contain objects is obtained. The average highest overlap rate ABO obtained by comparing the focus boxes generated by the algorithm with the area where the real situation is located is used to measure. For each determined category C, each real situation is represented as G is the target area where the object is located, and 1 is a sub-element of the set L of calculated position hypotheses. Then the expression of ABO is Equation 12:
[0150]
[0151] The calculation method of the overlap rate is Equation 13:
[0152]
[0153] Calculate the respective ABO for each category in turn, and then use the average value MBAO of the ABO of all categories for evaluation;
[0154] In the early stage of the algorithm, the similarity calculation method and the types of feature calculation methods are set in advance to limit the number of focus boxes obtained finally. The initial size threshold of the damaged area is set to 250 and 400, and finally about 200 focus boxes are obtained.
[0155] Object detection takes into account both the multi-level and diversity of the prediction area. First, the original segmentation area is obtained through image segmentation based on graph representation, and the areas with large differences are segmented. Then, classification and merging are performed according to the approximation calculation method of multiple indicators, and the areas with high approximation are merged. Through multiple data iterations, the final set of focus boxes is obtained.
[0156] (II) Location of the real position of the damaged area
[0157] The intersection over union of focus (IoU) is used to calculate the overlap rate between the detected focus box and the true calibration data, and to measure the accuracy of object detection. Given the focus box to be detected G1{(minx1, miny1)(maxx1, maxy1)} and the true position G2{(minx2, miny2)(maxx2, maxy2)}, assume that G1 and G2 have an intersection G u {(minx, miny)(maxx, maxy)}, where the specific values are shown in Equations 14 to 16:
[0158] minx = max(minx1, minx2)
[0159] miny = max(miny1, miny2)
[0160] maxx = min(maxx1, maxx2)
[0161] maxy = min(maxy1, maxy2) Equation 14
[0162] S(G u ) = (maxx - minx)(maxy - miny) Equation 15
[0163]
[0164] where S(G u ), S(G1), and S(G2) are the areas of the intersection G u and G1 and G2 respectively, and G u = G1 ∩ G2. If G1 and G2 do not intersect, then minx > maxx or miny > maxy. The graphical illustration is as Figure 2 .
[0165] Meanwhile, the critical value of IoU is determined based on prior knowledge and the relative relationship between the image size and the target size. In this application, the critical value is set to 0.6.
[0166] For focus area overlap and region merging, a large number of candidate boxes obtained from selective search are processed in three steps:
[0167] Step 1: Generate a large number of candidate windows through a sliding window or a candidate object generation scheme for selective search;
[0168] Step 2: Classify the candidate windows generated above using a classifier generated by a support vector machine, and assign each window a score as a subsequent evaluation criterion;
[0169] Step 3: Fuse the detection results to obtain the final detection result;
[0170] The obtained data includes the position information and scores of the focus boxes screened by the classifier. The focus boxes are fused on the same target. Suppose there are 8 focus boxes that frame the object to be measured, namely A, B, C, D, E, F, G, a total of 8 focus boxes. Set the intersection-over-union (IoU) threshold to a. The specific steps are as follows:
[0171] Step 1: Sort all the boxes in ascending order of the scores given by the classifier;
[0172] Step 2: Starting from the box A with the highest score, calculate its IoU with the other remaining boxes. If the ratio of the IoU to the area of A is greater than the given threshold a, then put that box into group G A , and the members within the group all have a large overlap with A; if there is no such box, then A forms a group by itself, and the remaining boxes B - G are grouped in the same way to obtain other groups (if any, denoted as G B -G G );
[0173] Step 3: For each set obtained after grouping, calculate the optimal regression box for each group, denoted as T, T = {T1, T2,..., T n}, where t n is the optimal regression box for the nth group.
[0174] Step 4: During the grouping calculation, the boxes with high IoU are clustered towards a specific target, and there may still be a situation where there are multiple boxes on a single target. To solve this problem, use a gradient IoU threshold that decreases step by step with the number of iterations, and then repeat steps 1, 2, and 3: re-group T according to the boxes in T and calculate the optimal solution, and finally completely separate the detection results to obtain a one-to-one corresponding result. Figure 3 This is the overall algorithm flow chart.
[0175] II. Damaged building retrieval process
[0176] It includes two parts. The training process uses the training dataset and the validation dataset to generate a network with classification using the Caffe framework with the help of a graphics card and CuDNN acceleration. The detection process uses the model obtained from the training process to detect the targets in the image to be measured, and compares the detection results with the true values pre-annotated manually to obtain the duplicate detection rate and the recall rate. The specific flow chart is as Figure 4 .
[0177] (I) Training process
[0178] 1) Convert the image dataset from the image format to a high-performance database to facilitate the efficient reading of data by the Caffe framework, and at the same time generate file labels;
[0179] 2) Calculate the mean file using the Caffe toolset;
[0180] 3) Select a network model and modify the model definition file;
[0181] 4) Develop corresponding training strategies;
[0182] 5) Select a training mode, randomly generate a weight initialization model, further train on the initialized model, and fine-tune the training of the model.
[0183] During the training process, update the propagation weights through error backpropagation. At the same time, use the momentum factor to smooth the weights to accelerate convergence. Use Dropout to randomly make the weights of some hidden layer nodes in the network not work. The non-working part of the nodes is temporarily considered not part of the network structure, but the weights are retained and not updated this time, and will be updated when the next wave of sample data is input.
[0184] (2) Detection process
[0185] After the training process ends, a detection model based on damaged residential buildings is obtained. The detection process uses this model to perform target detection on the image to be measured, which specifically includes the following steps:
[0186] 1) Use the selective search algorithm to generate a large number of focus boxes on the image of the area to be measured, record and retain the positions of the focus boxes;
[0187] 2) The trained detection model makes a category determination for the image data inside each focus box. The retained data are the features that are judged to be damaged residential building categories. Those judged to be other feature types such as landslides, factory buildings, and intact residential buildings are directly excluded from the results first;
[0188] 3) The content inside the remaining focus boxes is judged to be damaged residential buildings, but there are a large number of overlapping situations on the same target. Cross-use and focus on the intersection over union and the improved non-maximum suppression method to group and fuse all the remaining focus boxes, and eliminate redundant boxes by considering the area and the model score respectively. Finally, ensure that there is only one focus box for each target;
[0189] 4) Finally, calculate the recall rate and precision rate of the target detection, and calculate them separately for different models and different focus box production methods.
[0190] III. Detection Results and Analysis of Damaged Buildings Based on UAV Remote Sensing
[0191] (1) Target detection results
[0192] First, the image to be tested in area A is used to obtain the focus frames to be tested at multiple scale levels using the test area generation algorithm. Then, the classifier obtained after 50,000 iterations detects the data in all frames, determines the probability that the content of the focus frame is a damaged residential building, and labels it with a category. A group of focus frames with different sizes and scores that are determined to be damaged residential buildings are obtained, and then non-maximum suppression and focus intersection are used for fusion and elimination. Focus frames with low scores and large overlap ratios with higher-scoring areas are given priority for elimination or fusion, reducing the overall number of iterations and the amount of calculation. Finally, after multiple fusion and elimination, an optimal solution for the current detection of damaged residential houses is obtained.
[0193] The system evaluation indicators include: accuracy, recall, F value (weighted harmonic mean of accuracy and recall), and E value (weighted mean of accuracy and recall). In the final accuracy evaluation, these four indicators are used to evaluate the model detection under different experimental strategies.
[0194] In the experiment, two different transfer learning modes were used to train GoogleNet and AlexNet. The accuracy of model fine-tuning was better. The specific results of the four evaluation indicators are shown in Figure 5 . Overall, GoogLeNet's model fine-tuning is better than AlexNet's model fine-tuning. There are two reasons for choosing model fine-tuning. First, training networks such as AlexNet requires millions or even tens of millions of data such as imagenet, places (2.5 million image samples, a total of 205 image categories), places2 (10 million image samples, a total of 400 image categories), etc., while ordinary computer vision tasks cannot obtain data of the same order of magnitude. Second, because the features of the pre-trained model itself are sufficiently generalized, transfer learning can be performed and directly applied to another computer vision task. Between developing a new model and fine-tuning the model, model fine-tuning is chosen to complete the current task, because developing a new convolutional neural network classification model requires accumulating training experience and intuition, as well as a large amount of computing resources to try different network structures. By comparing the overall accuracy of the two networks, it is found that although the structure of AlexNet is not as complex as GoogLenet, it can also achieve better performance, which shows that AlexNet's own feature learning ability is still very good. This also shows that when performing tasks related to computer vision, it is not necessary to blindly pursue the latest and most complex models, but to fully consider the target focus and matching degree of the task.
[0195] from Figure 5It can be seen that the fine-tuning of the GoogLeNet model is generally better than that of the AlexNet in terms of accuracy and recall, with a gap of about 8 percentage points. This indicates that the performance of the object detection in this experiment obtained by fine-tuning the GoogLeNet is better than that obtained by fine-tuning the AlexNet. At the same time, the training time of the GoogLeNet fine-tuning model is twice as fast as that of the AlexNet fine-tuning model. In response to this phenomenon, it shows that the field of computer vision is indeed evolving towards better methods. In the process of model training in this experiment, various factors such as Dropout, ReLU activation function, LocalResponseNormalization, and Overlapping on the fitting of the model itself are also comprehensively considered. After many attempts, the detection accuracy shown in the table is finally obtained.
[0196] (2) Comparative analysis of different transfer learning modes
[0197] For the damaged residential building detection of this application, the dataset is much smaller than the original dataset, but there are significant differences in content. Training only a linear classifier may have better performance. Because the datasets are different, it is theoretically better to train the classifier starting from the activation function in the front.
[0198] Model fine-tuning targets multiple factors of the same model, the size of the new data, and the similarity between the new dataset and the original dataset. After selecting the network models of GoogLeNet and AlexNet, combined with the dataset of this application itself, two methods of random initialization and model fine-tuning are used for data training. The detection results of the GoogLeNet model weight initialization training and model fine-tuning are as Figure 6 and Figure 7 shown.
[0199] From Figure 7 Overall, looking at the histogram of the training results of the weight initialization and model fine-tuning of the GoogLeNet model in Area A and Area B, after model fine-tuning, GoogLeNet performs better than before, specifically manifested in the improvement of accuracy, reflected in the reduction of misdetected objects, and also a certain improvement in recall. The overall trend of Fl has a certain increase, and the value of E1 has a certain decrease.
[0200] (3) Comparative analysis of different training dataset sizes
[0201] The quality and scale of the dataset have a significant impact on the final fitting effect of the model, which is also the most time-consuming part of the entire experimental process. The impact of the dataset is much greater than that of the network structure and model fine-tuning. The optimization of the network structure and the adjustment of model parameters aim to prevent the weakening of the model's learning effect or overfitting. However, a poor data sample will be amplified during multiple learning, iteration, propagation, and feature synthesis of the model, directly resulting in too low accuracy of the model and making it unusable at all.
[0202] If the dataset is too small, overfitting is likely to occur because the deep model itself has not yet generalized the deep features in the image to a certain extent and reaches the node of model fitting. This will make the model too strict and not easily applicable to general scenarios. The principle of feature learning of the deep model is to abstract the features of the image deeply and then combine and rearrange the abstract features according to certain rules, presenting as a high-dimensional feature vector. This kind of feature vector represents a certain type of feature. After learning, it will continuously fit the optimal solution, and the detection process is actually a process of matching deep feature vectors. The computer itself doesn't know the real content of the current image, but through image feature comparison, it can judge the similarity between the features in the model and the features to be detected, and then give the detection result.
[0203] Three scales are involved in the experiment. Taking Area A as an example, the impact of different dataset sizes on the detection effect is studied. The selected model is obtained by fine-tuning the GoogLeNet model. The specific data and detection results are shown in Figure 8 , where 1 represents the initial sample size, 8 and 88 respectively represent the expansion ratios, and at the same time, the size of the negative samples has also changed accordingly. The three situations adopt the control variable method, and performance tests are carried out under other parameters and the same environment.
[0204] When the sample ratio is 1, both the accuracy and recall rate are relatively low, unable to meet the detection requirements of the task. Correspondingly, the weighted harmonic mean of the accuracy and recall rate and the weighted average of the accuracy and recall rate lose their original meanings. When the sample ratio increases to 8 times, it can be seen that both the accuracy and recall rate have increased significantly, but still do not meet the requirements of object detection. When the sample ratio increases to 88 times, both the accuracy and recall rate reach good values. At this time, the weighted harmonic mean and the weighted average also reflect that the model has relatively good performance at this time. In the target of damaged residential buildings detected in this application, the feature differences in different regions are relatively obvious. Therefore, although target samples at the same scale are extracted from multiple regions, there are still certain limitations. Within a certain range, the larger the dataset size, the better, because only with a large amount of data can the target features be generalized better and the model can find the targets to be detected in more complex environments.
[0205] (4) Detection performance comparison with other feature extraction methods
[0206] A commonly used method in the field of image processing is to extract the color histogram of the image itself. This method is only used to reflect the proportion of a single color in the image, and the distribution of the color itself in the color space is not the focus of this method. Although this method cannot deeply understand the objects in the image, it is more suitable for describing images whose own color distribution is too complex to be separated. RGB is a widely used color space, but there is a certain gap between the expression of the RGB color space and people's subjective perception in vision. Therefore, this application uses the HSV space to conduct this performance comparison analysis experiment. Converting H (hue), S (saturation), and V (value) respectively in space is a method of quantifying the color space. The specific experimental results are shown in Figure 9 .
[0207] In terms of precision, recall, and F-value, both AlexNet and GoogLeNet are much better than the color histogram model, and both have about twice the performance of the latter. Although there is a gap between GoogLeNet and AlexNet, it is within 10%. It can be seen that the damaged building detection model based on has better detection accuracy than traditional image processing methods. Considering the training cost and model fitting time comprehensively, GoogLeNet is more suitable to be selected as the original model for transfer learning than AlexNet.
[0208] The detection effect of the GoogLeNet fine-tuning model is better than that of the AlexNet fine-tuning model, and the GoogLeNet fine-tuning model is better than the GoogLeNet weight initialization model. At a certain scale, the larger the scale of the dataset, the better.
Claims
1. A method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images, characterized in that: Feature learning of damaged buildings in disaster images, using two training modes: full weight initialization and fine-tuning training, to detect damaged buildings based on GoogLeNet and AlexNet models. The deeper fine-tuning training network model has better detection performance under large-scale data sets; 1) Damaged building target feature extraction based on UAV remote sensing: GoogLeNet and AlexNet are used to extract damaged building target features to obtain more generalized damaged building target features; 2) Generation of damaged areas based on regional prediction: Using the regional prediction method, a large number of prediction areas are generated. When using UAV remote sensing for target detection, only the prediction areas are detected and classified to ensure the diversity and multi-scale characteristics of the prediction areas and improve the retrieval efficiency. 3) Accurate positioning of target detection results based on regional prediction: A method based on regional prediction is used to regress and fuse the focus frames of the damaged building targets detected. The JBJ focus intersection and union ratio and improved BKE method are used to fuse and eliminate the detection results with high overlap on the same target in UAV remote sensing target detection.
2. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 1 is characterized in that: Damaged area generation based on regional prediction: The damaged area is detected on the image to be tested, including two points: one is the possible orientation of the target to be tested, and prediction is made to improve the detection efficiency; the other is the real position of the target to be tested, that is, multiple adjacent damaged areas are judged, the edge selection is weakened, and the accuracy of the selection is improved; when the high-level image features of the detection, that is, the damaged area target is fixed, the whole image to be tested or the extracted part is used as the input element of the recognition model to obtain the category corresponding to all the input damaged area images and the probability of being judged as this category.
3. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 1 is characterized in that: Generation of damaged areas based on geological disaster images: First, based on the image segmentation represented by the graph, the image to be tested is split up to obtain a certain number of specific areas with special properties, and then the damaged area targets are extracted from them. In the image segmentation, the segmentation boundaries between different areas are determined. The boundary establishment is based on the judgment between specific areas under the graph representation. The greedy algorithm is used to screen and then segment the image, and the details of the areas with relatively large detail changes are selectively ignored, while the details of the areas with small detail changes are retained, so that the diverse areas are gathered, and the specific areas are more consistent visually; the efficiency of segmenting multiple areas in time is linearly related to the number of pixels in the original image; Graph representation expresses the image in a general sense as a graph in graph theory, which is composed of a number of given points and the lines between the points. Suppose there is a graph G = (V, E), and each pixel in the image is regarded as v i ∈V, every two adjacent pixels form an edge e i =(v i ,v j )∈E, the difference in pixel color values constitutes the weight of E w(e i ), w as v i and v j The core indicator for measuring the difference between them is that the weight is positively correlated with the pixel points. G = (V, E) is divided into a series of non-intersecting C∈M, where M is the set of subgraphs G'. All elements in M are independent of each other and do not intersect with each other. i to gather; Assume that G has been simplified to the minimum spanning tree MST(C,E), the internal difference Int(C) of C obtained by segmentation is the maximum weight of E contained in the current C, as shown in Formula 1: The difference between C1 and C2 is calculated by the weight of the smallest edge connecting their vertices. When there is no connecting edge between C1 and C2, Dif(C1, C2) = ∞, as shown in Formula 2: The boundary between two segmented regions is determined by comparing Dif with the minimum segmentation internal difference MInt, including equations 3 and 4: MInt(C1,C2)=min(lnt(C1)+τ(C1),lnt(C2)+τ(C2)) Equation 4 If the boundary assumption is correct, C1 and C2 are segmented, otherwise they are merged, where the critical value function τ controls the difference between the two. The critical value function is associated with the region size to better control the definition of the segmentation region boundary, see formula 5: τ(C)=k / |C| Equation 5 |C| refers to the number of pixels, and k is a parameter adjusted according to the image size. The larger k is, the larger the element is. In the implementation process, Gaussian smoothing is performed on the image to remove some noise with minimal changes to the image, and the overall effect of the image remains unchanged.
4. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 3 is characterized in that: The steps of the overall algorithm are divided into: Step 1: For E∈G, sort them in ascending order according to W; Step 2: S[0] is the initialization of segmentation, that is, each pixel v is an independent segmentation area C; Step 3: q = 1, 2, 3, ..., m (m is the number of edges), select an o[q] (v i ,v j ), if v i and v j They belong to two partitioned regions C respectively and do not intersect each other. If o[q](v i ,v j )<MInt, merge the two regions. If the judgment is no, continue to perform the repeated operation; Step 4: All the Cs obtained in the end i ∈S=S[m] is the segmentation set that is neither too fine nor too coarse; After obtaining the preliminary segmentation set, the set still has a large cardinality. The approximation of the original segmented area in the set is calculated. Two multivariate entropy strategies are used to enhance the multivariate entropy of the damaged area of the image, including color space multivariate entropy and approximation calculation multivariate entropy. The former is applied to the generation of the original area set in image segmentation, while the latter is applied to the merging of the original area.
5. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 4 is characterized in that: In the color space multivariate entropy, multiple color spaces including RGB, grayscale, Lab, normalized RG channel plus grayscale, HSV, normalized RGB, H, S, V are used to ensure that the original damaged area is generated with high diversity, and multiple color spaces are used to consider the image scenes with lighting conditions; In the approximation calculation of multivariate entropy, the specific calculation process is as follows: (1) Color similarity: S color (r i ,r j ) is used to calculate r in the area to be merged i and r j The color similarity of each region is obtained as a color distribution histogram When the color channel is 1, n = 25; if there are 3 color channels, n = 75; then the obtained color histogram is normalized, see formula 6: After the color approximation is calculated, i and r j Merge and get r t , the new region r t The color distribution histogram calculation method is shown in formula 7: C t =(size(r i )×C i +size(r j )×C j ) / size(r t ) Equation 7 The denominator in the formula is r t The new size, size(r t )=size(r i )+size(r j ); (2) Texture approximation: S texture (r i ,r j ) is used to calculate the texture similarity between regions. For each region, 8 directions of each color channel are taken, and Gaussian smoothing with a variance of 1 is performed. A histogram of 10 bins is obtained for each color of each channel. The color channel is normalized to 3, and a vector of n = 8 * 3 * 10 = 240 dimensions is obtained. The method for calculating the texture features of the new area is similar to the method for calculating the color approximation, see Equation 8: Like color, texture is transmitted; (3) Size approximation: S size (r i ,r j ) is used to calculate the size similarity between regions. The size is the number of pixels contained in the region. By calculating the ratio of the sum of the sizes of two regions to the total number of pixels, it is determined whether the region is small and whether it should be merged as soon as possible. For details, see formula 9: Where size(im) represents the total number of pixels of the original image; (4) Filling approximation: S fill (r i ,r j ) is used to calculate the filling approximation between regions, measuring whether there is an intersection or inclusion relationship between two regions. If so, they are merged first to avoid blank areas in the output results. The calculation obtains r i and r j Minimum outsourcing focus area BB ij , size(BB ij ) is smaller, the higher the filling approximation is, that is, the greater the possibility that the two areas intersect or contain each other, the smaller the number of pixels in the smallest enclosing area minus r i and r j The sum of the points is compared with the total number of pixels to determine r i and r j Whether they are mutually inclusive or intersecting relationships, see formula 10 for details: Finally, all four calculation methods are combined together, and the structure is as shown in formula 11: S(r i ,r j ) = a1S color (r i ,r j ) + a2S texture (r i ,r j ) + a3S size (i, r j ) + α4S fill (r i ,r j ) Equation 11 Among them, a i ∈{0, 1} means that the four approximation principles of color, texture, size, and filling are all optional; The original damaged area and the approximate degree of the damaged area are generated by image segmentation, and finally a bunch of focus frames that may contain objects are obtained. The image samples are trained with positive and negative samples, and the focus frames generated by the algorithm are compared with the areas where the real situation is located to obtain the average maximum overlap rate ABO. For each determined category C, each real situation is expressed as G is the target area where the object is located, 1 is the sub-element of the calculated position hypothesis set L, and the expression of ABO is formula 12: The overlap ratio is calculated as follows: The ABO is calculated for each category in turn, and then the MBAO, the average ABO value of all categories, is used for evaluation.
6. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 5 is characterized in that: In the early stage of the algorithm, the similarity calculation method and the type of feature calculation method are set to limit the number of focus frames obtained in the end. The size critical values of the initial damaged area are set to 250 and 400, and finally about 200 focus frames are obtained; Target detection takes into account the multi-level and diversity of the prediction area. First, the original segmented area is obtained through image segmentation based on graph representation, and the areas with large differences are separated. Then, they are classified and merged according to the similarity calculation method of multiple indicators, and the areas with high similarity are merged. After multiple data iterations, the final focus frame set is obtained.
7. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 1 is characterized in that: The real position of the damaged area is located by using the intersection-over-joint ratio (JBJ) to calculate the overlap rate between the detection focus frame and the real calibration data, and to measure the accuracy of target detection. The focus frame to be detected G1 (minx1, miny1) (maxx1, maxy1)} and the real position G2 {(minx2, miny2) (maxx2, maxy2)} are set. Assume that G1 and G2 have an intersection G u {(minx, miny)(maxx, maxy)}, where the specific values are shown in equations 14 to 16: minx=max(minx1,minx2) miny=max(miny1,miny2) maxx=min(maxx1,maxx2) maxy=min(maxy1,maxy2) Formula 14 S(G u )=(maxx-minx)(maxy-miny) Formula 15 Where S(G u ), S(G1), S(G2) are the intersection sets G u The area of G1 and G2, G u =G1∩G2, if G1 and G2 do not intersect, there exists minx>maxx or miny>maxy; Meanwhile, the critical value of JBJ is determined by prior knowledge and the relative relationship between the image size and the target size. In this application, the critical value is set to 0.
6.
8. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 7 is characterized in that: For focus area overlap and area merging, processing a large number of candidate boxes obtained by selective search is divided into three steps: Step 1: Generate a large number of candidate windows through a sliding window or selective search candidate target generation scheme; Step 2: Use the classifier generated by the support vector machine to classify the candidate windows generated previously, and give each window a score as a subsequent evaluation criterion; Step 3: Fuse the detection results to obtain the final detection results; The data obtained includes the position information and scores of the focus frames filtered by the classifier. The focus frames are fused on the same target. Assume that there are 8 focus frames selected to the object to be tested, namely A, B, C, D, E, F, and G, a total of 8 focus frames. The intersection-over-union ratio critical value is set to a. The specific steps are as follows: Step 1: Sort all the boxes from small to large according to the scores given by the classifier; Step 2: From the highest-scoring box A, calculate its intersection-and-union ratio with the remaining boxes. If the ratio of the intersection-and-union ratio to the area of A is greater than the given critical value a, then put the box into group G. A , the members of the group all have a large overlap with A; If there is no such box, then A is a group by itself, and the remaining boxes BG are grouped in the same way to obtain other groups (if they exist, recorded as G B -G G ); Step 3: For each set obtained after grouping, calculate the optimal regression box of each group, denoted as T, T = {T1, T2, ..., T n }, t n is the optimal regression box for the nth group. Step 4: Group calculation: The boxes with high IoU ratio are clustered towards a specific target. There may still be multiple boxes on a single target. To solve this problem, the gradient IoU ratio critical value is used, which is gradually reduced with the number of iterations. Then steps 1, 2, and 3 are repeated: T is regrouped according to the box situation in T, and the optimal solution is calculated. Finally, the detection results are completely separated to obtain one-to-one corresponding results.
9. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 1, characterized in that: Damaged Building Training Process: 1) Convert the image dataset from image format to a high-performance database to facilitate the Caffe framework to read the data efficiently and generate file tags; 2) Calculate the mean file using the Caffe toolset; 3) Select the network model and modify the model definition file; 4) Develop corresponding training strategies; 5) Select the training mode, randomly generate weights to initialize the model, further train on the initialized model, and fine-tune the model. During the training process, the error back propagation is used to update the propagation weights. At the same time, the momentum factor is used to smooth the weights to accelerate convergence. Dropout is used to randomly make the weights of some hidden layer nodes in the network non-working. The non-working nodes are temporarily determined not to be part of the network structure, but the weights are retained and will not be updated this time. They will be updated when the next wave of sample data is input.
10. The method for retrieval and assessment of damaged buildings using UAV remote sensing disaster images according to claim 1, characterized in that: Damaged building detection process: After the training process is completed, a detection model based on damaged residential buildings is obtained. The detection process uses this model to perform target detection on the image to be tested. Specifically, it includes the following steps: 1) Generate a large number of focus frames on the image of the area to be measured using a selective search algorithm, and record and retain the positions of the focus frames; 2) The trained detection model determines the category of the image data within each focus frame. The retained data are those that are identified as damaged residential buildings. Other types of objects, such as landslides, factory buildings, and intact residential buildings, are directly removed from the results. 3) The contents inside the remaining focus frames are judged as damaged residential buildings, but there are a lot of overlaps on the same target. All the remaining focus frames are grouped and merged using the intersection-over-union and improved non-maximum suppression methods. The area and model score are considered to eliminate redundant frames, and finally there is only one focus frame on each target. 4) Finally, the recall and precision of target detection are calculated for different models and different focus frame production methods.