A water system information extraction algorithm based on FasterR-CNN
Through the water system information extraction algorithm based on FasterR-CNN, combined with AlexNet transfer learning and morphological algorithm processing, the problem of low accuracy and efficiency of water system information extraction in the existing technology is solved, and the rapid and accurate identification of river network water systems is achieved.
Patent Information
- Application Number
- CN202111413901.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-25
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-11-25
AI Technical Summary
The prior art is difficult to adapt to diversified monitoring scenarios in the extraction of river network water system information. Image noise interference leads to low accuracy and efficiency. Algorithm parameters require manual intervention, low accuracy, and difficult to achieve rapid, quantitative and refined recognition.
The water system information extraction algorithm based on FasterR-CNN is used to pre-process remote sensing image data, establish a sample database, conduct deep learning of the FasterR-CNN network, combine AlexNet for transfer learning and morphological algorithm processing, and use the F1 score index to reduce the error detection rate.
It realizes rapid identification of the types, locations and widths of water systems, improves the accuracy and efficiency of water system information extraction, and is suitable for diversified water system monitoring scenarios.
Smart Images

Figure CN114140698B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning, and in particular, relates to a water system information extraction algorithm based on FasterR-CNN. Background Art
[0002] The river network is a very important part of the topography and hydrological characteristics, and research in this area is also of great significance to maintaining the stability of the ecological environment and the conservation of water and soil.
[0003] At present, the algorithms and methods for extracting river network water system information mostly use vision-based water system detection methods, such as adaptive threshold segmentation, edge method, wavelet analysis, etc., and there are still the following problems: it is difficult to apply a variety of water system monitoring and identification scenarios; image noise interference leads to low accuracy and efficiency; the algorithm has many parameters and requires manual intervention, which results in low accuracy. On the other hand, the resolution of a single water system image is high, and the amount of image processing data is large. The current method is difficult to achieve rapid, quantitative, and refined identification of water systems. Summary of the invention
[0004] In view of the above problems existing in the prior art, the object of the present invention is to provide a water system information extraction algorithm based on FasterR-CNN.
[0005] In order to solve the above problems, the technical solution adopted by the present invention is as follows:
[0006] A water system information extraction algorithm based on FasterR-CNN, characterized by comprising the following steps:
[0007] Step 1: Preprocess the remote sensing image data for training and establish a sample database;
[0008] Step 2: Use the data in the sample database to perform deep learning on the FasterR-CNN network and complete the algorithm training;
[0009] Step 3: Use the trained FasterR-CNN network to detect the missed detection rate and calculate the F1 score indicator;
[0010] Step 4: Use FasterR-CNN to calculate the actual collected data, obtain the training sample set, use AlexNet for transfer learning, and use AlexNet to locate the crack skeleton;
[0011] Step 5: Based on the water system classification results obtained in step 4, use the morphological algorithm to calculate the shape and width of the water system;
[0012] Step 6: Reduce the false positive rate of the final result through the F1 score indicator and output the final result.
[0013] The pretreatment in step 1 comprises the following steps:
[0014] Step 1: Extract image features through ResNet residual network and generate convolution feature map;
[0015] Step 2: Use the convolutional feature map to classify the water system through the region proposal network RPN.
[0016] The F1 score index is obtained by the following formula:
[0017]
[0018] Where: TP: true positive, FP: false positive, FN: false negative.
[0019]
[0020] The steps of the morphological algorithm are as follows:
[0021] Step 1: Extract the water system skeleton processed by Faster R-CNN and the pixel with the minimum gray value in each column processed by CNN in the bounding box area to form the initial basic skeleton of the water system;
[0022] Step 2: Use a 3*3 sliding window to traverse the image in step 1. If there is only one pixel inside the sliding window, filter out the pixel in this area.
[0023] Step 3: Filter out the connected areas with smaller areas through the area filter, and perform the connection operation of adjacent connected domains in the remaining connected areas;
[0024] Step 4: Skeletonize the water system morphology to obtain a complete continuous skeleton.
[0025] Step 5: Repeat steps 3 and 4 to obtain the width of the water system.
[0026] The AlexNet neural network is composed of five convolutional layers and three fully connected layers, with a total depth of eight layers. The output of the last fully connected layer is sent to the Softmax layer.
[0027] Compared with the existing technology, the present invention adopts the FasterR-CNN algorithm to quickly identify the type, location and width of the water system, and uses AlexNet transfer learning to locate the crack skeleton for the extracted water system area. Finally, the morphological algorithm is used to extract the river morphology and calculate its length and width. In order to solve the problem of low missed detection rate and high single false positive rate of FasterR-CNN, the F1 score indicator is introduced to reduce its false positive rate, which is suitable for application scenarios with more diversified water systems in reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a data object terrain range map;
[0029] Figure 2 This is the overall flow chart of this algorithm;
[0030] Figure 3 This is the architecture diagram of FasterR-CNN;
[0031] Figure 4 Figure 2 shows the actual usage of the region proposal network RPN.
[0032] Figure 5 This is a diagram of the AlexNet training process;
[0033] Figure 6 This is the AlexNet neural network model diagram;
[0034] Figure 7 This is the skeleton diagram of the water system after extraction. DETAILED DESCRIPTION
[0035] The present invention is further described below in conjunction with specific embodiments.
[0036] Example 1
[0037] This embodiment uses the terrain of Chongli District to demonstrate the algorithm.
[0038] The terrain range of the data collection object is as follows Figure 1 As shown, the study area is located in Chongli District, which belongs to Zhangjiakou City, Hebei Province. The geographical location is 40°47′-41°17′N, 114°17′-115°34′E. The area is 2334 square kilometers, the altitude is 813-2174m, and the maximum height difference is 1361m. Field observation data show that the average daily maximum temperature in the area is 12°, the average daily minimum temperature is -2°, the annual average water volume is 488 mm, the total precipitation is 1.13 billion cubic meters, the average annual runoff of surface water in the whole area is 42.9 mm, and the annual runoff volume is 100.69 million cubic meters. 80% of the territory is mountainous, and the forest coverage rate is as high as 52.38%. The territory is northwestern Hebei mountainous, mostly in the direction of North China-Southwest and East-West.
[0039] This embodiment uses the elevation data provided by ASTER, and performs preprocessing such as data mosaicking, coordinate conversion, and data clipping on the downloaded data to obtain a DEM original image that is consistent with the size of the study area. ASTER GDEM is one of the most commonly used surface digital elevations and is widely used in water system information extraction research.
[0040] Due to the influence of spatial resolution and system errors, some erroneous "depressions" will be generated during the generation of DEM data, which will cause incorrect water flow direction in the depressions in the DEM when extracting the river flow direction, thereby affecting the accuracy of river network extraction. Therefore, it is necessary to "fill depressions" on the original DEM to eliminate the influence of depressions. However, not all depressions are caused by data errors. Some depressions are the reflection of real terrain. Based on this theory, a reasonable filling threshold should be set when "filling depressions" on the original DEM to obtain the final initial data.
[0041] The overall flow chart of this algorithm is as follows: Figure 2 As shown in the figure, there are various means of collecting water system surface images in the existing technology. UAVs, professionally configured cameras, and remote sensing image data can all achieve fast, non-contact, and high-quality data collection. First, the collected data is preprocessed, the water system images are cropped and segmented, and then the image size is unified and the image data is enhanced. Then, the water system is classified according to its type, and the images are annotated using special software to establish a corresponding sample database; after the algorithm starts running, the sample data for training is input into the constructed Faster R-CNN network, and the entire network is deeply learned by adjusting the parameters; finally, the trained network is tested.
[0042] After training the network, the Faster R-CNN model is tested to detect the false detection rate of its images, and the precision rate P (Precession) and recall rate R (Recall) are introduced for algorithm evaluation. The definitions of each indicator are as follows:
[0043]
[0044] In the formula: TP: True Positive, FP: False Positive, FN: False Negative. According to the above formula, we can get the missed detection rate R-1 and the false detection rate P-1.
[0045] Finally, the F1 score indicator is introduced to reduce the false positive rate. The calculation method of the F1 score indicator is as follows:
[0046]
[0047] For the identified water system, after the algorithm determines its bounding box, it further refers to the AlexNet transfer learning in CNN and combines it with morphological algorithms to perform subsequent processing on the water system, extract the morphology of the water system, and obtain the length and width of the water system in the image of the water system, making the water system information more accurate and rich.
[0048] The FasterR-CNN algorithm mentioned above is derived from CNN and is an advanced version of the target detection algorithm R-CNN and Fast R-CNN. Its architecture is as follows Figure 3 As shown, first input the picture and extract the picture features through ResNet. The advantage of ResNet is that the use of residual network can improve the recognition accuracy by increasing the depth, and the more layers of the network, the richer the features extracted at different levels. Then, on the generated convolutional feature map (Feature Map), the region proposal network RPN is used to generate candidate windows and judge whether it is the foreground (target / water system). If the relevant area is classified as the foreground, the water system is classified through the fully connected layer, and the window coordinates are fine-tuned through border regression. Finally, a rectangular box is used on the output image to mark the specific water system category and give its classification confidence.
[0049] The region proposal network RPN belongs to the category of FCN (fully convolutional networks). Its purpose is to obtain high-quality water system candidate region boxes on the image and replace the previous selective search method so that the entire network can be trained end-to-end. Figure 4 As shown in the figure, first, a 3×3 convolution is continued on the feature map (generally 60×40 pixels in size and 256 dimensions in depth) output by the feature extraction network (ResNet, VGG, etc.), with the size and depth unchanged, to combine the surrounding information and obtain a convolution feature map. Then, each pixel on the feature map is mapped to the original image to generate 9 types of anchor boxes (Anchor). At the same time, two convolution operations (convolution kernel is 1×1) are performed synchronously on the convolution feature map, one of which is classification (divided into positive samples: target; negative samples: background), and each pixel outputs an 18-dimensional vector (9 anchor boxes × 2 classification scores); the other is regression, and the output is a 36-dimensional vector (9 anchor boxes × 4 coordinate values), so the final number of anchor boxes is: 60×40×9≈21600. Then, the following steps are used to filter out suitable positive and negative samples:
[0050]
[0051] Where IOU (Intersection over Union) is the interaction ratio, which is calculated as the ratio of the intersection and union of the areas of the predicted anchor box A and the corresponding actual anchor box B. Anchor boxes with IOU>0.6 are considered positive samples, and those with IOU<0.4 are considered negative samples, and the rest are discarded.
[0052] Then, non-maximum suppression (NMS) is used to remove the anchor box with the largest overlap rate with the true value, leaving only the candidate box with the largest prediction probability value as the final prediction result, and find the best detection target position.
[0053] Finally, Rank sorting is performed, and the remaining anchor boxes are sorted from high to low according to confidence. At most, the first 128 anchor boxes are taken as positive samples for final training, and 128 negative samples are provided. Finally, the screened candidate boxes are substituted into the training network for training.
[0054] After that, training was carried out. The water system samples used for training came from the actual water system pictures collected and the downloaded digital elevation data. The original picture size was 1024X1000 pixels. After slicing, the small sample picture size was 32X32 pixels. Then the data was enhanced (rotated, etc.) to obtain the training sample set, which contained 10118 water system samples and 9742 non-water system samples. Then, the prediction training network AlexNet was used to train the network for transfer learning. The training process is shown in the figure below: Figure 5 shown.
[0055] AlexNet neural network model Figure 6 As shown in the figure, the network consists of five convolutional layers and three fully connected layers, with a total depth of eight layers. The output of the last fully connected layer is sent to the Softmax layer, as shown below:
[0056]
[0057] Softmax can limit the output to the range of (0,1), thereby ensuring that the neurons are activated and producing a distribution covering multiple categories of labels. Compared with traditional neural networks, the Alex Net model has the following advantages:
[0058] 1. Use the ReLU activation function (ReLU(x) = max(x, 0)). If the input is not less than 0, the gradient of ReLU is always 1.
[0059] 2. Augment the dataset to suppress fitting;
[0060] 3. Use the Drop Out method to suppress overfitting;
[0061] 4. Use local response normalization layer to enhance generalization.
[0062] The input of the Alex Net model is 32*32*3. The pathological image used in this paper is a three-channel RGB image, and the dimension meets the requirements of the model. After the pathological image is resized, it is sent to the network and first reaches the convolution layer. After the first layer of convolution operation, the result is pooled and finally input to the second layer after standardization. The subsequent operations of the second to fifth layers are similar to those of the first layer and will not be repeated. The output result of the fifth layer of the network will be sent to the subsequent fully connected layer. The outputs of the sixth and seventh layers are both vectors of length 1000. The network finally uses the Softmax classifier to obtain the final classification result.
[0063] After the classification is completed, the morphological algorithm is used to extract the water system morphology. The Faster R-CNN technology is used to process the basic skeleton of the Chongli District water system in the form of anchor frames. However, since the remote sensing data is affected by the mountain terrain, there is a lot of noise and large errors, which makes some water system features have low resolution on the grayscale boundary of the background pixel. Therefore, the morphological algorithm is introduced to deeply process the rough skeleton of the water system after CNN processing, and restore the water system morphology through the minimum grayscale value in the bounding box area. The main steps are as follows:
[0064] Step 1. Extract the pixel with the minimum gray value in each column of the water system skeleton (CNN processing) within the bounding box area (Faster R-CNN processing) to form the initial basic skeleton of the water system;
[0065] Step 2. Due to the influence of the surrounding geographical environment, the first step extraction result is prone to produce salt and pepper noise. This paper uses a 3*3 sliding window to traverse the image. If there is only one pixel inside the sliding window, the pixel in this area is filtered out;
[0066] Step 3. Use the area filter to filter out the unconnected areas with smaller areas, and perform the connection operation on the adjacent connected domains in the remaining connected areas;
[0067] Step 4. Skeletonize the water system morphology to facilitate the subsequent extraction of the width of the water system. The skeleton obtained here is a complete and continuous type, which is different from the discontinuous skeleton obtained in the previous steps;
[0068] Step 5. Combine steps 3 and 4 to extract the width of the water system. This paper uses the method of traversing the water system skeleton and takes the normal length of the tangent of each point of the skeleton as the width of the water system at that point, such as Figure 7 As shown in the figure, 30 points of the water system width were selected along the water system direction for actual testing. The results showed that the calculated width and the actual width were within 1m.
[0069] Finally, we introduce the F1 score indicator in the previous formula to reduce the false detection rate of the algorithm and further improve the recognition accuracy of the algorithm. We determine the corresponding water system pixel area and confidence threshold according to the maximum value of the F1 score to reduce the false detection rate, which is used to adapt to the actual scenarios of water system diversity. The above combines the precision and recall rate into a score value, assuming that the precision and recall rate are equally important. The value range of F1 is [0,1], and the larger the value, the higher the corresponding recognition accuracy.
Claims
1. A water system information extraction algorithm based on FasterR-CNN, characterized in that: The following steps are involved: Step 1: Preprocess the remote sensing image data for training and establish a sample database; Step 2: Use the data in the sample database to perform deep learning on the FasterR-CNN network and complete the algorithm training; Step 3: Use the trained FasterR-CNN network to detect the missed detection rate and calculate the F1 score indicator; Step 4: Use FasterR-CNN to calculate the actual collected data, obtain the training sample set, use AlexNet for transfer learning, and use AlexNet to locate the crack skeleton; Step 5: Based on the water system classification results obtained in step 4, use the morphological algorithm to calculate the shape and width of the water system; Step 6: Reduce the false positive rate of the final result through the F1 score indicator and output the final result; The steps of the morphological algorithm are as follows: Step 1: Extract the water system skeleton processed by Faster R-CNN and the pixel with the minimum gray value in each column processed by CNN in the bounding box area to form the initial basic skeleton of the water system; Step 2: Use a 3*3 sliding window to traverse the image in step 1. If there is only one pixel inside the sliding window, filter out the pixel in this area. Step 3: Filter out the connected areas with smaller areas through the area filter, and perform the connection operation of adjacent connected domains in the remaining connected areas; Step 4: Skeletonize the water system morphology to obtain a complete continuous skeleton; Step 5: Repeat steps 3 and 4 to obtain the width of the water system.
2. The water system information extraction algorithm based on FasterR-CNN according to claim 1 is characterized in that: The pretreatment in step 1 comprises the following steps: Step 1: Extract image features through ResNet residual network and generate convolution feature map; Step 2: Use the convolutional feature map to classify the water system through the region proposal network RPN.
3. The water system information extraction algorithm based on FasterR-CNN according to claim 1 is characterized in that: The AlexNet neural network is composed of five convolutional layers and three fully connected layers, with a total depth of eight layers. The output of the last fully connected layer is sent to the Softmax layer.
4. The water system information extraction algorithm based on FasterR-CNN according to claim 1 is characterized in that: The F1 score index is obtained by the following formula: Where: TP: true positive, FP: false positive, FN: false negative; 。
Citation Information
Patent Citations
Automatic identification method for welding seam ultrasonic TOFD-D scanning defect type based on deep learning
CN107451997A
Vehicle-mounted video target detection method based on deep learning
CN109977812A