Water conservancy facility directional detection and labeling method based on remote sensing image

By applying deep convolutional neural networks and feature pyramid networks in remote sensing images, the directional frame detection and labeling of water conservancy facilities targets is achieved, and the problem of insufficient efficiency and accuracy of water conservancy facilities target detection in the existing technology is solved, and efficient and accurate identification and labeling effect is achieved.

CN120182811APending Publication Date: 2025-06-20CHINA YANGTZE POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510185088.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has failed to effectively apply the directional frame detection method in the target detection of water conservancy facilities, resulting in insufficient efficiency and accuracy of identification and labeling of water conservancy facilities in remote sensing images.

Method used

Using a deep convolutional neural network-based method, through data preprocessing, feature extraction, feature pyramid network training and neural network layer design, the directional box detection and labeling of water conservancy facilities targets in remote sensing images is realized. Specific steps include adjusting the image size, using improved residual network structure for feature extraction, designing feature pyramid networks and neural network layers for directional frame detection and labeling.

Benefits of technology

It realizes efficient and accurate identification and labeling of water conservancy facilities targets in remote sensing images, and has a more accurate frame selection range and higher recognition accuracy than traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182811A_ABST
    Figure CN120182811A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of remote sensing image recognition, and particularly provides a water conservancy facility directional detection and labeling method based on a remote sensing image, and the method comprises the steps: adjusting the size of a visible light remote sensing image of a water conservancy facility; a pre-trained deep convolutional neural network is used as a backbone network of the model, water conservancy facility visual features in the input remote sensing image are extracted, and the visual features are transmitted to a feature pyramid network from the backbone network; a neural network is designed at the tail end of the deep convolutional neural network, the foreground and the background of each water conservancy facility category are distinguished, and regression is carried out on the center, the length, the width and the rotation angle of the marked rectangle; and combining all candidate orientation frames predicted by the network through non-maximum suppression based on the orientation frames, and outputting a final orientation frame detection result of the algorithm. According to the method, the fine boundary extraction of the remote sensing target with any angle distribution in the visible light wide remote sensing image is facilitated, and the detection efficiency and the marking precision of water conservancy facilities such as dams and observation stations can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing image recognition, and particularly relates to a method for directional detection and annotation of water conservancy facilities based on remote sensing images. Background Art

[0002] As a basic task in the field of computer vision, object detection refers to the computer technology of obtaining the category and location information of objects of interest from images. In the early stage of the development of computer vision, researchers mostly used manually designed visual features and certain image window selection strategies to determine the most likely positions of objects. In the past decade, deep learning and related upstream and downstream strategies have gradually dominated the object detection task and become the de facto algorithm foundation. The detection of specific objects in remote sensing images is widely used in various fields such as emergency disaster reduction, national land surveying and mapping, and urban planning.

[0003] In recent years, the attention of researchers in this field has gradually shifted from the detection of horizontal bounding boxes (HBB) of objects to the detection of oriented bounding boxes (OBB) of objects. In the object detection task, the location information of an object is generally annotated using a horizontal box or an oriented box. The location information annotated by a horizontal box contains four dimensions, namely the abscissa of the center point of the horizontal box, the ordinate of the center point, and the length and width information of the horizontal box. The purpose of oriented box detection is also to find the location of the object of interest and annotate it. However, the difference is that these two pairs of line segments are no longer parallel to the horizontal and vertical axes of the image. The location information annotated by an oriented box contains five dimensions. In addition to the above center point, length, and width information, an additional angle information is required to describe the angle between the bounding box and the image coordinate axes. Even in a dense scene, the oriented box rarely overlaps with surrounding objects, which is especially suitable for object detection tasks with high aspect ratios in remote sensing images, such as water conservancy facilities like river embankments and dams. However, in the existing technology, the oriented box detection method has not been applied to the object detection of water conservancy facilities. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for directional detection and annotation of water conservancy facilities based on remote sensing images, so as to achieve efficient and accurate recognition and annotation of water conservancy facility objects in remote sensing images.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is: a method for directional detection and annotation of water conservancy facilities based on remote sensing images, comprising the following steps: Step 1: Data preprocessing: Adjust the size of the visible light remote sensing image of the water conservancy facility so that the image size does not exceed a predetermined value; Step 2: Deep convolutional neural network training. Use a pre-trained deep convolutional neural network as the backbone network of the model to extract the visual features of water conservancy facilities in the remotely sensed images processed in Step 1. Step 3: Feature pyramid network training: Transfer the visual features at 1 / 8, 1 / 16, and 1 / 32 resolutions from the backbone network to the feature pyramid network. Step 4: Design a neural network layer at the end of the deep convolutional neural network to distinguish the foreground and background of each water conservancy facility category, and perform regression on the center, length, width, and rotation angle of the annotation rectangle for oriented bounding box detection. Step 5: Put all the predicted rectangular boxes of the model together, and use a preset threshold to obtain all the predicted boxes above the confidence level. Then, apply non-maximum suppression to the qualified predicted boxes to obtain the final image annotation boxes.

[0006] In a preferred solution, in Step 1, when adjusting the size of the visible light remotely sensed image of water conservancy facilities, if the image size is larger than the predetermined value, the image will be scaled while maintaining the aspect ratio of the image. Then, zero pixels will be filled in the right and bottom edges of the scaled image so that both the length and width of the image can be divisible by 32.

[0007] In a preferred solution, in Step 2, the backbone network of the model adopts an improved residual network structure, and the improved residual network structure introduces spatial projection on the basis of the original residual network structure.

[0008] In a preferred solution, in Step 4, when distinguishing the foreground and background of each water conservancy facility category, according to the preset water conservancy facility feature labels, the weight information of the feature labels is incorporated into the end neural network layer.

[0009] In a preferred solution, a neural network layer is designed at the end of the deep convolutional neural network, with six convolutional layers. Among them, three neural network layers are used for the detection of the target contour, which are the target contour detection layers, and the other three neural network layers are used for the detection of the rectangular position, which are the rectangular detection layers.

[0010] In a preferred solution, the loss function is calculated on all six convolutional layers and added together to obtain the overall loss function value of the model. On all six neural network layers, the cross-entropy loss is independently used for each predicted category.

[0011] In a preferred solution, the calculation of the loss function on the six convolutional layers includes: For the target contour detection layer, use the cross-entropy loss function to perform gradient backpropagation, and the expression is as follows: ; Among them, represents all grids set as positive samples, represents all grids set as negative samples, K represents the number of categories of the target to be detected, λ is a preset weight; ∈{0, 1} represents i the confidence at On each rectangle detection layer, the following regression loss function is used to regress the remaining four parameters: ; Among them, G + represents the grid set as a positive sample, , respectively represent the predicted values of the offsets of the center position of the target box corresponding to grid i on the x-axis and y-axis, , respectively represent the predicted values of the short side and long side of the target box corresponding to grid i; t , respectively represent the current time and the predicted time.

[0012] In a preferred solution, the overall loss function value of the model is calculated using a joint loss function, and the expression is as follows: ; Among them, C represents the set of all contour detection layers, R represents the set of all rectangle detection layers, h represents the current detection layer, is a self-balanced angular loss function, represents the corresponding Laplace penalty function.

[0013] In a preferred solution, the definitions of and in the joint loss function are as follows: The aspect ratio information is used to set a dynamic weight for the angle of the prediction box, ; The expression of ; Among them, represents the lengths of the short side and long side of the target box corresponding to grid i, γ is a preset constant; represents the angle between the long side of the rectangle box and the horizontal axis, , Indicates the predicted result; Penalty term The expression is as follows: .

[0014] In a preferred embodiment, in step 4, when performing regression on the center, length, width, and rotation angle of the labeled rectangle, the natural embedding of the real projective space in the two-dimensional Euclidean space is used to represent the rotation angle, that is, directly predicting the coordinates represented by two real numbers (p, q) , and using the angle represented by the coordinates to represent the rotation angle of the oriented bounding box.

[0015] A method for directional detection and annotation of water conservancy facilities based on remote sensing images provided by the present invention has the following beneficial effects: 1. The present invention can realize the directional detection and efficient annotation of water conservancy facilities with high aspect ratios in large-scale visible light remote sensing images. Through deep learning methods, the detection and directional marking methods of target markers in remote sensing images are realized, and the box selection range is more accurate than traditional marking boxes.

[0016] 2. The present invention improves the weight ratio of target markers by introducing the category parameters of water conservancy facilities, and at the same time proposes an optimization algorithm for high aspect ratio targets, making the target recognition and annotation of water conservancy facilities more efficient and accurate.

[0017] 3. The present invention proposes an adaptive balanced angle loss, enabling the model to make a better trade-off between low aspect ratio objects and high aspect ratio objects.

[0018] 4. The advantage of step 1 of the present invention is that through specific filling methods and size adjustment strategies, combined with the requirements of water conservancy facility detection, the adaptability of the input data is optimized, and the model performance is improved.

[0019] 5. The advantage of step 2 of the present invention lies in the improvement of the existing residual network. For the water conservancy facility detection task, the feature extraction process is optimized, and the feature expression ability is enhanced by combining spatial projection.

[0020] 6. The advantage of step 3 of the present invention is that through the specific design of the feature pyramid, combined with the multi-scale characteristics of water conservancy facilities, the feature transfer and fusion methods are optimized.

[0021] 7. The advantage of step 4 of the present invention is that through the regression strategy and adaptive loss design, aiming at the particularity of water conservancy facilities, the deficiencies of traditional methods in angle processing are solved, and the detection accuracy and robustness are improved. Brief Description of the Drawings

[0022] The present invention will be further described below with reference to the drawings and embodiments: Figure 1 This is the flowchart of the method of the present invention; Figure 2 This is the value curve of the weights of different water conservancy facilities; Figure 3 This is the recognition and annotation effect of the dam body and pump house in the remote sensing image; Figure 4 This is the recognition and annotation effect of the riverbank dike and bridge in the remote sensing image. Specific implementation manners

[0023] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0024] A method for directional detection and annotation of water conservancy facilities based on remote sensing images includes the following steps: In Step 1 and Step 5, no trainable parameters are used, and all calculations of the deep convolutional neural network are performed in Steps 2 to 4.

[0025] Step 1: Adjust the size of the visible light remote sensing image of the water conservancy facilities so that the image size is not greater than a predetermined value, making it more suitable for parallel computing in the computer graphics card.

[0026] As Figure 1 shown, this step is the preprocessing link of the method, and preprocesses the input visible light remote sensing image.

[0027] Specifically, when adjusting the size of the visible light remote sensing image of the water conservancy facilities, if the image size is greater than the predetermined value, the image will be scaled while maintaining the aspect ratio of the image. Then, some zero pixels will be filled in the right and bottom edges of the scaled image so that the length and width of the image can be divisible by 32.

[0028] Step 2: Training of the deep convolutional neural network. Use a pre-trained deep convolutional neural network as the backbone network of the model to extract the visual features of the water conservancy facilities in the remote sensing image processed in Step 1.

[0029] The backbone network used in this step largely determines the speed, performance, convergence situation, etc. of the model.

[0030] In this embodiment, the backbone network of the model adopts an improved residual network structure (Improved Residual Networks, iResNet). The improved residual network structure introduces spatial projection on the basis of the original residual network structure to improve the network learning accuracy without increasing the model complexity and the number of model parameters.

[0031] Step 3: Feature Pyramid Network training. Transfer the visual features at 1 / 8, 1 / 16, and 1 / 32 resolutions from the backbone network to the Feature Pyramid Network.

[0032] The feature vectors at each position in the above resolutions respectively represent an 8x8, 16x16, and 32x32 pixel grid. In this step of the present invention, the dimensions of all feature vectors are unified and sent to the next step for regression.

[0033] Step 4: Foreground and background classification and oriented bounding box detection: Design neural network layers at the end of the deep convolutional neural network to distinguish between the foreground and background of each water conservancy facility category, and regress the center, length, width, and rotation angle of the labeled rectangle.

[0034] When distinguishing between the foreground and background of each water conservancy facility category, according to the preset water conservancy facility feature labels, incorporate the weight information of the feature labels into the end neural network layer to enhance the recognition accuracy of the network for water conservancy facilities.

[0035] The oriented bounding box detection method has three neural network layers at the end of the model for detecting the target contour, and another three neural network layers at the end of the model for detecting the rectangle position, that is, six convolutional layers are set, where three neural network layers are for detecting the target contour, which are the target contour detection layers, and the other three neural network layers are for detecting the rectangle position, which are the rectangle detection layers.

[0036] Calculate the loss function on each of the six convolutional layers and add them up to obtain the overall loss function value of the model.

[0037] On all six neural network layers, use the cross-entropy loss independently for each predicted category.

[0038] For the target contour detection layer, use the cross-entropy loss for gradient backpropagation, and the expression is as follows: ; Where, represents all grids set as positive samples, represents all grids set as negative samples, K represents the number of categories of the targets to be detected, λ is a preset weight, which is related to the preset types of target water conservancy facilities in the dataset. The larger the value, the higher the weight for detecting and recognizing such water conservancy facilities.

[0039] On each rectangle detection layer, use the following regression loss function to regress the remaining four parameters: ; Among them, G + represents the grid set as the positive sample, , respectively represent the predicted values of the offsets of the center position of the target box corresponding to grid i on the x-axis and y-axis, , respectively represent the predicted values of the short side and long side of the target box corresponding to grid i; t , respectively represent the current time and the predicted time.

[0040] The overall loss function value of the model is calculated using a joint loss function, and the expression is as follows: ; Among them, C represents the set of all contour detection layers, R represents the set of all rectangle detection layers, h represents the current detection layer, is a self-balanced angular loss function, represents and its corresponding Laplace penalty function.

[0041] The present invention uses a self-balanced angular loss function to better handle the importance of angular information in targets with different aspect ratios, that is, the problem of angular degradation. At the same time, a Laplace penalty term is added to ensure that the angular coordinates converge near the unit circle. For targets with a high aspect ratio, a small error in the angular information may cause a large deviation in the position of the oriented box; conversely, when the aspect ratio is low, for example, close to 1, the deviation of the angular information does not significantly affect the accuracy of the prediction box. For some targets close to a circle, it is even difficult to uniquely determine the angle of its oriented box. The width projection curve of targets with a high aspect ratio is closer to a trigonometric function with a period of π, while the curve corresponding to targets with a low aspect ratio has a smaller amplitude and more peaks and valleys.

[0042] Since water conservancy facilities such as dams and embankments are all targets with a high aspect ratio, a small error in the angular information may cause a large deviation in the position of the oriented box. Therefore, the present invention sets a dynamic weight for the angle of the prediction box using the aspect ratio information , and the expression is as follows: ; Among them, represents the lengths of the short side and long side of the target box corresponding to grid i, γ is a preset constant.

[0043] Such as Figure 2As shown, when the target has different aspect ratios, the equation will assign different weights to the loss function of the angle, thereby successively separating the importance of the angle information. When the aspect ratio of the target is larger, the loss function of the angle will be assigned a larger weight. The final prediction result of the model will also have a smaller tolerance for the angle deviation of such targets. Conversely, for targets close to a square or a circle, the loss function of the angle will be assigned a smaller weight. Then, for the angle deviation of such targets, the model will ultimately have a greater tolerance.

[0044] When regressing the center, length, width, and rotation angle of the annotated rectangle, the natural embedding of the real projective space in the two-dimensional Euclidean space is used to represent the rotation angle, that is, directly predict the coordinates represented by two real numbers (p, q) , and use the angle represented by the coordinates to represent the rotation angle of the oriented bounding box. More precisely, for a given angle θ , this study predicts cos 2θ and sin 2θ values in the model.

[0045] The angle loss function used consists of this adaptive weight and the L1 loss function. In the joint loss function is expressed as: The expression of is as follows: where represents the angle between the long side of the rectangle and the horizontal axis, , represents the prediction result.

[0046] To constrain the two predicted parameters (p,q) near the natural embedding described above to avoid their divergence to the two-dimensional Euclidean space, the present invention adopts the following penalty term , and the expression is as follows: .

[0047] Step 5: Put all the rectangle bounding boxes predicted by the model together, and use a preset threshold to obtain all the predicted bounding boxes above the confidence level. Then, for the predicted bounding boxes that meet the requirements, through a non-maximum suppression based on skew-IoU, the final image annotation bounding boxes are obtained.

[0048] The present invention proposes a simple and efficient method for detecting oriented bounding boxes. This method uses topology to solve the periodicity of angles and proposes a self-balanced angular loss function to trade off between objects with different aspect ratios. In the detection of low aspect ratio objects, the angle is not important for the bounding box, and in some cases, it is even difficult to define its rotation angle. On the contrary, in the detection of high aspect ratio objects in water conservancy target detection, the angle information plays a crucial role and has a decisive impact on the quality of the detection results. By utilizing the aspect ratios of different objects, the present invention proposes an adaptive balanced angular loss, enabling the model to make a better trade-off between low aspect ratio objects and high aspect ratio objects. For the rotation angle of each oriented bounding box, the present invention naturally embeds it into a two-dimensional Euclidean space for regression, thus avoiding overly redundant designs. Figure 3 The recognition and annotation effects of the dam body and pump house in the remote sensing image are shown. Figure 4 The recognition and annotation effects of the riverbank dike and bridge in the remote sensing image are shown. As can be seen from the figure, the innovative method described in the present invention has clear target recognition of water conservancy facilities and accurate and effective rotation angles.

[0049] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. The embodiments in this application and the features in the embodiments can be arbitrarily combined with each other without conflict. The protection scope of the present invention should be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A method for directional detection and labeling of water conservancy facilities based on remote sensing images, characterized in that: The following steps are involved: Step 1: Data preprocessing: adjust the size of the visible light remote sensing image of the water conservancy facilities so that the image size is not larger than the predetermined value; Step 2: Deep convolutional neural network training, using the pre-trained deep convolutional neural network as the backbone network of the model to extract the visual features of water conservancy facilities in the remote sensing image processed in step 1; Step 3: Feature Pyramid Network Training: Transfer visual features at 1 / 8, 1 / 16, and 1 / 32 resolutions from the backbone network to the feature pyramid network; Step 4: Design a neural network layer at the end of the deep convolutional neural network to distinguish the foreground and background of each water conservancy facility category, and regress the center, length, width and rotation angle of the labeled rectangle to perform oriented box detection; Step 5: Put all the rectangular boxes predicted by the model together, and use the pre-set threshold to get all the predicted boxes above the confidence level. Then, the predicted boxes that meet the requirements are suppressed by a non-maximum value to get the final image annotation box.

2. According to the method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 1, it is characterized in that: In step 1, when the size of the visible light remote sensing image of the water conservancy facilities is adjusted, if the image size is greater than a predetermined value, the image will be scaled while maintaining the aspect ratio of the image, and then the right and bottom edges of the scaled image will be filled with zero pixels so that the length and width of the image can be divided by 32.

3. According to the method for directional detection and labeling of water conservancy facilities based on remote sensing images in claim 1, it is characterized in that: In the step 2, the backbone network of the model adopts an improved residual network structure, and the improved residual network structure introduces spatial projection on the basis of the original residual network structure.

4. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 1 is characterized in that: In step 4, when distinguishing the foreground and background of each water conservancy facility category, the weight information of the feature label is integrated into the terminal neural network layer according to the preset water conservancy facility feature label.

5. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 4 is characterized in that: A neural network layer is designed at the end of the deep convolutional neural network, and six convolutional layers are set, of which three neural network layers are used for detecting target contours, namely target contour detection layers, and the other three neural network layers are used for detecting rectangular positions, namely rectangle detection layers.

6. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 5 is characterized in that: The loss function is calculated on each of the six convolutional layers and added together to obtain the overall loss function value of the model. On all six neural network layers, cross entropy loss is used independently for each predicted category.

7. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 6 is characterized in that: The calculation of the loss function on the six convolutional layers includes: For the target contour detection layer, the cross entropy loss function is used To perform gradient back propagation, the expression is as follows: ; in, represents all grids that are set as positive samples, represents all grids that are set as negative samples, K Indicates the number of categories of targets that need to be detected. λ is a pre-set weight; ∈{0,1} means i Confidence level at At each rectangle detection layer, the remaining four parameters are regressed using the following regression loss function: ; Among them, G + represents the grid that is set as a positive sample, , They represent the predicted values ​​of the offset of the center position of the target box corresponding to grid i on the x-axis and y-axis respectively. , Respectively represent the predicted values ​​of the short and long sides of the target box corresponding to grid i; t , Represent the current time and predicted time respectively.

8. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 7 is characterized in that: The overall loss function value of the model is calculated using the joint loss function, and the expression is as follows: ; Among them, C represents the set of all contour detection layers, R represents the set of all rectangle detection layers, h represents the current detection layer, is the self-balancing angle loss function, Represents and its corresponding Laplace penalty function.

9. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 8, characterized in that: The joint loss function and is defined as follows: Use the aspect ratio information to set a dynamic weight for the angle of the predicted box , the expression is as follows: ; The expression is as follows: ; in, Represents the length of the short and long sides of the target box corresponding to grid i, γ is a preset constant; Represents the angle between the long side of the rectangle and the horizontal axis, , Indicates the predicted result; Penalty The expression is as follows: 。 10. The method for directional detection and labeling of water conservancy facilities based on remote sensing images according to claim 1, characterized in that: In step 4, when regressing the center, length, width and rotation angle of the labeled rectangle, the natural embedding of the real projective space in the two-dimensional Euclidean space is used to represent the rotation angle, that is, directly predicting the coordinates represented by the two real numbers (p, q) , the angle represented by the coordinates is used to represent the angle of rotation of the orientation frame.