Intelligent power grid infrastructure fine identification method based on remote sensing image

By combining deep learning and domain knowledge of power infrastructure, optimizing target recognition strategies and network parameters, the problem of accuracy and low efficiency of power infrastructure recognition in remote sensing images is solved, and high accuracy and high efficiency recognition effects are achieved.

CN120182807AInactive Publication Date: 2025-06-20冯昌存

Patent Information

Application Number
CN202510124759.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has limitations in the automatic identification of power infrastructure in remote sensing images, and the fusion characteristics of multi-source remote sensing sensors are not fully utilized. The target recognition is time-consuming and the feature design and understanding are insufficient, resulting in low recognition accuracy and efficiency.

Method used

Combining the domain knowledge of deep learning and power infrastructure, the fusion target recognition strategies FasterR-CNN and ResNet are optimized and integrated, the network parameters are modified to adapt to power infrastructure detection tasks, and the rough screening of target recognition is completed through deep learning, and the domain knowledge characteristics are combined for fine screening, including water surface extraction, target reclassification and spatial relationship calculation.

Benefits of technology

It significantly improves the accuracy and recall rate of power infrastructure target recognition, reduces identification time, improves identification efficiency, and can more accurately identify power infrastructure in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182807A_ABST
    Figure CN120182807A_ABST
Patent Text Reader

Abstract

According to the intelligent power grid infrastructure fine recognition method based on the remote sensing image, deep learning and field knowledge of the power infrastructure are combined, target recognition of the power infrastructure in the multi-scale remote sensing image is achieved, target recognition strategies Faster R-CNN and ResNet are optimized and fused according to the characteristics of the power infrastructure in the remote sensing image, and the target recognition accuracy of the power infrastructure in the remote sensing image is improved. Network parameters are modified to adapt to a detection task of the power infrastructure, coarse screening of target recognition is completed based on deep learning, and the recall rate of a detection result is improved by reducing the confidence coefficient of target judgment so as to obtain more possible targets; and on the basis of the rough detection result of the electric power infrastructure, analyzing domain knowledge features of various electric power infrastructures for fine screening of electric power infrastructure detection, including water surface extraction, target reclassification, spatial relationship calculation and the like. Experiments show that the target identification accuracy of the electric power infrastructure is greatly improved, and meanwhile, the better recall rate can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to a method for identifying power grid facilities in remote sensing images, and particularly to an intelligent power grid infrastructure precise identification method based on remote sensing images, belonging to the technical field of remote sensing image recognition. Background Art

[0002] Due to the development of sensor technology and aerospace technology, the quantity and quality of remote sensing images have been greatly improved. Therefore, it is possible to obtain a large range of interested targets by means of remote sensing images, and solve problems existing in the current intelligent power grid, such as a serious lack of location information of power infrastructure and insufficient methods for obtaining power infrastructure. Currently, in power planning and design, a large amount of ground measurement and on-site investigation are still required, and the cost of information acquisition is high, the efficiency is low, and the difficulty is great, which obviously cannot meet the demand for obtaining geographical information of the intelligent power grid. Compared with traditional measurement methods, remote sensing technology can collect multi-topic geographical information on the ground over a large area and quickly, effectively overcoming defects such as high cost and long cycle of traditional ground measurement, and providing reliable GIS support for the construction of the intelligent power grid.

[0003] With the development of remote sensing technologies with high spectral, high spatial, and high temporal resolutions, the accuracy and precision of obtaining remote sensing geographical information have been greatly improved, effectively reducing spatial data errors, making it possible to obtain high-precision spatial information of power infrastructure. By means of various remote sensing technology means such as visible light, infrared, hyperspectral, and SAR, important power energy infrastructure information and other related topic information are obtained. Through the automatic identification of power infrastructure in remote sensing images, the geographical location information of power infrastructure can be automatically obtained, which is also the key to improving the production efficiency and quality of links such as planning, design, and construction. Realizing the automatic identification of power infrastructure in remote sensing images can provide high-quality geographical location information for the construction of the power energy infrastructure Internet, which is of great significance for the construction of the energy Internet.

[0004] The target recognition in remote sensing images is to realize the intelligent extraction of target information by combining remote sensing image processing, geoscience analysis, GIS (Geographic Information System), pattern recognition, and artificial intelligence technologies with the support of computer technology. There have been some achievements in the prior art on the automatic identification of power energy infrastructure based on remote sensing images, but there are limitations: one is that the fusion characteristics of multi-source remote sensing sensors are not fully utilized. Since the characteristics of different sensors have strong complementary properties, comprehensively using multi-sensors, such as infrared, hyperspectral, visible light, SAR and other sensors for target recognition can greatly improve the accuracy of target recognition; the second is that target recognition is time-consuming. Accurate target recognition requires a large amount of calculation, and advanced computing systems such as high-performance technology and cloud computing should be fully utilized; the third is that there are deficiencies in the design and understanding of features. Traditional features such as spectrum, shape, texture, and structure are not sufficient for accurate target category determination, and there is still much room for improvement.

[0005] The problems to be solved in the identification of power grid infrastructure in remote sensing images of the prior art and the key technical difficulties of this application include: Compared with the object recognition in natural images, the power infrastructure in remote sensing images has more complex scene content. First, due to the characteristics of remote sensing images themselves, they have a large amount of information and complex content. Due to factors such as sensors and imaging conditions, there are differences in the attitude, size, and resolution of remote sensing images. Although remote sensing images have very intuitive visual expression capabilities, it is easy for the human eye to distinguish various power infrastructures such as cooling towers, large power towers, and hydropower stations in optical remote sensing images. However, for computers, it is not an easy task to automatically identify these facilities in a large number of remote sensing images. Second, due to the complex characteristics of power infrastructure itself. The surrounding environmental conditions of power infrastructure are complex, and there are many types of itself. For example, the types of transmission towers can be divided into many types according to scale or material, which greatly increases the difficulty of object recognition. In addition, for some power facilities, when observing their shapes from the air, they are not much different from other civilian buildings, and it is very difficult for computers to automatically identify and discover such subtle differences. Third, there is the problem of small target detection, which is more serious in the detection task of medium and small transmission towers. Fourth, due to the problem that the sizes of the same ground object targets are inconsistent in images of different scales, it increases the detection difficulty of such ground objects. For example, a thermal power station includes a cooling tower, a transmission tower, and a substation. When we need to detect a small transmission tower, we hope to obtain higher-resolution data. However, in ultra-high-resolution remote sensing images, the sizes of the cooling tower and the substation will become very large, thus affecting the detection of the cooling tower and the substation.

[0006] Combining the above, the detection of power grid infrastructure in remote sensing images can be summarized into the following two key problems:

[0007] (1) Feature extraction of power grid infrastructure in remote sensing images

[0008] Complex power infrastructure targets such as large power stations have the characteristics of spatial distribution structure. The recognition and extraction of such targets belong to the recognition of complex scenes, and it is impossible to achieve using traditional object- and pixel-based methods. It is necessary to adopt a method of extracting and fusing multi-level features of comprehensive pixel-object-scene. By comprehensively expressing multi-level collaborative information such as pixel-object-scene, the features of complex power facility targets can be quickly extracted, and then the recognition and extraction of complex targets can be realized.

[0009] The objects in the image are composed of multiple pixels, and the objects are in a specific scene. Therefore, the objects have the characteristics of multi-layer features. The high-level features of the image are a combination of low-level features. The higher the level, the more abstract the features. Since most power infrastructure objects have a natural hierarchical structure, effective feature extraction methods need to be used to gradually abstract the low-level features of basic elements into more advanced features. That is, from the pixels that make up the power facility object, the pixels form straight lines, the straight lines form textures, the textures form objects, and the objects form scenes, extracting layer by layer from low to high, and making full use of the characteristics of the power infrastructure object itself to realize the recognition and extraction of the object.

[0010] In addition to the characteristics of the object itself, there are also relationship characteristics between objects. In the real world, objects are in a specific scene, and there are certain spatial relationships between objects. Various types of context information such as inside the object, in the neighborhood around the object, between objects and between the object and the scene where it is located can greatly enrich the feature expression of the object and effectively improve the accuracy in the case of small object recognition and occlusion. Using the relevant ground objects and environmental context information around the object as auxiliary information to construct a context model for object detection to improve the detection accuracy in the case of small objects and occlusion. Context information can eliminate the uncertainty and ambiguity in object recognition and make up for the recognition problems caused by factors such as resolution or occlusion, especially improving the detection accuracy in the case of small objects and occlusion.

[0011] A method of automatic learning based on deep learning algorithms is to learn from a large amount of data, mine the most effective features in the large amount of data, establish a robust object feature expression, fully obtain the associations between data, establish a powerful classifier, and realize the rough screening of power infrastructure object recognition.

[0012] (2) Feature extraction of power infrastructure based on domain knowledge

[0013] For different power infrastructures, extract their unique ground object features for fine screening of object recognition. Specifically, for transmission towers, object recognition is based on deep features; for hydropower stations, the method of image classification is used to extract the upstream and downstream scene features and consider the relationship between the hydropower station and the transmission tower; for substations, the initially detected substation objects are re-classified by images and combined with their spatial relationship with the surrounding transmission towers. For thermal power stations, pay attention to their unique cooling towers and their spatial logical relationship with the surrounding transmission towers;

[0014] Although deep learning has a certain degree of invariance to geometric transformations, deformations, illuminations, etc. of objects, it can effectively overcome some difficulties in object recognition. More importantly, it can adaptively construct feature descriptions driven by training data and has a certain degree of flexibility and generalization ability. However, the deep learning model that learns through training data-driven treats each type of object as a separate individual for detection, thus easily ignoring important information between objects, such as the spatial relationship and distance relationship between various power infrastructure facilities. It is necessary to combine the domain knowledge of power infrastructure to extract more implicit object features and further refine the object recognition results based on deep learning methods to achieve a higher detection accuracy rate. Summary of the Invention

[0015] In response to the need to build a smart grid to achieve energy resource sharing, this application combines deep learning and the domain knowledge of power infrastructure to achieve object recognition of power infrastructure in multi-scale remote sensing images. According to the characteristics of power infrastructure in remote sensing images, it optimizes and integrates the object recognition strategies Faster R-CNN and ResNet, and modifies the network parameters to adapt to the detection task of power infrastructure. Based on deep learning, it completes the rough screening of object recognition, and improves the recall rate of the detection results by reducing the confidence level of object determination to obtain more possible objects; then, based on the rough detection results of power infrastructure, it analyzes the domain knowledge characteristics of various power infrastructure facilities, and establishes object features based on grid domain knowledge for different types of power infrastructure facilities for the fine screening of power infrastructure detection, including water surface extraction, object reclassification, spatial relationship calculation, etc. Experiments show that the object recognition method that combines deep learning and the domain knowledge of power infrastructure has greatly improved the object recognition accuracy rate of power infrastructure and can also ensure a better recall rate.

[0016] To achieve the above technical effects, the technical solutions adopted in this application are as follows:

[0017] Intelligent grid infrastructure fine recognition method based on remote sensing images, which combines deep learning and domain knowledge of power infrastructure to achieve target recognition of power infrastructure in multi-scale remote sensing images: According to the characteristics of power infrastructure in remote sensing images, optimize and fuse the target recognition strategies FasterR-CNN and ResNet, and modify the network parameters to adapt to the detection task of power infrastructure. Based on the target recognition of deep learning, complete the rough screening of target recognition, and improve the recall rate of detection results by reducing the confidence of target determination to obtain more possible targets; Then, based on the rough detection results of power infrastructure, analyze the domain knowledge characteristics of various types of power infrastructure, and extract target characteristics that can be applied to the fine screening of target recognition, specifically including: characteristic information of substations at different levels, characteristic information of transmission towers at different levels and types, characteristic information of the cooling towers of thermal power plants, distinguishing characteristic information of hydropower stations and reservoirs, and spatial logical relationship characteristics between various types of power infrastructure. For different categories of power infrastructure, establish target characteristics based on grid domain knowledge for the fine screening of power infrastructure detection. The specific methods include water surface extraction, target reclassification, and spatial relationship calculation;

[0018] (1) Unsupervised training and weakly supervised learning of deep learning models: For the need of fine recognition of grid infrastructure, collect and organize training data for target recognition, and annotate the training data under unified data standards to establish a remote sensing image library of power infrastructure for target recognition;

[0019] (2) Rough screening method for power infrastructure recognition: Improve the target recognition strategy by fusing FasterR-CNN and ResNet. According to the characteristics of power infrastructure in remote sensing images, modify the network parameters to adapt to the detection task of power infrastructure. Based on the target recognition of the neural convolutional neural network, complete the rough screening of the final target recognition, and improve the recall rate of detection results by reducing the confidence of target determination to obtain more possible targets;

[0020] (3) Fine screening method for power infrastructure recognition: Analyze the domain knowledge characteristics of various types of power infrastructure, and extract target characteristics that can be applied to the fine screening of target recognition. Specifically for hydropower stations, use the method of deep learning image segmentation to extract the upstream and downstream scene characteristics, judge whether there is a water surface in its upstream and downstream environment, and combine the electric towers around the hydropower station for auxiliary judgment; For substations, perform image classification on the results of rough target screening, and combine their spatial relationship with the surrounding transmission towers; For thermal power plants, pay attention to their unique cooling towers and their spatial logical relationship with the surrounding transmission towers and substations; Combine the domain knowledge of power infrastructure to improve the target recognition effect in this field.

[0021] Preferably, for the rough screening of targets in remote sensing images of power grid facilities: Based on the TensorFlow deep learning framework, a convolutional neural network model for substation detection in remote sensing images is constructed. Based on the 101-layer ResNet of the network model, deep learning adopts a cross-computation structure of convolutional layers and downsampling layers to integrate the low-level features of the image, further obtain the high-level features of the image, extract the self-features and internal features of the region box and the annotation box, and obtain the effective information in the remote sensing image. The power infrastructure that needs to be subject to rough target screening includes: hydropower stations, substations, and thermal power stations;

[0022] Fuse the network model ResnNet and the target recognition framework FasterR-CNN to reduce the confidence of the output of the target candidate box to achieve a higher recall rate and realize the preliminary recognition of the target;

[0023] Suppose F(x) represents the mapping function of a block that only contains two or three layers, x is the input of the block, and F(x) is the output of the block. Assume that they have the same dimension. During the training process, it is hoped that an ideal h(x) can be fitted by modifying the w and b in the network, which is an ideal mapping function from the input to the output. The goal is to modify the w and b in F(x) to approximate h(x), use F(x) to approximate h(x)-x, and finally the output of the bolck changes from F(x) to F(x)+x. After the change, the goal is to approximate the training function F(x) to h(x)-x;

[0024] The residual network specifically opens up channels, making the optimization goal become H(x)-x. H(x)-x represents the difference between the output and the input, rather than the original fitted output H(x). Here, x is the input and H(x) is the original expected mapping output of a certain layer. The gradient from the deep layer can directly pass through unobstructed to the upper layer, enabling the effective training of the parameters of the shallow network.

[0025] Preferably, generate target candidate boxes based on RPN: Use the corresponding relationship to map the points on the feature map to the original image, and generate many windows at the positions on the original image. The scale of the windows is fixed. Then calculate the IOU between the windows and the target ground truth. According to the division rules of positive and negative samples, give it positive and negative labels and conduct learning. Train a network RPN. RPN makes three fixes for the scale windows: the scale change is fixed, the scale ratio change is fixed, and the sampling method is fixed, reducing the complexity of target recognition. The sampling method of RPN is to sample only on the corresponding ROI of each point on the feature map in the original image. The input of RPN is an image of any size, and it outputs some target candidate boxes, and each candidate box has a target existence score and confidence information;

[0026] Specifically, after the image input network, it successively passes through a series of convolutional layers and relu activation function layers, and finally obtains a feature map, and outputs a feature map of 51*39*256 dimensions. The feature map is used for the selection of subsequent alternative targets, and the coordinates at this time can still be mapped to the original image. Multiple regional targets are predicted through the scale of the target anchor and the set ratio. According to the sizes of various types of power infrastructure in the remote sensing images of different resolutions in the training data, the scale of the target anchor is modified. Specifically, the target anchors for detecting transmission towers are set to 2, 4, 8, 16, and the target anchors for hydropower stations, thermal power stations, and substations are 4, 8, 16, 32, and the ratio remains unchanged. More alternative target boxes of different scales will be generated at the reference points.

[0027] Preferably, positive and negative samples are divided: First, for each manually calibrated target ground truth region, retain the target anchor with the largest overlap ratio with it as the positive sample, ensuring that each target ground truth corresponds to at least one positive sample target anchor. For the remaining target anchors, if the overlap with the manually calibrated region is greater than 0.7, it is divided into the positive sample, and if the overlap with the manually calibrated region is less than 0.3, it is divided into the negative sample. Each real target box corresponds to multiple positive sample target anchors, but each positive sample target anchor only represents one real target. In addition, target anchors that do not meet the overlap area or cross the image boundary are discarded;

[0028] In actual training, 256 target anchors are randomly sampled in each image and put into the calculation of the loss of a small batch of images, ensuring that the ratio of positive samples to negative samples is 1:1. When the number of positive samples is less than 128, negative samples are automatically supplemented for sampling.

[0029] Preferably, a rough screening dataset for power infrastructure recognition is established: The target library is a remote sensing target image library, including: remote sensing images of thermal power stations, remote sensing images of hydropower stations, remote sensing images of transmission towers, and remote sensing images of substations;

[0030] Large and medium-sized transmission towers are located in the suburbs or the wild. Compared with other power infrastructure, the background is simple and independent. For images affected by sunlight, transmission towers in remote sensing images at different times have different shadows; in addition, the background types of transmission towers are diverse, including forests, farmlands, and deserts; the difficulty in identifying transmission towers lies in their small size, which is difficult to distinguish in images with low resolution. In Google images of 1 to 2 meters, they occupy 30 to 60 pixels, and in remote sensing images of 0.6 meters, they occupy 60 to 120 pixels;

[0031] In a remote sensing image with a resolution of 5 meters, a hydropower station occupies 150 - 200 pixels of the image. In a remote sensing image with a resolution of 2 meters, it occupies 375 - 500 pixels in the image. Compared with other power infrastructure, a hydropower station has a large floor area, and there must be water surfaces in the upper and lower environments. Coupled with the special line features of the hydropower station, it is relatively easy to identify in remote sensing images. The problem in the identification of hydropower stations lies in the distinction from similar ground objects, including dams and some similar bridges;

[0032] Compared with hydropower stations, thermal power stations have more complex background content and greater identification difficulty. However, the power generation principle and power generation system of thermal power determine that they have a special structure of a cooling tower. From the perspective of remote sensing images, all thermal power stations contain cooling towers. The cooling tower is used as the key feature for the preliminary identification of thermal power stations. In a remote sensing image with a resolution of 1 to 2 meters, the cooling tower occupies 60 - 80 pixels, and in a remote sensing image with a resolution of 0.6 meters, it occupies 120 - 160 pixels;

[0033] In a remote sensing image with a resolution of 5 meters, a complete substation occupies 300 - 500 pixels. In a remote sensing image with a resolution higher than 5 meters, it is difficult to retain the integrity of the substation in an image with a size of 2000 * 2000 pixels and will be cut into image blocks;

[0034] The samples are divided into training data and test data at a ratio of 4:1, and the data is augmented through rotation transformation, resampling, color transformation, and mirror transformation.

[0035] Preferably, the rough screening process for power infrastructure identification: Analyze the characteristics of various power infrastructures and set the parameters of the deep learning model. During the rough screening of power infrastructure, according to the sizes of various power infrastructure targets at different resolutions, set the target anchors for training, with sizes of 2, 4, 8, 16, 32. GPU acceleration calculation is used during the training process;

[0036] The models for identifying transmission towers and the models for identifying hydropower stations, thermal power stations, and substations adopt different training parameters. The learning rate is set to 0.001, and the image batch size for calculation is 256. Considering the size of the transmission tower in the image, the target anchors are set as: 2, 4, 8, 16, corresponding to the sizes of the transmission tower in the remote sensing image being 32, 64, 128, 256 pixels; Additionally, considering the amount of training data, the maximum number of training iterations is set to 80,000 times, and the range of the target anchors is set as: 4, 8, 16, 32. Due to the different amounts of data, the maximum number of training iterations is 100,000 times;

[0037] During the model training process, a model pre - trained on the ImageNet dataset is used to initialize the weights of the deep learning. In the rough screening of power infrastructure targets, reduce the confidence value of determining a ground object as a certain category, and reduce the confidence level to 0.3.

[0038] Preferably, prior knowledge in the field of power grid infrastructure:

[0039] 1) Characteristics of transmission towers: Transmission towers appear small in remote sensing images. When the fusion degree of the tower and the background environment is high, the detection effect of the transmission tower is poor. The target recognition of transmission towers adopts a target recognition method based on deep learning, specifically using a target recognition strategy that combines ResNet and Faster R-CNN;

[0040] 2) Characteristics of hydropower stations: The imaging target of hydropower stations is large in images. It is necessary to combine the domain knowledge unique to hydropower stations. There must be a water surface environment upstream and downstream of the hydropower station. Deep learning is used again to classify the images, extract the water surface in a certain scene, and perform fine screening on the rough screening results of the hydropower station. In addition, within a certain range of large and medium-sized hydropower stations, there will be a certain number of transmission towers. The spatial and quantitative relationships between the hydropower station and the transmission towers are used to assist the fine screening of the hydropower station;

[0041] 3) Characteristics of substations: The site of the substation is close to the road and the access road is short. There are a certain number of transmission towers within a certain range of the substation. The fine screening of the substation first uses the spatial logical relationship between the substation and the transmission towers. Secondly, the substation covers a large area and appears as a regular rectangular block in remote sensing images. The rough screening results of the substation are classified by images, and the deep neural convolutional network will be used again for image classification;

[0042] 4) Characteristics of thermal power plants: In the rough screening stage of thermal power plants, the cooling tower is used as the key feature for the identification of thermal power plants. There are a certain number of transmission towers around the thermal power plant, and there is a substation set within a certain distance range. Combining the domain knowledge of thermal power plants, in the fine screening stage of thermal power plant detection, the spatial and quantitative relationships between the thermal power plant and the transmission towers and substations are used.

[0043] Preferably, feature extraction based on power grid domain knowledge: In addition to the geometric information, color information, texture information, and context information features of power infrastructure, for different power infrastructure, extract their unique ground object features for fine screening of target recognition. Specifically, for hydropower stations, use the method of deep learning image segmentation to extract the scene features of their upstream and downstream, judge whether there is a water surface in their upstream and downstream environments, and combine the transmission towers around the hydropower station for auxiliary judgment; for substations, classify the images of the rough screening results of the target and combine their spatial relationships with the surrounding transmission towers; for thermal power plants, pay attention to their unique cooling towers and their spatial logical relationships with the surrounding transmission towers and substations;

[0044] Establish a classifier applicable to the fine screening of power infrastructure target recognition based on the specific domain knowledge of various power infrastructures. Calculate the scores for the final target fine screening according to the domain knowledge of various power infrastructures, and set the threshold for judging the scores of various targets to 0.8;

[0045] 1) Fine screening of hydropower station recognition results

[0046] The fine screening of hydropower station target recognition needs to combine the target recognition results of transmission towers and the water surface extraction results. The extraction of the water surface adopts an image semantic segmentation method based on deep learning. Specifically, a porous full convolutional neural network is used. When using deep learning for image segmentation tasks, two bottlenecks are encountered. One is the information loss caused by downsampling, which is solved by the method of dilated convolution; the other is the inaccurate edges caused by spatial invariance, which is solved by fully connected CRF;

[0047] The fine screening process of hydropower station target recognition is as follows:

[0048] First, taking the rough screening results of the hydropower station target as the center, calculate the distance L from all transmission towers to the center of the hydropower station. The center coordinates of the hydropower station are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is formula 1:

[0049]

[0050] According to prior knowledge, when L is less than 1 kilometer, the number of transmission towers within the hydropower station range increases by 1. The number of power towers is represented by N, and the influence value S(A N ) generated by the number of power towers within a certain range of the hydropower station, and the calculation formula of S(A N ) is as formula 2:

[0051]

[0052] When there are no transmission towers within the range of L, the value of S(A N ) is 0; when the number of transmission towers N is greater than 0 and less than 10, the value of S(A N ) is 0.5; when the number of transmission towers N is greater than 10, the value of S(A N ) takes 1, and the calculation formula is as follows:

[0053]

[0054] In formula 3 and formula 4, the score of the hydropower station is S(B), and the influence value of the water surface on the hydropower station is S(W). When there is no water environment in the image, S(W) = 0. When there is a water surface in the image, s(W) = 1, S(B,A N) represents the score of a certain target being determined as a hydropower station under the influence of the number of electric towers, and S(B, W) represents the score of a certain target being determined as a hydropower station under the influence of water surface conditions;

[0055] In summary, the final score calculation formula for the possible targets of a hydropower station is as shown in Equation 5:

[0056]

[0057] S(B, A N , W) represents the score of the final refined screening of the hydropower station. The larger S(B, A N , W) is, the greater the possibility that the target is a hydropower station;

[0058] 2) Refined screening of substation identification results

[0059] First, use the spatial logical relationship between the substation and the transmission tower. Secondly, since the substation occupies a large area, perform image classification on the results of the rough screening. Image classification uses the deep neural convolutional network again. In this application, the target boxes obtained from the rough screening are used as the input images, and deep learning is used to perform image classification on the alternative target boxes. GoogleNet is used as the classification network model, and the dense block structure is used to approximately obtain the sparse structure, achieving the effect of improving performance without increasing the computational complexity;

[0060] The process of refined screening for substation target recognition is as follows: Based on the rough screening results of substation target recognition, a series of alternative target areas that may be substations are obtained, and each area corresponds to a score S(C1) determined as a substation. Then, the possible areas are input into the GoogleNet model for image classification. The model will output the category and score of the possible areas. If the category is determined to be a substation, the output score is taken as S(C2). Finally, considering the number of electric towers around the substation, use S(A N ) to represent the influence value of the electric towers within a certain range of the substation on the substation;

[0061] First, taking the rough screening results of the substation target as the center, calculate the distance L between all transmission towers and the center of the substation. The center coordinates of the substation are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is Equation 6:

[0062]

[0063] According to prior knowledge, when L is less than 500 meters, the number of transmission towers within the range of the hydropower station increases by 1. The number of electric towers is represented by N. The influence value S(A N ) generated by the electric towers within a certain range of the substation on the substation, and the calculation formula of S(A N ) is as

[0064] Formula 7:

[0065]

[0066] When there is no transmission tower within the range L, then the value of P(A N ) is 0; when the number N of transmission towers in the current year is greater than 0 and less than 10, the value of S(A N ) is 0.5; when the number N of transmission towers in the current year is greater than 10, the value of S(A N ) takes 0.5, and the calculation formula is as follows:

[0067]

[0068] In Formula 8 and Formula 9, S(C) represents the score for determining the target as a substation. The larger S(C) is, the greater the possibility that the target is a substation. S(C1) is the score of the substation in the rough screening stage of target recognition, and S(C2) is the score for classifying the alternative target box as a substation. The influence value of the power tower on the substation is S(A N ), and S(C,A N ) represents the score for determining a certain target as a substation under the condition that there are a certain number of transmission towers within the range.

[0069] Preferably, establish a refined screening dataset for power infrastructure detection:

[0070] 1) Water surface extraction: The refined screening of the target of power infrastructure is based on the rough screening of target recognition. The dataset for refined screening of the target is the detection result of rough screening of target recognition. In addition to the data for rough screening of targets such as hydropower stations, thermal power stations, and substations, the data for refined screening of the target includes water surface segmentation data and substation data for scene classification;

[0071] Divide all samples into training data and test data according to 4:1 for model training of water surface extraction;

[0072] 2. Image classification: Divide all samples into training data and test data according to 4:1. When collecting images of substations, not only the consistency of the distribution of training data should be considered, but also the images of substations in all situations should be ensured;

[0073] The data information for model training of image classification includes a total of 8 categories including substations. In the preparatory work for training, the data is augmented, including rotating, mirroring, and translating the data, and the training data is normalized to the same size.

[0074] Preferably, the refined screening process for power infrastructure detection:

[0075] 1) Transmission tower target recognition: The target recognition of the transmission tower integrates the target recognition strategies of Faster R-CNN and ResNet. The model training for identifying the transmission tower is implemented with the help of the TensorFlow framework;

[0076] According to the size of the transmission tower in the image, the target anchor sizes for training are set to 2, 4, 8, and 16. The training environment is the Linux system, and the GPU is used during the training process. The initial learning rate is 0.001, the maximum number of training times is 80,000 times, and the learning rate decreases once every 20,000 iterations;

[0077] During the model training process, the model pre-trained on the ImageNet dataset is used to initialize the weights of the deep learning, which can distribute the already learned model parameters to the new model, accelerating and optimizing the learning efficiency of the model;

[0078] 2) Water surface extraction: The learning rate parameter used for training the deep learning model for water surface segmentation is 0.0001, the batchsize parameter is 4, the momentum parameter is 0.9, the weight decay is 0.0005, and the maximum number of iterations is 50,000 times. Fine-tuning is performed on the pre-trained model;

[0079] 3) Image classification: Based on the Caffe deep learning framework, the GoogLeNet network model is used for the model training of image classification. The learning rate is 0.01. Different from the model for object detection, for the model for image classification, during training, the Batchsize of the training data and the Batchsize of the validation data are different, which are 64 and 32 respectively;

[0080] In addition to the substation images as the positive samples for image classification, during the training process, another 7 types of ground object samples, namely farmland, water surface, desert, forest, oil tank, airplane, and residential area, are added as negative samples to participate in the training. In the classification results, only the classification results of the substation are taken.

[0081] Compared with the prior art, the innovation points and advantages of this application are as follows:

[0082] (1) In response to the need to build a smart grid for energy resource sharing, this application combines deep learning and domain knowledge of power infrastructure to achieve the target recognition of power infrastructure in multi-scale remote sensing images. According to the characteristics of power infrastructure in remote sensing images, the target recognition strategies FasterR-CNN and ResNet are optimized and integrated, and the network parameters are modified to adapt to the detection task of power infrastructure. Based on deep learning-based target recognition, the rough screening of target recognition is completed, and the recall rate of detection results is improved by reducing the confidence of target determination to obtain more possible targets. Then, based on the rough detection results of power infrastructure, the domain knowledge characteristics of various types of power infrastructure are analyzed, and the target characteristics applicable to the fine screening of target recognition are extracted, including the characteristic information of substations at different levels, the characteristic information of transmission towers at different levels and types, the characteristic information of the cooling towers of thermal power plants, the distinguishing characteristic information of hydropower stations and reservoirs, and the spatial logical relationship characteristics between various types of power infrastructure. For different types of power infrastructure, target characteristics based on grid domain knowledge are established for the fine screening of power infrastructure detection. The specific methods include water surface extraction, target reclassification, spatial relationship calculation, etc. Experiments show that the target recognition method combining deep learning and domain knowledge of power infrastructure significantly improves the accuracy of power infrastructure target recognition and can also ensure a relatively good recall rate.

[0083] (2) This application proposes unsupervised training and weakly supervised learning of deep learning models. For the need of fine recognition of grid infrastructure, the training data for target recognition are collected and sorted, and the training data are labeled under a unified data standard to establish a remote sensing image library of power infrastructure for target recognition. A rough screening method for power infrastructure recognition is proposed, and the target recognition strategy integration of FasterR-CNN and ResNet is improved. According to the characteristics of power infrastructure in remote sensing images, the network parameters are modified to adapt to the detection task of power infrastructure. Based on the target recognition of convolutional neural networks, the rough screening of the final target recognition is completed, and the recall rate of detection results is improved by reducing the confidence of target determination to obtain more possible targets. This application makes full use of the fusion characteristics of multi-source remote sensing sensors, comprehensively uses multi-sensors, such as infrared, hyperspectral, visible light, SAR and other sensors for target recognition to improve the accuracy of target recognition; makes full use of advanced computing systems such as high-performance technology and cloud computing to reduce the time-consuming of target recognition; through in-depth analysis, design and understanding of features, makes full use of features such as spectrum, shape, texture, structure, etc. for accurate target category determination, and cleverly combines the prior knowledge of the grid infrastructure field, so that the recognition speed and accuracy of the grid infrastructure in remote sensing images are greatly improved.

[0084] (3) This application proposes a refined screening method for power infrastructure identification. By analyzing the domain knowledge characteristics of various power infrastructures, target features applicable to refined screening of target identification are extracted. Specifically, for hydropower stations, a method based on deep learning image segmentation is used to extract the scene features upstream and downstream, determine whether there is a water surface in the upstream and downstream environments, and combine with the power towers around the hydropower station for auxiliary judgment; for substations, image classification is performed on the results of rough screening of the target, and combined with its spatial relationship with the surrounding transmission towers; for thermal power stations, attention is paid to its unique cooling towers and its spatial logical relationship with the surrounding transmission towers and substations; by combining the domain knowledge of power infrastructure, the target identification effect in this field is improved. GPU and CUDA acceleration are used, along with optimization of scene perception and refined identification of real-time power grid infrastructure, to establish a real-time perception and power target detection intelligent system applicable to UAV remote sensing, including software and hardware. The system has good stability, high spatial perception efficiency, fast target detection speed, high accuracy, and preferably solves the identification and monitoring of UAV remote sensing avoiding intelligent power grid infrastructure. Description of the Drawings

[0085] Figure 1 It is a flowchart of rough screening for power infrastructure target identification by integrating ResNet and FasterR-CNN.

[0086] Figure 2 It is an example diagram of some training samples of transmission towers.

[0087] Figure 3 It is an example diagram of some training samples of hydropower stations.

[0088] Figure 4 It is an example diagram of some training samples of thermal power stations.

[0089] Figure 5 It is an example diagram of some training samples of substations.

[0090] Figure 6 It is a schematic diagram of the transformation methods and parameters of training samples.

[0091] Figure 7 It is a diagram of the correct detection result of rough screening of hydropower stations.

[0092] Figure 8 It is a schematic diagram of the wrong detection of misjudging an oil tank as a cooling tower in rough screening of thermal power stations.

[0093] Figure 9 It is a schematic diagram of the classifier for refined screening of power infrastructure target identification.

[0094] Figure 10 It is a schematic diagram of calculating the refined screening score of target identification taking a hydropower station as an example. Detailed Implementation Manner

[0095] The technical solution of the intelligent power grid infrastructure fine recognition method based on remote sensing images provided by this application will be further described below in conjunction with the accompanying drawings, so that those skilled in the art can better understand this application and be able to implement it.

[0096] Remote sensing image target recognition is to process image features and recognize relevant targets in the image. The process of target recognition is actually a discrimination process of different image objects by combining multiple pattern recognition algorithms and target feature libraries with the support of a target library. Due to the development of sensor technology and aerospace technology, the quantity and quality of remote sensing images have been greatly improved. This intelligent information processing method is crucial for the understanding of remote sensing images.

[0097] In response to the need to build an intelligent power grid to achieve energy resource sharing, this application combines deep learning and domain knowledge of power infrastructure to achieve target recognition of power infrastructure in multi-scale remote sensing images.

[0098] This application optimizes and integrates the target recognition strategies FasterR-CNN and ResNet according to the characteristics of power infrastructure in remote sensing images, and modifies the network parameters to adapt to the detection task of power infrastructure. The target recognition based on deep learning completes the rough screening of target recognition, and improves the recall rate of the detection results by reducing the confidence of target determination to obtain more possible targets; then, based on the rough detection results of power infrastructure, the domain knowledge characteristics of various types of power infrastructure are analyzed, and the target features that can be applied to the fine screening of target recognition are extracted, specifically including: characteristic information of substations of different levels, characteristic information of transmission towers of different levels and types, characteristic information of the cooling towers of thermal power plants, discriminative characteristic information between hydropower stations and reservoirs, and spatial logical relationship characteristics between various types of power infrastructure. For different categories of power infrastructure, target features based on power grid domain knowledge are established for the fine screening of power infrastructure detection. The specific methods include water surface extraction, target reclassification, spatial relationship calculation, etc. Experiments show that the target recognition method combining deep learning and domain knowledge of power infrastructure can improve the accuracy of power infrastructure target recognition and at the same time ensure a better recall rate.

[0099] (1) Unsupervised training and weakly supervised learning of deep learning models

[0100] For the need of fine recognition of power grid infrastructure, collect and organize training data for target recognition, annotate the training data under a unified data standard, and establish a remote sensing image library of power infrastructure for target recognition.

[0101] (2) Rough screening method for power infrastructure recognition.

[0102] Improve the target recognition strategy by integrating Faster R-CNN and ResNet. According to the characteristics of power infrastructure in remote sensing images, modify the network parameters to adapt to the detection task of power infrastructure. Based on the object recognition of the neural convolutional neural network, complete the rough screening of the final target recognition, and improve the recall rate of the detection results by reducing the confidence of target determination to obtain more possible targets.

[0103] (III) Fine screening method for power infrastructure recognition.

[0104] Analyze the domain knowledge characteristics of various power infrastructures and extract target features that can be applied to the fine screening of target recognition. Specifically, for hydropower stations, use the method of deep learning image segmentation to extract the scene features upstream and downstream of them, judge whether there is water surface in the upstream and downstream environments, and combine the electric towers around the hydropower stations for auxiliary judgment; for substations, perform image classification on the results of rough target screening and combine their spatial relationship with the surrounding transmission towers; for thermal power stations, pay attention to their unique cooling towers and their spatial logical relationship with the surrounding transmission towers and substations; combine the domain knowledge of power infrastructure to improve the target recognition effect in this field.

[0105] I. Rough screening of targets in remote sensing images of deep learning power grid facilities

[0106] Based on the TensorFlow deep learning framework, construct a convolutional neural network model for substation detection in remote sensing images. Based on ResNet with 101 layers of the network model, deep learning adopts a cross-computation structure of convolutional layers and downsampling layers to integrate the low-level features of the image and further obtain the high-level features of the image. Extract the self-features and internal features of the region box and the annotation box to obtain the effective information in the remote sensing image. The power infrastructures that need to be roughly screened for targets include: hydropower stations, substations, and thermal power stations.

[0107] Integrate the network model ResnNet and the target recognition framework Faster R-CNN, reduce the confidence of the output of the target alternative box to achieve a higher recall rate, and realize the preliminary recognition of the target. The flow chart of the rough screening of power infrastructure targets integrating ResNet and Faster R-CNN is as Figure 1 shown:

[0108] Suppose \(F(x)\) represents the mapping function of a block that only contains two or three layers. \(x\) is the input of the block, and \(F(x)\) is the output of the block. Assume they have the same dimension. During the training process, it is hoped that by modifying \(w\) and \(b\) in the network, an ideal \(h(x)\) can be fitted, which is an ideal mapping function from the input to the output. The goal is to modify \(w\) and \(b\) in \(F(x)\) to approximate \(h(x)\), use \(F(x)\) to approximate \(h(x)-x\). Finally, the output of the block changes from \(F(x)\) to \(F(x)+x\). After the change, the goal is to approximate the training function \(F(x)\) to \(h(x)-x\).

[0109] The residual network specifically opens up channels, making the optimization goal become \(H(x)-x\). \(H(x)-x\) represents the difference between the output and the input, rather than the original fitting output \(H(x)\). Here, \(x\) is the input and \(H(x)\) is the original expected mapping output of a certain layer. The gradient from the deep layer can directly pass through unobstructed to the upper layer, enabling the effective training of the parameters of the shallow network.

[0110] In the problem of power infrastructure target recognition in remote sensing images, in addition to requiring a certain recognition accuracy, it is more necessary to achieve a higher recall rate. The recognition accuracy and recall rate of power infrastructure are more important than the recognition speed. The target recognition algorithm based on the alternative region has experienced a process from RCNN to SPP-Net, then from SPP-Net to FastR-CNN, and then to FasterR-CNN, gradually realizing end-to-end training, achieving a faster detection speed and better detection effects.

[0111] First, input an image of any size into the deep learning network ResNet. After the forward propagation of the CNN to the last shared convolutional layer, on the one hand, a feature map for the input of the RPN network is obtained, and on the other hand, continue the forward propagation to the specific convolutional layer to generate a higher-dimensional feature map. The feature map input to the RPN network obtains region alternative boxes and region scores by the RPN network, and the non-maximum suppression algorithm is used for the region scores to output the first \(N\) ( \(N = 300\)) region alternative boxes. Then, input the high-dimensional feature map and the region alternative boxes into the ROI pooling layer to extract the features of the region alternative boxes. Finally, input the features of the region alternative boxes into the fully connected layer and output the classification score of this region and the target box after regression.

[0112] (1) Generate target alternative boxes based on RPN

[0113] Using the correspondence relationship, map the points on the feature map to the original image, generate many windows at the positions of the original image, with the scale of the windows fixed. Then calculate the IOU between the windows and the target ground truth (manual annotation), and give it positive and negative labels according to the division rules of positive and negative samples for learning. Train a network RPN. RPN makes three fixations on the scale windows: the scale change (three scales) is fixed, the scale ratio change is fixed, and the sampling method is fixed to reduce the complexity of target recognition. The sampling method of RPN is to sample only on the corresponding ROI in the original image for each point on the feature map. The input of RPN is an image of any size, and it outputs some target candidate boxes, and each candidate box has a target existence score and confidence information.

[0114] Specifically, after the image is input into the network, it successively passes through a series of convolutional layers and relu activation function layers, and finally obtains a feature map, and outputs a feature map of 51*39*256 dimensions. The feature map is used for the subsequent selection of candidate targets, and the coordinates at this time can still be mapped to the original image. Multiple regional targets are predicted through the scale of the target anchor and the set ratio. According to the sizes of various types of power infrastructure in remote sensing images of different resolutions in the training data, modify the scale of the target anchor. Specifically, the target anchors for detecting transmission towers are set to 2, 4, 8, 16, and the target anchors for hydropower stations, thermal power stations, and substations are 4, 8, 16, 32, and the ratio remains unchanged. More target candidate boxes of different scales will be generated at the reference points.

[0115] (II) Divide positive and negative samples

[0116] First, for each manually calibrated target ground truth region, retain the target anchor with the largest overlap ratio with it as the positive sample, ensuring that each target ground truth corresponds to at least one positive sample target anchor. For the remaining target anchors, if the overlap with the manually calibrated region is greater than 0.7, it is divided into the positive sample, and if the overlap with the manually calibrated region is less than 0.3, it is divided into the negative sample. Each true target box corresponds to multiple positive sample target anchors, but each positive sample target anchor represents only one true target. In addition, target anchors that do not meet the overlap area or cross the image boundary are discarded.

[0117] In actual training, randomly sample 256 target anchors in each image and put them into the calculation of the loss of a small batch of images to ensure that the ratio of positive samples to negative samples is 1:1. When the number of positive samples is less than 128, negative samples are automatically supplemented for sampling.

[0118] (III) Define the loss function

[0119] For each target anchor, first connect a softmax binary classifier at the back, set 2 score outputs, respectively representing the score of being an object and the score of not being an object P i, and then connect the regression output of a bounding box, representing the 4 coordinate positions t of this target anchor. i , to determine the overall loss function of the RPN.

[0120] (4) Establish a rough screening dataset for power infrastructure recognition

[0121] The target library is a remote sensing target image library, including: remote sensing images of thermal power plants, remote sensing images of hydropower plants, remote sensing images of transmission towers, and remote sensing images of substations;

[0122] The data used for rough screening of target recognition are all manually collected from Google Maps images and Tianditu remote sensing images, including 4 types of power infrastructure, namely 256 hydropower plant images, 196 thermal power plant images, and 164 substation images; among them, there are 261 hydropower stations, 607 cooling towers, and 195 substations. All the collected data are manually labeled with the true values of the targets under a unified data annotation specification to facilitate the calculation of the accuracy of target recognition.

[0123] Large and medium-sized transmission towers are located in the suburbs or the wild. Compared with other power infrastructure, the background is simple, independent, and affected by sunlight. Transmission towers in remote sensing images at different times have different shadows; in addition, the background types of transmission towers are diverse, including forests, farmlands, and deserts. The difficulty in recognizing transmission towers lies in their small size, which is difficult to distinguish in images with low resolution. In Google images with a resolution of 1 to 2 meters, they occupy 30 to 60 pixels, and in remote sensing images with a resolution of 0.6 meters, they occupy 60 to 120 pixels. There are a total of 627 data samples of transmission towers, including 1519 transmission towers;

[0124] In remote sensing images with a resolution of 5 meters, hydropower plants occupy 150 - 200 pixels in the image, and in remote sensing images with a resolution of 2 meters, they occupy 375 - 500 pixels in the image. Compared with other power infrastructure, hydropower plants cover a large area, and there must be water surfaces in the upper and lower environments. Coupled with the special line features of hydropower plants, they are relatively easy to recognize in remote sensing images. The problem in recognizing hydropower plants lies in the distinction from similar ground objects, including dams and some similar bridges; there are a total of 256 data samples of hydropower plants, including 261 hydropower stations.

[0125] Compared with hydropower plants, thermal power plants have more complex background content and greater recognition difficulty. However, the power generation principle and power generation system of thermal power determine the special structure of cooling towers. From the perspective of remote sensing images, all thermal power plants contain cooling towers. The cooling towers are used as the key features for the preliminary recognition of thermal power plants. In remote sensing images with a resolution of 1 to 2 meters, the cooling towers occupy 60 to 80 pixels, and in remote sensing images with a resolution of 0.6 meters, they occupy 120 to 160 pixels; there are 196 data samples of thermal power plants, including 607 cooling towers.

[0126] In a remote sensing image with a resolution of 5 meters, a complete substation occupies 300 to 500 pixels. In a remote sensing image with a resolution higher than 5 meters, it is difficult to retain the integrity of a substation in an image with a size of 2000 * 2000 pixels and will be cut into image blocks; there are 164 data samples of substations, including 195 substations.

[0127] Figure 2 , Figure 3 , Figure 4 , Figure 5 Examples of some samples corresponding to transmission towers, hydropower stations, thermal power stations, and substations respectively. The samples are divided into training data and test data in a ratio of 4:1, and the data is augmented through rotation transformation, resampling, color transformation, and mirror transformation. The variation parameters of the training samples are as Figure 6 shown.

[0128] (V) Coarse Screening Process for Power Infrastructure Identification

[0129] Compared with ordinary target recognition, the power infrastructure in remote sensing images has more complex scene content, which brings difficulties to the identification of power infrastructure. Each type of power infrastructure has special characteristics. Analyze the characteristics of various types of power infrastructure and set the parameters of the deep learning model. During the coarse screening of power infrastructure, the aim is to achieve a higher recall rate and avoid missing targets. According to the sizes of various power infrastructure targets at different resolutions, set the training target anchors with sizes of 2, 4, 8, 16, and 32. The training environment for the experiment is the Linux system, and GPU acceleration computing is used during the training process.

[0130] The models for identifying transmission towers and the models for identifying hydropower stations, thermal power stations, and substations use different training parameters. The learning rate is set to 0.001, the image batch size for calculation is 256. Considering the size of the transmission tower in the image, set the target anchors as: 2, 4, 8, 16, corresponding to the sizes of the transmission towers in the remote sensing image being 32, 64, 128, and 256 pixels. Additionally, considering the amount of training data, set the maximum number of training iterations to 80,000 times. The range of the target anchors is set to: 4, 8, 16, 32. Due to the different amounts of data, the maximum number of training iterations is 100,000 times.

[0131] During the model training process, use a model pre-trained on the ImageNet dataset to initialize the weights of the deep learning. To improve the recall rate of the target recognition results, during the coarse screening of power infrastructure targets, reduce the confidence value for determining a ground object as a certain class. Reduce the confidence level to 0.3.

[0132] (V) Analysis of Coarse Screening Results for Power Infrastructure Identification

[0133] The recognition accuracy of the transmission towers in the model is 89.26%, and the recall rate is 85.69%. To ensure the recognition accuracy of transmission towers, in the target recognition of transmission towers, the confidence value determined for the candidate bounding boxes as transmission towers has not been reduced and remains 0.5. Currently, for the target recognition of transmission towers, the existing problem is the missed detection of transmission towers. The reason for this situation may be the characteristics of multiple categories of transmission towers, resulting in differences among transmission towers in remote sensing images. Additionally, affected by lighting, the shadow directions of transmission towers are difficult to be consistent, leading to differences in targets. The missed-detected transmission towers are generally small and the imaging is not clear. For transmission tower targets that are small (less than 20 pixels in the image) and have poor imaging effects, the model is still unable to perform target recognition on them.

[0134] The recognition accuracies of hydropower stations, thermal power stations, and substations are 85.03%, 88.75%, and 77.73% respectively, and the recall rates of the recognition of these three types of targets are 92.34%, 98.06%, and 90.89% respectively. In the experiment, more attention is paid to the recall rate of target recognition in the target rough screening stage, and the accuracy of the target recognition results is emphasized in the target fine screening stage.

[0135] For hydropower stations, there is a situation of misjudgment of similar targets, such as dams and bridges. Especially after reducing the confidence value determined for the target, the false alarms in the target recognition results will increase. Although the recall rate of target recognition will increase, the accuracy of target recognition will decrease. Directly using a deep learning model to detect hydropower stations is prone to misjudgment of similar ground objects, and it is difficult to consider all similar ground objects as negative samples for training in the experiment. Figure 7 The correct detection result diagrams of hydropower stations are listed.

[0136] For thermal power stations, the problem is the misjudgment of similar ground objects, such as oil tanks. Some images containing oil tanks were collected in the experiment, and it was found that the rough detection model misjudged some oil tanks as condensation towers, thus affecting the recognition effect of thermal power stations. Some of the incorrect detection results (misjudging oil tanks as condensation towers) are as Figure 8 shown. In subsequent experiments, it can be considered to add ground objects such as oil tanks to participate in the training of the model, and it can be verified through experiments whether this method is effective.

[0137] For substations, the existing problems are that the size of substations changes too much in images with different resolutions, and often an incomplete substation is detected. Additionally, there is also a phenomenon of missed detection in the detection results of substations. The reasons for these problems may be that the data distribution of substations is relatively complex, and it is difficult to have a unified standard in the annotation stage of substation targets. In addition, the complex characteristics of substations themselves and inconsistent imaging conditions are also important reasons affecting the recognition effect of substations.

[0138] Generally speaking, there are still many problems in the results of the rough screening of target recognition. The misjudgment of similar object targets is relatively serious. When the accuracy of target recognition is guaranteed, the recall rate of target recognition will be affected. Therefore, the method of directly using deep learning to solve the problem of power infrastructure target recognition in remote sensing images cannot well meet the actual application requirements. However, the rough screening of power infrastructure target recognition can be completed by using the target recognition method of deep learning. From the experimental results, the recall rates of the detection results of various power infrastructures are relatively high, and the correct targets need to be screened out from the detection results with high recall rates. Therefore, it is necessary to further refine the results of the rough screening of target recognition by combining the domain knowledge of power infrastructure.

[0139] II. Fine Screening of Targets Based on Grid Domain Knowledge

[0140] (1) Prior Knowledge in the Field of Grid Infrastructure

[0141] 1. Characteristics of Transmission Towers

[0142] Transmission towers have small images in remote sensing images. When the fusion degree of the tower and the background environment is high, the detection effect of the transmission tower is poor. The target recognition of the transmission tower adopts the target recognition method based on deep learning, specifically using the target recognition strategy that combines ResNet and FasterR-CNN.

[0143] 2. Characteristics of Hydropower Stations

[0144] Hydropower stations have large target images in the images. However, the results of the rough screening of hydropower station target recognition show that it is difficult to distinguish similar objects directly using the detection algorithm of deep learning, and the accuracy of hydropower station target recognition is not high. Therefore, it is necessary to combine the unique domain knowledge of hydropower stations. There must be a water surface environment upstream and downstream of the hydropower station. Use deep learning to classify the image again to extract the water surface in a certain scene and refine the rough screening results of the hydropower station. In addition, within a certain range of large and medium-sized hydropower stations, there will be a certain number of transmission towers. Use the spatial relationship and quantity relationship between the hydropower station and the transmission tower to assist the fine screening of the hydropower station.

[0145] 3. Characteristics of Substations

[0146] The site of the substation is close to the road, and the access road is short. There are a certain number of transmission towers within a certain range of the substation. For the fine screening of the substation, first, use the spatial logical relationship between the substation and the transmission tower. Second, the substation covers a large area and is mostly a regular rectangular block in remote sensing images. Classify the rough screening results of the substation, and the image classification will use the deep neural convolutional network again.

[0147] 4. Characteristics of Thermal Power Stations

[0148] In the rough screening stage of a thermal power station, the condensation tower is used as a key feature for identifying the thermal power station. There are a certain number of transmission towers around the thermal power station, and a substation is set within a certain distance range. Combining the domain knowledge of thermal power stations, in the refined screening stage of thermal power station detection, the spatial and quantitative relationships between the thermal power station, transmission towers, and substations are utilized.

[0149] (2) Feature extraction based on power grid domain knowledge

[0150] In addition to the geometric, color, texture, and context information features of power infrastructure, for different power infrastructure, extract their unique ground object features for use in the refined screening of target recognition. Specifically, for hydropower stations, use a method based on deep learning image segmentation to extract the upstream and downstream scene features, determine whether there is a water surface in its upstream and downstream environment, and combine the power towers around the hydropower station for auxiliary judgment; for substations, perform image classification on the results of the target rough screening and combine their spatial relationship with the surrounding transmission towers; for thermal power stations, pay attention to their unique condensation towers and the spatial logical relationship with the surrounding transmission towers and substations.

[0151] Based on the unique domain knowledge of various types of power infrastructure, establish a classifier suitable for the refined screening of power infrastructure target recognition. The domain knowledge of various types of power infrastructure to be combined is as Figure 9 shown. According to the domain knowledge of various types of power infrastructure, calculate the scores for the final target refined screening, and set the threshold for judging the scores of various targets to 0.8.

[0152] 1. Refined screening of hydropower station recognition results

[0153] The refined screening of hydropower station target recognition needs to combine the target recognition results of transmission towers and the water surface extraction results. The extraction of the water surface uses a method of deep learning-based image semantic segmentation. Specifically, a porous hole fully convolutional neural network is used. When using deep learning for image segmentation tasks, two bottlenecks are encountered. One is the information loss caused by downsampling, which is solved by the method of dilated convolution; the other is the inaccurate edges caused by spatial invariance, which is solved by fully connected CRF.

[0154] Based on the target recognition results of transmission towers and the water surface extraction results, the process of refined screening of hydropower station targets: The specific process is to perform water surface extraction, target recognition of transmission towers, and rough screening of hydropower station targets on the image to be detected simultaneously.

[0155] The process of refined screening of hydropower station target recognition is as follows:

[0156] First, centered on the target rough screening result of the hydropower station, calculate the distance L from all transmission towers to the center of the hydropower station. The center coordinates of the hydropower station are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is Equation 1:

[0157]

[0158] According to prior knowledge, when L is less than 1 kilometer, the number of transmission towers within the scope of the hydropower station increases by 1. The number of electric towers is represented by N, and the influence value S(A N ) generated by the number of electric towers within a certain range of the hydropower station is calculated as follows. The calculation formula of S(A N ) is Equation 2:

[0159]

[0160] When there are no transmission towers within the range of L, the value of S(A N ) is 0; when the number of transmission towers N in the current year is greater than 0 and less than 10, the value of S(A N ) is 0.5; when the number of transmission towers N in the current year is greater than 10, the value of S(A N ) is 1. The calculation formula is as follows:

[0161]

[0162] In Equations 3 and 4, the score of the hydropower station is S(B), and the influence value of the water surface on the hydropower station is S(W). When there is no water environment in the image, S(W) = 0. When there is a water surface in the image, S(W) = 1. S(B, A N ) represents the score of a certain target being determined as a hydropower station under the influence of the number of electric towers, and S(B, W) represents the score of a certain target being determined as a hydropower station under the influence of the water surface condition;

[0163] In summary, the final score calculation formula for the possible target of a certain hydropower station is Equation 5:

[0164]

[0165] S(B, A N , W) represents the score of the final refined screening of the hydropower station. The larger S(B, A N , W) is, the greater the possibility that the target is a hydropower station.

[0166] 2. Refined screening of substation identification results

[0167] Firstly, utilize the spatial logical relationship between the substation and the transmission tower. Secondly, since the substation occupies a large area, perform image classification on the results of the rough screening. For image classification, use the deep neural convolutional network again. In this application, take the target bounding boxes obtained from the rough screening as the input images, and use deep learning to perform image classification on the alternative target bounding boxes. Employ GoogleNet as the network model for classification, and utilize the dense block structure to approximately obtain the sparse structure, achieving the effect of performance improvement without increasing the computational complexity.

[0168] The process of fine screening for the target recognition of the substation is as follows: Based on the results of the rough screening for substation target recognition, obtain a series of alternative target regions that may be the substation, and each region corresponds to a score S(C1) determined to be the substation. Then, input the possible regions into the GoogleNet model for image classification. The model will output the category and score of the possible regions. If the category is determined to be the substation, take the output score as S(C2). Finally, consider the number of power towers around the substation, and use S(A N ) to represent the influence value of the power towers within a certain range of the substation on the substation;

[0169] Firstly, taking the results of the rough screening of the substation target as the center, calculate the distance L between all transmission towers and the center of the substation. The center coordinates of the substation are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is Formula 6:

[0170]

[0171] According to prior knowledge, when L is less than 500 meters, the number of transmission towers within the range of the hydropower station increases by 1. The number of power towers is represented by N. The influence value S(A N ) generated by the power towers within a certain range of the substation is as follows. The calculation formula of S(A N ) is as

[0172] Formula 7:

[0173]

[0174] When there are no transmission towers within the range of L, the value of P(A N ) is 0; when the number N of transmission towers in that year is greater than 0 and less than 10, the value of S(A N ) is 0.5; when the number N of transmission towers in that year is greater than 10, the value of S(A N ) is taken as 0.5. The calculation formula is as follows:

[0175]

[0176] In Equations (8) and (9), S(C) represents the score of the target determined to be a substation. The larger S(C) is, the greater the possibility that the target is a substation. S(C1) is the score of the substation in the rough screening stage of the target, and S(C2) is the score of the alternative target box classified as a substation. The influence value of the electric tower on the substation is S(A N ), and S(C,A N ) represents the score of a certain target determined to be a substation that meets the condition of having a certain number of transmission towers within a range.

[0177] 3. Fine Screening of Thermal Power Station Recognition Results

[0178] There are a certain number of transmission towers around thermal power stations, and there are substations set within a certain distance range. For the fine screening of thermal power stations, their spatial relationships with transmission towers and substations are considered.

[0179] The fine screening process of thermal power station targets is as follows:

[0180] There is more than one cooling tower in a thermal power station. The average value of the scores of all targets determined to be cooling towers is obtained as S(D). The calculation formula of S(D) is as shown in Equation (10). Then, the influence of the existence of transmission towers and substations around the thermal power station on the thermal power station is considered;

[0181]

[0182] The process of fine screening for the recognition of thermal power station targets is as follows: First, taking the rough screening results of thermal power station targets as the center, calculate the distance L between the centers of all transmission towers and the substation. The center coordinates of the substation are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is as shown in Equation (11):

[0183]

[0184] According to prior knowledge, when L is less than 500 meters, the number of transmission towers within the range of this thermal power station increases by 1. The number of electric towers is represented by N. The influence value of the transmission tower on the thermal power station within a certain range of the thermal power station is S(A N ), and the calculation formula of S(A N ) is as shown in Equation (12):

[0185]

[0186] When there are no transmission towers within the range of L, the value of S(A N ) is 0; when the number N of transmission towers is greater than 0 and less than 10, the value of S(A N ) is 0.5; when the number N of transmission towers is greater than 10, S(A N) takes a value of 0.5. There is more than one condenser tower in a thermal power station. S(A N ) takes the maximum. The calculation formula is as follows:

[0187]

[0188] There is more than one condenser tower in a thermal power station. S(D) is the average score of all objects determined to be condenser towers. In Equations 13, 14, and 15, S(D,A N ) represents the score of the object targeted at a thermal power station when the number condition of transmission towers within a certain range is met; S(D,C) represents the score of the object targeted at a thermal power station under the condition that there is a substation; S(D,A N ,C) represents the score of an object determined to be a thermal power station when the conditions of the existence of both transmission towers and a substation are met.

[0189] (III) Establishing a refined screening dataset for power infrastructure detection

[0190] 1. Water surface extraction

[0191] The refined screening of power infrastructure targets is based on the rough screening of target recognition. The dataset for refined screening is the detection results of rough screening of target recognition. In addition to the data used for rough screening of hydropower stations, thermal power stations, and substation targets, the data for refined screening includes water surface segmentation data and substation data (tile data with a size of 256*256) for scene classification;

[0192] The data for training the porous hole convolutional neural network model for water surface extraction is sourced from GF-2 satellite data with a resolution of 0.8 meters. There are a total of 1572 images, each with a size of 500*500 pixels. All samples are divided into training data and test data at a ratio of 4:1. A part of the training data samples is listed for model training of water surface extraction.

[0193] 2. Image classification

[0194] The samples for substation image classification are images with a size of 224*224, totaling 1606, from Google Images and Tianditu. In the experiment, all samples are divided into training data and test data at a ratio of 4:1. When collecting substation images, not only the consistency of the training data distribution should be considered, but also the substation images in all situations should be ensured.

[0195] The data information for training the image classification model includes a total of 8 categories including substations. In the preparatory work for training, the data is augmented, including rotating, mirroring, and translating the data, and the training data is normalized to the same size.

[0196] (IV) Refined screening process for power infrastructure detection

[0197] 1. Transmission tower target recognition

[0198] The target recognition of the transmission tower integrates the target recognition strategy of FasterR-CNN and ResNet, and the model training for identifying the transmission tower is implemented with the help of the TensorFlow framework.

[0199] According to the size of the transmission tower in the image, the target anchor sizes for training are set to 2, 4, 8, and 16. The training environment for the experiment is the Linux system, and the NVIDIA Titan 1080 GPU is used during the training process. The initial learning rate is 0.001, the maximum number of training times is 80,000 times, and the learning rate decreases once every 20,000 iterations.

[0200] During the model training process, the model pre-trained on the ImageNet dataset is used to initialize the weights of the deep learning, which can distribute the already learned model parameters to the new model, accelerating and optimizing the learning efficiency of the model.

[0201] 2. Water surface extraction

[0202] The learning rate parameter used for training the deep learning model for water surface segmentation is 0.0001, the batchsize parameter is 4, the momentum parameter is 0.9, the weight decay is 0.0005, the maximum number of iterations is 50,000 times, and fine-tuning is performed on the pre-trained model, and the model accuracy reaches 92.3%.

[0203] 3. Image classification

[0204] Based on the Caffe deep learning framework, the GoogLeNet network model is used for the model training of image classification, where the learning rate is 0.01. Different from the model for object detection, for the model for image classification, the Batchsize of the training data and the Batchsize of the validation data are different during training, which are 64 and 32 respectively.

[0205] In addition to the substation images as the positive samples for image classification, during the training process, 7 types of ground object samples including farmland, water surface, desert, forest, oil tank, airplane, and residential area are additionally added as negative samples to participate in the training. In the classification results, only the classification results of the substation are taken in this application.

[0206] (V) Analysis of the refined screening results of power infrastructure detection

[0207] 1. Water surface extraction results

[0208] The water surface extraction results of 4 experimental images are given. The results are binary images, with white representing the water surface and black representing non-water surface. For the requirements, an accuracy of 92% is sufficient as the basis for judging the water surface scene. Therefore, it is feasible to apply the results of water surface extraction using DeepLab to the judgment of the water surface environment.

[0209] 2. Calculation of the refined screening score for target recognition - taking a hydropower station as an example

[0210] Figure 10 Three image examples for target recognition are given. After rough screening of the targets, it is found that there are targets judged as hydropower stations in all three images. However, it is obvious that Figure 10 (a) and (b) are both cases of misjudged targets. Therefore, it is necessary to screen these three images again by combining the domain knowledge of hydropower stations.

[0211] The specific calculation results are as follows: The target recognition strategy that fuses FasterR-CNN and ResNet, Figure 10 (a) The score of the object judged as a hydropower station is S Ba = 90.29%; Figure 10 (b) The score of the object judged as a hydropower station is S Bb = 85.52%; Figure 10 (c) The score of the object judged as a hydropower station is S Bc = 99.95%.

[0212] For the water surface extraction situation, Figure 10 (a) and (c) both have water surfaces, while there is no water surface in figure (b). Therefore Figure 10 (a) The score based on the water surface is S Wa = 1; Figure 10 (b) The score based on the water surface is S Wb = 1; Figure 10 (c) The score of the water surface is S Wc = 0;

[0213] For the information of electric towers, Figure 10 (a) The number of electric towers is 0, so the score of the number of electric towers is S ANa = 0; Figure 10 (b) The number of transmission towers is 3, so the score of the number of electric towers is S ANb = 0.5; Figure 10 (b) The number of transmission towers is more than 10, so the score of the number of its electric towers is S ANc = 1.

[0214] Based on the above conditions, it is calculated that Figure 10 (a) The score of the target is 67.72%, Figure 10The score of the target in (b) is 53.45%, Figure 10 and the score of the target in (c) is 99.95%. It can be seen from the calculation results that Figure 10 for the target with a very high initial rough screening score in (a), after the refined screening, the score decreased by 22.57%. Figure 10 The decrease in (b) is even greater, by 32.07%. And Figure 10 the initial rough screening result in (c) was already very high, and after the refined screening, the score was maintained. Theoretically, the higher the score, the greater the similarity to the detection target. According to the set score threshold of 0.8, Figure 10 (a) and (b) are both less than this threshold, so they do not appear in the detection results in the end. From the experimental results, the scores of the targets similar to hydropower stations are suppressed in the end, while the true hydropower station targets maintain a high recognition rate, thus showing the feasibility of the method adopted in this application.

[0215] 3. Refined Screening Results of Power Infrastructure

[0216] After the refined screening of power infrastructure, the accuracy rates of the detection results of hydropower stations, thermal power stations, and substations have all increased. Among them, the target recognition accuracy rate of hydropower stations has increased from the original 85.03% to 89.48%, an increase of 4.45%. The accuracy rate of thermal power stations has increased from 88.75% to 92.73%, an increase of 3.98%. The accuracy rate of substations has increased from 77.73% to 81.51%, an increase of 3.78%. At the same time, the recall rates of the detection results of these three types of power infrastructure are preserved. Thus, it can be seen that the detection method of power infrastructure adopted in this application is effective.

[0217] Separate result diagrams of the detection of some hydropower stations, thermal power stations, and substations are listed. For hydropower stations, the result of water surface extraction is used to judge whether there is a water surface around the hydropower station, so ground objects similar to hydropower stations on land, such as some bridges, can be excluded. Combining with the detection results of transmission towers, false results of similar ground objects such as dams, water surface roads, and bridges can be filtered. For the detection results of substations, there are often some incomplete substations, but only partial results of substations, so there is a great deal of uncertainty in the detection results. Classifying the ground objects similar to the detection results of substations again can improve the accuracy rate of the substation detection results. In the detection of the cooling towers of thermal power stations, it is easy to misjudge some gray oil tanks as cooling towers. Although smoke can be used as a feature of cooling towers, sometimes there is no smoke in the cooling towers. Therefore, using the characteristics of transmission towers and substations around thermal power stations can better exclude similar ground objects such as oil tanks.

[0218] III. Experimental Environment

[0219] (1) Model training environment

[0220] Due to the large dataset training scale of deep learning, a graphics processing unit (GPU) was used to accelerate the training of a ResNet model with 240,000 iterations in the experiment. In the experiment, TensorFlow framework was used for deep learning training, and GPU and CUDA acceleration were adopted under the Linux server.

[0221] (2) Object recognition environment

[0222] The object recognition environment was carried out on a computer under Windows, and GPU acceleration was also used for processing. All algorithms involved in object recognition were implemented in Python. Objects were detected in images with 1000 - 2000 pixels, and the time to detect one image was 7 - 9 seconds.

Claims

1. A smart grid infrastructure precise identification method based on remote sensing images, characterized in that: Combine deep learning with domain knowledge of power infrastructure to achieve target recognition of power infrastructure in multi-scale remote sensing images: According to the characteristics of power infrastructure in remote sensing images, optimize the fusion target recognition strategy FasterR-CNN and ResNet, and modify the network parameters to adapt to the detection task of power infrastructure. Complete the rough screening of target recognition based on deep learning target recognition, and improve the recall rate of detection results by reducing the confidence of target judgment to obtain more possible targets; Then, based on the rough detection results of power infrastructure, the domain knowledge features of various types of power infrastructure are analyzed to extract target features that can be used for precise target identification and screening, including: feature information of substations of different levels, feature information of transmission towers of different levels and types, feature information of condensing towers of thermal power stations, distinguishing feature information of hydropower stations and reservoirs, and spatial logical relationship features between various types of power infrastructure. For different types of power infrastructure, target features based on power grid domain knowledge are established for precise screening of power infrastructure detection. Specific methods include water surface extraction, target reclassification, and spatial relationship calculation. (1) Unsupervised training and weakly supervised learning of deep learning models: To meet the needs of accurate identification of power grid infrastructure, we collect and organize training data for target identification, annotate the training data under unified data standards, and establish a remote sensing image library of power infrastructure for target identification. (2) Rough screening method for power infrastructure identification: Improve the target recognition strategy by integrating FasterR-CNN and ResNet. According to the characteristics of power infrastructure in remote sensing images, modify the network parameters to adapt to the detection task of power infrastructure. The target recognition based on neural convolutional neural network completes the rough screening of the final target recognition. By reducing the confidence of target judgment, the recall rate of detection results is improved to obtain more possible targets. (3) Precision screening method for power infrastructure identification: Analyze the domain knowledge characteristics of various types of power infrastructure and extract target features that can be used for precision screening of target identification. Specifically, for hydroelectric power stations, a method based on deep learning image segmentation is used to extract the scene features of their upstream and downstream, determine whether there is water in their upstream and downstream environments, and combine the power towers around the hydroelectric power station to assist in the judgment; for substations, perform image classification on the results of rough target screening and combine their spatial relationship with the surrounding transmission towers; for thermal power stations, pay attention to their unique condensing towers and their spatial logical relationship with the surrounding transmission towers and substations; combine the domain knowledge of power infrastructure to improve the target recognition effect in this field.

2. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Deep learning for rough screening of targets in remote sensing images of power grid facilities: Based on the Tensorflow deep learning framework, a convolutional neural network model for substation detection in remote sensing images is constructed. Based on the 101-layer ResNet network model, deep learning uses a cross-calculation structure of convolutional layers and downsampling layers to integrate low-level features of the image, further obtain high-level features of the image, extract the inherent features and internal features of the region box and the annotation box, and obtain effective information in the remote sensing image. The power infrastructure that needs rough screening of targets includes: hydroelectric power stations, substations, and thermal power stations; The fusion network model ResnNet and the target recognition framework FasterR-CNN reduce the output confidence of the target candidate box to achieve a higher recall rate and realize the preliminary recognition of the target; Assume that F(x) represents a mapping function of a block with only two or three layers, x is the input of the block, and F(x) is the output of the block. Assume that they have the same dimension. During the training process, we hope to modify w and b in the network to fit an ideal mapping function h(x) from input to output. The goal is to modify w and b in F(x) to approximate h(x), and use F(x) to approximate h(x)-x. Finally, the output of bolck changes from F(x) to F(x)+x. The goal after the change is to make the training function F(x) approximate h(x)-x. The residual network specially opens up a channel so that the optimization target becomes H(x)-x, where H(x)-x represents the difference between the output and the input, rather than the original fitted output H(x), where x is the input and H(x) is the original expected mapping output of a certain layer. The gradient from the deep layer can pass directly and unimpeded to the previous layer, allowing the shallow network parameters to be effectively trained.

3. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Generate target candidate boxes based on RPN: Use the correspondence relationship to map the points on the feature map to the original image, and generate many windows at the position of the original image. The scale of the window is fixed, and then the IOU between the window and the target truth value is calculated. According to the division rule of positive and negative samples, it is given positive and negative labels for learning and training a network RPN. RPN fixes three scale windows: fixed scale change, fixed scale ratio change and fixed sampling method to reduce the complexity of target recognition. The sampling method of RPN is to sample only the corresponding ROI of each point in the feature map in the original image. The input of RPN is an image of any size, and some target candidate boxes are output, and each candidate box has a target existence score and confidence information; Specifically, after the image is input into the network, it passes through a series of convolutional layers and relu excitation function layers in sequence, and finally a feature map is obtained, and a 51*39*256-dimensional feature map is output. The feature map is used for the subsequent selection of candidate targets, and the coordinates at this time can still be mapped to the original image. Multiple regional targets are predicted by the scale of the target anchor and the set ratio. The scale of the target anchor is modified according to the size of various types of power infrastructure in remote sensing images of different resolutions in the training data. The target anchors used to detect transmission towers are set to 2, 4, 8, and 16, and the target anchors used for hydroelectric power stations, thermal power stations, and substations are 4, 8, 16, and 32. The ratios remain unchanged, and more target candidate boxes of different scales will be generated on the reference point.

4. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Divide positive and negative samples: First, for each manually calibrated target true value area, retain the target anchor with the largest overlap ratio and use it as the positive sample, ensuring that each target true value corresponds to at least one positive sample target anchor. For the remaining target anchors, if the overlap with the manually calibrated area is greater than 0.7, they are classified as positive samples. If the overlap with the manually calibrated area is less than 0.3, they are classified as negative samples. Each true target box corresponds to multiple positive sample target anchors, but each positive sample target anchor only represents one real target. In addition, target anchors that do not meet the overlapping area or cross the image boundary are discarded; In actual training, 256 target anchors are randomly sampled in each image and put into the loss calculation of a small batch of images to ensure that the ratio of positive samples to negative samples is 1:

1. When the number of positive samples is less than 128, negative samples are automatically sampled.

5. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Establish a rough screening data set for power infrastructure identification: the target library is a remote sensing target image library, including: remote sensing images of thermal power stations, remote sensing images of hydroelectric power stations, remote sensing images of transmission towers, and remote sensing images of substations; Large and medium-sized transmission towers are located in the suburbs or in the wild. Compared with other power infrastructure, the background is simple and independent. In the images illuminated by sunlight, the transmission towers in remote sensing images at different times have different shadows. In addition, there are many types of backgrounds for transmission towers, including forests, farmlands, and deserts. The difficulty in identifying transmission towers lies in their small size and difficulty in distinguishing them in low-resolution images. They occupy 30 to 60 pixels in Google images of 1 to 2 meters and 60 to 120 pixels in remote sensing images of 0.6 meters. In a remote sensing image with a resolution of 5 meters, a hydroelectric power station occupies 150-200 pixels, and in a remote sensing image with a resolution of 2 meters, it occupies 375-500 pixels. Compared with other power infrastructure, a hydroelectric power station occupies a large area, and there must be water above and below it. In addition, the special line features of a hydroelectric power station make it relatively easy to identify in remote sensing images. The problem in identifying a hydroelectric power station lies in distinguishing it from similar landforms, including dams and some similar bridges. Compared with hydroelectric power stations, thermal power stations have more complex background content and are more difficult to identify. However, the power generation principle and power generation system of thermal power generation determine the special structure of the condensing tower. From the remote sensing image, it can be seen that thermal power stations all have condensing towers. The condensing tower is used as the key feature for the preliminary identification of thermal power stations. The condensing tower occupies 60 to 80 pixels in the remote sensing image of 1 to 2 meters, and 120 to 160 pixels in the remote sensing image of 0.6 meters. In a remote sensing image with a resolution of 5 meters, a complete substation occupies 300 to 500 pixels. In a remote sensing image with a resolution higher than 5 meters, it is difficult to preserve the substation in its entirety in an image of 2000*2000 pixels and it will be cut into image blocks. The samples are divided into training data and test data at a ratio of 4:1, and the data are expanded through rotation transformation, resampling, color transformation, and mirror transformation.

6. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Power infrastructure identification and rough screening process: Analyze the characteristics of various types of power infrastructure, set the parameters of the deep learning model, and during the rough screening of power infrastructure, set the training target anchors according to the sizes of various types of power infrastructure targets at different resolutions. The sizes are 2, 4, 8, 16, and 32. The training process uses GPU accelerated computing. The model used to identify transmission towers and the model used to identify hydropower stations, thermal power stations, and transmission stations use different training parameters. The learning rate is set to 0.001, and the image batch size used for calculation is 256. Considering the size of the transmission tower in the image, the target anchors are set to: 2, 4, 8, 16, which corresponds to the size of the transmission tower in the remote sensing image is 32, 64, 128, 256 pixels. In addition, considering the amount of training data, the maximum number of training iterations is set to 80,000 times, and the range of the target anchor is set to: 4, 8, 16, 32. Due to the different data amounts, the maximum number of training iterations is 100,000 times. During the model training process, the model pre-trained on the ImageNet dataset was used to initialize the weights of deep learning. In the rough screening of basic power targets, the confidence value of judging the ground objects as a certain category was reduced to 0.

3.

7. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Prior knowledge in the field of power grid infrastructure: 1) Transmission tower features: Transmission towers are small in remote sensing images. When the towers are highly integrated with the background environment, the detection effect of the transmission towers is poor. The target recognition of transmission towers adopts a target recognition method based on deep learning, specifically a target recognition strategy that integrates ResNet and FasterR-CNN. 2) Characteristics of hydropower stations: Hydropower stations are large imaging targets in images. We need to combine the unique domain knowledge of hydropower stations. There must be water surface environments upstream and downstream of hydropower stations. We use deep learning to classify images again, extract the water surface in a certain scene, and perform fine screening on the rough screening results of hydropower stations. In addition, within a certain range of large and medium-sized hydropower stations, there will be a certain number of transmission towers. The spatial relationship and quantitative relationship between hydropower stations and transmission towers are used to assist in the fine screening of hydropower stations. 3) Substation characteristics: The substation is located close to the highway, the access road is short, and there are a certain number of transmission towers within a certain range of the substation. The fine screening of substations must first use the spatial logical relationship between substations and transmission towers. Secondly, substations occupy a large area and are mostly regular rectangular blocks in remote sensing images. The coarse screening results of substations are used for image classification, and the image classification will again use the deep neural convolutional network; 4) Characteristics of thermal power stations: In the coarse screening stage of thermal power stations, condensing towers are used as key features for identifying thermal power stations. There are a certain number of transmission towers around thermal power stations, and substations are set up within a certain distance. Combined with the domain knowledge of thermal power stations, in the fine screening stage of thermal power station detection, the spatial and quantitative relationships between thermal power stations and transmission towers and substations are used.

8. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Feature extraction based on knowledge of power grid domain: In addition to the geometric information, color information, texture information, and context information features of power infrastructure, for different power infrastructure, its unique ground features are extracted for precise screening of target recognition. Specifically, for hydroelectric power stations, a method based on deep learning image segmentation is used to extract the scene features of its upstream and downstream, to determine whether there is water surface in its upstream and downstream environment, and to assist in the judgment by combining the electric towers around the hydroelectric power station; for substations, image classification is performed on the results of rough screening of targets, and combined with its spatial relationship with surrounding transmission towers; for thermal power stations, attention is paid to its unique condensing tower, as well as its spatial logical relationship with surrounding transmission towers and substations; Based on the unique domain knowledge of various types of power infrastructure, a classifier suitable for precise identification and screening of power infrastructure targets is established. Based on the domain knowledge of various types of power infrastructure, the final target precise screening score is calculated, and the threshold for judging the score of each type of target is set to 0.8; 1) Precision screening of hydropower station identification results The precise screening of target identification of hydropower stations requires the combination of target identification results of transmission towers and water surface extraction results. The water surface is extracted using an image semantic segmentation method based on deep learning, specifically a multi-hole fully convolutional neural network. When using deep learning for image segmentation tasks, two bottlenecks are encountered. One is the information loss caused by downsampling, which is solved by the atrous convolution method. The other problem is that the edges are not accurate enough due to spatial invariance, which is solved by fully connected CRF; The target identification and fine screening process of hydropower station is as follows: First, taking the target rough screening result of the hydropower station as the center, calculate the distance L from all transmission towers to the center of the hydropower station. The center coordinates of the hydropower station are (X0, Y0), and the center coordinates of the transmission tower are (X1, Y1). The distance calculation formula is formula 1: According to prior knowledge, when L is less than 1 km, the number of transmission towers within the hydropower station increases by 1. The number of towers is represented by N. The impact value generated by the number of towers within a certain range of the hydropower station is S(A N ), S(A N ) is calculated as in Formula 2: When there is no transmission tower within the range of L, then S(A N ) is 0; when the number of transmission towers N is greater than 0 and less than 10, S(A N ) is 0.5; when the number of transmission towers N is greater than 10, S(A N ) is set to 1, and the calculation formula is as follows: In equations 3 and 4, the score of the hydroelectric power station is S(B), and the impact of the water surface on the hydroelectric power station is S(W). When there is no water environment in the image, S(W) = 0, and when there is a water surface in the image, s(W) = 1, S(B, A N ) represents the score of a target being determined as a hydroelectric power station under the influence of the number of towers, and S(B,W) represents the score of a target being determined as a hydroelectric power station under the influence of water surface conditions; In summary, the final score calculation formula for a possible target of a hydropower station is as follows: S(B,A N ,W) represents the final score of the hydropower station fine screening, S(B,A N ,W) is larger, the more likely it is that the target is a hydropower station; 2) Fine screening of substation identification results First, the spatial logical relationship between the substation and the transmission tower is used. Secondly, the substation occupies a large area. The rough screening results are used for image classification. The image classification uses the deep neural convolutional network again. This application uses the rough screened target box as the input image, uses deep learning to classify the candidate target box, uses GoogleNet as the classification network model, and uses the dense block structure to approximate the sparse structure, so as to achieve the effect of improving performance without increasing the computational complexity. The process of substation target identification fine screening is as follows: Based on the coarse screening results of substation target identification, a series of candidate target areas that may be substations are obtained, and the area corresponds to a score S(C1) that is judged as a substation. Then, the possible area is input into the GoogleNet model for image classification. The model will output the category and score of the possible area. If the category is judged as a substation, the output score is taken as S(C2). Finally, considering the number of towers around the substation, S(A N ) represents the impact of the towers within a certain range of the substation on the substation; First, taking the target rough screening result of the substation as the center, calculate the distance L between all transmission towers and the substation center. The coordinates of the substation center are (X0, Y0), and the coordinates of the transmission tower center are (X1, Y1). The distance calculation formula is Equation 6: According to prior knowledge, when L is less than 500 meters, the number of transmission towers within the hydropower station increases by 1, and the number of towers is represented by N. The impact value of the towers within a certain range of the substation on the substation is S (A N ),S(A N ) is calculated as in Formula 7: When there is no transmission tower within the range of L, then P(A N ) is 0; when the number of transmission towers N is greater than 0 and less than 10, S(A N ) is 0.5; when the number of transmission towers N is greater than 10, S(A N ) is taken as 0.5, and the calculation formula is as follows: In equations 8 and 9, S(C) represents the score of the target being judged as a substation. The larger S(C) is, the greater the possibility that the target is a substation. S(C1) is the score of the substation in the target rough screening stage, S(C2) is the score of the candidate target box classified as a substation, and the influence value of the tower on the substation is S(A N ), S(C,A N ) indicates the score at which a target is judged to be a substation if a certain number of transmission towers are present within the range.

9. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Establishing a precision screening dataset for power infrastructure detection: 1) Water surface extraction: The target fine screening of power infrastructure is based on the target coarse screening. The target fine screening data set is the detection result of the target coarse screening. In addition to the data used for the coarse screening of hydropower stations, thermal power stations and substations, the target fine screening data includes water surface segmentation data and substation data for scene classification; All samples are divided into training data and test data at a ratio of 4:1 for model training of water surface extraction; 2) Image classification: All samples are divided into training data and test data at a ratio of 4:

1. When collecting substation images, not only the consistency of the distribution of training data should be considered, but also substation images in all situations should be considered; The data information used for training the image classification model includes 8 categories including substations. In the preparation for training, the data is expanded, including rotation, mirroring, and translation, and the training data is normalized to the same size.

10. The method for accurately identifying smart grid infrastructure based on remote sensing images according to claim 1, characterized in that: Power infrastructure testing and screening process: 1) Transmission tower target recognition: The target recognition of transmission towers integrates the target recognition strategies of FasterR-CNN and ResNet. The model training for identifying transmission towers is implemented with the help of Tensorflow framework; According to the size of the transmission tower in the image, the target anchor size for training is set to 2, 4, 8, and 16. The training environment is the Linux system. The training process uses GPU. The initial learning rate is 0.001, the maximum number of training times is 80,000, and the learning rate decreases every 20,000 iterations. During the model training process, the model pre-trained on the ImageNet dataset is used to initialize the weights of deep learning, which can distribute the learned model parameters to the new model first, speeding up and optimizing the learning efficiency of the model; 2) Water surface extraction: The deep learning model used for water surface segmentation was trained with a learning rate parameter of 0.0001, a batch size parameter of 4, a momentum parameter of 0.9, a weight decay of 0.0005, a maximum number of iterations of 50,000, and fine-tuned on the pre-trained model; 3) Image classification: Based on the Caffe deep learning framework, the Googlenet network model is used for model training of image classification, where the learning rate is 0.

01. Different from the model used for target detection, the batch size of the training data and the batch size of the validation data are different during training, which are 64 and 32 respectively. In addition to substation images as positive samples for image classification, seven types of ground features, including farmland, water surface, desert, forest, oil tank, airplane, and residential area, are added as negative samples to participate in the training during the training process. In the classification results, only the classification results of substations are taken.

Citation Information

Patent Citations

  • Target Recognition of Thermal Power Station in Remote Sensing Image

    CN109344774A

  • Distribution line inspection data identification and analysis method based on image identification

    CN113536944A

  • Monocular visual odometer scale recovery method based on optimized angular point screening road surface points

    CN117557592A

Cited By

  • Agricultural product quality identification method and system based on multi-modal fusion and feature enhancement

    CN120599546A