Power transmission line pin defect identification algorithm
By using the improved Faster R-CNN algorithm and YOLOv5 deep learning network, combined with FPN and cascade detection models, the problem of low accuracy in identifying pin defects in transmission lines is solved, achieving efficient and accurate pin defect detection and saving computing resources.
Patent Information
- Application Number
- CN202510832927.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies have the problem of low recognition accuracy in the identification of pin defects in transmission lines. Traditional methods have poor robustness, the end-to-end model consumes human resources and accumulates recognition errors, and ignoring regional information affects the final accuracy.
An improved Faster R-CNN algorithm is adopted, combined with the Feature Pyramid Network (FPN) and the YOLOv5 deep learning network. Through data preprocessing, backbone network, neck network, region proposal network and ROI head network, prior information is integrated to build a cascade detection model, extract the context area information where the pin is located, reduce background interference, and improve recognition accuracy.
Without increasing data annotation, the accuracy and efficiency of pin defect recognition are improved, computing resources are saved, and the prior information of the pin installation position is fully utilized to achieve accurate positioning and recognition of the pin.
Smart Images

Figure CN120747737A_ABST
Abstract
Description
[0001] The present invention relates to the field of power transmission line pin defect detection, and in particular to a power transmission line pin defect recognition algorithm, which solves the problem of low accuracy in power transmission line pin defect recognition in current methods. Background Art
[0002] Pins, as fasteners, connect and secure key components in power towers. Any defects can affect the safe and stable operation of the power system. Due to the small size of pins and the large number of hardware similar in shape to them, traditional target detection technologies are less effective in identifying defective pins in practical applications. Therefore, this paper identifies pin defects and proposes a pin defect cascade detection model based on Faster R-CNN and Feature Pyramid Network (FPN) that integrates prior information. Current methods for detecting pin defects in transmission lines include the following:
[0003] 1. Use traditional image processing operators for manual feature extraction. Commonly used feature extraction operators include Canny operator, SIFT (Scale Invariant Feature) operator, HOG (HIStogram of Oriented Gradients) operator, LBP (Local Binary Patterns) operator, etc.
[0004] 2. Use a non-end-to-end power pole tower pin defect detection model to detect transmission line pins.
[0005] 3. Use an end-to-end pin defect detection model to detect pin defects on transmission lines. The existing path planning method for substation oil sampling robot operations has the following shortcomings:
[0006] Existing methods for identifying pin defects in transmission lines have the following disadvantages:
[0007] 1. Using traditional image processing operators for manual feature extraction requires a lot of manual template design work, has poor robustness, is sensitive to the appearance and size changes of pins, and has great application limitations.
[0008] 2. Using a non-end-to-end power tower pin defect detection model to inspect transmission line pins reduces information loss during feature sampling, avoids interference from complex background information, and introduces prior information into the model. However, training the multi-stage model requires multiple rounds of data annotation, which consumes a significant amount of human resources. Furthermore, the cascaded model means that recognition errors in each stage are propagated to the next stage, affecting the model's ultimate recognition accuracy.
[0009] 3. An end-to-end pin defect detection model is used to detect pin defects in transmission lines. Residual units are used to construct different forms of feature pyramids to fuse high-level semantic information and low-level semantic information. This makes up for the problem of weak detection ability of small-sized pins due to the loss of sampling information. However, it ignores the regional information where the pins are located, which also helps to enhance the feature expression ability, and the ability to utilize information is not strong.
[0010] In view of the above shortcomings, the present invention proposes a transmission line pin defect recognition algorithm based on the Faster R-CNN algorithm to improve the accuracy and efficiency of transmission line pin defect recognition. Summary of the Invention
[0011] The purpose of this invention is to provide a power transmission line pin defect recognition algorithm that uses an improved Faster R-CNN algorithm to achieve higher recognition accuracy. Figure 1 shown.
[0012] To achieve the above object, the present invention provides the following solutions:
[0013] like Figure 2 As shown in the figure, a transmission line pin defect recognition algorithm uses the FasterR-CNN algorithm to achieve accurate recognition of transmission line pin defects, including the following steps:
[0014] 1) Data Preprocessing: Given the large size of the original aerial inspection images and the small size of the pins, and due to limited computing resources, the baseline model needs to downsample the original images to a smaller size if it directly uses the original aerial images for pin defect detection. This results in the loss of the already limited pin pixel information, making it difficult for the baseline model to extract effective features for pin defect recognition, resulting in extremely poor pin defect recognition performance. Therefore, the present invention uses a uniform cropping algorithm to crop the original aerial inspection images into multiple sub-images, which are then used to train the baseline model.
[0015] 2) Backbone network: The backbone network is the basic feature extractor of the target detection model, which is used to extract semantic features from the original image and output feature maps. The backbone network of this project is a convolutional neural network, which is mainly composed of a stack of convolution layers, pooling layers and activation functions. The backbone network of the present invention uses ResNet-50 pre-trained on the ImageNetM dataset. The network structure configuration of ResNet-50 consists of 5 stages. The first stage consists of a 7*7 convolution kernel with a stride of 2 plus a 3*3 maximum pooling layer with a stride of 2. The network structures of the second to fifth stages are basically the same, mainly composed of a stack of residual units. The residual units adopt the design ideas of identity mapping and shortcut connection, which can solve the problem of network performance degradation as the number of network layers deepens, and also help solve the problems of gradient disappearance and gradient explosion.
[0016] 3) Neck network: The neck network is also called the feature fusion network. Its main function is to fuse the original feature maps generated by the backbone network. In the original feature map, high-level features are downsampled many times, with a high degree of abstraction and rich high-level semantic information, but a lot of detailed information is lost. Low-level features are downsampled less times and have rich detailed information, such as edge contours, color, and texture. By fusing high-level semantic information with low-level detail information in a "top-down" or "bottom-up" manner, a feature map with richer semantic information can be obtained, which can improve the detection capability of the model. The present invention uses FPN as the neck network in the baseline model to improve the detection capability of small-sized pins. The fused feature maps P2-P6 are obtained through lateral connection and top-down operations.
[0017] 4) Region Proposal Network: The essence of the Region Proposal Network is a sliding window-based unclassified object detector. Its workflow is to first densely lay anchor boxes on the feature map to generate candidate regions, then determine whether the region contains an object and predict the coordinate offset;
[0018] 5) ROI Head Network: The ROI Head Network performs fine-grained classification and coordinate adjustment of the regions of interest (ROIs) selected by the region proposal network. First, the ROI Head Network extracts feature information of the ROI region from the corresponding feature map. ROI features are then transformed to a uniform size using ROIAlign. These features are then fed into two fully connected layers, FC1 and FC2, for feature transformation to obtain a 1024-dimensional feature vector. Finally, the feature vector passes through the classification and regression heads to obtain the predicted category and relative coordinate offset.
[0019] 6) Recognition result restoration: In order to restore the recognition result of the model on the sub-image to the original image, the absolute coordinates of the upper left corner of each sub-image in the original image are recorded during the cropping process of the original image, which is recorded as (x pic ,y pic ); The position of the prediction box in the sub-image is expressed as the coordinates of the upper left corner and the lower right corner (x l ,y l , x r ,y r ); then the coordinates of the predicted box in the sub-image after being restored to the original image can be expressed as (x pic +x l ,y pic +y l , x pic +x r ,y pic +y l ).
[0020] Furthermore, FPN is introduced as a feature fusion module; the Faster R-CNN baseline model usually uses a uniform cropping algorithm to preprocess the original image; however, a large number of sub-images will be generated, resulting in a waste of computing resources, and the prior information that pins are often installed in fixed positions is ignored; to address the above problems, the present invention designs a pin defect cascade detection model that integrates prior information and introduces FPN as a feature fusion module. This can extract the structural area where the pin is located, reduce the interference of complex background on the model, help the model better model contextual information, thereby improving the model's recognition effect on defective pins, and can reduce training and testing time and save computing resources.
[0021] Furthermore, a pin defect cascade detection model integrating prior information was constructed. The model mainly consists of two-level networks; the first-level network is the pin structure area extraction network, which takes the original aerial inspection image as input and uses YOLOv5 to locate and identify the pins in it. The pin recognition at this stage does not distinguish between normal pins and defective pins; then the first-level network uses the density space clustering-based pin structure area extraction algorithm on the YOLOv5 prediction box to extract the context area where the pins are located; the second-level network is the defective pin recognition network, and the network structure is the same as the baseline model, that is, it uses the FasterR-CNN+FPN structure; it is responsible for locating and accurately classifying pins in the pin structure area extracted by the first-level network.
[0022] Furthermore, we utilized the improved YOLOv5 deep learning network. YOLOv5 is a one-stage detection model that is an improvement on YOLOv3 and YOLOv4. The official code provides five models: YOLOv5n, YOLOv5s, YOLOv5m, YOLOv5l, and YOLOv5x, with increasing depth and width. Considering the role of YOLOv5 in the first-level network and the size of the original aerial inspection images, this paper selected the largest YOLOv5x as the pin localization model to ensure the first-level network's recall rate for pins.
[0023] YOLOv5 in the primary network and FasterR-CNN in the secondary network are trained separately. YOLOv5 is first trained using the original training set. The primary network then extracts the pin context structure regions from the original training set to form a sub-image training set. FasterR-CNN is then trained using the sub-image training set. However, the XML label files used for YOLOv5 training do not need to be annotated separately. A simple Python script can be used to convert "nail_good" and "nail_bad" in the original label files to "nail". There is no difference between the two except for the label category.
[0024] YOLOv5 only needs to identify whether the current sample is a pin or background, and does not need to perform fine-grained classification of pins. This design is intended to reduce the learning difficulty of YOLOv5 and ensure that YOLOv5 can achieve a good pin recall rate even when using original aerial inspection images as input.
[0025] Furthermore, the cascade model trains the secondary network by extracting the contextual information of the pins. The secondary network's pin structure region extraction algorithm based on density spatial clustering can be roughly divided into the following four steps:
[0026] Step 1: Get the center of the target frame; Represents the target frame set in the original aerial inspection image; where N represents the total number of target frames in an aerial inspection image, g i ={x1 i ,y1 i ,x2 i ,y2 i} represents the i-th target box, (x1 i ,y1 i ) and (x2 i ,y2 i ) represent the coordinates of the upper left corner and lower right corner of the target frame respectively; based on G, the coordinate set of the center point of the target frame can be calculated Where p i ={x1 i ,y1 i} represents the center point of a target box;
[0027] Step 2: Clustering; Use the DBSCAN algorithm to cluster the center point coordinate set P and obtain the cluster set Where M represents the number of clusters obtained after division, which is related to the distribution of the target box of each image, the ∈ neighborhood of the DBSCAN algorithm and the MinPts hyperparameter setting, and is not fixed. i is the set of target boxes contained in cluster i; the DBSCAN algorithm can divide the pins installed on the same hardware structure into a cluster, which is very helpful for the subsequent extraction of the context structure area where the pins are located;
[0028] Step 3: Calculate cluster centers; calculate the cluster center set based on the coordinates of the target box center point in each cluster Where cp i =(cx i ,cy i ) represents the coordinates of the center point of cluster i, cx i Indicates the horizontal coordinate cy of the center point irepresents the vertical coordinate of the center point; the horizontal and vertical coordinates of the cluster center point are the mean of the horizontal and vertical coordinates of the center points of the target boxes in the cluster, as shown in formulas (1) and (2), where n represents the number of target boxes contained in the cluster, x i ,y i Respectively represent the horizontal and vertical coordinates of the i-th target box in the cluster;
[0029]
[0030] Step 4: Generate subgraph boundaries; define the subgraph boundary set as The i-th subgraph boundary b i ={cx i ,cy i ,w i ,h i}, it takes the center point of cluster i (cx i ,cy i ) as its own center point, and set the width w i High f i The definition is as follows:
[0031] w i =max(1024,(max{x2|g∈G i}-min{x1|g∈G i})) (3)
[0032] h i =max(1024,(max{y2|g∈G i}-min{y1|g∈G i})) (4)
[0033] The cascade model trains the secondary network by extracting the contextual area information where the pin is located. This not only reduces the interference of complex background on the model but also helps the model model prior context information and achieve accurate positioning of the pin.
[0034] The advantage of the present invention is that the method proposed in the present invention can alleviate the challenge of resolution and make full use of the prior information that the pins are installed in fixed positions without the need for additional data annotation, thereby improving the accuracy of pin recognition while saving computing power. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is the application scenario of the present invention, transmission line pin defect identification;
[0036] Figure 2 This is the baseline model network structure diagram of the Faster R-CNN+FPN mode provided in this application;
[0037] Figure 3This is a structural diagram of the pin defect cascade detection model that integrates prior information provided in this application;
[0038] Figure 4 Diagram of the improved YOLOv5 deep learning architecture provided for this application. DETAILED DESCRIPTION
[0039] The technical solutions in the embodiments of the present invention will be described in detail below with reference to the accompanying drawings and implementation examples.
[0040] The purpose of the present invention is to provide a transmission line pin defect recognition algorithm based on the Faster R-CNN algorithm, and a higher convergence speed is obtained by using an improved Faster R-CNN algorithm design.
[0041] The transmission line pin defect recognition algorithm includes the following steps:
[0042] 1) Data Preprocessing: Given the large size of raw aerial inspection images and the small size of pins, and limited computing resources, the baseline model, if used directly for pin defect detection, would need to downsample the original images to a smaller size. This results in the loss of information about the pin pixels, which is already limited, making it difficult for the baseline model to extract effective features for pin defect recognition, resulting in extremely poor pin defect recognition performance. Therefore, the present invention uses a uniform cropping algorithm to crop the raw aerial inspection images into multiple sub-images, which are then used to train the baseline model.
[0043] 2) Backbone network: The backbone network is the basic feature extractor of the target detection model, which is used to extract semantic features from the original image and output feature maps. The backbone network of this project is a convolutional neural network, which is mainly composed of a stack of convolution layers, pooling layers and activation functions. The backbone network of the present invention uses ResNet-50 pre-trained on the ImageNetM dataset. The network structure configuration of ResNet-50 consists of 5 stages. The first stage consists of a 7*7 convolution kernel with a step size of 2 plus a 3*3 maximum pooling layer with a step size of 2. The network structures of the second to fifth stages are basically the same, mainly composed of a stack of residual units. The residual unit adopts the design ideas of identity mapping and shortcut connection, which can solve the problem of network performance degradation as the number of network layers deepens, and at the same time helps to solve the problems of gradient disappearance and gradient explosion.
[0044] 3) Neck network (Neck): The neck network is also called the feature fusion network. Its main function is to perform feature fusion on the original feature map generated by the backbone network. In the original feature map, high-level features are downsampled many times, with a high degree of abstraction and rich high-level semantic information, but a lot of detail information is lost; while low-level features are downsampled few times and have rich detail information, such as edge contours, colors, textures, etc. By fusing high-level semantic information and low-level detail information using a "top-down" or "bottom-up" path, a feature map with richer semantic information can be obtained, which can improve the detection capability of the model. The present invention uses FPN as the neck network in the baseline model to improve the detection capability of small-sized pins. The fused feature maps P2-P6 are obtained through lateral connection and top-down operations.
[0045] 4) Region Proposal Network (RPN): The essence of the RPN is a sliding window-based category-free object detector. Its workflow is to first densely lay anchor boxes on the feature map to generate candidate regions, then determine whether the region contains the target and predict the coordinate offset.
[0046] 5) ROI Head Network: The ROI Head Network performs fine-grained classification and coordinate adjustment of the region of interest (ROI) selected by the region proposal network. First, the ROI Head Network extracts feature information of the ROI region from the corresponding feature map. Then, ROIAlign is used to transform the ROI features to a uniform size. These features are then fed into two fully connected layers (FC1 and FC2) for feature transformation, resulting in a 1024-dimensional feature vector. Finally, the feature vector passes through the classification and regression heads to obtain the predicted category and relative coordinate offset.
[0047] 6) Recognition result restoration: In order to restore the recognition result of the model on the sub-image to the original image, the absolute coordinates of the upper left corner of each sub-image in the original image are recorded during the cropping process of the original image, which is recorded as (x pic ,y pic The position of the prediction box in the sub-image is expressed as the coordinates of the upper left corner and the lower right corner (x l ,y l , x r ,y r ). Then the coordinates of the predicted box in the sub-image after being restored to the original image can be expressed as (x pic +x l ,y pic +y l , x pic +x r ,ypic +y l ).
[0048] It should be emphasized that the embodiments described in the present invention are illustrative rather than restrictive. Therefore, the present invention includes but is not limited to the embodiments described in the specific embodiments. Any other embodiments derived by those skilled in the art based on the technical solutions of the present invention also fall within the scope of protection of the present invention.
Claims
1. A transmission line pin defect recognition algorithm, using the FasterR-CNN algorithm, achieves accurate recognition of transmission line pin defects, including the following steps: 1) Data Preprocessing: Given the large size of the original aerial inspection images and the small size of the pins, and due to limited computing resources, the baseline model needs to downsample the original images to a smaller size if it directly uses the original aerial images for pin defect detection. This results in the loss of the already limited pin pixel information, making it difficult for the baseline model to extract effective features for pin defect recognition, resulting in extremely poor pin defect recognition performance. Therefore, the present invention uses a uniform cropping algorithm to crop the original aerial inspection images into multiple sub-images, which are then used to train the baseline model. 2) Backbone network: The backbone network is the basic feature extractor of the target detection model, which is used to extract semantic features from the original image and output feature maps. The backbone network of this project is a convolutional neural network, which is mainly composed of a stack of convolution layers, pooling layers and activation functions. The backbone network of the present invention uses ResNet-50 pre-trained on the ImageNetM dataset. The network structure configuration of ResNet-50 consists of 5 stages. The first stage consists of a 7*7 convolution kernel with a stride of 2 plus a 3*3 maximum pooling layer with a stride of 2. The network structures of the second to fifth stages are basically the same, mainly composed of a stack of residual units. The residual units adopt the design ideas of identity mapping and shortcut connection, which can solve the problem of network performance degradation as the number of network layers deepens, and also help solve the problems of gradient disappearance and gradient explosion. 3) Neck network: The neck network, also known as the feature fusion network, is mainly used to fuse the original feature maps generated by the backbone network. In the original feature map, high-level features are downsampled many times, with a high degree of abstraction and rich high-level semantic information, but a lot of detailed information is lost. In contrast, low-level features are downsampled less times and are rich in detailed information, such as edge contours, color, and texture. By fusing high-level semantic information with low-level detailed information in a "top-down" or "bottom-up" manner, a feature map with richer semantic information can be obtained, which can improve the detection capability of the model. The present invention uses FPN as the neck network in the baseline model to improve the detection capability of small-sized pins. The fused feature maps P2-P6 are obtained through lateral connection and top-down operations. 4) Region Proposal Network: The essence of the Region Proposal Network is a sliding window-based unclassified object detector. Its workflow is to first densely lay anchor boxes on the feature map to generate candidate regions, then determine whether the region contains an object and predict the coordinate offset; 5) ROI Head Network: The ROI Head Network performs fine-grained classification and coordinate adjustment of the regions of interest (ROIs) selected by the region proposal network. First, the ROI Head Network extracts feature information of the ROI region from the corresponding feature map. ROI features are then transformed to a uniform size using ROIAlign. These features are then fed into two fully connected layers, FC1 and FC2, for feature transformation to obtain a 1024-dimensional feature vector. Finally, the feature vector passes through the classification and regression heads to obtain the predicted category and relative coordinate offset. 6) Recognition result restoration: In order to restore the recognition result of the model on the sub-image to the original image, the absolute coordinates of the upper left corner of each sub-image in the original image are recorded during the cropping process of the original image, which is recorded as (x pic ,y pic ); The position of the prediction box in the sub-image is expressed as the coordinates of the upper left corner and the lower right corner (x l ,y l , x r ,y r ); then the coordinates of the predicted box in the sub-image after being restored to the original image can be expressed as (x pic +x l ,y pic +y l , x pic +x r ,y pic +y l ).
2. A power transmission line pin defect recognition algorithm according to claim 1, characterized in that ,FPN is introduced as the feature fusion module; The Faster R-CNN baseline model typically uses a uniform cropping algorithm to preprocess the original image; however, this generates a large number of sub-images, resulting in a waste of computing resources, and ignores the prior information that pins are often installed in fixed positions. To address the above issues, the present invention designs a pin defect cascade detection model that integrates prior information and introduces FPN as a feature fusion module. This can extract the structural area where the pin is located, reduce the interference of complex background on the model, and help the model better model contextual information, thereby improving the model's recognition of defective pins. It can also reduce training and testing time and save computing resources.
3. The power transmission line pin defect recognition algorithm according to claim 1 is characterized in that , a pin defect cascade detection model integrating prior information was constructed. The model mainly consists of two-level networks; the first-level network is the pin structure area extraction network, which takes the original aerial inspection image as input and uses YOLOv5 to locate and identify the pins in it. The pin recognition at this stage does not distinguish between normal pins and defective pins; then the first-level network uses the pin structure area extraction algorithm based on density space clustering to extract the context area where the pins are located on the prediction box of YOLOv5; the second-level network is the defective pin recognition network, and the network structure is the same as the baseline model, that is, the FasterR-CNN+FPN structure is used; it is responsible for locating and accurately classifying pins on the pin structure area extracted by the first-level network.
4. A power transmission line pin defect recognition algorithm according to claim 1, characterized in that ,Using the improved YOLOv5 deep learning network, YOLOv5 is a one-stage detection model, which is improved on the basis of YOLOv3 and YOLOv4; the official code provides five models: YOLOv5n, YOLOv5s, YOLOv5m, YOLOv5l and YOLOv5x, and their depth and width gradually increase; considering the role of YOLOv5 in the first-level network and the size of the original aerial inspection image, this paper selects the largest YOLOv5x as the pin positioning model to ensure the recall rate of the first-level network for pins; YOLOv5 in the primary network and FasterR-CNN in the secondary network are trained separately. YOLOv5 is first trained using the original training set. The primary network then extracts the pin context structure regions from the original training set to form a sub-image training set. FasterR-CNN is then trained using the sub-image training set. However, the XML label files used for YOLOv5 training do not need to be annotated separately. A simple Python script can be used to convert "nail_good" and "nail_bad" in the original label files to "nail". There is no difference between the two except for the label category. YOLOv5 only needs to identify whether the current sample is a pin or background, and does not need to perform fine-grained classification of pins. This design is intended to reduce the learning difficulty of YOLOv5 and ensure that YOLOv5 can achieve a good pin recall rate even when using original aerial inspection images as input.
5. The power transmission line pin defect recognition algorithm according to claim 1 is characterized in that ,The cascade model trains the secondary network by extracting the contextual area information where the pins are located. The secondary network's pin structure area extraction algorithm based on density space clustering can be roughly divided into the following four steps: Step 1: Get the center of the target frame; Represents the target frame set in the original aerial inspection image; where N represents the total number of target frames in an aerial inspection image, g i ={x1 i ,y1 i ,x2 i ,y2 i } represents the i-th target box, (x1 i ,y1 i ) and (x2 i ,y2 i ) represent the coordinates of the upper left corner and lower right corner of the target frame respectively; based on G, the coordinate set of the center point of the target frame can be calculated Where p i ={x1 i ,y1 i } represents the center point of a target box; Step 2: Clustering; Use the DBSCAN algorithm to cluster the center point coordinate set P and obtain the cluster set Where M represents the number of clusters obtained after division, which is related to the distribution of the target box of each image, the ∈ neighborhood of the DBSCAN algorithm and the MinPts hyperparameter setting, and is not fixed. i is the set of target boxes contained in cluster i; the DBSCAN algorithm can divide the pins installed on the same hardware structure into a cluster, which is very helpful for the subsequent extraction of the context structure area where the pins are located; Step 3: Calculate cluster centers; calculate the cluster center set based on the coordinates of the target box center point in each cluster Where cp i =(cx i ,cy i ) represents the coordinates of the center point of cluster i, cx i Indicates the horizontal coordinate cy of the center point i represents the vertical coordinate of the center point; the horizontal and vertical coordinates of the cluster center point are the mean of the horizontal and vertical coordinates of the center points of the target boxes in the cluster, as shown in formulas (1) and (2), where n represents the number of target boxes contained in the cluster, x i ,y i Respectively represent the horizontal and vertical coordinates of the i-th target box in the cluster; Step 4: Generate subgraph boundaries; define the subgraph boundary set as The i-th subgraph boundary b i ={cx i ,cy i ,w i ,h i }, it takes the center point of cluster i (cx i ,cy i ) as its own center point, and set the width w i High f i The definition is as follows: w i =max(1024,(max{x2|g∈G i }-min{x1|g∈G i })) (3) h i =max(1024,(max{y2|g∈G i }-min{y1|g∈G i })) (4) The cascade model trains the secondary network by extracting the contextual area information where the pin is located. This not only reduces the interference of complex background on the model but also helps the model model prior context information and achieve accurate positioning of the pin.