Identification processing method for power grid terminal defects
Through the combined application of priority screening module, multi-scale feature priority processing module, ordinary coded information supplement module and multi-level token processing module, the calculation burden and stability problems of DETR algorithm in the identification of defects in power grid terminals is solved, and efficient defect detection is achieved.
Patent Information
- Application Number
- CN202510350524.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-08
AI Technical Summary
The existing DETR algorithms have heavy computational burdens in the identification of power grid terminal defects and have high dependence on query selection stability, resulting in unstable detection performance and redundant.
The priority screening module, multi-scale feature priority processing module, ordinary encoding information supplement module and multi-level token processing module are adopted to combine relative position embedding and absolute position embedding through priority value confidence screening query, and multiple single feature processing modules are used for multiple transformations and fusions to generate the coded features used by the decoder.
It realizes a good balance between computing efficiency and detection performance of the power grid terminal defect identification model, improves the accuracy and speed of detection, and reduces the cost and risk of manual inspection.
Smart Images

Figure CN120278969A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and particularly relates to a method for identifying and processing grid terminal defects. Background Art
[0002] Methods similar to DETR have significantly improved detection performance in an end-to-end manner. The mainstream two-stage frameworks of these algorithms perform intensive self-attention and select a part of the queries for sparse cross-attention, which has been proven effective in improving detection performance but also introduces a heavy computational burden and a high dependence on stable query selection. The two-stage selection strategy will cause scale bias and redundancy due to the mismatch between the queries selected in the two-stage initialization and the objects.
[0003] The identification of grid terminal defects using image recognition technology has significant advantages. Through efficient image processing algorithms, the power system can achieve automated and real-time monitoring, accurately identify various defects such as insulator cracks, broken wire strands, and equipment corrosion. This method not only improves the speed and accuracy of detection, reduces the cost and risk of manual inspection, but also can discover potential problems at an early stage, thereby improving the reliability and safety of the power system and ensuring the stability of power supply. Summary of the Invention
[0004] The present invention provides a method for identifying and processing grid terminal defects, aiming to propose a priority screening module, a multi-scale feature priority processing module, a general coding information supplement module, and a multi-level token processing module. The priority screening module screens queries through priority value confidence, retains the most discriminative queries, the multi-scale feature priority processing module further processes the screened queries to overcome scale bias, the general coding information supplement module fuses relative position embedding and absolute position embedding, combines with unselected queries, and increases the background information of unselected queries. The multi-level token processing module fuses adjacent hierarchical features, utilizes multiple single-feature processing modules, performs multiple transformations and fusions on features, and generates the coding features used by the decoder. The combined application of multiple modules ensures a good balance between the computational efficiency and detection performance of the grid terminal defect recognition model.
[0005] The present invention aims to propose a priority screening module, a multi-scale feature priority processing module, a general coding information supplement module, and a multi-level token processing module, and provides a method for identifying and processing grid terminal defects, including the following steps: S1. Production of the grid terminal defect dataset: Collect grid terminal defect pictures, including various defect types, and perform defect annotation on each grid terminal defect picture to form a grid terminal defect dataset; S2. Construct a multi-scale feature generation module, input the grid terminal defect pictures, and obtain three-scale features; S3. Construct a priority screening module, input three scale features, and retain the queries corresponding to the priority values with confidence levels higher than the priority value; S4. Construct a multi-scale feature priority processing module, perform attention operations on the priority queries, and keep other queries unchanged; S5. Construct an ordinary coding information supplement module, fuse relative embeddings and absolute embeddings, and combine them with the unselected queries; S6. Construct a multi-level token processing module, which includes multiple single-feature processing modules and outputs encoded features; S7. Construct a power grid terminal defect recognition model, including input, multi-scale feature generation module, priority screening module, multiple multi-scale feature priority processing modules, ordinary coding information supplement module, multi-level token processing module, decoder, and output; S8. Train and detect the power grid terminal defect recognition model. Use the power grid terminal defect dataset to train the power grid terminal defect recognition model. After training, input the new power grid terminal picture to be detected into the power grid terminal defect recognition model to obtain the detection result, and the detection result includes the defect location and type in the power grid terminal picture.
[0006] Preferably, in step S1, for multiple defect types, including insulator defects, conductor defects, cable defects, pole defects, and power equipment defects, insulator defects are divided into cracks, breakages, fouling, and discharge marks, conductor defects are divided into broken strands, abrasions, and corrosion, cable defects are divided into outer sheath breakages and aging, pole defects are divided into tilts and rust, and power equipment defects are divided into oil leakage of power equipment, shell breakages, and loose wiring.
[0007] Preferably, in step S2, for the multi-scale feature generation module, input the power grid terminal defect picture I, I ∈ R H ×W×3 , where H, W, and 3 represent the height, width, and channels of the power grid terminal defect picture I. Input the power grid terminal defect picture I into the backbone network to obtain three scale features F l , l ∈ {1, 2, 3}, H l 、W l and C represent the height, width, and channels of F l .
[0008] Preferably, in step S3, for the priority screening module, input the three scale features F l obtained from the multi-scale feature generation module, l ∈ {1, 2, 3}. For the query at position (i, j) in the l-scale feature corresponding to the position c in the power grid terminal defect picture I, c = (x, y), s l represents the downsampling scale for forming the l-scale feature, and the query The priority value is Query through the priority value confidence Filter, only retain the queries corresponding to the priority values higher than the priority value confidence The setting range of the priority value confidence is from 0.5 to 1, and the calculation method of the priority value is D Bbox represents the true bounding box, D Bbox ∈(x, y, w, h), where discau(c, D Bbox ) represents the relative distance between the query point c and the target center, Δx and Δy represent the distances between the query point and the target center, and w and h represent the width and height of the target respectively.
[0009] Preferably, in step S4, for the multi-scale feature priority processing module, for the l-th scale feature in the t-th multi-scale feature priority processing module, only those higher than v t w l The queries are called priority queries, and attention operations are performed on the priority queries, while other queries remain unchanged, where v t and w l are two screening confidences, 1 ≤ t ≤ T, 1 ≤ l ≤ T, T is the number of multi-scale feature priority processing modules, L is the number of scale features, here L is 3, and the specific implementation formula is where q i represents the i-th query, pos i represents the position encoding of the i-th query, Ω t is the set of priority queries in the t-th multi-scale feature priority processing module, q is the set of all queries in 3 scales, pos is the position encoding corresponding to the query, and SA represents the self-attention mechanism.
[0010] Preferably, in step S5, for the ordinary coding information supplement module, information is supplemented to the queries not selected by the multi-scale feature priority processing module to enhance their representation ability. Specifically, relative embedding and absolute embedding information are fused and combined with the unselected queries. The relative embedding formula is Interp represents the interpolation operation, r represents the row embedding, c represents the column embedding, represents the outer product of the row embedding and the column embedding, b l represents the background embedding, represents converting the embedding vector of the initial dimension n×n to the target dimension H through the interpolation operation l ×W l ; The absolute embedding formula is b (i,j)= Concat(r(i), c(j)), where Concat represents the concatenation operation, and r(i) and c(j) are the embedding vectors of the row and column respectively, b (i,j) represents the absolute background embedding; then fuse the relative embedding and the absolute embedding, b fused = αb l + (1 - α)b (i,j) , where α is a learnable weight parameter used to balance the importance of the relative embedding and the absolute embedding, b fused represents the fused relative embedding and absolute embedding; then combine b fused with the unselected query, q′ i = q i + b fused , q i represents the unselected query, and q′ i represents the combined query.
[0011] Preferably, in step S6, for the multi-level token processing module, first concatenate the tokens f l and f h of adjacent scale features, and generate the initial fused feature through a convolution operation That is UP represents the upsampling operation, Concat represents the concatenation operation, Conv represents the convolution operation. Input the initial fused feature into the main branch and the sub-branch. In the main branch, use multiple cascaded single-feature processing modules for feature processing. The number of single-feature processing modules is N. The processing process of the single-feature processing module is as follows, GC represents group convolution, β is a learnable weight parameter, + represents element-wise addition, and R represents the ReLU activation function, is the output feature of the nth single-feature processing module, 1 ≤ n ≤ N, and serves as the input feature of the (n + 1)th single-feature processing module. The output feature of the Nth single-feature processing module is In the sub-branch, according to obtain FC represents the fully connected layer, and R represents the ReLU activation function; finally, and are added element-wise to obtain f O , f O is the encoded feature, f OIt is the output of the multi-level token processing module. Since the multi-scale feature generation module outputs three-scale features, the three-scale features input to the multi-level token processing module are called low-scale features, medium-scale features, and high-scale features here. Then, the low-scale features and medium-scale features will be input to the multi-level token processing module to obtain medium-scale encoded features, and the medium-scale features and high-scale features will be input to the multi-level token processing module to obtain high-scale encoded features, that is, 2 multi-level token processing modules are used. Finally, the low-scale features, medium-scale encoded features, and high-scale encoded features will be input to the subsequent decoder.
[0012] Preferably, in step S7, for the power grid terminal defect recognition model, input the power grid terminal picture into the multi-scale feature generation module to generate three different-scale features, and input them into the priority screening module. Query screening is performed through the priority value confidence level among the three scales, and then the screened queries are input into multiple cascaded multi-scale feature priority processing modules. The priority queries are screened multiple times and self-attention calculation is performed. Then, in the ordinary coding information supplement module, the fused relative embedding and absolute embedding information are combined with the unselected queries, and then the low-scale features, medium-scale encoded features, and high-scale encoded features are obtained through 2 multi-level token processing modules and input into the decoder, thereby outputting the detection results. The detection results include the defect positions and types in the power grid terminal picture.
[0013] Compared with the prior art, the present invention has the following technical effects: The technical solution provided by the present invention proposes a priority screening module, a multi-scale feature priority processing module, an ordinary coding information supplement module, and a multi-level token processing module. The priority screening module screens queries through the priority value confidence level and retains the most discriminative queries. The multi-scale feature priority processing module further processes the screened queries to overcome scale deviation. The ordinary coding information supplement module fuses relative position embedding and absolute position embedding, combines them with the unselected queries, and increases the background information of the unselected queries. The multi-level token processing module fuses adjacent-level features, uses multiple single-feature processing modules, performs multiple transformations and fusions on the features, and generates the encoded features used by the decoder. The combined application of multiple modules ensures a good balance between the computational efficiency and detection performance of the power grid terminal defect recognition model. Brief Description of the Drawings
[0014] Figure 1 It is the flowchart of power grid terminal defect recognition provided by the present invention.
[0015] Figure 2 It is the structural diagram of the multi-scale feature generation module provided by the present invention.
[0016] Figure 3 It is the structural diagram of the multi-level token processing module provided by the present invention.
[0017] Figure 4 It is the structure diagram of the power grid terminal defect recognition model provided by the present invention.
[0018] Figure 5 It is the pole inclination recognition diagram of the power grid terminal defect provided by the present invention. Specific implementation manners
[0019] The present invention aims to propose a recognition and processing method for power grid terminal defects, and proposes a priority screening module, a multi-scale feature priority processing module, an ordinary coding information supplement module and a multi-level token processing module. The priority screening module screens and queries through the priority value confidence level, and retains the most discriminative queries. The multi-scale feature priority processing module further processes the screened queries to overcome the scale deviation. The ordinary coding information supplement module fuses the relative position embedding and the absolute position embedding, combines with the unselected queries, and increases the background information of the unselected queries. The multi-level token processing module fuses the features of adjacent levels, uses multiple single-feature processing modules, performs multiple transformations and fusions on the features, and generates the coding features used by the decoder. The combined application of multiple modules ensures a good balance between the computational efficiency and the detection performance of the power grid terminal defect recognition model.
[0020] Please refer to Figure 1 As shown, a recognition and processing method for power grid terminal defects in an embodiment of the present application: S1. Produce a power grid terminal defect data set, collect 500 pictures of power grid terminal defects, including various defect types, and perform defect annotation on each picture of power grid terminal defects to form a power grid terminal defect data set; S2. Construct a multi-scale feature generation module, input a single picture of power grid terminal defects, and obtain three-scale features; S3. Construct a priority screening module, input three-scale features, and retain the queries corresponding to the priority values higher than the priority value confidence level; S4. Construct a multi-scale feature priority processing module, perform an attention operation on the priority queries, and keep other queries unchanged; S5. Construct an ordinary coding information supplement module, fuse the relative embedding and the absolute embedding, and combine with the unselected queries; S6. Construct a multi-level token processing module, including 5 single-feature processing modules, and output coding features; S7. Construct a power grid terminal defect recognition model, including an input, 1 multi-scale feature generation module, 1 priority screening module, 7 multi-scale feature priority processing modules, 1 ordinary coding information supplement module, 2 multi-level token processing modules, a DETR decoder and an output; S8. Train the power grid terminal defect recognition model and perform detection. Use the power grid terminal defect dataset to train the power grid terminal defect recognition model. After training, input the new power grid terminal image to be detected into the power grid terminal defect recognition model to obtain the detection result, which includes the defect location and type in the power grid terminal image.
[0021] Further, in step S1, for multiple defect types, including insulator defects, conductor defects, cable defects, pole defects, and power equipment defects, insulator defects are divided into cracks, breakages, dirt, and discharge marks, conductor defects are divided into broken strands, abrasions, and corrosion, cable defects are divided into outer sheath breakages and aging, pole defects are divided into tilts and rusts, and power equipment defects are divided into power equipment oil leakage, shell breakages, and loose connections.
[0022] Further, in step S2, for the multi-scale feature generation module, as Figure 2 shown, input the power grid terminal defect image I, I ∈ R 640×640×3 , 640, 640, and 3 represent the height, width, and channels of the power grid terminal defect image I. Input the power grid terminal defect image I into the backbone network to obtain three-scale features F l , l ∈ {1, 2, 3}, H l , W l , and C represent the height, width, and channels of F l .
[0023] Further, in step S3, for the priority screening module, input the three-scale features F l obtained from the multi-scale feature generation module, l ∈ {1, 2, 3}. For the query of the position (i, j) in the l-scale feature corresponding to the position c in the power grid terminal defect image I, c = (x, y), s l represents the downsampling scale for forming the l-scale feature. The priority value of the query is Perform query filtering through the priority value confidence. Only retain the queries corresponding to the priority values higher than the priority value confidence. The priority value confidence setting interval is 0.75, and the priority value calculation method is DB represents the true bounding box, D box ∈ (x, y, w, h), where discau(c, DB Bbox ) represents the relative distance between the query point c and the target center. box Δx and Δy represent the distances between the query point and the target center, and w and h represent the width and height of the target respectively.
[0024] Further, in step S4, for the multi-scale feature prioritization module, for the l-th scale feature in the t-th multi-scale feature prioritization module, only queries higher than v t w l are called prioritized queries. The prioritized queries are subjected to an attention operation, while other queries remain unchanged, where v t and w l are two screening confidence levels, 1 ≤ t ≤ T, 1 ≤ l ≤ L, T is the number of multi-scale feature prioritization modules, here T is 7, and L is the number of scale features, here L is 3. The specific implementation formula is where q i represents the i-th query, pos i represents the position encoding of the i-th query, Ω t is the set of prioritized queries in the t-th multi-scale feature prioritization module, q is the set of all queries in the 3 scales, pos is the position encoding corresponding to the query, and SA represents the self-attention mechanism.
[0025] Further, in step S5, for the ordinary encoding information supplement module, information is supplemented to the queries not selected by the multi-scale feature prioritization module to enhance their representation ability. Specifically, relative embedding and absolute embedding information are fused and combined with the unselected queries. The relative embedding formula is Interp represents the interpolation operation, r represents the row embedding, c represents the column embedding, represents the outer product of the row embedding and the column embedding, b l represents the background embedding, represents converting the embedding vector of the initial dimension n×n to the target dimension H l ×W l ; the absolute embedding formula is b (i,j) = Concat(r(i), c(j)), Concat represents the concatenation operation, r(i) and c(j) are the row and column embedding vectors respectively, and b (i,j) represents the absolute background embedding; then the relative embedding and the absolute embedding are fused, b fused = αb l + (1 - α)b (i,j) , α is a learnable weight parameter used to balance the importance of relative embedding and absolute embedding, and b fused represents the fused relative embedding and absolute embedding; then b fused is combined with the unselected queries, q′ i = q i + b fused , q i represents the unselected queries, and q′ i represents the combined queries.
[0026] Further, in step S6, for the multi-level token processing module, its structure is as Figure 3 shown. First, the tokens f l and f h of adjacent scale features are concatenated and an initial fusion feature is generated through a convolution operation That is UP represents the upsampling operation, Concat represents the concatenation operation, Conv represents the convolution operation. The initial fusion feature is input into the main branch and the sub-branch. In the main branch, 5 cascaded single-feature processing modules are used for feature processing. The processing process of the single-feature processing module is GC represents group convolution, β is a learnable weight parameter, + represents element-wise addition, and the + in Figure 3 corresponds to the inside, R represents the ReLU activation function, is the output feature of the nth single-feature processing module, 1 ≤ n ≤ 5, and serves as the input feature of the (n + 1)th single-feature processing module. The output feature of the 5th single-feature processing module is In the sub-branch, according to obtain FC represents the fully connected layer, R represents the ReLU activation function; finally, and are element-wise added to obtain f O f O is the encoded feature, f O is the output of the multi-level token processing module. Since the multi-scale feature generation module outputs 3 scale features, the three scale features input into the multi-level token processing module here are called the low-scale feature, the middle-scale feature, and the high-scale feature. Then the low-scale feature and the middle-scale feature will be input into the multi-level token processing module to obtain the middle-scale encoded feature, and the middle-scale feature and the high-scale feature will be input into the multi-level token processing module to obtain the high-scale encoded feature. That is, 2 multi-level token processing modules are used. Finally, the low-scale feature, the middle-scale encoded feature, and the high-scale encoded feature will be input into the subsequent decoder.
[0027] Further, in step S7, for the power grid terminal defect recognition model, its structure is as Figure 4As shown, the input power grid terminal picture is fed into the multi-scale feature generation module to generate features of three different scales, which are then input into the priority screening module. Queries are screened through priority value confidence among the three scales, and then the screened queries are input into seven cascaded multi-scale feature priority processing modules. Priority queries are screened multiple times and self-attention calculation is performed. Then, in the ordinary coding information supplement module, the fused relative embedding and absolute embedding information are combined with the unselected queries. Subsequently, low-scale features, medium-scale coding features, and high-scale coding features are obtained through two multi-level token processing modules and input into the DETR decoder, thereby outputting the detection results, which include the defect positions and types in the power grid terminal picture.
[0028] Furthermore, an NVIDIA RTX 3090 GPU (24GB) is used, and the power grid terminal defect recognition model is trained using the AdamW optimizer with a weight decay of 0.0001 and an initial learning rate of 0.00001. The batch size for each GPU is set to 2.
[0029] Furthermore, the recognition effect of pole inclination in power grid terminal defects is as Figure 5 shown. Figure 5 The height and width of [the picture] are 640 and 640. It can be seen that the power grid terminal defect recognition model can well locate the pole inclination defect positions and accurately classify them.
[0030] The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the creative concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A method for identifying and processing defects in power grid terminals, characterized in that, It includes the following steps: S1. Production of the power grid terminal defect dataset. Collect power grid terminal defect pictures, which contain multiple defect types. Perform defect annotation on each power grid terminal defect picture to form the power grid terminal defect dataset; S2. Construct a multi-scale feature generation module. Input the power grid terminal defect pictures to obtain three-scale features; S3. Construct a priority screening module. Input the three-scale features and retain the queries corresponding to the priority values with a confidence level higher than the priority value; S4. Construct a multi-scale feature priority processing module. Perform attention operations on the priority queries and keep other queries unchanged; S5. Construct an ordinary coding information supplement module. Integrate relative embedding and absolute embedding and combine them with the unselected queries; S6. Construct a multi-level token processing module, which contains multiple single-feature processing modules and outputs encoded features; S7. Construct a power grid terminal defect recognition model, including input, multi-scale feature generation module, priority screening module, multiple multi-scale feature priority processing modules, ordinary coding information supplement module, multi-level token processing module, decoder and output; S8. Train and detect the power grid terminal defect recognition model. Use the power grid terminal defect dataset to train the power grid terminal defect recognition model. After training, input the new power grid terminal pictures to be detected into the power grid terminal defect recognition model to obtain the detection results, and the detection results include the defect positions and types in the power grid terminal pictures.
2. The identification and processing method for grid terminal defects according to claim 1, characterized in that, In step S1, for multiple defect types, including insulator defects, conductor defects, cable defects, pole defects, and power equipment defects, insulator defects are divided into cracks, breakages, dirt, and discharge traces, conductor defects are divided into broken strands, abrasions, and corrosion, cable defects are divided into outer skin breakages and aging, pole defects are divided into tilts and rusts, and power equipment defects are divided into power equipment oil leakage, shell breakages, and loose wiring.
3. The identification and processing method for grid terminal defects according to claim 1, wherein In step S2, for the multi-scale feature generation module, the input is the power grid terminal defect image I, where I ∈ R H×W×3 , and H, W, and 3 represent the height, width, and channels of the power grid terminal defect image I. The power grid terminal defect image I is input into the backbone network to obtain three scale features F l , where l ∈ {1, 2, 3}, H l , W l , and C represent the height, width, and channels of F l .
4. A method for identifying and processing defects in a power grid terminal according to claim 1, characterized in that, In step S3, for the priority screening module, the three scale features F obtained from the multi-scale feature generation module are input l , l∈{1,2,3}, for the query at position (i,j) in the l-scale feature Corresponding to the position c in the grid terminal defect image I, c = (x, y), s l Represents the downsampling scale to form l-scale features, query The priority value is Query by priority value confidence Filter to keep only queries corresponding to priority values with higher confidence than the priority value The priority value confidence setting interval is 0.5 to 1, and the priority value is calculated as follows: D Bbox represents the true bounding box, D Bbox ∈(x,y,w,h), where discau(c,D Bbox ) represents the relative distance between the query point c and the target center, Δx and Δy represent the distances between the query point and the center of the object, and w and h represent the width and height of the object, respectively.
5. The identification and processing method for grid terminal defects according to claim 1, characterized in that In step S4, for the multi-scale feature prioritization module, for the l-th scale feature in the t-th multi-scale feature prioritization module, only queries higher than v t w l are called prioritized queries. The prioritized queries are subjected to an attention operation, and other queries remain unchanged, where v t and w l are two screening confidence levels, 1 ≤ t ≤ T, 1 ≤ l ≤ L, T is the number of multi-scale feature prioritization modules, L is the number of scale features, here L is 3, and the specific implementation formula is where q i represents the i-th query, pos i represents the position encoding of the i-th query, Ω t is the set of prioritized queries in the t-th multi-scale feature prioritization module, q is the set of all queries in the 3 scales, pos is the position encoding corresponding to the query, and SA represents the self-attention mechanism.
6. A method for identifying and processing defects in a power grid terminal according to claim 1, characterized in that In step S5, for the general coding information supplement module, information is supplemented for the queries not selected by the multi-scale feature prioritization module to enhance their representation ability. Specifically, the relative embedding and absolute embedding information are fused and combined with the unselected queries. The relative embedding formula is Interp represents the interpolation operation, r represents the row embedding, and c represents the column embedding. represents the outer product of the row embedding and the column embedding, and b l represents the background embedding, and Interp (n,n) →(H l , W l ) represents converting the embedding vector of the initial dimension n×n to the target dimension H through the interpolation operation l ×W l ; The absolute embedding formula is b (i,j) = Concat(r(i), c(j)), where Concat represents the concatenation operation, and r(i) and c(j) are the embedding vectors of the row and column respectively, and b (i,j) represents the absolute background embedding; Then fuse the relative embedding and the absolute embedding, b fused = αb l + (1 - α)b (i,j) , where α is a learnable weight parameter used to balance the importance of the relative embedding and the absolute embedding, and b fused represents the fused relative embedding and absolute embedding; then combine b fused with the unselected query, q' i = q i + b fused , where q i represents the unselected query and q' i represents the combined query.
7. A method for identifying and processing defects in a power grid terminal according to claim 1, characterized in that, In step S6, for the multi-level token processing module, first, the tokens f l and f h of adjacent scale features are concatenated and an initial fusion feature is generated through a convolution operation That is UP represents the upsampling operation, Concat represents the concatenation operation, Conv represents the convolution operation. The initial fusion feature is input into the main branch and the sub-branch. In the main branch, multiple cascaded single-feature processing modules are used for feature processing. The number of single-feature processing modules is N. The processing process of the single-feature processing module is GC represents group convolution, β is a learnable weight parameter, + represents element-wise addition, and R represents the ReLU activation function is the output feature of the nth single-feature processing module, 1 ≤ n ≤ N, and serves as the input feature of the (n + 1)th single-feature processing module. The output feature of the Nth single-feature processing module is In the sub-branch, according to obtain FC represents the fully connected layer, and R represents the ReLU activation function; finally, and are element-wise added to obtain f O f O is the encoded feature, f O is the output of the multi-level token processing module. Since the multi-scale feature generation module outputs three scale features, the three scale features input into the multi-level token processing module are called the low-scale feature, the middle-scale feature, and the high-scale feature here. Then the low-scale feature and the middle-scale feature will be input into the multi-level token processing module to obtain the middle-scale encoded feature, and the middle-scale feature and the high-scale feature will be input into the multi-level token processing module to obtain the high-scale encoded feature. That is, 2 multi-level token processing modules are used. Finally, the low-scale feature, the middle-scale encoded feature, and the high-scale encoded feature will be input into the subsequent decoder 8. A method for identifying and processing defects in a power grid terminal according to claim 1, characterized in that, In step S7, for the power grid terminal defect recognition model, input the power grid terminal pictures into the multi-scale feature generation module to generate three different-scale features, and input them into the priority screening module. Perform query screening through the priority value confidence level at the three scales, and then input the screened queries into multiple cascaded multi-scale feature priority processing modules. Screen out the priority queries multiple times and perform self-attention calculations. Then, in the ordinary coding information supplement module, combine the integrated relative embedding and absolute embedding information with the unselected queries, and then obtain low-scale features, medium-scale encoded features, and high-scale encoded features through two multi-level token processing modules and input them into the decoder, so as to output the detection results, and the detection results include the defect positions and types in the power grid terminal pictures.