Power transmission line defect target tracking method and system

By introducing attention mechanism and dynamic template update mechanism into the dual-branch twin network, combined with the online migration strategy, the identification and tracking problems of transmission line defect targets in complex scenarios are solved, and the goal tracking effect with high accuracy and adaptability is achieved.

CN120147236APending Publication Date: 2025-06-13STATE GRID JIANGSU ELECTRIC POWER CO LTD NANJING POWER SUPPLY COMPANY +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510186506.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and stably track transmission line defect targets in complex scenarios, especially under conditions such as occlusion, light fluctuations and background noise.

Method used

A dual-branch twin network with improved attention mechanism is adopted to dynamically update template images, use target confidence to improve template update methods, and use online migration strategies to iterate the model parameters, so as to achieve accurate identification and stable tracking of small goals.

Benefits of technology

Effectively deal with occlusion, lighting fluctuations and background noise in complex scenarios, achieve stable, continuous tracking and accurate identification of dynamic targets, and improve the accuracy and adaptability of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147236A_ABST
    Figure CN120147236A_ABST
Patent Text Reader

Abstract

A power transmission line defect target tracking method is characterized by comprising the following steps: acquiring a continuous image sequence of a power transmission line, and sending a reference frame image and a to-be-detected frame image into a double-branch twin network; constructing an attention branch in a feature extraction module of the double-branch twin network, wherein a channel attention module, a space attention module, a time attention module and a template updating module are sequentially connected in series on the attention branch; extracting target features in the reference frame image by using the feature extraction module, and updating a target detection template according to the target features; and identifying the defect target of the power transmission line from the to-be-detected frame by using a target detection module. According to the method, variable scenes such as shielding, illumination fluctuation and background noise are effectively processed, and stable and continuous tracking and accurate recognition of a dynamic target are completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems, and more specifically, to a method and system for tracking defective targets on transmission lines. Background Art

[0002] As an important part of the power system, the safe and stable operation of transmission lines is crucial for ensuring the national economy and social life. However, affected by natural environment and human factors, transmission lines often face various potential safety hazards, such as broken strands of conductors, pollution of insulators, deformation of tower frames, etc. Therefore, it is necessary to conduct regular inspections and maintenance on transmission lines to ensure the reliability and safety of power supply.

[0003] Traditional inspections of transmission lines mostly rely on manual work, which is not only inefficient but also poses a great safety risk. With the development of technologies, the application of technologies such as unmanned aerial vehicles, machine vision, and artificial intelligence provides new solutions for the detection and maintenance of transmission lines. By using an unmanned aerial vehicle equipped with a high-definition camera to conduct aerial photography of transmission lines and combining image processing and deep learning technologies, automatic detection and identification of defective targets on transmission lines can be achieved. This method not only improves the efficiency and accuracy of detection but also greatly reduces the safety risk of personnel.

[0004] The technologies for detecting and tracking defective targets on transmission lines mainly involve knowledge in fields such as image processing, pattern recognition, and machine learning. By preprocessing the collected image data, extracting key features, and then using a trained deep learning model to analyze these features, accurate identification and positioning of various defective targets on transmission lines can be achieved. In addition, tracking technologies can dynamically monitor the detected defective targets, update their status in real time, and provide a basis for maintenance decisions. The combined use of these technologies makes the inspection work of transmission lines more intelligent and automated, greatly improving the management efficiency and response speed of the power system. With the development of technologies such as big data and cloud computing, the technologies for detecting and tracking defective targets on transmission lines are also constantly advancing. By using big data analysis, historical inspection data can be deeply mined to identify potential fault patterns and risk points, and then preventive measures can be taken in advance to avoid accidents. The cloud computing platform provides powerful data storage and computing capabilities, enabling complex image processing and analysis tasks to be efficiently completed in the cloud, greatly improving the speed and scale of data processing. However, defective targets on transmission lines are usually collected by high-speed moving line inspection equipment. Static defects on transmission lines are difficult to collect, and dynamic defects are difficult to track. So far, there is still no method for accurately identifying, stably tracking, and continuously following defective targets.

[0005] In view of the above problems, there is an urgent need for a method and system for tracking defective targets on transmission lines. Summary of the Invention

[0006] To address the deficiencies in the existing technologies, the present invention provides a method and system for tracking defective targets in transmission lines. The dual-branch Siamese network improved by an attention mechanism is used to dynamically update the template image. The template update method is improved using the target confidence, and an online migration strategy is adopted to iteratively improve the model parameters, thereby achieving accurate recognition of small targets.

[0007] The present invention adopts the following technical solutions.

[0008] In a first aspect of the present invention, there is provided a method for tracking defective targets in transmission lines, the method comprising the following steps: collecting a continuous image sequence of a transmission line, and sending a reference frame image and a frame image to be detected into a dual-branch Siamese network; constructing an attention branch in the feature extraction module of the dual-branch Siamese network, and sequentially connecting a channel attention module, a spatial attention module, a temporal attention module, and a template update module on the attention branch; using the feature extraction module to extract the target features in the reference frame image, and updating the target detection template according to the target features; and using the target detection module to identify the defective targets in the transmission line from the frame to be detected.

[0009] Preferably, collecting a continuous image sequence of a transmission line and sending a reference frame image and a frame image to be detected into a dual-branch Siamese network includes: the dual-branch Siamese network includes a feature extraction module, a target detection module, and a target extraction module; wherein, both the feature extraction module and the target detection module are dual-branch networks constructed by an appearance branch and a semantic branch; and in the semantic branch of the feature extraction module, there is also an attention branch.

[0010] Preferably, constructing an attention branch in the feature extraction module of the dual-branch Siamese network includes: inputting the initial feature map φ(z) of the reference frame image z extracted by the semantic branch into the channel attention module to obtain a first weighted feature map φ′(z); inputting the first weighted feature map φ′(z) into the spatial attention module to obtain a second weighted feature map φ″(z); and inputting the second weighted feature map φ″(z) into the temporal attention module to obtain a third weighted feature map φ″′(z).

[0011] Preferably, using the feature extraction module to extract the target features in the reference frame image and updating the target detection template according to the target features includes: inputting the third weighted feature map φ″′(z) into the module update module to dynamically update the current template image φ template (z - 1) to dynamically update the current template image φ template (z).

[0012] Preferably, using the template image φ template (z - 1) obtained last time to dynamically update the current template image φtemplate (z), including: the current template image is:

[0013] φ template (z) = φ template (z - 1)+(1 - λ)φ″′(z)

[0014] where z - 1 is the previous frame image of the reference frame image z,

[0015] λ is the template update weight, which is adjusted according to the target confidence score ConfidenceScore(φ″′(z), x).

[0016] Preferably, using the target detection module, the transmission line defect target is identified from the frame to be detected, including: using similarity calculation, the updated current template image φ template (z) and the feature map φ(x) output by the target detection module are subjected to similarity analysis to obtain a similarity response map.

[0017] Preferably, the training method of the double - branch Siamese network is: constructing the loss function of the double - branch Siamese network by using the weighted sum of the classification loss, regression loss, and attention weight loss; adopting the online transfer learning strategy, and updating the pre - selected training parameters of the double - branch Siamese network to:

[0018]

[0019] where is the training parameter vector of the double - branch Siamese network, is the nth parameter trained at time t,

[0020] η is the learning rate, which is a constant,

[0021] L(θ 1 , θ 2 , …, θ n ) = λ 1 L cls +λ 2 L reg +λ 3 L att is the loss function, where L cls represents the classification loss, L reg represents the regression loss, L att represents the attention weight loss, λ 1, λ 2 , λ 3 are the weight coefficients;

[0022] is the function L(θ 1 , θ 2 , …, θn ) the gradient in the direction of any parameter θ n ;

[0023] Calculate the loss function of the double-branch Siamese network after parameter update. When the tracking accuracy is met, end the training.

[0024] In the second aspect of the present invention, there is provided a transmission line defect target tracking system using the method in the first aspect of the present invention; the system includes an acquisition module, a construction module, an update module, and an identification module; wherein, the acquisition module is configured to acquire a continuous image sequence of the transmission line and send the reference frame image and the image to be detected into the double-branch Siamese network; the construction module is configured to construct an attention branch in the feature extraction module of the double-branch Siamese network, and a channel attention module, a spatial attention module, a temporal attention module, and a template update module are sequentially connected in series on the attention branch; the update module is configured to extract the target feature in the reference frame image by using the feature extraction module and update the target detection template according to the target feature; the identification module is configured to identify the transmission line defect target from the image to be detected by using the target detection module.

[0025] In the third aspect of the present invention, there is provided a terminal, including a processor and a storage medium; the storage medium is used for storing instructions; the processor is configured to operate according to the instructions to execute the steps of the method in the first aspect of the present invention.

[0026] In the fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the method in the first aspect of the present invention are implemented.

[0027] The beneficial effects of the present invention are as follows. Compared with the prior art, in a transmission line defect target tracking method and system in the present invention, the template image is dynamically updated by a double-branch Siamese network improved by an attention mechanism, the template update method is improved by using the target confidence, and an online migration strategy is adopted to iteratively improve the model parameters, so as to complete the accurate recognition of small targets. The present invention effectively processes variable scenes such as occlusion, light fluctuation, and background noise, and completes the stable, continuous tracking and accurate recognition of dynamic targets.

[0028] The beneficial effects of the present invention further include:

[0029] 1. The present invention addresses the challenges in the field of video object tracking, such as object occlusion, morphological changes, or background interference, and proposes a novel tracking method. This method utilizes an enhanced dual-branch Siamese network structure, significantly improving the recognition and discrimination ability of object features. In addition, a specific loss function is used to train the deep network model to achieve high-accuracy tracking and positioning of objects within a video sequence. The method of the present invention is not only applicable to the detection of transmission line defects but can also be extended to object tracking tasks in other fields, such as UAV tracking, autonomous driving vehicles, etc.

[0030] 2. The present invention integrates a temporal attention mechanism and a dynamic object template update mechanism. This mechanism utilizes time series prediction and feature degradation suppression strategies and provides an effective fusion of the first-frame image features of the video sequence with high-confidence features in subsequent frames, thereby achieving precise object recognition, stable tracking, and continuous update. The dynamic template update strategy can merge the features of the starting frame of the video with high-confidence features in subsequent frames, effectively reducing the possibility of object position offset. The introduced continuous adaptive learning mechanism enables the model to continuously calibrate the temporal attention model, ensuring the accuracy of long-term tracking and the adaptability to newly emerging object types.

[0031] 3. The temporal attention mechanism introduces advanced time series prediction techniques and feature degradation suppression algorithms to significantly improve the model's prediction ability in time and its ability to handle rapid dynamic changes. An improved convolutional long short-term memory network (ConvLSTM) is applied to analyze the time series features of the object. ConvLSTM can process time series data with spatial correlation and can be used to learn the motion pattern of the object between consecutive frames and predict its state in future frames. Combining with the optical flow analysis technique in image processing, the optical flow algorithm Farneback is applied to estimate the motion of the object. The pixel-level motion between consecutive frames is calculated to generate an optical flow field. The optical flow field not only provides information about the motion speed and direction of the object but can also be used to assist feature matching and data augmentation. Through enhanced data augmentation techniques, the model can be exposed to more diverse data during the training phase, thereby improving its generalization ability and its ability to handle unknown situations.

[0032] 4. The attention mechanism model for channels is integrated into the dual-branch network to assign weights to the key feature channels, which is used to improve the model's recognition ability for the key channel features. The spatial attention mechanism model is added to assign weights to the key spatial positions of the image, enhancing the perception ability for the key spatial positions of the image. The channel attention, spatial attention, and temporal attention mechanisms assign corresponding weights to different feature channels and spatial regions, and concentrate on processing the information crucial for tracking. This process continues until the end of the video sequence. Tests on the OTB2013 and OTB2015 standard data sets show that compared with the existing leading technologies, the present invention has achieved a 7.6% improvement in performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic flowchart of a method for tracking defective targets on a transmission line according to the present invention;

[0034] Figure 2 It is a schematic diagram of the module structure of a system for tracking defective targets on a transmission line according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0035] To make the objectives, technical solutions and advantages of the present invention clearer and more correct, the technical solutions of the present invention will be described in detail below through multiple specific embodiments. It should be noted that the embodiments adopted by the present invention are only used to explain the present invention and do not limit the content of the present invention.

[0036] In the first aspect of the present invention, it relates to a method for tracking defective targets on a transmission line, and the method includes the following steps 1 to 4.

[0037] Step 1, collect a continuous image sequence of the transmission line, and send the reference frame image and the image to be detected into the dual-branch Siamese network.

[0038] Preferably, collecting a continuous image sequence of the transmission line and sending the reference frame image and the image to be detected into the dual-branch Siamese network includes: the dual-branch Siamese network includes a feature extraction module, a target detection module, and a target extraction module; among them, both the feature extraction module and the target detection module are dual-branch networks constructed by an appearance branch and a semantic branch; in the semantic branch of the feature extraction module, there is also an attention branch.

[0039] The present invention processes the transmission line video, which consists of n frames of images, in the order of the first frame, the second frame... the nth frame. In the feature extraction module, the first frame image of the video is processed as the reference template for tracking. In the target detection module, the Tth frame image waiting to be tracked is processed. The first frame image is marked as z, and the Tth frame image is marked as x. In addition, a network φ for feature extraction is integrated for feature recognition and importance weighting in the channel, spatial, and temporal dimensions.

[0040] In one embodiment, first, by performing high-frequency component analysis on the depth features of the target in the video sequence, a high-pass filter is used to process the depth features of the target to extract the edge and contour information of the target, so as to identify and suppress feature degradation caused by rapid movement or emergencies.

[0041] Pairs of images z, such as the first frame of the video and x, are taken, and the current frame of the target being tracked is used as the input. In one embodiment, both the feature extraction module and the target detection module adopt the ResNet50 architecture, which includes 49 convolutional layers and 1 fully connected layer. These convolutional layers are distributed in four main residual modules, and each residual module is composed of several residual units containing 3 convolutional layers. The sliding window parameters of the convolution in the network are diverse. The first convolutional layer uses a 7×7 convolutional kernel and sets the stride to 2, aiming to obtain preliminary image features and reduce the spatial dimension of the feature map. In the subsequent residual modules, most 3×3 convolutional kernels use a stride of 1 to maintain the size of the feature map, while some convolutional layers for spatial downsampling use a 2×2 sliding window and set the stride to 2. With such a feature extraction design, the convolutional layers of ResNet50 extract multi-scale features from residual modules at different depths, and combine the adaptive feature fusion technology to optimize the performance of the model input features, laying a solid foundation for defect detection and tracking in the complex background of transmission lines.

[0042] Step 2, construct an attention branch in the feature extraction module of the dual-branch Siamese network, and sequentially connect a channel attention module, a spatial attention module, a temporal attention module, and a template update module on the attention branch.

[0043] Constructing an attention branch in the feature extraction module of the dual-branch Siamese network includes: inputting the initial feature map φ(z) of the reference frame image z extracted by the semantic branch into the channel attention module to obtain the first weighted feature map φ′(z); inputting the first weighted feature map φ′(z) into the spatial attention module to obtain the second weighted feature map φ″(z); inputting the second weighted feature map φ″(z) into the temporal attention module to obtain the third weighted feature map φ″′(z).

[0044] The initial feature map is the access position of the attention branch on the feature extraction module. Multi-scale feature extraction is performed through a specific network, and the initial feature map is obtained by weighting with an enhanced triple attention module. Subsequent frame images also undergo feature extraction and calculate the similarity response map with the dynamically updated first frame feature map to determine the most accurate position of the target in consecutive frames.

[0045] In the channel attention mechanism, the input feature map φ(z) is globally average pooled to obtain p(z) = GAP(φ(z)), which is then fed into a fully connected layer to adjust the feature dimension f(z) = FC(p(z)), where FC represents the fully connected layer. The Sigmoid function generates the channel weights α = sigmoid(f(z)). Finally, the weighted feature map φ′(z) = φ(z) ⊙ α is generated, where ⊙ represents element-wise multiplication.

[0046] In one embodiment, the features of the image are processed using Global Average Pooling to obtain a reduced feature vector of dimension 1×1×C. The feature dimension is adjusted via a series of fully connected layers and corresponding activation functions, including dimension increase and decrease. This operation not only increases the non-linear fitting ability of the model but also reduces the parameters and computational amount of the model, while retaining the 1×1×C dimension of the feature vector. The Sigmoid function is used to generate the specific weight α for each channel, and the feature values of each channel are multiplied by the corresponding weight α to weight the input image features φ(z) and highlight the crucial feature channels.

[0047] In the spatial attention mechanism, the convolutional layer generates the spatial weights s(z) = Conv(φ′(z)). The Softmax function normalizes the spatial weights β = softmax(s(z)). Finally, the weighted feature map φ″(z) = φ′(z) ⊙ β is generated.

[0048] In spatial attention, the first frame image of the video is input into the spatial attention mechanism network to obtain the weight values at each position in the feature map. These weighted first frame features are compared with the current tracking frame to achieve template matching for identifying and locating the tracking target. The specific position of the tracking target is determined based on the highest response value obtained during the matching process.

[0049] In the temporal attention mechanism, the state of the prediction target in future frames is predicted timeseriesprediction(φ″(z), x). The output of the ConvLSTM and the optical flow information are input into the temporal attention module, and the temporal attention weight γ = TimeAttention(φ″(z), x) is calculated through a fully connected layer and an activation function. The temporal attention weight γ is subjected to an element-wise multiplication operation with the depth feature map φ″(z) to generate the weighted feature map φ″′(z) = φ″(z) ⊙ γ.

[0050] In one embodiment, the feature map generated by the target detection module is input into the temporal attention module, and the optical flow matrix between the reference frame image and the image to be detected is calculated. Assuming that the image gradient is constant and the local optical flow is constant, between the two frames of images, the neighborhood of the feature points is searched to construct an objective function to solve the position movement of the target. The objective function e(i) is:

[0051]

[0052] In the formula, i is any pixel point in image z and image x, and A and b are image parameters;

[0053] Δi ∈ I is an adjacent point on the neighborhood I of pixel i, and d is the moving distance between image z and image x.

[0054] For a two-dimensional grayscale image, the grayscale value of a pixel point is a two-dimensional variable function f(i x , i y ). Then, with the feature point as the center, a local coordinate system is constructed, and the image is binomially expanded, then:

[0055] f(i x , i y ) = i T Ai + b T i + c

[0056] A is a 2×2 symmetric matrix, b is a 2×1 matrix, and i is a two-dimensional column vector.

[0057] Then the moving mode of image z to image x is:

[0058]

[0059] In the formula, A 1 = A 2 , and b 1 and b 2 are respectively the parameters of the grayscale functions of image z and image x.

[0060] The feature point can be obtained by calculating the region of interest. In the region of interest, a point with continuous grayscale with other surrounding pixel points is a feature point. Usually, the neighborhood of the feature point is a square region centered on the feature point with a size of 2n + 1.

[0061] In another embodiment, in the temporal attention mechanism, the Farneback algorithm can also be used to calculate the optical flow fields I(i x , i y, t), obtain the target motion information, and in the same way, input the optical flow field and depth features in consecutive images into the ConvLSTM network. Then, through the ConvLSTM network, learn the temporal dynamics of the target and predict the state of the target in future frames timeseriesprediction(φ″(x - 1), x). Input the output of the ConvLSTM and the optical flow information into the temporal attention module, and calculate the temporal attention weight γ = TimeAttention(φ″(z), x) through a fully connected layer and an activation function. Perform an element-wise multiplication operation on the temporal attention weight γ and the depth feature map φ″(z) to generate a weighted feature map φ″′(z) = φ″(z) ⊙ γ.

[0062] In the ConvLSTM network, the optical flow field not only provides information about the motion speed and direction of the target, but can also be used to assist in feature matching and data augmentation. Through enhanced data augmentation techniques, the model can be exposed to more diverse data during the training phase, thereby improving its generalization ability and the ability to handle unknown situations.

[0063] In the above process, increase the diversity of training samples through techniques such as illumination change simulation, motion blur simulation, and background change simulation, and improve the generalization ability of the model.

[0064] In addition, the method also supports using deep learning models (such as CNN) to extract the feature representation of the target and combine the optical flow information to predict the moving trend of the target. When the target temporarily disappears or is occluded, the model can predict the possible position of the target through the learned historical information.

[0065] Step 3, use the feature extraction module to extract the target features in the reference frame image, and update the target detection template according to the target features.

[0066] Preferably, using the feature extraction module to extract the target features in the reference frame image and update the target detection template according to the target features includes: input the third weighted feature map φ″′(z) into the module update module to dynamically update the current template image φ template (z - 1) to dynamically update the current template image φ template (z).

[0067] The present invention introduces a dynamic template update strategy based on the object state confidence. Calculate the target confidence score: c = ConfidenceScore(φ″′(z), x). The confidence score can be calculated in ways such as similarity, IOU, etc.

[0068] By analyzing the confidence score and appearance change of the tracked target in real time, dynamically update the tracking template. Use the previously obtained template image φ template (z - 1) to dynamically update the current template image φtemplate (z), including: the current template image is:

[0069] φ template (z) = φ template (z - 1)+(1 - λ)φ″′(z)

[0070] where z - 1 is the previous frame image of the reference frame image z

[0071] λ is the template update weight, which is adjusted according to the target confidence score ConfidenceScore(φ″′(z), x). λ is dynamically adjusted according to the confidence score c to ensure the timeliness and accuracy of the tracking model.

[0072] Thus, the present invention supports the use of an online learning algorithm to update the model weight by means of stochastic gradient descent to ensure the accuracy of long - term tracking and the adaptability to newly emerging target types.

[0073] Step 4, use the target detection module to identify the transmission line defect target from the frame to be detected.

[0074] Using similarity calculation, perform similarity analysis on the updated current template image φ template (z) and the feature map φ(x) output by the target detection module to obtain a similarity response map.

[0075] Through similarity analysis, the position of the target in the current frame can be determined, which helps to maintain the stability of tracking in complex situations such as when the target appearance changes or is occluded. The result of similarity analysis can be used to evaluate the update effect of the template image. If the similarity analysis shows a decline in tracking performance, the template update strategy can be adjusted to change the update frequency or weight. The result of similarity analysis can also be used to evaluate the confidence of the target. High similarity indicates that the target is successfully tracked, while low similarity may indicate tracking failure or target loss. Similarity analysis provides a feedback mechanism that can be used to adjust and optimize the tracking algorithm. By analyzing the change in similarity, problems in the tracking process can be identified and corresponding adjustments can be made.

[0076] The updated current template image and the features extracted from the current frame are subjected to a convolution operation to achieve feature matching and generate a similarity response map. This process is expressed as:

[0077] r(x) = ConvMatch(φ″′(z), φ(x))

[0078] The training method of the double - branch Siamese network is:

[0079] Construct the loss function of the double - branch Siamese network using the weighted sum of the classification loss, regression loss, and attention weight loss;

[0080] The Classification Loss is used to measure the performance of the model in the classification task, that is, the difference between the predicted class by the model and the true class. The Classification Loss can help the model distinguish the target and the background, or identify different target types. The Regression Loss is used to measure the performance of the model in the regression task, that is, the difference between the continuous value predicted by the model and the true value, and is usually used to predict the position or size of the target. The above two loss functions can be calculated in the common way.

[0081] The Attention Weight Loss is used to measure the performance of the model in attention allocation, that is, the difference between the attention weights allocated by the model and the ideal weights. The Attention Weight Loss can help the model learn how to correctly focus on the target and ignore the background.

[0082] In one embodiment, the loss function is:

[0083] L = λ 1 L cls + λ 2 L reg + λ 3 L att

[0084] where, L cls represents the Classification Loss, L reg represents the Regression Loss, L att represents the Attention Weight Loss, λ 1, λ 2 , λ 3 are weight coefficients.

[0085] Adopt the online transfer learning strategy to update the training parameters of the preselected dual-branch Siamese network to:

[0086]

[0087] In the formula, is the training parameter vector of the dual-branch Siamese network, is the nth parameter trained at time t,

[0088] η is the learning rate, which is a constant,

[0089] L = λ 1 L cls + λ 2 L reg + λ 3 L att is the loss function, where, L cls represents the Classification Loss, L reg represents the Regression Loss, Latt Denote the attention weight loss as λ 1, λ 2 , λ 3 is the weight coefficient

[0090] is the gradient of the function L(θ 1 , θ 2 , …, θ n ) in the direction of any parameter θ n ;

[0091] Calculate the loss function of the double-branch Siamese network after parameter update. When the tracking accuracy is met, end the training.

[0092] The classification loss is used to measure the performance of the model in the classification task, that is, the difference between the category predicted by the model and the true category. The classification loss can help the model distinguish the target and the background, or identify different target types.

[0093] The regression loss is used to measure the performance of the model in the regression task, that is, the difference between the continuous value predicted by the model and the true value, and is usually used to predict the position or size of the target.

[0094] The attention weight loss is used to measure the performance of the model in attention allocation, that is, the difference between the attention weight assigned by the model and the ideal weight. The attention weight loss can help the model learn how to correctly focus on the target and ignore the background.

[0095] Through this process, fine-tune the network parameters to achieve more stable tracking accuracy. By continuously iteratively optimizing the loss function, a fine-tuned network model is finally formed and ready for actual target tracking tasks.

[0096] The present invention uses online learning to extract sequence image features in the whole tracking process, and updates the model parameters by using the transfer learning mechanism, so that the model can self-calibrate during actual deployment and operation, adapt to new defect types and environmental changes, and ensure the adaptive ability and long-term stability of the system. Use the trained deep model to track and locate the target in the video sequence. Even under variable and complex conditions, this method can maintain stable tracking of the target, avoid position drift, and ensure the stability and accuracy of the tracking result.

[0097] In a second aspect of the present invention, there is provided a transmission line defect target tracking system. The system is implemented by using the method according to the first aspect of the present invention, and includes an acquisition module, a construction module, an update module, and an identification module. The acquisition module is configured to acquire a continuous image sequence of the transmission line and send the reference frame image and the frame to be detected into a dual-branch Siamese network. The construction module is configured to construct an attention branch in the feature extraction module of the dual-branch Siamese network, and a channel attention module, a spatial attention module, a temporal attention module, and a template update module are sequentially connected in series on the attention branch. The update module is configured to extract the target features in the reference frame image by using the feature extraction module and update the target detection template according to the target features. The identification module is configured to identify the transmission line defect target from the frame to be detected by using the target detection module.

[0098] In a third aspect of the present invention, there is provided a terminal, including a processor and a storage medium. The storage medium is configured to store instructions. The processor is configured to operate according to the instructions to execute the steps of the method according to the first aspect of the present invention.

[0099] In a fourth aspect of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored. The computer program, when executed by a processor, implements the steps of the method according to the first aspect of the present invention.

[0100] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that there are still modifications or equivalent replacements that can be made to the specific embodiments of the present invention within the technical solutions of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for tracking a defect target of a transmission line, characterized in that: The method comprises the following steps: Collect continuous image sequences of the transmission line, and send the reference frame image and the frame image to be detected into the double-branch twin network; An attention branch is constructed in the feature extraction module of the dual-branch Siamese network, and a channel attention module, a spatial attention module, a temporal attention module and a template update module are sequentially connected in series on the attention branch; Utilizing the feature extraction module to extract target features in the reference frame image, and updating the target detection template according to the target features; The target detection module is used to identify the transmission line defect targets from the frames to be detected.

2. A method for tracking transmission line defects according to claim 1, characterized in that: The method of collecting a continuous image sequence of the power transmission line and sending the reference frame image and the frame image to be detected into a double-branch twin network includes: The dual-branch twin network includes a feature extraction module, a target detection module and a target extraction module; wherein, The feature extraction module and the target detection module are both double-branch networks constructed by an appearance branch and a semantic branch; The semantic branch in the feature extraction module also includes an attention branch.

3. A method for tracking a power transmission line defect target according to claim 2, characterized in that: The constructing of an attention branch in the feature extraction module of the dual-branch Siamese network comprises: The semantic branch extracts the initial feature map φ(z) of the reference frame image z and inputs it into the channel attention module to obtain the first weighted feature map φ′(z); Input the first weighted feature map φ′(z) into the spatial attention module to obtain the second weighted feature map φ″(z); The second weighted feature map φ″(z) is input into the temporal attention module to obtain the third weighted feature map φ″′(z).

4. A method for tracking a power transmission line defect target according to claim 3, characterized in that: The step of extracting target features from the reference frame image using the feature extraction module and updating the target detection template according to the target features includes: The third weighted feature map φ″′(z) is input into the module update module to utilize the template image φ obtained last time template (z-1) Dynamically update the current template image φ template (z).

5. A method for tracking a transmission line defect target according to claim 4, characterized in that: The template image φ obtained last time is used template (z-1) Dynamically update the current template image φ template (z) including: The current template image is: φ template (z)=φ template (z-1)+(1-λ)φ″′(z) Wherein, z-1 is the previous frame image of the reference frame image z, λ is the template update weight, which is adjusted with the target confidence score CpmfodemceScore(φ″′(z),x).

6. A method for tracking a transmission line defect target according to claim 5, characterized in that: The method of using the target detection module to identify the power transmission line defect target from the frame to be detected includes: Using similarity calculation, the updated current template image φ template (z) performing similarity analysis with the feature map φ(x) output by the target detection module to obtain a similarity response map.

7. A method for tracking a power transmission line defect target according to claim 6, characterized in that: The training method of the dual-branch twin network is: The loss function of the two-branch twin network is constructed by using the weighted sum of classification loss, regression loss, and attention weight loss; Using the online transfer learning strategy, the training parameters of the pre-selected two-branch twin network are updated to: In the formula, is the training parameter vector of the two-branch twin network, is the nth parameter trained at time t, η is the learning rate, which is a constant. L(θ1,θ2,…,θ n )=λ1L cls +λ2L reg +λ3L att is the loss function, where L cls represents the classification loss, L reg represents the regression loss, L att represents the attention weight loss, λ 1, λ2,λ3 are weight coefficients; is the function L(θ1,θ2,…,θ n ) for any parameter θ n Directional gradient; The loss function of the two-branch twin network after parameter update is calculated, and the training is terminated when the tracking accuracy is met.

8. A transmission line defect target tracking system using the method according to any one of claims 1 to 7; characterized in that: The system includes a collection module, a construction module, an update module and an identification module; wherein, The acquisition module is used to acquire a continuous image sequence of the power transmission line, and send the reference frame image and the frame image to be detected into the double-branch twin network; The construction module is used to construct an attention branch in the feature extraction module of the dual-branch Siamese network, and the channel attention module, the spatial attention module, the temporal attention module and the template update module are sequentially connected in series on the attention branch; The updating module is used to extract the target features in the reference frame image by using the feature extraction module, and update the target detection template according to the target features; The identification module is used to identify the power transmission line defect target from the frame to be detected by using the target detection module.

9. A terminal comprising a processor and a storage medium; characterized in that: The storage medium is used to store instructions; The processor is configured to operate according to the instructions to execute the steps of the method according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Deep learning-based intelligent control method and system for explosion-proof safety cabinet

    CN120335312A

  • Method and system for on-line detection and marking of cable surface flaws based on machine vision

    CN122694261A