Robot inspection system and method based on task instruction
By introducing Transformer, Capsule Networks and Kalman Filter into the robot inspection system, combined with multi-source sensor data, the existing inspection system has solved the shortcomings in defect identification accuracy and adaptability, and achieved high-precision, strong robustness and flexible adaptation inspection results.
Patent Information
- Application Number
- CN202411802074.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing robot inspection system has shortcomings in defect identification accuracy and adaptability, and faces the problem of frequent false alarms and omissions, and is difficult to meet the high-frequency special inspection needs of core key equipment such as transformers.
A robot inspection system based on task instructions is adopted, data is collected using multi-source sensors, global shared features are extracted through Transformer, local structural features are extracted in combination with Capsule Networks, and continuous tracking and prediction of device status is achieved through Kalman Filter.
It improves the accuracy and reliability of defect identification, enhances the migration and generalization capabilities of the model, can effectively adapt to different inspection tasks, and meets the needs of high-frequency special inspections.
Smart Images

Figure CN119942278A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent inspection technology, and in particular to a robot inspection system and method based on task instructions. Background Art
[0002] As an important infrastructure for national economic and social development, the power system plays an irreplaceable role in the normal operation of the economy and society and the improvement of people's quality of life. Transmission lines and substations are the most critical components of the power system, and their safe and stable operation directly affects the reliability of the power grid. However, since power equipment is exposed to harsh environments such as high voltage, high temperature, and humidity for a long time, defects such as insulator flashover, tower tilt, and excessive line sag are very likely to occur, which seriously threaten the safety of the power system [1].
[0003] Traditional power inspection mainly relies on manual visual inspection of lines and stations on a regular basis, which has problems such as insufficient number of maintenance personnel, high labor intensity, and high missed inspection rate [2]. With the continuous increase in the mileage and voltage level of transmission lines, the manual inspection mode can no longer meet the growing demand for power grid operation and maintenance. To this end, various power companies have introduced intelligent means such as robots and drones to assist in inspection, aiming to reduce the manpower burden and improve the quality of work [3]. However, the existing robot inspection system still has deficiencies in defect recognition accuracy and adaptability, and faces the problem of frequent false alarms and missed alarms [4]. At the same time, the demand for high-frequency special inspections of core key equipment such as transformers has not been effectively met. Therefore, there is an urgent need for an intelligent perception method that is high-precision, robust, and can flexibly adapt to different inspection tasks.
[0004] In recent years, deep learning technology has made breakthrough progress in the field of computer vision, providing new ideas for the research and development of intelligent inspection systems.
[0005] Transformer is a deep learning model based on the self-attention mechanism. Compared with the classic recurrent neural network, Transformer can better capture more dependencies, and has the advantages of high computational efficiency and easy parallelization. Recently, researchers have tried to expand it to computer vision tasks. Vision Transformer has shown excellent performance in benchmark tests such as image classification and object detection, reflecting its powerful feature representation ability. Inspired by this, the present invention intends to introduce Transformer into intelligent inspection to extract global semantic features shared between different tasks.
[0006] Capsule Networks is a type of neural network structure that can encode the spatial position and posture of local features. Unlike scalar neurons, a capsule is a group of neurons in the form of vectors, and each capsule represents the instantiation parameters of a specific object or its parts. Through a dynamic routing algorithm, shallow capsules can activate high-level capsules from bottom to top, thereby modeling the hierarchical relationship between the whole and the local.
[0007] In the power system, the equipment state has obvious time continuity and stability. However, most of the existing visual detection methods use static modeling, ignoring the correlation information between the previous and next frames. Kalman Filter is a Bayesian filter that can recursively estimate the state of a dynamic system and is widely used in target tracking, navigation and positioning and other fields. The basic idea is to use the system model to predict the current state distribution, and then correct the predicted distribution in combination with the observed value to obtain the posterior state distribution. With the help of Kalman Filter, the present invention can realize the continuous evaluation of the current operating status and future development trends of power equipment, detect abnormal signs early and take preventive measures. Summary of the invention
[0008] In view of the problems existing in the existing inspection system, the present invention proposes a robot inspection system and method based on task instructions.
[0009] On the one hand, the present invention proposes a robot inspection method based on task instructions, and the method specifically comprises the following steps:
[0010] Step 1. Use multi-source sensors to collect image data of transmission line components and perform preprocessing.
[0011] Step 1.1 Use multi-source sensors such as visible light cameras, infrared thermal imagers, and ultraviolet imagers to collect image data of transmission line towers, insulators, conductors, and other components. raw , where X raw ={X visible , X infrared , X ultraviolet}. Among them, X visible Represents visible light image data; X infrared Represents infrared image data; X ultraviolet Represents UV image data.
[0012] Step 1.2: collect the original image data X raw Perform preprocessing, including image correction, noise removal, data enhancement, etc., to obtain the preprocessed image data X preprocessed , where X preprocessed =f preprocess (X raw ). Among them, fpreprocess (·) represents the preprocessing function.
[0013] Step 2. Aggregate the preprocessed multi-source heterogeneous image data into a unified data matrix.
[0014] The pre-processed visible light, infrared, and ultraviolet image data X preprocessed Converge into a unified data matrix X fused , where X fused =f fuse (X preprocessed ). Among them, f fuse (·) represents the data aggregation function.
[0015] Step 3. Use the Transformer self-attention mechanism to extract the global shared features of the image data.
[0016] Step 3.1 The aggregated data matrix X fused Divide the image into fixed-size blocks and flatten each block into a vector, denoted by x i .
[0017] Step 3.2: For the image block vector x i Perform linear transformation to obtain the query matrix Q, key matrix K and value matrix V:
[0018] Q=XW Q , K = XW K , V = XW V Formula 1;
[0019] Among them, W Q , W K , W V represents the learnable weight matrix; X represents the image block vector x i The matrix composed of.
[0020] Step 3.3 Calculate the self-attention matrix A:
[0021]
[0022] Among them, softmax(·) represents the softmax normalization function; d k Represents the dimension of the key matrix K.
[0023] Step 3.4 multiplies the self-attention matrix A by the value matrix V to get the output Z of the Transformer encoder:
[0024] Z=AV Formula 3;
[0025] Step 3.5: Embed the output Z of the Transformer encoder into F as a shared semantic featureglobal :
[0026] F global =Z Formula 4;
[0027] Step 4. Extract local structural features based on Capsule Network.
[0028] Step 4.1 Use the convolutional neural network CNN to extract the primary features of the image block and obtain the primary capsule u i .
[0029] Step 4.2 Aggregate primary capsules u through a dynamic routing algorithm i , get the high-level capsule v j :
[0030]
[0031] Among them, W ij represents the transformation matrix from the i-th primary capsule to the j-th high-level capsule; c ij Represents the coupling coefficient from the i-th primary capsule to the j-th high-level capsule.
[0032] Step 4.3 Repeat step 4.2 to get the top-level capsule {v1, v2, ..., v m}.
[0033] Step 4.4: Concatenate the top-level capsules to get the local structural features embedded in F local :
[0034] F local =Concat(v1, v2, …, v m ) Formula 6;
[0035] Among them, Concat(·) represents the concatenation operation.
[0036] Step 5. Fusion of global features extracted by Transformer and local features extracted by Capsule Network.
[0037] The global feature F extracted by Transformer global And the local feature F extracted by Capsule Network local Splicing to get the fusion feature F fused :
[0038] F fused =Concat(F global , F local ) Formula 7;
[0039] Step 6. Evaluate the current state of the device by combining the current feature estimate with the historical state information.
[0040] Step 6.1 Fusion feature F fused Input classifier f cls (·) (such as MLP), get the current state estimate y of the device t :
[0041] y t =f cls (F fused ) Formula 8;
[0042] Among them, y t Represents the estimated state of the device at time t.
[0043] Step 6.2 Estimate the current state y t Input Kalman Filter, combined with historical state estimation, to obtain the corrected state estimation
[0044]
[0045] in, represents the prediction of the state at time t using the state estimate at time t-1; K t represents Kalman gain; H t Represents the observation matrix.
[0046] Step 7. Dynamically adjust the model according to different task instructions and generate targeted inspection strategy reports.
[0047] According to different task instructions (such as daily inspections, special inspections, etc.), dynamically adjust data input, model weights and output forms to generate targeted inspection strategies and reports.
[0048] On the other hand, the present invention also proposes a robot inspection system based on task instructions. Figure 2 As shown in the figure, the system consists of the following key modules:
[0049] Data acquisition module: This module consists of multi-source sensors, including visible light cameras, infrared thermal imagers, and ultraviolet imagers. These sensors are installed on drones, robots, or handheld devices to collect image data of transmission line towers, insulators, conductors, and other components. The types of data collected include visible light images, infrared thermal images, and ultraviolet images, which reflect the appearance, temperature distribution, and corona discharge of the equipment, respectively. The data acquisition module also includes environmental sensors for collecting environmental parameters such as temperature, humidity, and wind speed, as well as GPS receivers for recording the geographic location information of the image.
[0050] Data preprocessing module: This module receives the raw image data from the data acquisition module and performs a series of preprocessing operations on it. The preprocessing steps include image correction, noise removal, and data enhancement. The image correction step eliminates image distortion and distortion and improves the geometric accuracy of the image through distortion correction and geometric calibration. The noise removal step uses image filtering algorithms such as median filtering and wavelet transform to suppress noise interference in the image and improve image quality. The data enhancement step expands the diversity of the training data set and improves the robustness of the model through operations such as image rotation, flipping, and scaling. The preprocessed image data will be passed to the feature extraction module.
[0051] Feature extraction module: This module contains two submodules: global feature extraction and local feature extraction. The global feature extraction submodule uses the Transformer self-attention mechanism to extract features with global semantic information from the entire image. Transformer calculates the attention weights between image blocks and builds long-range dependencies between blocks to obtain a global representation of the image. The local feature extraction submodule uses Capsule Network to extract local structural features of the image. Capsule Network uses a dynamic routing algorithm to adaptively aggregate low-level capsules to form high-level capsules, thereby modeling the hierarchical relationship between the whole and the local. The extraction of global features and local features is carried out in parallel, and finally the two features are fused to form a complete image representation.
[0052] State evaluation module: This module receives the fused features output by the feature extraction module and evaluates the current state of the equipment. State evaluation is divided into two stages: current state estimation and state timing optimization. In the current state estimation stage, the fused features are input into a classifier (such as a multilayer perceptron) to predict the current state of the equipment, such as normal, abnormal, or defects of different levels. In the state timing optimization stage, the current state estimation results are input into the Kalman Filter together with the historical state estimation results, and the current state is corrected through the state space model. The Kalman Filter recursively estimates the posterior probability distribution of the state, smoothes and optimizes the state estimation results in the time dimension, and improves the accuracy and stability of the evaluation.
[0053] Task planning module: This module dynamically adjusts the system's data input, model weights, and output forms based on the task instructions input by the user, and generates targeted inspection strategies and reports. Task instructions can be different types such as daily inspections and special inspections, and each task corresponds to different inspection priorities and requirements. The task planning module will selectively call the data acquisition module to collect specific types of data based on the task type, adjust the weights of Transformer and CapsuleNetwork in the feature extraction module, highlight task-related features, and customize the evaluation criteria and decision thresholds of the state evaluation module. The inspection results will be presented in different forms according to the task type, such as defect type, location, severity, etc., and a corresponding inspection report will be generated.
[0054] Human-computer interaction module: This module provides a graphical interface for users to interact with the system. Users can enter task instructions, select inspection areas and equipment types, and view inspection results and reports through this interface. The human-computer interaction module also provides data visualization functions, which intuitively display equipment status, defect distribution and other information in the form of charts, heat maps, etc., so that users can quickly understand the operation of the power system. In addition, the module also supports user feedback and annotation. Users can confirm, correct or supplement the inspection results. This feedback information can be used to optimize the model and improve system performance.
[0055] The present invention has the following beneficial technical effects:
[0056] (1) Transformer is introduced to uniformly model different inspection objects and task objectives, learn universal global semantic features, and enhance the model's migration and generalization capabilities.
[0057] (2) Capsule Networks are used to characterize the local structured information of key components, overcome the risk of missed detection caused by interference factors such as occlusion and deformation, and improve the accuracy and reliability of defect identification.
[0058] (3) Use Kalman Filter to continuously track and predict changes in the status of power equipment, understand the development laws of equipment health status, and meet the needs of high-frequency special inspections of key areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flow chart of a robot inspection method based on task instructions provided by the present invention.
[0060] Figure 2 A structure diagram of a robot inspection system based on task instructions provided by the present invention DETAILED DESCRIPTION
[0061] The method proposed by the present invention is further explained below in conjunction with specific embodiments. Figure 1 As shown in the figure, the overall idea of the method of the present invention is: first input the multi-source heterogeneous data of the equipment to be inspected, including visible light, infrared, ultraviolet images, etc.; then send it to Transformer to extract shared features with global semantics, and at the same time model local structured information through Capsule Networks; then cascade the two features, and give the current state evaluation through the defect discriminator; finally, fuse the historical state through Kalman Filter to generate the health profile prediction of the equipment. The whole process is driven by task instructions, and the data input, model components and output forms are flexibly adjusted according to different instruction types.
[0062] The input of the method of the present invention comes from a variety of sensors and data sources, mainly including:
[0063] (1) Visible light images: Use high-definition visible light cameras carried by drones or robots to photograph line components such as towers, insulators, and conductors, as well as station equipment such as transformers and reactors, to obtain visible light images that reflect their appearance and structure. Considering the differences in shooting angles and lighting conditions, image correction, enhancement, and other preprocessing are required.
[0064] (2) Infrared thermal imaging: The radiation heat map of the equipment is collected by handheld or airborne infrared thermal imagers to reflect its surface temperature distribution. It can be used to detect defects such as overheating of wires and joints and partial discharge of transformers. Due to the influence of factors such as imaging distance and ambient temperature, the thermal map needs to be calibrated and noise filtered.
[0065] (3) Ultraviolet imaging: Using the power frequency electric field excitation characteristics, ultraviolet imaging is used to detect the corona discharge phenomenon on the surface and surroundings of high-voltage equipment. This steady-state radiation can be used as a basis for diagnosing defects such as insulation aging and metal foreign matter. In order to suppress the interference of the environmental background, filtering and background reduction are usually used to improve the signal-to-noise ratio.
[0066] (4) Other data: Environmental parameters such as temperature, humidity, and wind speed provided by meteorological sensors are used to analyze the impact of external conditions on the occurrence and development of defects. Structured data such as equipment nameplates and operation records are used to obtain background information such as equipment attributes and historical status.
[0067] Transformer is a sequence modeling method based on the self-attention mechanism. Different from time-recursive models such as RNN, Transformer introduces a self-attention mechanism to model the interdependence between elements in a sequence. Given an input sequence of length n Transformer first linearly transforms it into three matrices: query matrix Q, key matrix K, and value matrix V:
[0068] Q=XW Q , K = XW K , V = XW V Formula 1;
[0069] Where W Q , W K , W V is the learnable weight matrix.
[0070] Then the attention matrix A is obtained by calculating the normalized dot product of Q and K:
[0071]
[0072] Among them, softmax(·) is the normalization function, d k is the dimension of K, used to scale the dot product result. Each element a in A ij represents the attention weight from position i to j.
[0073] Finally, multiply A and V to get the output sequence Z:
[0074] Z=AV Formula 3;
[0075] Z is a weighted sum representation of the input X, where the weights A are adaptively assigned through the self-attention mechanism. This allows the Transformer to flexibly model dependencies between different positions.
[0076] The present invention utilizes Transformer to extract shared semantic features in power inspection data.
[0077] Specifically, for a set of heterogeneous inputs (such as visible light, infrared, and ultraviolet images), the present invention first divides them into fixed-size blocks (16×16) and flattens each block into a vector. Then, position embedding is added to all vectors to tell their spatial position in the image. These vectors are then concatenated into a long sequence and input into the Transformer encoder. Through multi-layer self-attention calculations, the Transformer can adaptively aggregate semantic dependencies between different blocks and extract global features. The present invention uses the output Z of the last layer as a shared semantic embedding, denoted by F global :
[0078] F global = Transformer(Concat(X visible , X infrared , X ultraviolet )) Formula 4;
[0079] Where X visible , X infrared , Xultraviolet They represent visible light, infrared, and ultraviolet image block sequences respectively, and Concat(·) is the sequence concatenation function.
[0080] Although Transformer can effectively mine global semantics, it lacks explicit modeling of local structured features. For targets such as transmission line towers and substation equipment, the spatial topology and subordinate relationships of their internal components often contain discriminative information. However, Transformer's self-attention calculation "flattens" and mixes the features of different blocks, which is prone to losing such structural clues. Therefore, the present invention further introduces Capsule Networks to model local features.
[0081] Capsule is a group of vector neurons, whose activity values represent the multi-dimensional attributes of a visual entity, such as position, size, direction, etc. Intuitively speaking, if scalar neurons are compared to detecting whether a visual concept appears, then capsules correspond to detecting the way it appears (i.e. posture). Capsule Networks self-organizes the activation transfer from shallow capsules to deep capsules through a dynamic routing algorithm, thereby modeling the hierarchical relationship between the whole and the part.
[0082] Let the output of the u-th Capsule in the i-th layer be u i , the input of the vth Capsule in the jth layer is s j The two are transformed by the matrix W ij connect:
[0083]
[0084] in is u i For j The prediction vector of .
[0085] s j By aggregating all prediction vectors get:
[0086]
[0087] where c ij for u i to j The coupling coefficient satisfies c ij Updated by iterative dynamic routing algorithm:
[0088]
[0089] where v j Yes jThe output vector after nonlinear transformation (usually squash function). Formula (7) is the softmax(·) normalization; Formula (8) converts b ij Updated to With v j Intuitively, if u i Prediction With v j If they are consistent, then increase c ij , that is, it is more inclined to i The information is passed to s j .
[0090] For the extracted image blocks, the present invention first uses CNN to generate its primary capsule representation. Then, through multi-layer capsules and dynamic routing algorithms, capsules with strong correlations are adaptively aggregated to establish a hierarchical relationship between local components. The present invention concatenates the output of the highest-level capsule as a local structured feature embedding, denoted as F local :
[0091] F local =Concat(v1, v2, …, v m ) Formula 9;
[0092] Where m is the number of capsules in this layer.
[0093] Finally, the present invention converts the local feature F local And the global feature F extracted by Transformer global Cascade and send to the classifier (MLP) to determine the current status of the device:
[0094] y t =f cls (Concat(F global , F local )) Formula 10;
[0095] where f cls is the classifier, y t is the state estimate at time t.
[0096] The state of power equipment usually has a certain degree of continuity and stability on a time scale. That is, the current state not only depends on instantaneous multi-source data, but is also closely related to the previous historical state. In order to mine this kind of timing constraint information, the present invention further introduces Kalman Filter to track the equipment state.
[0097] Let the state variable at time t be x t , the observed variable is z t The state space equation is:
[0098] x t =F t-1 x t-1 +B t-1 u t-1 +w t-1 Formula 11;
[0099] z t =H t x t +v t Formula 12;
[0100] where F t is the state transfer matrix, B t is the control matrix, u t is the control variable, H t is the observation matrix, are process noise and observation noise respectively, and they obey Gaussian distribution with mean zero.
[0101] Assume that the posterior estimate of the state at time t-1 is known and its covariance matrix P t-1|t-1 . Then the state prediction at time t is:
[0102]
[0103] The predicted state covariance matrix is:
[0104]
[0105] When we observe z at time t t When available, the Kalman gain K is calculated by t :
[0106]
[0107] The predicted value is corrected with the observation error to obtain the state posterior estimate:
[0108]
[0109] The corrected covariance matrix is:
[0110] P t|t =(IK t H t ) t|t-1 Formula 17;
[0111] by and P t|t As the prior for the next moment, repeat the above prediction and update process.
[0112] In the method of the present invention, the state estimation y output by Transformer and Capsule Network is t Considered as observation z t , the real state of the device is the hidden variable x t The result of equation (10) is sent to the Kalman Filter through the optional state transfer and observation model, and the historical state information of the device is integrated to output the state posterior estimate after filtering correction. This estimation comprehensively considers the current multi-source data-driven diagnostic results and the timing constraints of historical status, and can more robustly and comprehensively reflect the health level of the equipment.
[0113] In the actual power inspection process, the degree of interest in different equipment parts and different defect types varies depending on the task. For example, daily inspections focus on comprehensive screening, while special inspections focus on specific hidden dangers. The two differ in data source selection, analysis focus, etc. In order to adapt to different task requirements, the present invention introduces a task instruction-driven human-computer interaction mechanism.
[0114] Specifically, the present invention predefines a series of structured task instructions, covering common inspection scenarios, such as "daily inspection", "special inspection of wire dancing", "special inspection of pole tower tilt", "special inspection of insulator contamination", etc. Each instruction corresponds to several parameters, such as target equipment type, type of defect of concern, source of visual data, etc. Given the instructions and parameters, the system adaptively adjusts the data input pipeline and selects relevant multi-source data for aggregation. At the same time, constraints are imposed on the weights of each component of the model (such as Transformer, Capsule Network) according to the instructions to focus on specific patterns of interest. For example, the wire dancing detection task will select infrared data as the key input to guide the Capsule to extract the contour structure features of the wire.
[0115] The presentation form of the inspection results will also be dynamically generated according to the instructions. For daily inspections, the system outputs global information such as the overall health score of the equipment and the probability of defects of each component; for special inspections, the system focuses on analyzing the severity and development trend of a given type of defect and gives a quantitative analysis report. All results are converted into text descriptions that are easy for engineers to understand through the natural language generation module, and organized into structured inspection logs in accordance with power industry specifications.
[0116] By introducing task instructions, the method of the present invention can flexibly adapt to different inspection requirements, balance versatility and pertinence, and greatly improve the practical value of the system.
[0117] To verify the effectiveness of the method of the present invention, a large-scale multi-source heterogeneous data set was collected in a real power system. The data set covers the main equipment of transmission lines and substations, including towers, insulators, conductors, lightning arresters, transformers, etc. Visible light, infrared, and ultraviolet images are collected for each device, and environmental parameters and device metadata are recorded at the same time.
[0118] Active learning strategies are used to expand the annotation of small sample defects in a targeted manner to improve the balance of data categories. The final dataset meets the research use standards in terms of defect type coverage, scenario diversity, and professional quality.
[0119] The present invention randomly divides the data set into training set, validation set and test set in a ratio of 7:2:1. According to the distribution characteristics of infrared and ultraviolet data, the present invention explores a variety of data enhancement methods to further expand the richness of training samples.
[0120] This paper uses indicators such as precision, recall, F1 value, and intersection over union (IoU) to evaluate model performance:
[0121]
[0122]
[0123] TP, FP, and FN represent the number of true positive, false positive, and false negative pixels, respectively. Intuitively speaking, precision reflects the accuracy of the detection results, recall reflects the completeness of the detection results, and F1 value is the harmonic average of the two. IoU measures the overlap between the detection results and the true value, and is more sensitive to the fineness of defect location.
[0124] Considering the safety-oriented characteristics of power inspection, the present invention gives priority to the recall rate indicator to control the risk of missed detection. When the recall rate is equivalent, the precision and IoU are weighed to take into account the false alarm rate and positioning accuracy. The present invention compares the proposed method with the existing SOTA method in various indicators, and analyzes the statistical significance of the performance difference through significance test.
[0125] In order to understand the effectiveness of each component of the model, the present invention conducted a series of experiments. The baseline model uses CNN to extract features and MLP for classification. On this basis, the present invention gradually adds modules such as Transformer (for global features), CapsuleNetwork (for local features), and Kalman Filter (for time series fusion) to examine the performance trend. The configuration of each model and the corresponding performance are shown in Table 1. It can be seen that:
[0126] (1) Adding either Transformer or Capsule Network alone can significantly improve model performance, verifying the complementarity of global semantic features and local structural features. Combining the two can further improve the effect and achieve the best result.
[0127] (2) Compared with using only the current frame information, incorporating the previous state can significantly improve the detection recall rate, thanks to the correction of missed defects by KalmanFilter. At the same time, Kalmax Filter can also alleviate false positives to a certain extent, reflecting the gain of timing constraints on stability.
[0128] Table 1:
[0129]
[0130] The full model is compared with the following SOTA methods:
[0131] (1) FPN: Feature Pyramid Network, a representative multi-scale object detection framework;
[0132] (2) Mask R-CNN: Instance segmentation model based on region proposal and ROIAlign;
[0133] (3) PSPNet: Pyramid scene parsing network, a multi-scale semantic segmentation framework;
[0134] (4) DeepLab v3+: Dilated convolutional semantic segmentation network, which enhances the capture of fine-grained and multi-scale information.
[0135] The performance comparison of each model on the test set is shown in Table 2. It can be seen that the method of the present invention is significantly better than the existing SOTA in all evaluation indicators. Compared with the detection and segmentation framework that focuses on visual features, the method of the present invention fully considers the field characteristics of power inspection, introduces timing constraints and multi-source fusion, and strengthens the defect characterization capabilities from both the data and model levels. In addition, the task instruction driven mechanism also makes the method of the present invention more flexible and efficient in adapting to different needs. Combining the characteristics of high recall, high precision, and strong generalization, the method of the present invention not only meets the safety standards of the power industry but also takes into account practicality, and has the potential for actual implementation.
[0136] Table 2:
[0137]
[0138] In order to examine the adaptability of the method of the present invention to new scenarios, the present invention collected test images at three other substations to evaluate the generalization of the model. These images are different from the training data in terms of equipment type, operating environment, imaging conditions, etc., which can better test the robustness of the method. The average performance of each model on three external test sets is shown in Table 3. It can be seen that the method of the present invention still maintains a leading position in all indicators, and the advantages are particularly significant. This is attributed to the fact that the present invention adopts powerful feature extractors such as Transformer, supplemented by strategies such as data enhancement and active learning, which enhances the model's ability to represent and generalize complex scenes. In contrast, traditional methods such as FPN and PSPNet have dropped significantly on external test sets, and their generalization capabilities need to be improved.
[0139] Table 3:
[0140]
[0141] Aiming at the key problems in intelligent power inspection, the present invention proposes a task instruction driven method integrating Transformer, Capsule Networks and Kalman Filter. Starting from the two levels of data and model, the method makes full use of multi-source heterogeneous information and timing constraints to build a defect characterization model for complex scenarios. On this basis, the task instruction driven mechanism is further introduced to enable the system to flexibly adapt to different inspection requirements. Experiments in a real power grid environment show that the method of the present invention is significantly superior to existing methods in terms of defect detection accuracy, recall rate and other indicators, and has good generalization ability, showing practical application value.
Claims
1. A robot inspection method based on task instructions, characterized in that: The method steps are as follows: Step 1. Collect image data of power transmission line components using multi-source sensors and perform preprocessing; Step 2. Aggregate the preprocessed multi-source heterogeneous image data into a unified data matrix; Step 3. Use the Transformer self-attention mechanism to extract the global shared features of the image data; Step 4. Extract local structural features based on Capsule Network; Step 5. Fusion of global features extracted by Transformer and local features extracted by Capsule Network; Step 6. Evaluate the current state of the device by combining the current feature estimate with the historical state information; Step 7. Dynamically adjust the model according to different task instructions and generate targeted inspection strategy reports.
2. A robot inspection method based on task instructions according to claim 1, characterized in that: The step 1 comprises: Step 1.1 Use visible light camera, infrared thermal imager, ultraviolet imager multi-source sensor to collect image data of transmission line towers, insulators, and conductor components. raw , where X raw ={X visible , X infrared , X ultraviolet }. Among them, X visible Represents visible light image data; X infrared Represents infrared image data; X ultraviolet represents the UV image data, Step 1.2: collect the original image data X raw Perform preprocessing, including image correction, noise removal, and data enhancement, to obtain the preprocessed image data X preprocessed , where X prprocessed =f preprocess (X raw ). Among them, f preprocess (·) represents the preprocessing function.
3. The method according to claim 1, characterized in that: The step 2 specifically includes: The pre-processed visible light, infrared, and ultraviolet image data X preprocessed Converge into a unified data matrix X fused , where X fused =f fuse (X preprocessed ). Among them, f fuse (·) represents the data aggregation function.
4. The robot inspection method based on task instructions according to claim 1 is characterized in that: The step 3 specifically includes: Step 3.1 The aggregated data matrix X fused Divide the image into fixed-size blocks and flatten each block into a vector, denoted by x i , Step 3.2: For the image block vector x i Perform linear transformation to obtain the query matrix Q, key matrix K and value matrix V: Q = XW Q ,K = XW K ,V = XW V Formula 1; Among them, W Q , W K , W V represents the learnable weight matrix; X represents the image block vector x i The matrix composed of Step 3.3 Calculate the self-attention matrix A: Among them, softmax(·) represents the softmax normalization function; d k represents the dimension of the key matrix K, Step 3.4 multiplies the self-attention matrix A by the value matrix V to get the output Z of the Transformer encoder: Z=AV Formula 3, Step 3.5 embeds the output Z of the Transformer encoder into Fglobal as a shared semantic feature: F global =Z Formula 4.
5. The robot inspection method based on task instructions according to claim 1 is characterized in that: The step 4 specifically includes: Step 4.1 Use the convolutional neural network CNN to extract the primary features of the image block and obtain the primary capsule u i , Step 4.2 Aggregate primary capsules u through a dynamic routing algorithm i , get the high-level capsule v j , Among them, W ij represents the transformation matrix from the i-th primary capsule to the j-th high-level capsule; c ij represents the coupling coefficient from the i-th primary capsule to the j-th high-level capsule, Step 4.3 Repeat step 4.2 to get the top-level capsule {v1, v2, …, v m }, Step 4.4: Concatenate the top-level capsules to get the local structural features embedded in F local : F local =Concat(v1, v2, …, v m ) Formula 6; Among them, Concat(·) represents the concatenation operation.
6. The robot inspection method based on task instructions according to claim 1 is characterized in that: The step 5 specifically includes: The global feature F extracted by Transformer global And the local feature F extracted by Capsule Network local Splicing to get the fusion feature F fused : F fused =Concat(F global , F local ) Formula 7.
7. The robot inspection method based on task instructions according to claim 1 is characterized in that: The step 6 includes: step 6.1: fusion feature F fused Input classifier f cls (·) (such as MLP), get the current state estimate y of the device t : y t =f cls (F fused ) Formula 8; Among them, y t represents the estimated state of the device at time t, Step 6.2 Estimate the current state y t Input Kalman Filter, combined with historical state estimation, to obtain the corrected state estimation in, represents the prediction of the state at time t using the state estimate at time t-1; K t represents Kalman gain; H t Represents the observation matrix.
8. A robot inspection system based on task instructions, characterized in that: The system comprises a data acquisition module, a data preprocessing module, a feature extraction module, a state evaluation module, a task planning module, and a human-computer interaction module; the system is used to execute the method according to any one of claims 1 to 7.
9. A robot inspection system based on task instructions as claimed in claim 8, characterized in that: The feature extraction module includes two submodules: global feature extraction and local feature extraction; the global feature extraction submodule adopts the Transformer self-attention mechanism to extract features with global semantic information from the entire image range; Transformer calculates the attention weights between image blocks and builds long-range dependencies between blocks to obtain a global representation of the image. The local feature extraction submodule uses Capsule Network to extract local structural features of the image. Capsule Network uses a dynamic routing algorithm to adaptively aggregate low-level capsules to form high-level capsules, modeling the hierarchical relationship between the whole and the local. The extraction of global features and local features is carried out in parallel, and finally the two features are fused to form a complete image representation.
10. A robot inspection system based on task instructions as claimed in claim 8, characterized in that: The state evaluation module receives the fused features output by the feature extraction module and evaluates the current state of the device; the state evaluation is divided into two stages: current state estimation and state timing optimization; in the current state estimation stage, the fused features are input into the classifier to predict the current state of the device; In the state timing optimization stage, the current state estimation result and the historical state estimation result are input into the Kalman Filter together, and the current state is corrected through the state space model; the Kalman Filter smoothes and optimizes the state estimation result in the time dimension by recursively estimating the posterior probability distribution of the state.
Citation Information
Patent Citations
Artwork classification method and system based on multi-modal fusion
CN114170460A
Malicious URL detection and classification method based on capsule neural network
CN116471096A
Intelligent photovoltaic hot spot fault detection method and system, medium, equipment and terminal
CN117095311A
Power transmission line inspection method based on unmanned aerial vehicle and CNN-BiLSTM model
CN117541534A
Distribution line unmanned aerial vehicle inspection image target detection and hidden danger intelligent identification method
CN117726958A
Cited By
Method and system for identifying defects of power equipment based on multi-modal information fusion
CN121211207A
Cable defect monitoring and risk grading method based on multi-modal deep fusion
CN121350851A
A cable defect monitoring and risk grading method based on multi-modal deep fusion
CN121350851B