A method, system and computer device for detecting and deploying a power inspection robot

By applying knowledge distillation and model pruning technology in power inspection, and optimizing the power inspection robot detection model, the problems of high frequency of lidar use and low data processing efficiency are solved, an efficient and economical inspection plan is achieved, and the real-time and accuracy of inspection tasks are ensured.

CN119910662BActive Publication Date: 2025-06-13HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510397569.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-06-13
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

The power inspection tasks use lidar frequently, resulting in waste of resources in equipment procurement costs, maintenance costs and operation training. At the same time, it is necessary to quickly obtain on-site data in emergency situations. It is difficult for the existing technology to reduce the frequency of lidar usage and improve data processing efficiency while ensuring the quality of inspection.

Method used

The power inspection robot detection and deployment method based on knowledge distillation and model pruning is adopted. Through the joint training of the teacher model and the student model, combined with focus distillation, relational distillation and global distillation losses, the student model is optimized to reduce hardware requirements, and the optimized model is converted into ONNX format for TensorRT deployment.

Benefits of technology

The detection accuracy of low-wire harness radar is achieved close to that of high-wire harness radar, while reducing the computing burden, simplifying the difficulty of network deployment, and improving real-time processing speed, which not only ensures patrol accuracy but also effectively reduces costs, and ensures the real-time and efficient patrol tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119910662B_ABST
    Figure CN119910662B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and computer device for detecting and deploying a power inspection robot. The three-dimensional point cloud data of high and low wire harnesses of the road scene to be detected is obtained and input into a teacher model and a student model respectively to extract BEV features. The focus module is used to perform distillation processing on the three-dimensional features of the teacher and student models, create foreground and background masks, and calculate the loss of foreground features. The focus relationship distillation module is introduced to extract the features of nine corner points at the foreground position and calculate the Gaussian similarity, improving the balance of the model in learning foreground features of different categories. The global distillation module is applied to improve the overall performance. After the model training is completed, model pruning is carried out to optimize the model performance. The pruned and optimized model is converted into the ONNX format and deployed and accelerated on TensorRT to further improve the inference efficiency. It effectively solves the problems of low detection accuracy of low wire harness radar and difficulties in actual deployment of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of electric power inspection target detection, and in particular relates to an electric power inspection robot detection and deployment method, system and computer equipment. Background Art

[0002] In recent years, LiDAR technology has shown wide application potential in the field of power inspection and has become an important means to improve inspection accuracy and efficiency. Power inspection tasks mainly cover the inspection of transmission lines, substations and other facilities, and LiDAR, with its high-precision three-dimensional data acquisition capabilities, can effectively identify the status and location of power equipment in complex environments, improving the reliability of inspections. However, since power inspection tasks are usually only carried out during specific periods of time and at a low frequency, overly expensive LiDAR equipment often leads to a waste of resources, including high equipment procurement costs, maintenance costs, and operator training burdens.

[0003] At the same time, power inspection also places high demands on real-time performance, especially in emergency situations, where inspectors need to quickly obtain on-site data and make decisions. Therefore, how to reduce the frequency of using LiDAR while ensuring the quality of inspections and achieve efficient and fast data processing has become a problem that the power industry needs to solve urgently. Summary of the invention

[0004] In view of the above technical problems, the present invention provides a method, system and computer equipment for detecting and deploying an electric power inspection robot.

[0005] The technical solution adopted by the present invention to solve the technical problem is:

[0006] A method for detecting and deploying an electric power inspection robot, the method comprising the following steps:

[0007] S100: In an open road environment, use a camera to collect two-dimensional image data of the road to be inspected, obtain corresponding three-dimensional point cloud data through high-beam laser radar and low-beam laser radar respectively, and save the calibration file and the true value annotation file;

[0008] S200: inputting the high-beam 3D point cloud data into the teacher model, inputting the low-beam 3D point cloud data into the student model, respectively extracting the 3D features of the high-beam 3D point cloud and the low-beam 3D point cloud through the 3D backbone network in the model, and respectively inputting the 3D features into the 2D backbone network for conversion to obtain BEV features, and then obtaining the prediction results, wherein the teacher model is trained separately for the high-beam point cloud to obtain pre-weights for sharing by the two models;

[0009] S300: Read the ground truth annotation file through the focal distillation module to obtain the foreground three-dimensional coordinates, project the foreground three-dimensional coordinates onto the two-dimensional BEV coordinates, and set masks to distinguish foreground features and background features. Use the foreground mask and background mask to weight the BEV feature losses of the teacher model and the student model to obtain the focal loss;

[0010] S400: Extract the nine corner point features of the prediction results of the teacher model and the student model through the focal relationship distillation module, calculate the Gaussian similarity of the corner points of the teacher model and the student model, and perform contrast distillation to obtain the relationship distillation loss;

[0011] S500: Calculate the difference between the BEV features of each sample of the teacher model and the student model through the global distillation module to obtain the global distillation loss;

[0012] S600: Train the student model based on the focal loss, relationship distillation loss, global distillation loss, and calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove unimportant filters according to a certain proportion, reconstruct the network, and retrain the pruned student model to obtain the optimized student model;

[0013] S700: Convert the optimized student model to the ONNX format, deploy and accelerate it on TensorRT, and test the obtained real-time input data.

[0014] Preferably, S200 includes:

[0015] S210: Input the high beam three-dimensional point cloud data into the teacher model, and input the low beam three-dimensional point cloud data into the student model. Divide the point cloud space of a sample into cylindrical grids through the three-dimensional backbone network in the teacher and student models. The points in the sample are assigned to the corresponding cylinders according to their spatial coordinates, and the cylinders without points are regarded as empty cylinders;

[0016] S220: Assume that the number of non-empty cylinders in the sample is P, and at the same time limit the maximum number of points in each cylinder to N. If the number of points in a cylinder is less than N, it is filled with 0. If it exceeds N, sample N points from the points in the cylinder and encode each point in the cylinder, where the representation of each point includes the coordinates of the point and the reflection intensity;

[0017] S230: After obtaining the cylinder representation tensor of the point cloud, first process the points in each cylinder, transform the feature dimension of each point from the original dimension to the target dimension through a multi-layer perceptron, and then perform max pooling on all the points in each cylinder to obtain the feature vector of each cylinder, that is, the three-dimensional feature;

[0018] S240: Finally, the 2D backbone networks in the teacher model and the student model expand the 3D features of each cylinder into pseudo-image features according to the position of the cylinder in space. By flattening the number of cylinders into the height and width of a 2D image, the feature map is finally transformed into an image-like form to obtain the BEV features. The position of the target object in the BEV features is the prediction result.

[0019] Preferably, S300 includes:

[0020] S310: Read the ground truth annotation file, calculate the mask region for each foreground position R, and project the foreground 3D coordinates into the 2D BEV domain:

[0021] ;

[0022] ;

[0023] ;

[0024] ;

[0025] Where , , and are the bounding box ranges after projecting the 3D positions into 2D. and are determined by the valid range of the lidar point cloud. and are the voxel sizes of the KITTI dataset base configuration file. represents the downsampling times when the voxel network extracts deep features;

[0026] S320: Set a binary mask Mask to segment the background features and foreground features:

[0027] ;

[0028] Where r represents the foreground box, and i and j are the horizontal and vertical coordinates of the feature map respectively. If and , then = 1, otherwise 0;

[0029] S330: Normalize the weights of each foreground position and calculate the focal loss:

[0030] ;

[0031] ;

[0032] Among them, N is the batch size, and M is the number of ground truth boxes or ROIs. and represent the features at the foreground positions in the BEV features of the m-th sample in the n-th batch for the teacher features and the student features, respectively.

[0033] Preferably, S400 includes:

[0034] S410: Combine the , , and to form four vertices, calculate the midpoints of the sides formed by the four vertices to obtain four midpoints, then connect the diagonals of the vertices pairwise, and calculate the center point of the bounding box;

[0035] S420: Collect the features of these nine corner points in the BEV feature maps of the student model and the teacher model to form a 9*9 Gaussian similarity of the teacher model and a Gaussian similarity of the student model :

[0036] ;

[0037] where and are the features of the teacher model corner points and respectively, and are the features of the student model corner points and respectively, is the standard deviation of the Gaussian function;

[0038] S430: Calculate the relationship distillation loss through contrastive distillation, specifically:

[0039] .

[0040] Preferably, S500 is specifically:

[0041] ;

[0042] where and represent the BEV features of the teacher-student models, calculate the difference through the L2 norm, and use it as the global distillation loss.

[0043] Preferably, S600 includes:

[0044] S610: Count the number of filters in the student model and assign an index to each filter;

[0045] S620: Zero out each filter one by one, test its impact on the detection accuracy, and record the accuracy change. In this way, process all filters one by one, and finally generate a TXT file containing the importance of each filter. Among them, the smaller the accuracy change value, the lower the importance of the filter;

[0046] S630: Read and sort the importance of the filters in the TXT file in ascending order, and update the corresponding filter indices at the same time;

[0047] S640: Remove the filters with lower importance according to the set pruning rate;

[0048] S650: Adjust the channel structure of the student model according to the number of remaining filters, and retrain it to optimize the weights on this basis to achieve the best performance.

[0049] Preferably, S700 includes:

[0050] S710: Reconstruct the PointPillar model and convert it to the ONNX format. During this process, the PointPillar model is divided into two parts, and finally these two parts are spliced together; among them, the first part includes the PFE module and the Scatter module, which are responsible for extracting the features of the pillars, and the second part contains the Backbone module and the Densehead module, which are responsible for extracting the deep features of the pillars and making predictions;

[0051] S720: Adjust or use custom operators to customize the operators according to the ONNX support library to make the ONNX format compatible with all operators in the PointPillar model;

[0052] S730: Optimize the ONNX model, convert the ONNX model to the TensorRT acceleration format engine file, and use the engine file to test the input data to complete the deployment.

[0053] A power inspection robot detection and deployment system, including a data acquisition module, a BEV feature extraction module, a focus distillation module, a focus relationship distillation module, a global distillation module, a model pruning and training module, and a model deployment module;

[0054] The data acquisition module is used to collect two-dimensional image data of the road to be detected using a camera in an open road environment, obtain the corresponding three-dimensional point cloud data through a high-beam lidar and a low-beam lidar respectively, and save the calibration file and the ground truth annotation file;

[0055] The BEV feature extraction module is used to input the high-line beam three-dimensional point cloud data into the teacher model and the low-line beam three-dimensional point cloud data into the student model. The three-dimensional features of the high-line beam three-dimensional point cloud and the low-line beam three-dimensional point cloud are respectively extracted through the three-dimensional backbone network in the model, and the three-dimensional features are respectively input into the two-dimensional backbone network for conversion to obtain the BEV features, and then the prediction results are obtained. Among them, the teacher model is trained separately for the high-line beam point cloud to obtain the pre-weights for sharing by the two models;

[0056] The focal distillation module is used to read the ground truth annotation file to obtain the foreground three-dimensional coordinates, project the foreground three-dimensional coordinates onto the two-dimensional BEV coordinates, and set masks to distinguish the foreground features and the background features, and use the foreground mask and the background mask to weight the BEV feature losses of the teacher model and the student model to obtain the focal loss;

[0057] The focal relationship distillation module is used to extract the nine corner point features of the prediction results of the teacher model and the student model, calculate the Gaussian similarity of the corner points of the teacher model and the student model and perform contrast distillation to obtain the relationship distillation loss;

[0058] The global distillation module is used to calculate the difference between the BEV features of each sample of the teacher model and the student model to obtain the global distillation loss;

[0059] The model pruning training module is used to train the student model based on the focal loss, the relationship distillation loss, the global distillation loss and the calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove the unimportant filters according to a certain proportion, reconstruct the network and retrain the pruned student model to obtain the optimized student model;

[0060] The model deployment module is used to convert the optimized student model into the ONNX format, deploy and accelerate it on TensorRT, and test the obtained real-time input data.

[0061] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the power inspection robot detection and deployment method are implemented.

[0062] The above power inspection robot detection and deployment method, system, and computer device design a lightweight model for power inspection robots based on knowledge distillation and model pruning and implement its deployment for detection. The lightweight model can reduce hardware requirements, enable the detection accuracy of low-beam lidar to approach that of high-beam lidar, and also reduce the computational burden. By adopting model pruning to simplify the network deployment difficulty, the real-time processing speed of the network is improved, ensuring both accuracy and effectively reducing costs, and ensuring the real-time and efficient nature of the inspection task. By optimizing the system design and data processing algorithm, the lightweight model provides a more flexible and economical solution for power inspection and will play an important role in future inspection tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 The flowchart of a method for detecting and deploying a power inspection robot according to an embodiment of the present invention;

[0064] Figure 2 The schematic diagram of the principle of a method for detecting and deploying a power inspection robot according to an embodiment of the present invention;

[0065] Figure 3 The flowchart of the operation of the focus distillation module according to an embodiment of the present invention;

[0066] Figure 4 The flowchart of the operation of the focus relationship distillation module according to an embodiment of the present invention;

[0067] Figure 5 The flowchart of pruning the student model according to an embodiment of the present invention;

[0068] Figure 6 The flowchart of deploying the student model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] In order to enable those skilled in the art of the present technology to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings.

[0070] In one embodiment, as Figure 1 and Figure 2 shown, a method for detecting and deploying a power inspection robot, the method includes the following steps:

[0071] S100: In an open road environment, use a camera to collect two-dimensional image data of the road to be detected, and respectively obtain corresponding three-dimensional point cloud data through a high-beam lidar and a low-beam lidar, and save the calibration file and the ground truth annotation file;

[0072] S200: Input the high-beam three-dimensional point cloud data into the teacher model, and input the low-beam three-dimensional point cloud data into the student model. Extract the three-dimensional features of the high-beam three-dimensional point cloud and the low-beam three-dimensional point cloud respectively through the three-dimensional backbone network in the model, and input the three-dimensional features into the two-dimensional backbone network for conversion to obtain BEV features, and then obtain the prediction results. Among them, the teacher model is trained separately for the high-beam point cloud to obtain pre-weights for sharing by the two models;

[0073] S300: Read the ground truth annotation file through the focal distillation module to obtain the foreground three-dimensional coordinates, project the foreground three-dimensional coordinates onto the two-dimensional BEV coordinates, and set masks to distinguish foreground features and background features. Use the foreground mask and background mask to weight the BEV feature losses of the teacher model and the student model to obtain the focal loss;

[0074] S400: Extract the nine corner features of the prediction results of the teacher model and the student model through the focal relationship distillation module, calculate the Gaussian similarity of the corners of the teacher model and the student model and perform contrast distillation to obtain the relationship distillation loss;

[0075] S500: Calculate the difference between the BEV features of each sample of the teacher model and the student model through the global distillation module to obtain the global distillation loss;

[0076] S600: Train the student model based on the focal loss, relationship distillation loss, global distillation loss and calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove unimportant filters according to a certain proportion, reconstruct the network and retrain the pruned student model to obtain the optimized student model;

[0077] S700: Convert the optimized student model to the ONNX format, deploy and accelerate it on TensorRT, and test the obtained real-time input data.

[0078] In one embodiment, S200 includes:

[0079] S210: Input the high-beam three-dimensional point cloud data into the teacher model, and input the low-beam three-dimensional point cloud data into the student model. Divide the point cloud space of a sample into cylinder grids through the three-dimensional backbone network in the teacher and student models. The points in the sample are assigned to the corresponding cylinders according to their spatial coordinates, and the cylinders without points are regarded as empty cylinders;

[0080] S220: Assume that the number of non-empty cylinders in the sample is P, and at the same time, limit the maximum number of points in each cylinder to N. If the number of points in a cylinder is less than N, it is filled with 0s. If it exceeds N, N points are sampled from the points in the cylinder, and each point in the cylinder is encoded. The representation of each point includes the coordinates of the point and the reflection intensity;

[0081] S230: After obtaining the cylinder representation tensor of the point cloud, first process the points in each cylinder. Through a multi-layer perceptron, the feature dimension of each point is transformed from the original dimension to the target dimension. Then, max pooling is performed on all the points in each cylinder to obtain the feature vector of each cylinder, that is, the three-dimensional feature;

[0082] S240: Finally, the two-dimensional backbone networks in the teacher model and the student model expand the three-dimensional feature of each cylinder into pseudo-image features according to the position of the cylinder in space. By flattening the number of cylinders into the height and width of a two-dimensional image, the feature map is finally transformed into an image-like form to obtain the BEV feature. The position of the target object in the BEV feature is the prediction result.

[0083] In one embodiment, as Figure 3 shown, S300 includes:

[0084] S310: Read the ground truth annotation file, calculate the mask region for each foreground position R, and project the foreground three-dimensional coordinates into the two-dimensional BEV domain:

[0085] ;

[0086] ;

[0087] ;

[0088] ;

[0089] Among them, , , and are the bounding box ranges after projecting the three-dimensional position into two dimensions. and are determined by the valid range of the lidar point cloud, which are 0 and -40 here respectively. and are the voxel sizes of the KITTI dataset base configuration file, which are set to 0.05 here. represents the downsampling times when the voxel network extracts deep features, which is set to 8;

[0090] S320: Set a binary mask Mask to segment the background features and foreground features:

[0091] ;

[0092] Where r represents the foreground box, and i and j are the horizontal and vertical coordinates of the feature map respectively. If and , then = 1, otherwise 0;

[0093] S330: Normalize the weights of each foreground position and calculate the focal loss:

[0094] ;

[0095] ;

[0096] Where N is the batch size, M is the number of ground truth boxes or ROIs, and represent the features at the foreground positions in the BEV features of the m-th sample in the n-th batch of the teacher features and the student features respectively.

[0097] Specifically, the foreground three-dimensional coordinates are obtained by reading the ground truth annotation file using the focal distillation module, the three-dimensional coordinates are projected onto the two-dimensional BEV coordinates, the weight 1 is assigned to the foreground feature positions within the coordinate range, and the weight 0 is assigned to other positions, i.e., the background feature positions, so as to effectively distinguish the foreground and background features, enhance the model's attention to the foreground pixels and suppress the background noise.

[0098] In one embodiment, as Figure 4 shown, S400 includes:

[0099] S410: The , , and obtained by S310 are combined into four vertices, the midpoints of the sides formed by the four vertices are calculated to obtain four midpoints, and then the diagonals connecting the vertices are connected pairwise to calculate the center point of the bounding box;

[0100] S420: Collect the features of these nine corner points in the BEV feature maps of the student model and the teacher model to form the 9*9 Gaussian similarity of the teacher model and the Gaussian similarity of the student model:

[0101] ;

[0102] Where and are the features of the teacher model corner points and respectively, and are the corner points of the student model and features, is the standard deviation of the Gaussian function;

[0103] S430: Calculate the relationship distillation loss by contrast distillation, specifically:

[0104] .

[0105] Specifically, since different objects occupy different-sized regions, the general model tends to align the features of large objects rather than small objects. Here, the focal relationship distillation module evenly selects nine key points for each object for alignment, which can effectively mitigate the above influence.

[0106] In one embodiment, S500 is specifically:

[0107] ;

[0108] wherein and represent the BEV features of the teacher and student models, calculate the difference through the L2 norm, and use it as the global distillation loss.

[0109] Specifically, local focal distillation enables the model to focus on the details of specific regions, but often ignores the semantic information of the overall scene. Therefore, introducing global distillation can help the model effectively learn global context information, thereby enhancing the robustness and overall understanding ability of the model.

[0110] In one embodiment, as Figure 5 shown, S600 includes:

[0111] S610: Count the number of filters in the student model and assign an index to each filter;

[0112] S620: Zero out each filter one by one, test its impact on the detection accuracy, and record the accuracy change. In this way, process all filters one by one, and finally generate a TXT file containing the importance of each filter. Among them, the smaller the accuracy change value, the lower the importance of the filter;

[0113] S630: Read and sort the importance of the filters in the TXT file in ascending order, and update the corresponding filter indices at the same time;

[0114] S640: Remove the filters with lower importance according to the set pruning rate; After multiple experiments, it is found that when the pruning rate is set to 40%, the model performs best in the balance of speed and accuracy and has the highest stability;

[0115] S650: Adjust the channel structure of the student model according to the number of remaining filters and retrain it to optimize the weights based on this to achieve the best performance.

[0116] Specifically, perform model pruning on the student model, calculate the importance of each filter in the model, and then remove unimportant filters at a certain ratio. After the pruning operation, obtain the pre-trained weights before pruning and extract the parameters of the remaining filters. The pruned model will be trained based on these remaining parameters to ensure that the remaining feature representations are further optimized. Although the structure of the model has changed, the training process after pruning is basically the same as that before pruning: still perform forward propagation through the input data, calculate the loss value and perform backpropagation, and finally use the loss functions (classification loss, regression loss, and orientation loss) for supervised training. During this process, the model will optimize and adjust the remaining filters according to the loss value to gradually improve the performance of the model. After training, the pruned model will be evaluated to verify whether its performance on the test set meets the expected goals. Reconstruct the network and retrain it to improve the operation speed of the model without reducing the model accuracy too much.

[0117] In one embodiment, as Figure 6 shown, S700 includes:

[0118] S710: Reconstruct the PointPillar model and convert it to the ONNX format. During this process, the PointPillar model is divided into two parts, and finally these two parts are spliced together; among them, the first part includes the PFE (Point Feature Encoding Module) module and the Scatter (Scatter Module) module, which are responsible for extracting the features of the pillars, and the second part contains the Backbone (Backbone Module) module and the Densehead (DenseHead Module) module, which are responsible for extracting the deep features of the pillars and making predictions; among them, the PFE module converts the original point cloud data into a meaningful feature representation, the Scatter module scatters the features extracted by the PFE module back into the sparse three-dimensional space, the Backbone module is the part of the neural network used to extract the high-level features of the input data, and the DenseHead module is used to generate prediction results from the deep features extracted by the Backbone module;

[0119] S720: Adjust or use custom operators to customize the operators according to the ONNX support library, so that the ONNX format is compatible with all operators in the PointPillar model; specifically, since the ONNX format is not fully compatible with some operators in the PointPillar model, these operators need to be rewritten or replaced. Specifically, some operators implemented in the original framework may not be directly convertible to ONNX, and it is necessary to adjust according to the ONNX support library or use custom operators to ensure the correct execution and optimization of the model after conversion;

[0120] S730: Optimize the ONNX model, convert the ONNX model into a TensorRT accelerated format engine file, and use the engine file to test the input data to complete the deployment. Specifically, use TensorRT to accelerate the reconstructed ONNX model to further improve the inference speed. Through TensorRT optimization techniques such as layer fusion, precision scaling, and dynamic tensors, optimize the model to ensure that the accelerated model can still maintain the original accuracy and robustness under different input conditions.

[0121] In the above power inspection robot detection and deployment method, the model fully combines the advantages of the high-beam lidar in detection accuracy and the characteristics of the lightweight model after pruning, effectively solving the challenges of the low-beam lidar in detection accuracy and network deployment. Specifically, the model uses the teacher model of the high-beam lidar to guide the student model of the low-beam lidar to learn, and further optimizes through model pruning to obtain the best effect of the lightweight model. Finally, the model is converted through the ONNX format and efficiently deployed on the Orin platform with the help of TensorRT, ensuring the dual optimization of performance and deployment.

[0122] A power inspection robot detection and deployment system, including a data acquisition module, a BEV feature extraction module, a focus distillation module, a focus relationship distillation module, a global distillation module, a model pruning training module, and a model deployment module,

[0123] The data acquisition module is used to collect two-dimensional image data of the road to be detected using a camera in an open road environment, obtain the corresponding three-dimensional point cloud data through a high-beam lidar and a low-beam lidar respectively, and save the calibration file and the ground truth annotation file;

[0124] The BEV feature extraction module is used to input the high-beam three-dimensional point cloud data into the teacher model and the low-beam three-dimensional point cloud data into the student model. The three-dimensional features of the high-beam three-dimensional point cloud and the low-beam three-dimensional point cloud are respectively extracted through the three-dimensional backbone network in the model, and the three-dimensional features are respectively input into the two-dimensional backbone network for conversion to obtain BEV features, and then the prediction results are obtained. Among them, the teacher model is trained separately for the high-beam point cloud to obtain pre-weights for sharing by the two models;

[0125] The focal distillation module is used to read the ground truth annotation file to obtain the foreground three-dimensional coordinates, project the foreground three-dimensional coordinates onto the two-dimensional BEV coordinates, and set masks to distinguish foreground features and background features, and use the foreground mask and background mask to weight the BEV feature losses of the teacher model and the student model to obtain the focal loss;

[0126] The focal relationship distillation module is used to extract the nine corner features of the prediction results of the teacher model and the student model, calculate the Gaussian similarity of the corners of the teacher model and the student model and perform contrast distillation to obtain the relationship distillation loss;

[0127] The global distillation module is used to calculate the difference between the BEV features of each sample of the teacher model and the student model to obtain the global distillation loss;

[0128] The model pruning training module is used to train the student model based on the focal loss, relationship distillation loss, global distillation loss and calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove unimportant filters according to a certain ratio, reconstruct the network and retrain the pruned student model to obtain the optimized student model;

[0129] The model deployment module is used to convert the optimized student model into the ONNX format, deploy and accelerate it on TensorRT, and test the obtained real-time input data.

[0130] For the specific limitations of the power inspection robot detection and deployment system, reference can be made to the limitations on the power inspection robot detection and deployment method in the above text, which will not be elaborated here. Each module in the above power inspection robot detection and deployment system can be implemented in whole or in part by software, hardware and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0131] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method for detecting and deploying a power inspection robot are implemented.

[0132] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0133] The above has introduced in detail a method, system, and computer device for detecting and deploying a power inspection robot provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the core idea of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for detecting and deploying a power inspection robot, characterized in that: The method comprises the following steps: S100: In an open road environment, use a camera to collect two-dimensional image data of the road to be inspected, obtain corresponding three-dimensional point cloud data through high-beam laser radar and low-beam laser radar respectively, and save the calibration file and the true value annotation file; S200: inputting the high-beam 3D point cloud data into the teacher model, inputting the low-beam 3D point cloud data into the student model, respectively extracting the 3D features of the high-beam 3D point cloud and the low-beam 3D point cloud through the 3D backbone network in the model, and respectively inputting the 3D features into the 2D backbone network for conversion to obtain BEV features, and then obtaining the prediction results, wherein the teacher model is trained separately for the high-beam point cloud to obtain pre-weights for sharing by the two models; S300: The focus distillation module is used to read the true value annotation file to obtain the foreground three-dimensional coordinates, project the foreground three-dimensional coordinates to the two-dimensional BEV coordinates, and set a mask to distinguish the foreground features from the background features. The foreground mask and the background mask are used to weight the BEV feature losses of the teacher model and the student model to obtain the focus loss. S400: extracting nine corner point features of the prediction results of the teacher model and the student model through the focus relationship distillation module, calculating the Gaussian similarity of the corner points of the teacher model and the student model, and performing comparative distillation to obtain the relationship distillation loss; S500: Calculate the difference between the BEV features of each sample of the teacher model and the student model through the global distillation module to obtain the global distillation loss; S600: Train the student model based on the focal loss, relational distillation loss, global distillation loss and calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove unimportant filters according to a certain ratio, reconstruct the network and retrain the pruned student model to obtain an optimized student model; S700: Convert the optimized student model to ONNX format, deploy and accelerate it on TensorRT, and test it on the acquired real-time input data.

2. The method according to claim 1, characterized in that S200 includes: S210: inputting high-beam 3D point cloud data into the teacher model, inputting low-beam 3D point cloud data into the student model, dividing the point cloud space of a sample into cylindrical grids through the 3D backbone networks in the teacher and student models, assigning points in the sample to corresponding cylinders according to their spatial coordinates, and cylinders without points are regarded as empty cylinders; S220: Assume that the number of non-empty cylinders included in the sample is P, and the maximum number of points in each cylinder is limited to N. If the number of points in a cylinder is less than N, it is padded with 0. If it exceeds N, N points are sampled from the points in the cylinder, and each point in the cylinder is encoded, wherein the representation of each point includes the coordinates of the point and the reflection intensity; S230: After obtaining the cylinder representation tensor of the point cloud, firstly process the points in each cylinder, transform the feature dimension of each point from the original dimension to the target dimension through a multi-layer perceptron, and then perform maximum pooling on all the points in each cylinder to obtain the feature vector of each cylinder, i.e., the three-dimensional feature; S240: Finally, the two-dimensional backbone network in the teacher model and the student model expands the three-dimensional features of each cylinder into a pseudo-image feature according to the position of the cylinder in space. By flattening the number of cylinders into the height and width of the two-dimensional image, the feature map is finally converted into an image-like form to obtain the BEV feature. The position of the target object in the BEV feature is the prediction result.

3. The method according to claim 2, characterized in that S300 includes: S310: Read the true value annotation file, calculate the mask area for each foreground position R, and project the foreground three-dimensional coordinates into the two-dimensional BEV domain: ; ; ; ; in, , , and is the bounding box range after the three-dimensional position is projected into two dimensions, and Determined by the effective range of the LiDAR point cloud, and The voxel size of the base profile for the KITTI dataset, Indicates the number of downsampling times when the voxel network extracts deep features; S320: Set a binary mask to separate background features and foreground features: ; Where r represents the foreground box, i and j are the horizontal and vertical coordinates of the feature map respectively. If and ,but =1, otherwise 0; S330: Normalize the weight of each foreground position and calculate the focal loss: ; ; Among them, N is the batch size, M is the number of true value boxes or ROIs, and It represents the features of the foreground position in the BEV features of the mth sample in the nth batch of the teacher features and the student features.

4. The method according to claim 3, characterized in that S400 includes: S410: obtained through S310 , , and Combine into four vertices, calculate the midpoint of the edge formed by the four vertices, get four midpoints, then connect the diagonals of the vertices two by two to calculate the center point of the bounding box; S420: Collect the features of these nine corner points in the BEV feature map of the student model and the teacher model to form a 9*9 teacher model Gaussian similarity Gaussian similarity with the student model : ; in and are the corner points of the teacher model and Features, and are the corner points of the student model and Features, is the standard deviation of the Gaussian function; S430: Compare the distillation calculation relationship distillation loss, specifically: 。 5. The method according to claim 4, characterized in that S500 is specifically: ; in and Represents the BEV features of the teacher-student model, and the difference is calculated by the L2 norm as the global distillation loss.

6. The method according to claim 5, characterized in that S600 includes: S610: Count the number of filters in the student model and assign an index to each filter; S620: setting filters to zero one by one, testing their influence on detection accuracy, and recording the change in accuracy. In this way, all filters are processed one by one, and finally a TXT file containing the importance of each filter is generated, wherein the smaller the value of the change in accuracy is, the lower the importance of the filter is; S630: Read and sort the importance of filters in the TXT file in ascending order, and update the corresponding filter index; S640: removing filters with lower importance according to a set pruning rate; S650: Adjust the channel structure of the student model according to the number of retained filters and retrain to optimize the weights on this basis to achieve the best performance.

7. The method according to claim 6, characterized in that S700 includes: S710: Reconstruct the PointPillar model and convert it to ONNX format. During this process, the PointPillar model is divided into two parts, which are finally spliced ​​together. The first part includes the PFE module and the Scatter module, which are responsible for extracting the features of the column. The second part includes the Backbone module and the Densehead module, which are responsible for extracting the deep features of the column and making predictions. S720: Adjust or customize operators using custom operators according to the ONNX support library to make the ONNX format compatible with all operators in the PointPillar model; S730: Optimize the ONNX model, convert the ONNX model into the TensorRT acceleration format engine file, use the engine file to test the input data, and complete the deployment.

8. A power inspection robot detection and deployment system, characterized in that: It includes data acquisition module, BEV feature extraction module, focus distillation module, focus relationship distillation module, global distillation module, model pruning training module and model deployment module. The data acquisition module is used to collect two-dimensional image data of the road to be inspected using a camera in an open road environment, obtain corresponding three-dimensional point cloud data through high-beam laser radar and low-beam laser radar, and save calibration files and true value annotation files; The BEV feature extraction module is used to input the high-beam 3D point cloud data into the teacher model and the low-beam 3D point cloud data into the student model, and respectively extract the 3D features of the high-beam 3D point cloud and the low-beam 3D point cloud through the 3D backbone network in the model, and respectively input the 3D features into the 2D backbone network for conversion to obtain BEV features, and then obtain the prediction results, wherein the teacher model is trained separately for the high-beam point cloud to obtain pre-weights for sharing by the two models; The focus distillation module is used to read the true value annotation file to obtain the foreground 3D coordinates, project the foreground 3D coordinates to the 2D BEV coordinates, set the mask to distinguish the foreground features from the background features, and use the foreground mask and the background mask to weight the BEV feature losses of the teacher model and the student model to obtain the focus loss; The focus relation distillation module is used to extract the nine corner point features of the prediction results of the teacher model and the student model, calculate the Gaussian similarity of the corner points of the teacher model and the student model, and perform comparative distillation to obtain the relation distillation loss; The global distillation module is used to calculate the difference between the BEV features of each sample of the teacher model and the student model to obtain the global distillation loss; The model pruning training module is used to train the student model based on the focal loss, relational distillation loss, global distillation loss and calibration file until the total loss value converges to obtain the trained student model; perform model pruning on the trained student model, calculate the importance of each filter in the trained student model, remove unimportant filters according to a certain ratio, reconstruct the network and retrain the pruned student model to obtain the optimized student model; The model deployment module is used to convert the optimized student model into ONNX format, deploy and accelerate it on TensorRT, and test the acquired real-time input data.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Knowledge distillation method and device for image classification model and computer equipment

    CN112232397A

  • Power transmission channel engineering vehicle target detection method and system based on knowledge distillation

    CN115131747A