Power Equipment Detection, Ranging and Warning Method Based on Artificial Intelligence and Monocular Vision
Through monocular vision technology based on artificial intelligence, combined with example segmentation and depth estimation, the low-cost problem of three-dimensional distance measurement and early warning of power equipment is solved, and the automatic measurement and early warning of safe distance between power equipment and operation and maintenance personnel is realized to avoid electric shock accidents.
Patent Information
- Application Number
- CN202210946142.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-08
AI Technical Summary
It is difficult for the existing technology to achieve low-cost, convenient and effective three-dimensional ranging and early warning around power equipment. Binocular vision and lidar have problems such as large size, large calculation volume and expensive in actual applications, and it is impossible to monitor and early warning the safe distance of staff in real time.
Using an artificial intelligence and monocular vision method, combined with instance segmentation, depth estimation, depth reconstruction and back projection technology, the three-dimensional distance prediction and safety warning between power equipment and operation and maintenance personnel is realized through the improved SOLOv2 model and Diversedepth network.
It realizes refined mask segmentation and automatic distance measurement between power equipment and operation and maintenance personnel, and can be promptly warned, avoid electric shock accidents and reduce casualties.
Smart Images

Figure CN115439741B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power equipment safety monitoring, and relates to a method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision. Background Art
[0002] The safety monitoring of staff is crucial in the operation and construction of power systems. Although there are generally warning signs, fences or anti-electric shock measures around live high-voltage equipment, this method is too passive and may not be able to give warnings in time due to reasons such as damage to protection measures or mental slack of staff. This will not only cause huge economic losses but also result in casualties. Therefore, it is particularly important to monitor the safety distance of the staff around substation equipment in real time and give warnings in time.
[0003] In recent years, the rapid development of computer vision technology and deep learning methods has greatly improved the efficiency of target recognition and detection, and has become a research hotspot in visual detection in the power field. However, most of the research on visual detection of power equipment mainly focuses on the detection of two-dimensional targets. Due to the lack of information in one dimension, it cannot be directly used for the detection and ranging of targets in three-dimensional space.
[0004] The current technologies for obtaining three-dimensional information of a scene are mainly realized through binocular vision and lidar (LiDAR). In actual power scenarios, binocular cameras have problems such as large volume, large amount of computation required for stereo matching, and the predicted distance being limited by the binocular baseline. When the texture of the image is not obvious enough, it is difficult to capture enough features for binocular matching; as an active measurement technology, LiDAR is not suitable for real-time monitoring tasks and is extremely expensive. Therefore, a low-cost, easy-to-deploy and convenient and effective three-dimensional ranging scheme is needed for the measurement and warning of the safety distance of live equipment. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision. By combining instance segmentation, depth estimation, depth reconstruction and back-projection technology, the automatic prediction of the three-dimensional distance between power equipment and operation and maintenance personnel is completed, and the refined power equipment and "person" target masks are segmented to realize the automatic measurement of the safety distance of live equipment and safety warning, further avoiding electric shock accidents and reducing casualties.
[0006] To solve the above technical problem, the technical solution adopted by the present invention is: a method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision, which includes the following steps:
[0007] S1, Acquisition: Acquire images of power equipment and operation and maintenance personnel, preprocess the obtained images, and form an effective image dataset with the extracted COCO-person data;
[0008] S2, Construction: Construct an improved SOLOv2 instance segmentation model for detecting and segmenting refined power equipment and personnel masks;
[0009] S3, Mapping: Based on the inverse projection transformation of the camera, use the Diversedepth network and the point cloud encoder network to predict depth values and virtual focal lengths, and obtain the 2D-3D mapping relationship;
[0010] S4, Prediction: Combine the 2D mask and the mapping relationship to obtain the 3D point cloud of the target, and complete the prediction of the minimum distance between power equipment and people according to the conversion ratio of virtual distance and actual distance;
[0011] S5, Judgment: Automatically judge whether personnel are in danger and give an alarm according to the regulations on the safe distance of high-voltage equipment in the standard;
[0012] S6, Warning: Based on the structural characteristics of energized equipment, use the predicted minimum distance and triangular geometric relationship to calculate the distance between the energized part of the equipment and people, and then judge whether personnel are in danger and give an alarm according to the regulations on the safe distance between people and energized bodies in the enterprise standards of China Southern Power Grid, so as to realize real-time ranging and safety warning of energized equipment.
[0013] In S1, preprocessing the images includes screening the original power equipment images, performing data augmentation operations such as mirror symmetry on the images, and using the EISeg tool to perform mask annotation on power equipment and "person" targets; The COCO-person data is composed of part of the data containing "person" targets separately stripped from the COCO_val2014 dataset, including images and labels.
[0014] In S2, improve and optimize the backbone network, neck, and detection head of the SOLOv2 model:
[0015] S2-1, Introduce three structures, Res2Net, ResNet-c, and ResNet-d, into the feature extraction network ResNet50 to optimize the model, so as to improve the model's feature extraction ability and reduce the computational amount; Among them, the ResNet-c and ResNet-d structures reduce the weight parameters of the network and the information loss of the feature map, and the Res2Net structure further expands the receptive field of the network layer;
[0016] S2-2. Replace the original feature fusion module FPN with a more performant weighted bidirectional feature pyramid network BiFPN structure to further enhance the model's ability to fuse multi-scale features;
[0017] S2-3. Introduce deformable convolution DCNv2 into the SOLO detection head. DCNv2 improves the model's ability to model the geometric deformation of objects, enabling the model to more accurately predict the regions of objects.
[0018] Use transfer learning to train the improved SOLOv2 model on the constructed dataset. The training loss function includes the following two types: classification loss and mask loss:
[0019] ;
[0020] Where:
[0021] ;
[0022] In the formula, L cls is the Focal loss function for semantic class classification, L mask is the mask prediction loss based on the Dice loss function, is the weight factor of the c -th class sample, used to balance the imbalance between positive and negative samples, with a value range of [0, 1], γ is the coefficient for adjusting the calculation weight of easy and hard samples, and its value is greater than or equal to 0, represents the predicted probability of the c -th class sample output using the Softmax activation function, n p is the number of positive samples, is the instance class score at the position of the original image ( i , j ), f is the indicator function, when > 0, f takes 1, otherwise takes 0, P k and G k respectively represent the predicted pixel matrix and the true pixel matrix of the mask k , L Dice represents the Dice loss function, and D ( p , q ) represents the Dice coefficient corresponding to the matrices p , q , and ( x , y ) represents the original image (i , j ), the corresponding feature map coordinates p x,y and q x,y are the predicted mask pixel value and the true mask pixel value for the position ( x , y ) in the feature map; after the model is trained, the test set is used to test the model.
[0023] In S3, the back-projection transformation of the camera is deduced from the following camera coordinate system conversion relationship:
[0024] ;
[0025] where, ( X w , Y w , Z w ) are the world coordinate system coordinates, ( X c , Y c , Z c ) are the camera coordinate system coordinates, ( x, y ) are the image coordinate system coordinates, ( u, v ) are the pixel coordinate system coordinates, the rotation matrix R and the translation vector T are the external parameters of the camera, f x , f y , u 0 , v 0 is the internal parameter of the camera, where, f is the camera focal length in mm, ( dx dy are the physical sizes of one pixel in the image coordinate axes x, y directions), f x =f / dx , it is the focal length of the camera on the x axis, f y =f / dy , it is the focal length of the camera on the y axis, ( u 0 , v 0) are the pixel coordinates corresponding to the camera optical center.
[0026] In monocular vision, it is assumed that the world coordinate system coincides with the camera coordinate system, that is, rotation ( R ) and translation ( T ) operations are not considered, and the external parameter matrix in the matrix operation is ignored M2. Then calculate the conversion formula between the world coordinates ( X w , Y w , Z w ) and the pixel coordinates ( u ,v ), that is, the back-projection transformation formula is:
[0027] ;
[0028] In the formula, Z c is the depth value (depth of field).
[0029] Based on the back-projection transformation formula, to complete the 2D-3D projection conversion, it is necessary to first obtain the depth information d p and the focal length f of the 2D image, use the big data-driven Diversedepth depth estimation model to predict the depth value of the 2D image, and combine the results of Diversedepth and the point cloud encoder network to further predict the virtual focal length f of the camera.
[0030] Automatically generate evenly distributed ranging points on the 2D mask, convert the ranging points into 3D coordinates in the virtual coordinate system according to the 2D-3D mapping relationship, and calculate the minimum virtual distance between the power equipment and the "person" target according to the following three-dimensional Euclidean distance formula:
[0031] .
[0032] In S4, the measurement of the actual distance between the power equipment and the operation and maintenance personnel includes the following steps:
[0033] S4-1. In each new live equipment monitoring scenario, select the height of the operation and maintenance personnel as the known reference quantity for the conversion of the actual distance and the virtual relative distance in this scenario, and calculate the fixed ratio K of the actual height of the personnel to the height in the virtual coordinate system,
[0034] ;
[0035] In the formula, dt act_h is the actual height of the operation and maintenance personnel, dt pse_h is the height of the personnel in the virtual coordinate system, and K is the fixed ratio between the two;
[0036] S4-2: Use the virtual distance between the live equipment and the operation and maintenance personnel obtained above to calculate its actual minimum distance. The calculation formula is as follows:
[0037] ;
[0038] wherein, dt pse is the virtual distance between the device and the personnel.
[0039] In S5, the safety status of the operation and maintenance personnel is directly judged in combination with the minimum distance between the power equipment and the person and the standard, and a warning is given in a timely manner;
[0040] In S6, the distance and angle information from the ranging point on the device to the energized part are measured, and the distance from the person to the energized part is calculated by using the triangular geometric relationship and the measured distance above; the safety status of the operation and maintenance personnel is judged according to the standards of the power grid enterprise and a warning is given.
[0041] The main beneficial effects of the present invention are as follows:
[0042] Through technologies such as instance segmentation, depth estimation, depth reconstruction, and inverse projection transformation, 3D ranging of the target in the 2D image is realized.
[0043] An optimized instance segmentation model is used to obtain the masks of the power equipment and the personnel. The model improves SOLOv2 from three aspects: feature extraction, feature fusion, and the ability to handle target geometric transformations:
[0044] Three structures, Res2Net, ResNet-c, and ResNet-d, are introduced into the feature extraction network ResNet50 to optimize the model, so as to improve the feature extraction ability of the model and reduce the calculation amount; among them, the ResNet-c and ResNet-d structures can reduce the weight parameters of the network and the information loss of the feature map, and the Res2Net structure can further expand the receptive field of the network layer;
[0045] The original feature fusion module FPN is replaced with a weighted bidirectional feature pyramid network (BiFPN) structure with better performance to further improve the model's ability to fuse multi-scale features;
[0046] Deformable Convolution (DCNv2) is introduced into the SOLO detection head. DCNv2 can improve the model's ability to model the geometric deformation of the target and make the model more accurately predict the region of the target.
[0047] A depth map and the camera focal length are obtained by using depth estimation and the point cloud encoder, and the 2D-3D mapping relationship is obtained by combining inverse projection.
[0048] With the help of the 2D mask, the 3D coordinates of the target can be calculated, and further the relative distance between the power transformer and the personnel in the virtual coordinate system can be obtained.
[0049] Under each new monitoring scenario, with reference to the actual height of a person, a fixed conversion ratio of the actual-virtual distance of the monitoring scenario is obtained, and the actual distance of the target is calculated.
[0050] The present invention can segment refined power equipment and "person" target masks, realize automatic ranging and safety warning between live equipment and operation and maintenance personnel, and further avoid electric shock accidents and reduce casualties. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The present invention will be further described below in conjunction with the drawings and embodiments.
[0052] Figure 1 It is a flowchart of the embodiment of the present invention;
[0053] Figure 2 It is a schematic diagram of the idea of the detection and ranging method of the embodiment of the present invention;
[0054] Figure 3 It is one of the backbone network structure diagrams of the embodiment of the present invention;
[0055] Figure 4 It is the training loss diagram of the improved SOLOv2 model of the embodiment of the present invention;
[0056] Figure 5 and Figure 6 It is the detection result diagram of Mask R-CNN of the embodiment of the present invention;
[0057] Figure 7 and Figure 8 It is the detection result diagram of Cascade Mask R-CNN of the embodiment of the present invention;
[0058] Figure 9 and Figure 10 It is the detection result diagram of the SOLOv2 model of the embodiment of the present invention;
[0059] Figure 11 It is the operation flow chart of the 3D ranging module of the embodiment of the present invention;
[0060] Figure 12 It is the pinhole camera imaging model diagram of the embodiment of the present invention;
[0061] Figure 13 It is the original image with pixel coordinate scales of the embodiment of the present invention;
[0062] Figure 14 It is the instance segmentation result diagram of the embodiment of the present invention;
[0063] Figure 15 It is the depth estimation effect diagram of the embodiment of the present invention;
[0064] Figure 16 Automatic ranging point selection diagram of the embodiment of the present invention;
[0065] Figure 17 Pseudo-laser point cloud diagram after depth reconstruction of the embodiment of the present invention;
[0066] Figure 18 Laser point cloud ranging diagram of the embodiment of the present invention;
[0067] Figure 19 Diagram of the position relationship between the energized part of the power equipment and people in the embodiment of the present invention; Detailed implementation manners
[0068] As Figures 1 to 19 In [reference], a power equipment detection, ranging and early warning method based on artificial intelligence and monocular vision includes the following steps:
[0069] S1, Acquisition: Acquire images of power equipment and operation and maintenance personnel, preprocess the acquired images, and form an effective image dataset with the extracted COCO-person data;
[0070] S2, Construction: Construct an improved SOLOv2 instance segmentation model for detecting and segmenting refined power equipment and personnel masks;
[0071] S3, Mapping: Based on the inverse projection transformation of the camera, use the Diversedepth network and the point cloud encoder network to predict the depth value and virtual focal length, and obtain the 2D-3D mapping relationship;
[0072] Preferably, as shown in [[reference]], it mainly consists of two parts: 2D instance segmentation and 3D ranging. Figure 2 Shown, it mainly consists of two parts: 2D instance segmentation and 3D ranging.
[0073] S4, Prediction: Combine the 2D mask and the mapping relationship to obtain the 3D point cloud of the target, and complete the prediction of the minimum distance between the power equipment and people according to the conversion ratio between the virtual distance and the actual distance;
[0074] S5, Judgment: Automatically judge whether the personnel are in danger and give an early warning according to the regulations on the safe distance of high-voltage equipment in the standard;
[0075] Preferably, the above standard preferably selects the Chinese national standard.
[0076] S6, Early warning: Based on the structural characteristics of the energized equipment, use the predicted minimum distance and the triangular geometric relationship to calculate the distance between the energized part of the equipment and people, and then judge whether the personnel are in danger and give an early warning according to the regulations on the safe distance between people and energized bodies in the enterprise standard of the Southern Power Grid, so as to realize real-time ranging and safety early warning of the energized equipment.
[0077] In a preferred solution, in S1, preprocessing the image includes screening the original power equipment images, performing data augmentation operations such as mirror symmetry on the images, and using the EISeg tool to perform mask annotation on the power equipment and "person" targets; the COCO-person data is composed of part of the data containing "person" targets separately stripped from the COCO_val2014 dataset, including images and labels.
[0078] Preferably, preprocessing the image includes screening the original power equipment images, performing data augmentation operations such as mirror symmetry on the images, and using the EISeg tool to perform mask annotation on the power equipment and "person" targets. The main power equipment selected is the power transformer.
[0079] In a preferred solution, in S2, the backbone network, neck, and detection head of the SOLOv2 model are improved and optimized:
[0080] S2-1, introducing three structures, Res2Net, ResNet-c, and ResNet-d, into the feature extraction network ResNet50 to optimize the model, so as to improve the model's feature extraction ability and reduce the computational amount; among them, the ResNet-c and ResNet-d structures reduce the weight parameters of the network and the information loss of the feature map, and the Res2Net structure further expands the receptive field of the network layer;
[0081] Preferably, the original residual network structure consists of a basic input block C1 and four consecutive residual blocks C2-C5. The improved backbone network is as Figure 3 shown: A. Introducing the ResNet-c and ResNet-d structures, replacing the original 7×7 convolutional kernel in C1 with three consecutive 3×3 convolutional kernels, changing the stride of the 1×1 convolutional kernel in the downsampling structure, and also adding an average pooling layer in path B. The introduction of ResNet-c and ResNet-d allows the model to make full use of feature information and reduce memory occupancy. B. Introducing the Res2Net structure with four feature groups into the optimized backbone network, replacing the 3×3 convolutional kernel of each residual block in C2-C5 with a hierarchical cascaded feature group convolution. The design of hierarchical convolution expands the receptive field of each network layer, enabling more refined capture of local and global features. In addition, the powerful feature retention mechanism of Res2Net allows the network to retain the integrity of information even under deeper network structures, and provides better robustness and generalization ability. The optimized backbone network is called Res2Net50-cd.
[0082] S2-2. Replace the original feature fusion module FPN with a more performant weighted bidirectional feature pyramid network BiFPN structure to further enhance the model's ability to fuse multi-scale features;
[0083] Preferably, the pyramid network is a Bi-directional feature pyramid network; BiFPN introduces a fast normalization fusion strategy to learn the weights of different input features, removes feature nodes that contribute little to the model, and uses a bidirectional cross-scale connection method to mitigate the problem of feature information loss caused by deepening the network layers, and further enhances the model's ability to fuse multi-scale features.
[0084] S2-3. Introduce deformable convolution DCNv2 into the SOLO detection head. DCNv2 improves the model's ability to model the geometric deformation of targets, enabling the model to more accurately predict the regions of targets.
[0085] Preferably, DCNv2 ensures the effective extraction of target information by the model and improves the model's ability to model the geometric deformation of targets. Adding deformable convolution to the detection head is not only used for feature extraction but also for generating offsets and weights to accurately predict the positions of targets and classify targets according to the position regions. In the same power scenario, the size ratio between "person" and power equipment varies greatly, and the scale-transformable deformable convolution of DCNv2 can better predict their positions.
[0086] In a preferred solution, transfer learning is used to train the improved SOLOv2 model on the constructed dataset. The loss function for training includes the following two types: classification loss and mask loss:
[0087] ;
[0088] Where:
[0089] ;
[0090] In the formula, L cls is the Focal loss function for semantic class classification, L mask is the mask prediction loss based on the Dice loss function, is the weight factor for the c -th class of samples, used to balance the imbalance between positive and negative samples, with a value range of [0, 1], γ is the coefficient for adjusting the calculation weights of easy and hard samples, and its value is greater than or equal to 0, represents the predicted probability of the c -th class of samples output using the Softmax activation function, np is the number of positive samples, is the instance category score at the position of the original image ( i , j ), f is the indicator function, when > 0, f it takes 1, otherwise it takes 0, P k and G k respectively represent the predicted pixel matrix and the true pixel matrix of the mask k , L Dice represents the Dice loss function, D ( p , q ) represents the Dice coefficient corresponding to the matrix p , q , ( x , y ) represents the coordinates of the feature map corresponding to the position of the original image ( i , j ); p x,y and q x,y are the predicted mask pixel value and the true mask pixel value at the position ([[]] x , y ) in the feature map; after the model is trained, the test set is used to test the model.
[0091] Preferably, the training process uses the SGD optimizer with a momentum of 0.9 and L2 regularization with a weight decay factor of 0.0001 to optimize and improve the training of the SOLOv2 model, which helps to prevent overfitting during training.
[0092] Preferably, according to the GPU performance and model characteristics used, the batch size is set to 2 and trained for 100 epochs.
[0093] Preferably, the training is decayed in two segments with a decay factor of 0.1, and the initial learning rate is 0.005. Bilinear interpolation is also used to scale the images to enhance the training images.
[0094] Preferably, the loss curve of the training is as Figure 4 shown, and the training process quickly reaches the convergence state.
[0095] Preferably, the qualitative comparison result graph of the improved SOLOv2 segmentation model and other models is shown.
[0096] Preferably, compare Figure 5 , Figure 7 , Figure 9, where the solid circles are the areas of power equipment blocked by the pipeline and the operation and maintenance personnel respectively. It can be seen that the Mask R-CNN and Cascade Mask R-CNN models identify the pipeline as a transformer. More unacceptably, since the head of the operation and maintenance personnel blocks the transformer, they also identify the head of "person" as part of the transformer. This is undoubtedly an unacceptable defect for the next step of ranging. The improved SOLOv2 can predict high-quality "person" and transformer masks to support the ranging of the 3DPOR module. Comparison sub Figure 6 , Figure 8 , Figure 10 , where the dashed box is the sign outside the transformer.
[0097] Preferably, the Mask R-CNN and Cascade Mask R-CNN models not only fail to well predict the masks of the edge part of the transformer, but also output the sign as a mask, while the improved SOLOv2 is not affected by the occlusion of the sign. It is verified that the 2D SOLO model for object segmentation is robust to the occlusion problem. When testing the entire test set, with the IOU thresholds of 0.5 and 0.75, the mean average precision (mAP) is 0.749 and 0.579 respectively. Under the strict IoU threshold, the mAP of the proposed method 50:95 reaches more than 53%.
[0098] In the preferred solution, the operation process of the 3D ranging module is as shown in 6. In S3, the back-projection transformation of the camera is deduced from the following camera coordinate system conversion relationship:
[0099] ;
[0100] where, ( X w , Y w , Z w ) are the coordinates in the world coordinate system, ( X c , Y c , Z c ) are the coordinates in the camera coordinate system, ( x, y ) are the coordinates in the image coordinate system, ( u, v ) are the coordinates in the pixel coordinate system, the rotation matrix R and the translation vector T are the external parameters of the camera, f x , f y , u 0 , v 0 is the internal parameter of the camera, where, fwhere the camera focal length is in mm, ( dx dy are the physical sizes of one pixel in the x, y directions of the image coordinate axes), f x =f / dx , which is the focal length of the camera on the x axis, f y =f / dy , which is the focal length of the camera on the y axis, ( u 0 , v 0) is the pixel point coordinate corresponding to the camera optical center.
[0101] In monocular vision, it is assumed that the world coordinate system coincides with the camera coordinate system, that is, rotation ( R ) and translation ( T ) operations are not considered, and the external parameter matrix M 2 in the matrix operation is ignored. Then the conversion formula between the world coordinates ( X w , Y w , Z w ) and the pixel coordinates ( u ,v ) in monocular vision is calculated, that is, the back-projection transformation formula is:
[0102] ;
[0103] In the formula, Z c is the depth value (depth of field).
[0104] In a preferred solution, based on the back-projection transformation formula, to complete the 2D-3D projection conversion, it is necessary to first obtain the depth information d p of the 2D image and the focal length f , use the big data-driven Diversedepth depth estimation model to predict the depth value of the 2D image, and combine the results of Diversedepth and the point cloud encoder network to further predict the virtual focal length f of the camera.
[0105] In a preferred solution, uniformly distributed ranging points are automatically generated on the 2D mask, the ranging points are converted into 3D coordinates in the virtual coordinate system according to the 2D-3D mapping relationship, and the minimum virtual distance between the power equipment and the "person" target is calculated according to the following three-dimensional Euclidean distance formula:
[0106] .
[0107] In a preferred solution, in S4, the measurement of the actual distance between the power equipment and the operation and maintenance personnel includes the following steps:
[0108] S4-1. In each new live equipment monitoring scenario, select the height of the operation and maintenance personnel as the known reference quantity for the conversion of the actual distance and the virtual relative distance in this scenario, and calculate the fixed ratio K between the actual height of the personnel and the height in the virtual coordinate system.
[0109] ;
[0110] In the formula, dt act_h is the actual height of the operation and maintenance personnel, dt pse_h is the height of the personnel in the virtual coordinate system, and K is the fixed ratio between the two;
[0111] S4-2: Use the virtual distance between the live equipment and the operation and maintenance personnel obtained above to calculate its actual minimum distance. The calculation formula is as follows:
[0112] ;
[0113] In the formula, dt pse is the virtual distance between the equipment and the personnel.
[0114] Preferably, various visualized results output in this step are as Figures 13 to 17 shown. According to the actual scenario requirements, formulate rules to automatically generate the 2D ranging points of the target: Select 800 non-edge points in the transformer mask and 20 upper points in the "person" mask as the ranging points. And 10 images with distance label information are selected to test this step. The results show that the maximum error rate of ranging is less than 9%, and the average error rate is 5.337%.
[0115] In a preferred solution, in S5, combine the minimum distance between the power equipment and the person and the standard to directly judge the safety status of the operation and maintenance personnel and give an early warning in a timely manner;
[0116] In S6, as Figure 19 shown, measure the distance and angle information from the ranging points on the measuring equipment to the live part, and use the trigonometric geometric relationship and the above-measured distance to calculate the distance from the person to the live part; make a judgment on the safety status of the operation and maintenance personnel according to the grid enterprise standard and give an early warning.
[0117] The above embodiments are only the preferred technical solutions of the present invention and should not be regarded as limitations on the present invention. In the present application, the embodiments and the features in the embodiments can be arbitrarily combined with each other without conflict. The protection scope of the present invention shall be the technical solutions recorded in the claims, including the equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, the equivalent replacement improvements within this scope are also within the protection scope of the present invention.
Claims
1. A method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision, characterized in that, It includes the following steps: S1, Acquisition: Acquire images of power equipment and operation and maintenance personnel, preprocess the acquired images, and form an effective image dataset with the extracted COCO-person data; S2, Construction: Construct an improved SOLOv2 instance segmentation model for detecting and segmenting refined power equipment and personnel masks; S3, Mapping: Based on the inverse projection transformation of the camera, use the Diversedepth network and the point cloud encoder network to predict depth values and virtual focal lengths, and obtain the 2D-3D mapping relationship; S4, Prediction: Combine the 2D mask and the mapping relationship to obtain the 3D point cloud of the target, and complete the prediction of the minimum distance between power equipment and people according to the conversion ratio of virtual distance and actual distance; S5, Judgment: Automatically judge whether the personnel are in danger and give an early warning according to the regulations on the safe distance of high-voltage equipment in the standard; S6, Early warning: Based on the structural characteristics of the energized equipment, use the predicted minimum distance and triangular geometric relationship to calculate the distance between the energized part of the equipment and people, and then judge whether the personnel are in danger and give an early warning according to the regulations on the safe distance between people and energized bodies in the enterprise standard of Southern Power Grid, so as to realize real-time ranging and safety early warning of energized equipment; In S2, the backbone network, neck, and detection head of the SOLOv2 model are improved and optimized: S2-1, Introduce three structures, Res2Net, ResNet-c, and ResNet-d, into the feature extraction network ResNet50 to optimize the model, so as to improve the model's feature extraction ability and reduce the computational amount; among them, the ResNet-c and ResNet-d structures reduce the weight parameters of the network and the information loss of the feature map, and the Res2Net structure further expands the receptive field of the network layer; S2-2, Replace the original feature fusion module FPN with a more performant weighted bidirectional feature pyramid network BiFPN structure to further improve the model's ability to fuse multi-scale features; S2-3, Introduce deformable convolution DCNv2 into the SOLO detection head. DCNv2 improves the model's ability to model the geometric deformation of the target and enables the model to more accurately predict the region of the target.
2. The method for detecting, ranging, and warning of power equipment based on artificial intelligence and monocular vision according to claim 1, characterized in that: In S1, preprocessing the images includes screening the original power equipment images, performing data augmentation operations such as mirror symmetry on the images, and using the EISeg tool to perform mask annotation on the power equipment and "person" targets; The COCO-person data is composed of part of the data containing "person" targets separately stripped from the COCO_val2014 dataset, including images and labels.
3. The method for detecting, ranging, and warning of power equipment based on artificial intelligence and monocular vision according to claim 1, characterized in that: Use transfer learning to train the improved SOLOv2 model on the constructed dataset, and the loss function for training includes the following two types of classification loss and mask loss: ; Where: ; Wherein, L cls is the Focal loss function for semantic class classification, L mask is the mask prediction loss based on the Dice loss function, is the weight factor for the c -th class of samples, used to balance the imbalance between positive and negative samples, and the value range is [0, 1], γ is the coefficient for adjusting the calculation weight of easy and difficult samples, and its value is greater than or equal to 0, represents the predicted probability of the c -th class of samples output by using the Softmax activation function, n p is the number of positive samples, is the original image ( i , j ) instance class score at the position, f is the indicator function, when > 0, f takes 1, otherwise takes 0, P k and G k respectively represent the predicted pixel matrix and the true pixel matrix of the mask k , L Dice represents the Dice loss function, D( p , q ) represents the Dice coefficient corresponding to the matrices p , q , ([[]] x , y ) represents the coordinates of the feature map corresponding to the position of the original image ( i , j ), p x,y and q x,y are the predicted mask pixel value and the true mask pixel value at the position ([[]] x , y ) in the feature map; after the model is trained, the test set is used to test the model.
4. The method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision according to claim 1 is characterized in that: In S3, the back-projection transformation of the camera is deduced from the following camera coordinate system conversion relationship: ; Among them, ( X w , Y w , Z w ) are the coordinates in the world coordinate system, ( X c , Y c , Z c ) are the coordinates in the camera coordinate system, ( x, y ) are the coordinates in the image coordinate system, ( u, v ) are the coordinates in the pixel coordinate system. The rotation matrix R and the translation vector T are the external parameters of the camera. f x , f y , u 0 , v 0 are the internal parameters of the camera. Among them, f is the camera focal length in mm. dx, dy are respectively the physical sizes of one pixel in the x, y directions of the image coordinate axes. f x = f / dx , which is the focal length of the camera on the x axis. f y = f / dy , which is the focal length of the camera on the y axis. ( u 0 , v 0) are the pixel coordinates corresponding to the camera optical center; Under monocular vision, it is assumed that the world coordinate system coincides with the camera coordinate system, that is, rotation R and translation T operations are ignored, and the external parameter matrix in matrix operations is M ignored; then the conversion formula between the world coordinates ( X w , Y w , Z w ) and pixel coordinates ( u, v ) under monocular vision is calculated. The inverse projection transformation formula is: ; Wherein, Z c is the depth value, i.e., the depth of field.
5. The method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision according to claim 1 is characterized in that: Based on the back-projection transformation formula, to complete the 2D-3D projection conversion, it is necessary to first obtain the depth information of the 2D image d p and the focal length f , use the big data-driven Diversedepth depth estimation model to predict the depth value of the 2D image, and combine the results of Diversedepth and the point cloud encoder network to further predict the virtual focal length of the camera f .
6. The method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision according to claim 1 is characterized in that: Uniformly distributed ranging points are automatically generated on the 2D mask, the ranging points are converted into 3D coordinates in the virtual coordinate system according to the 2D-3D mapping relationship, and the minimum virtual distance between the power equipment and the "person" target is calculated according to the following three-dimensional Euclidean distance formula: 。 7. The method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision according to claim 1 is characterized in that: In S4, the measurement of the actual distance between the power equipment and the operation and maintenance personnel includes the following steps: S4-1. In each new live equipment monitoring scenario, the height of the operation and maintenance personnel is selected as the known reference quantity for the conversion of the actual distance and the virtual relative distance in this scenario, and the fixed ratio K of the actual height of the personnel to the height in the virtual coordinate system is calculated. ; Wherein, dt act_h is the actual height of the operation and maintenance personnel, dt pse_h is the height of the personnel in the virtual coordinate system, and K is the fixed ratio between the two; S4-2: Using the virtual distance between the live equipment and the operation and maintenance personnel, calculate its actual minimum distance, and the calculation formula is as follows: ; Wherein, dt pse is the virtual distance between the device and the personnel.
8. The method for detecting, ranging and warning of power equipment based on artificial intelligence and monocular vision according to claim 1 is characterized in that: In S5, combining the minimum distance between the power equipment and the person and the standard, directly judge the safety status of the operation and maintenance personnel and give a warning in time; In S6, measure the distance and angle information from the ranging point on the measuring device to the live part, and use the trigonometric geometric relationship and the distance from the ranging point on the measuring device to the live part to calculate the distance from the person to the live part; judge the safety status of the operation and maintenance personnel according to the grid enterprise standard and give a warning.