Electric power engineering near-power operation supervision method based on multimode data target detection
Through the multi-mode data correlation and feature fusion method, combined with lidar, millimeter-wave radar and camera data, high-precision near-electric operation supervision in complex environments is achieved, solving the problems of weak radar semantic information capabilities and lidar sensitive to weather conditions in the existing technology.
Patent Information
- Application Number
- CN202510059341.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-23
AI Technical Summary
In the supervision of near-electric operation in complex environments, the radar semantic information capability is weak, and the lidar is sensitive to weather conditions, resulting in a decrease in detection accuracy.
Using a multimode data-based target detection method, data is collected through lidar, millimeter wave radar and camera, point clouds and images are correlated using transformation matrix and internal reference matrix, point cloud columns are divided for feature extraction and fusion, and target detection and distance comparison are combined with multimodal target detection network to achieve near-electric operation warnings.
It improves the target detection accuracy and robustness in complex environments, and enhances the efficiency and reliability of near-electric operation supervision.
Smart Images

Figure CN120032215A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of supervision of near-electrical operations, and in particular to a supervision method of near-electrical operations of electric power engineering projects based on multi-mode data target detection. Background Art
[0002] The safety work regulations for power construction have strict requirements on the safe distance of near-electrical operations at different voltage levels. In order to ensure that personnel and equipment working near electric power maintain a sufficient safe distance from surrounding live equipment, the work site needs to repeatedly plan the work area and conduct near-electrical operations control, which is inefficient.
[0003] In the prior art, there are solutions for achieving near-electric operation control based on radar target detection tasks. For near-electric operation control, it is necessary to accurately obtain the position of the target through radar and to perceive the target's motion state information. Lidar can provide accurate position measurement capabilities because it can better measure the shape of objects, but lidar is insensitive to motion and has weak ability to detect physical motion states. Single-mode lidar data cannot achieve near-electric operation control with high safety standards. To achieve better near-electric operation control, the bidirectional lidar and millimeter-wave radar fusion model combines lidar and millimeter-wave radar, which is more sensitive to motion, to detect targets. The bidirectional lidar and millimeter-wave radar fusion model encodes each mode into a bird's-eye view feature separately. Then, it uses query-based lidar to millimeter-wave radar height feature fusion and query-based lidar to millimeter-wave radar bird's-eye view feature fusion. In these two steps, we query and group lidar points and lidar bird's-eye view features close to each non-empty grid cell position on the millimeter-wave radar feature map. The grouped lidar raw points are aggregated to generate pseudo mmWave radar height features, while the grouped lidar bird's-eye view features are aggregated to generate pseudo mmWave radar bird's-eye view features. The generated pseudo mmWave radar height and bird's-eye view features are fused with the mmWave radar bird's-eye view features through splicing to enhance the mmWave radar features, and then the bidirectional lidar and mmWave radar fusion model performs the fusion of the mmWave radar bird's-eye view representation to the lidar bird's-eye view representation in a unified lidar and mmWave radar bird's-eye view representation. Finally, the bird's-eye view detection network consisting of the bird's-eye view backbone network and the detection head is applied to output the 3D object detection results.
[0004] However, the actual site of substation near-electrical operations has problems such as random areas, complex environments, and a large number of equipment types. In complex environments, semantic information is complex and radar semantic information capabilities are weak. In complex environments, the radar waves of the desired target will be interfered with by certain noises, affecting accuracy. In addition, near-electrical operations are supervised outdoors and will be affected by weather factors such as fog, rain, and snow. However, lidar shows significant sensitivity to weather conditions. In harsh scenarios, the data point cloud of the lidar will seriously decay and introduce a lot of noise. During the encoding process, the noise generated by weather factors is encoded layer by layer into the deep layer, which introduces a lot of noise that affects target detection into the fusion features. In the scheme of first extracting deep features and then fusing them, data degradation on either side has a great negative impact on target detection and affects the accuracy of near-electrical operation supervision. Summary of the invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a method for supervising near-electrical operations in electric power engineering based on multi-mode data target detection.
[0006] In a first aspect, the present invention provides a method for supervising near-electrical work of an electric power engineering project based on multi-mode data target detection, comprising: collecting data of a construction site through a laser radar, a millimeter-wave radar, and a camera; using a transformation matrix between the laser radar and the camera, a transformation matrix between the millimeter-wave radar and the camera, and a camera internal parameter matrix, respectively associating a sparse laser radar point cloud and a denoised millimeter-wave radar point cloud of the construction site with an image;
[0007] In the point cloud space of the two point cloud data, the sparse laser radar point cloud and the denoised millimeter wave radar point cloud are divided into a number of vertical point cloud columns, which have a set width and length and cover the entire point cloud space in height; within the range of each point cloud column, the sparse laser radar point cloud and the denoised millimeter wave radar point cloud are extracted to obtain the first feature and the second feature respectively; in each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature and the second fused feature;
[0008] The multimodal target detection network performs target detection based on the first fusion feature of the associated sparse lidar point cloud, the second fusion feature of the denoised millimeter-wave radar point cloud and the image, and then determines the distance between each target, compares the determined distance with the corresponding minimum distance standard between targets, and issues a near-electric operation warning when the corresponding minimum distance standard between targets is reached.
[0009] Furthermore, the laser radar point cloud is subjected to sparse filtering processing to obtain a sparse laser radar point cloud; the probability of each point cloud point in the millimeter wave radar point cloud representing the target is predicted by the point cloud segmentation network, and the points with a probability lower than a set threshold are filtered out, and the millimeter wave point cloud is denoised to obtain the denoised millimeter wave radar point cloud.
[0010] Furthermore, features of the sparse lidar point cloud and the denoised millimeter-wave radar point cloud are extracted within the range of each point cloud column to obtain the first feature and the second feature respectively, including: for the sparse lidar point cloud, in each point cloud column, the coordinates of the midpoint of the sparse lidar point cloud, the deviation of the point relative to the center of all the sparse lidar point cloud points in the point cloud column, and the reflection intensity of the point are extracted as the first feature; for the denoised millimeter-wave radar point cloud, in each point cloud column, the coordinates of the midpoint of the denoised millimeter-wave radar point cloud, the deviation of the point relative to the center of all the denoised millimeter-wave radar point cloud points in the point cloud column, the radar reflection cross-sectional area and Doppler information are extracted as the second feature.
[0011] Furthermore, in each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature and second fused feature, including:
[0012] For the first feature and the second feature with the same coordinate position, the deviation of the first feature relative to the center of all sparse lidar point cloud points in the point cloud column is interleaved and integrated into the second feature, and the deviation of the second feature relative to the center of all denoised millimeter-wave radar point cloud points in the point cloud column is interleaved and integrated into the first feature;
[0013] For the first feature and the second feature with the same coordinate position, the reflection intensities of all the laser radar points in the first feature within the point cloud column where the first feature is located are averaged and added to the second feature;
[0014] For the first feature and the second feature with the same coordinate position, the radar reflection cross-sectional area and Doppler information of the millimeter-wave radar in all the second features within the point cloud column where the second feature is located are averaged and added to the first feature.
[0015] Furthermore, the multimodal target detection network includes: a radar feature encoding network, an image feature encoding network, an image feature enhancement network, an enhanced feature fusion module and a detection head; the radar feature encoding network uses the first fusion feature and the second fusion feature Extract radar features; the image feature encoding network encodes the image into image features in a densely embedded form, the image feature enhancement network uses radar features to enhance image features, the enhanced feature fusion module fuses the enhanced image features and radar features, and then the detection head based on the region proposal network uses the fusion results to detect targets.
[0016] Furthermore, the radar feature encoding network converts the first fusion feature and the second fusion feature Stitching to get radar modal aggregation features The first fusion feature Second fusion feature and radar modal aggregation features Convolutional features are extracted through a convolutional layer with batch normalization and ReLU activation function. After each convolution, the convolutional features of the radar modal aggregation features are processed through convolution and Sigmoid activation function to generate weights. The weights are shared with the convolutional features of the first fusion features, the second fusion features, and the radar modal aggregation features weighted extracted in the three convolutions. After multiple layers of processing, the final convolutional features of the first fusion features, the second fusion features, and the radar modal aggregation features are concatenated to obtain the radar features.
[0017] Furthermore, the image feature enhancement network uses atrous convolution to expand the receptive field of radar features, and generates the spatial pattern of radar features through convolutional layers and Sigmoid activation functions S = Sigmoid (Conv (AtrousConv (F p ))), after broadcasting the spatial pattern of the radar feature along the channel dimension, the spatial pattern is combined with the image feature through element-by-element multiplication to generate the enhanced image feature: ⊙ is element-wise multiplication.
[0018] Furthermore, the enhanced feature fusion module combines the enhanced image features With radar signature F p Splicing, and then fusion through the convolution layer: Then, based on the fusion results, the convolution layer and Softmax are used to generate image feature weights and radar weights: W V , W F =Softmax(Conv(F)), based on image feature weights W V and radar weight W F Fusion enhanced image features With radar signature F p Get the final radar signature
[0019] Furthermore, a set of 3D anchor points are predefined at each position of the final radar feature; each anchor point contains the following attributes: the coordinates of the center point of the anchor point, the three-dimensional size of the anchor point and the direction of the anchor point; the region proposal network contains: a shared convolutional layer, the final radar feature is further extracted through the shared convolutional layer and then input into the classification branch and regression branch based on the multi-layer perceptron; wherein the classification branch outputs the probability of whether each anchor point contains the target, and uses the binary cross-entropy loss (Binary Cross-Entropy Loss) for supervision; the regression branch outputs the position offset of each anchor point relative to the true detection frame, wherein the position offset includes the coordinate offset of the center point of the anchor point, the three-dimensional size offset of the anchor point and the direction offset of the anchor point; the regression branch uses the smooth L1 loss Smooth L1 Loss for supervision; the anchor point is adjusted according to the classification score and the regression offset to generate the candidate region, and the candidate regions with high overlap are removed using non-maximum suppression to obtain the target detection result.
[0020] In a second aspect, the present invention provides a power engineering near-electrical operation supervision device based on multi-mode data target detection, comprising: at least one processing unit, the processing unit is connected to a storage unit, a camera, a laser radar and a millimeter-wave radar through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the power engineering near-electrical operation supervision method based on multi-mode data target detection is implemented.
[0021] The above technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0022] The present invention collects data of the construction site through a laser radar, a millimeter-wave radar and a camera; uses a transformation matrix between the laser radar and the camera, a transformation matrix between the millimeter-wave radar and the camera and a camera internal parameter matrix to respectively associate the sparse laser radar point cloud and the denoised millimeter-wave radar point cloud of the construction site with an image; in the point cloud space of the two point cloud data, the sparse laser radar point cloud and the denoised millimeter-wave radar point cloud are divided into a plurality of vertical point cloud columns, the point cloud columns have a set width and length, and the height covers the entire point cloud space; within the range of each point cloud column, the sparse laser radar point cloud and the denoised millimeter-wave radar point cloud are extracted to obtain a first feature and a second feature respectively; in each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain corresponding first fused features and second fused features; a multimodal target detection network performs target detection based on the first fused features of the associated sparse laser radar point cloud, the second fused features of the denoised millimeter-wave radar point cloud and the image, and then determines the distance between each target, compares the determined distance with the corresponding minimum distance standard between the targets, and when the corresponding minimum distance standard between the targets is reached, issues a near-electric operation warning. LiDAR has a short wavelength and is good at capturing the shape and structure of objects, so it has the ability to accurately measure the target position, but LiDAR is insensitive to motion and has a weak ability to capture dynamic motion information of the target. Millimeter-wave radar has a long wavelength, strong detection and resistance to weather interference, and is good at detecting the range and motion information of objects. Images have deeper semantic information. This application combines three types of data to perceive the situation of objects at the construction site, provides effective data for more robust and accurate target detection, and realizes near-electric operation control. This application first performs shallow feature fusion and then extracts deep features, which can effectively utilize the filtering capabilities of neural networks and reduce noise during feature extraction.
[0023] In each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature and the second fused feature. The interactive fusion includes: for the first feature and the second feature with the same coordinate position, the deviation of the first feature relative to the center of all sparse laser radar point cloud points in the point cloud column is interleaved and integrated into the second feature, and the deviation of the second feature relative to the center of all noise-reduced millimeter-wave radar point cloud points in the point cloud column is interleaved and integrated into the first feature; by exchanging the offsets of the first feature and the second feature with the same coordinate position, the respective geometric features are enriched. For the first feature and the second feature with the same coordinate position, the reflection intensity of all laser radar points in the first feature within the point cloud column where the first feature is located is averaged and added to the second feature; the reflection intensity of all first features within the point cloud column is fused into the corresponding second feature to enhance the quality of the second feature under normal conditions. For the first feature and the second feature with the same coordinate position, the radar reflection cross-sectional area and Doppler information of the millimeter-wave radar in all second features within the point cloud column where the second feature is located is averaged and added to the first feature. The Doppler information and radar cross-sectional area features of all second features in the point cloud column range are integrated into the corresponding first features to enhance the quality of the first features in harsh environments.
[0024] The image feature enhancement network uses radar features to enhance image features. The enhanced feature fusion module fuses the enhanced image features and radar features. Then the detection head based on the region proposal network uses the fusion results for target detection. In target detection, although image features have high-dimensional semantic information, they lack depth information and have low feature quality under harsh lighting conditions. Radar features contain rich spatial information, which can make up for the shortcomings of image features. Radar point clouds are sparse and difficult to provide rich semantic information. Images contain rich texture and semantic information, which can make up for the shortcomings of radar. In any harsh conditions, the final radar features can effectively provide data for target detection and enhance robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0027] Figure 1A flowchart of a method for supervising power engineering near-electrical operations based on multi-mode data target detection provided by an embodiment of the present invention;
[0028] Figure 2 In each point cloud column provided by the embodiment of the present invention, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature F c 1 and a flow chart of the second fusion feature;
[0029] Figure 3 An architecture diagram of a multimodal target detection network provided by an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of a radar feature coding network provided by an embodiment of the present invention;
[0031] Figure 5 A schematic diagram of an image feature enhancement network provided by an embodiment of the present invention;
[0032] Figure 6 A schematic diagram of a power engineering near-electrical operation monitoring device based on multi-mode data target detection provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0034] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0035] Example 1
[0036] like Figure 1 As shown, the technology of the present invention realizes a method for supervising power engineering near-line operations based on multi-mode data target detection, including:
[0037] Data from the construction site is collected through LiDAR, millimeter-wave radar and cameras. LiDAR has a short wavelength and is weak in resisting weather interference during detection. It is good at capturing the external structure of objects, so it has the ability to accurately measure the target position, but LiDAR is insensitive to motion and has a weak ability to capture dynamic motion information of the target. Millimeter-wave radar has a long wavelength and is strong in resisting weather interference during detection. It is good at detecting the range and motion information of objects. Images have deeper semantic information, but lack spatial information. This application combines three types of data to perceive the situation of objects at the construction site, provide effective data for more robust and accurate target detection, and realize near-electric operation control.
[0038] Perform sparse filtering on the lidar point cloud to obtain a sparse lidar point cloud is the nth point in the sparse lidar point cloud, with coordinates
[0039] The point cloud segmentation network is used to predict the probability of each point cloud point in the millimeter wave radar point cloud representing the target, and the points with probability lower than the set threshold are filtered out to reduce the noise of the millimeter wave point cloud to obtain the reduced noise millimeter wave radar point cloud. The mth point in the denoised millimeter-wave radar point cloud has coordinates
[0040] The transformation matrix between the laser radar and the camera, the transformation matrix between the millimeter-wave radar and the camera, and the camera internal parameter matrix are used to associate the sparse laser radar point cloud and the denoised millimeter-wave radar point cloud of the construction site with the image, respectively. Thus, a spatial mapping relationship between the image and the two point clouds is established.
[0041] In the point cloud space of the two point cloud data, the associated sparse lidar point cloud and denoised millimeter wave radar point cloud are divided into K vertical point cloud columns C = {c k |k=1,2,...K}, the point cloud column has a set width and length, and the height covers the entire point cloud space. k is the kth point cloud column in the point cloud space.
[0042] The features of the sparse LiDAR point cloud and the denoised millimeter-wave radar point cloud are extracted within the range of each point cloud column to obtain the first feature and the second feature respectively. For the sparse LiDAR point cloud, in each point cloud column, the coordinates of the midpoints of the sparse LiDAR point cloud are extracted, and the deviation of the point from the center of all the sparse LiDAR point cloud points in the point cloud column and the reflection intensity of the point are taken as the first feature. The first feature set is: For example, a point in a sparse lidar point cloud Belongs to point cloud column c k , then in the point cloud column c k, extract the coordinates of the points in the sparse lidar point cloud point Relative to point cloud column c k The deviation of the center of all sparse lidar point cloud points in and Point The reflection intensity To construct the first feature. For the denoised millimeter-wave radar point cloud, in each point cloud column, the coordinates of the midpoint of the denoised millimeter-wave radar point cloud, the deviation of the point relative to the center of all the denoised millimeter-wave radar point cloud points in the point cloud column, the radar reflection cross-sectional area and the Doppler information are extracted as the second feature; Belongs to point cloud column c k , then in the point cloud column c k , extract the coordinates of the midpoints in the denoised millimeter-wave radar point cloud point Relative to point cloud column c k The deviation of the center of all noise-reduced millimeter-wave radar point cloud points The radar reflection cross-sectional area and Doppler information are used to construct the second feature, and the second feature set is:
[0043] In each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature and the second fusion feature
[0044] Among them, Figure 2 As shown, the fusion includes:
[0045] For the first feature and the second feature with the same coordinate position, the deviation of the first feature relative to the center of all sparse lidar point cloud points in the point cloud column is interlaced and integrated into the second feature, and the deviation of the second feature relative to the center of all denoised millimeter-wave radar point cloud points in the point cloud column is interlaced and integrated into the first feature; by exchanging the offsets of the first feature and the second feature with the same coordinate position, their respective geometric features are enriched.
[0046] For the first feature and the second feature with the same coordinate position, the reflection intensities of all the lidar points in the first feature within the point cloud column where the first feature is located are averaged and added to the second feature; the reflection intensities of all the first features within the point cloud column are merged into the corresponding second feature to enhance the quality of the second feature under normal conditions.
[0047] For the first feature and the second feature with the same coordinate position, the radar reflection cross-sectional area and Doppler information of the millimeter-wave radar in all the second features within the point cloud column where the second feature is located are averaged and added to the first feature. The Doppler information and radar reflection cross-sectional area features of all the second features within the point cloud column are integrated into the corresponding first feature to enhance the quality of the first feature in harsh environments.
[0048] The multimodal target detection network performs target detection based on the first fused features of the associated sparse lidar point cloud, the second fused features of the denoised millimeter-wave radar point cloud and the image.
[0049] like Figure 3 As shown, the multimodal target detection network includes: a radar feature encoding network, an image feature encoding network, an image feature enhancement network, an enhanced feature fusion module and a detection head. The radar feature encoding network uses the first fusion feature and the second fusion feature Extract radar feature F p The image feature encoding network encodes the image into a densely embedded image feature V, the image feature enhancement network uses radar features to enhance image features, the enhanced feature fusion module fuses the enhanced image features with the radar features, and then the detection head based on the region proposal network uses the fusion results to perform target detection.
[0050] In the specific implementation process, Figure 4 As shown, the radar feature encoding network first fuses the feature and the second fusion feature Stitching to get radar modal aggregation features The first fusion feature Second fusion feature and radar modal aggregation features Convolutional features are extracted through convolutional layers with batch normalization and ReLU activation functions. After each convolution, the convolutional features of the radar modality aggregation features are processed through convolution and Sigmoid activation functions to generate weights. The weights are shared with the first fusion features, the second fusion features, and the convolutional features of the radar modality aggregation features extracted by the three convolutions. After multiple layers of processing, the final convolutional features of the first fusion features, the second fusion features, and the radar modality aggregation features are concatenated to obtain the radar feature F. p .
[0051] The image feature encoding network encodes the corresponding RGB image into image features V in the form of dense embedding. To reduce computational overhead, a pre-trained ResNet model is used and the weights are kept unchanged during training. Note that in the specific implementation, the image feature encoding network can be replaced with a stronger visual model such as the visual Transformer according to the computational budget.
[0052] The image feature enhancement network uses radar features to enhance image features, and the enhanced feature fusion module fuses the enhanced image features with radar features.
[0053] In target detection, although image features have high-dimensional semantic information, they lack spatial information and have low feature quality under harsh lighting conditions. Radar features contain rich spatial information and can make up for the shortcomings of image features. The image feature enhancement network uses radar features to explicitly predict the probability of target objects existing in different spatial positions, integrates them into image features, and gives image feature structural information to enhance image features. In the specific implementation process, Figure 5 As shown in the figure, the image feature enhancement network uses atrous convolution to expand the receptive field of radar features, and generates the spatial pattern of radar features through convolutional layers and Sigmoid activation functions S = Sigmoid (Conv (AtrousConv (F p ))), where Sigmoid(), Conv(), and AtrousConv() are sigmoid activation, convolution, and dilated convolution functions, respectively. After broadcasting the spatial pattern of the radar feature along the channel dimension, the spatial pattern is combined with the image feature through element-by-element multiplication to generate enhanced image features: ⊙ is the element-by-element multiplication. The image feature enhancement network enhances the image features through the spatial information of the radar features, makes up for the defect of the image lacking depth information, and improves the expression ability of the image features in spatial position, especially under poor lighting conditions.
[0054] Radar point clouds are sparse and difficult to provide rich semantic information. Images contain rich texture and semantic information, which can make up for the shortcomings of radar. The enhanced feature fusion module uses the enhanced image features to Enhance radar features. With radar signature F p Splicing, and then fusion through the convolution layer: Among them, Conv(), Concat() are convolution and splicing functions respectively, and then the convolution layer and Softmax are used to generate image feature weights and radar weights W based on the fusion results. V , W F =Softmax(Conv(F)), based on image feature weights W V and radar weight WF Fusion enhanced image features With radar signature F p Get the final radar signature The final radar features are used for subsequent target detection tasks.
[0055] The detection head of the region proposal network uses the fusion results for target detection. A set of 3D anchors is predefined at each position of the final radar feature; each anchor contains the following attributes: the coordinates of the center point of the anchor, the three-dimensional size of the anchor, and the direction of the anchor; the region proposal network contains: a shared convolutional layer, and the final radar feature is further extracted through the shared convolutional layer and input into the classification branch and regression branch based on the multi-layer perceptron. Among them, the classification branch outputs the probability of whether each anchor contains a target, and uses the binary cross-entropy loss for supervision. The regression branch outputs the position offset of each anchor relative to the true detection box, where the position offset includes the coordinate offset of the center point of the anchor, the three-dimensional size offset of the anchor, and the direction offset of the anchor; the regression branch uses the smooth L1 loss for supervision. The anchor is adjusted according to the classification score and the regression offset to generate candidate regions, and non-maximum suppression is used to remove candidate regions with high overlap.
[0056] After detecting the target, the distance between each target is determined, and the determined distance is compared with the corresponding minimum distance standard between the targets. When the corresponding minimum distance standard between the targets is reached, a near-electric operation warning is issued.
[0057] Example 2
[0058] See also Figure 6 As shown, an embodiment of the present invention provides a device for supervising near-electrical work of electric power engineering based on multi-mode data target detection, comprising: at least one processing unit, the processing unit is connected to a storage unit via a bus unit, the storage unit is a computer-readable storage medium, and can be used to store software programs, computer executable programs, and modules, such as the software programs, computer executable programs, and modules corresponding to a method for supervising near-electrical work of electric power engineering based on multi-mode data target detection in an embodiment of the present invention. The processing unit implements the above-mentioned method for supervising near-electrical work of electric power engineering based on multi-mode data target detection by running the software programs, computer executable programs, and modules stored in the storage unit, comprising:
[0059] The data of the construction site is collected through laser radar, millimeter wave radar and camera; the sparse laser radar point cloud and denoised millimeter wave radar point cloud of the construction site are respectively associated with the image using the transformation matrix between the laser radar and the camera, the transformation matrix between the millimeter wave radar and the camera and the camera internal parameter matrix;
[0060] In the point cloud space of two kinds of point cloud data, several vertical point cloud columns are divided from the sparse lidar point cloud and the noise-reduced millimeter-wave radar point cloud to be associated. The point cloud columns have a set width and length, and the height covers the entire point cloud space. Features of the sparse lidar point cloud and the noise-reduced millimeter-wave radar point cloud are extracted respectively to obtain a first feature and a second feature within the range of each point cloud column. In each point cloud column, the first feature and the second feature with the same coordinate position respectively fuse the features of the other modality to obtain corresponding first fusion features and second fusion features.
[0061] The multi-modal target detection network performs target detection based on the first fusion feature of the associated sparse lidar point cloud, the second fusion feature of the noise-reduced millimeter-wave radar point cloud and the image, and then determines the distances between the targets. The determined distances are compared with the minimum distance criteria between the corresponding targets. When the minimum distance criteria between the corresponding targets are met, a near-power operation warning is issued.
[0062] Certainly, the computer program stored in the storage unit of a power engineering near-power operation supervision device based on multi-modal data target detection provided by an embodiment of the present invention is not limited to the method operations described above, and can also execute the relevant operations in a power engineering near-power operation supervision method based on multi-modal data target detection provided by any embodiment of the present invention.
[0063] Embodiment 3
[0064] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which when executed, implements the power engineering near-power operation supervision method based on multi-modal data target detection, including:
[0065] Collect data of the construction site through a lidar, a millimeter-wave radar and a camera; use the transformation matrix between the lidar and the camera, the transformation matrix between the millimeter-wave radar and the camera, and the camera internal parameter matrix to associate the sparse lidar point cloud and the noise-reduced millimeter-wave radar point cloud of the construction site with the image respectively.
[0066] In the point cloud space of two kinds of point cloud data, several vertical point cloud columns are divided from the sparse lidar point cloud and the noise-reduced millimeter-wave radar point cloud. The point cloud columns have a set width and length, and the height covers the entire point cloud space. Features of the sparse lidar point cloud and the noise-reduced millimeter-wave radar point cloud are extracted respectively to obtain a first feature and a second feature within the range of each point cloud column. In each point cloud column, the first feature and the second feature with the same coordinate position respectively fuse the features of the other modality to obtain corresponding first fusion features and second fusion features.
[0067] The multimodal target detection network performs target detection based on the first fusion feature of the associated sparse lidar point cloud, the second fusion feature of the denoised millimeter-wave radar point cloud and the image, and then determines the distance between each target, compares the determined distance with the corresponding minimum distance standard between targets, and issues a near-electric operation warning when the corresponding minimum distance standard between targets is reached.
[0068] A computer-readable storage medium provided in an embodiment of the present invention stores a computer program which is not limited to the method operations described above, but can also execute related operations in a method for supervising near-electrical operations of electric power engineering based on multi-mode data target detection provided in any embodiment of the present invention.
[0069] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, structures or units, which can be electrical, mechanical or other forms.
[0070] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0071] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0072] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A method for supervising power engineering near-line operations based on multi-mode data target detection, characterized in that: include: Collect data from construction sites through LiDAR, millimeter-wave radar, and cameras; The sparse lidar point cloud and denoised millimeter-wave radar point cloud of the construction site are associated with the image using the transformation matrix between the lidar and the camera, the transformation matrix between the millimeter-wave radar and the camera, and the camera internal parameter matrix. In the point cloud space of the two point cloud data, the associated sparse laser radar point cloud and the denoised millimeter wave radar point cloud are divided into a number of vertical point cloud columns, each of which has a set width and length and covers the entire point cloud space in height; within the range of each point cloud column, the sparse laser radar point cloud and the denoised millimeter wave radar point cloud are extracted to obtain the first feature and the second feature respectively; in each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused feature and the second fused feature; The multimodal target detection network performs target detection based on the first fusion feature of the associated sparse lidar point cloud, the second fusion feature of the denoised millimeter-wave radar point cloud and the image, and then determines the distance between each target, compares the determined distance with the corresponding minimum distance standard between targets, and issues a near-electric operation warning when the corresponding minimum distance standard between targets is reached.
2. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 1 is characterized in that: The laser radar point cloud is subjected to sparse filtering processing to obtain a sparse laser radar point cloud; the probability of each point cloud point in the millimeter wave radar point cloud representing the target is predicted through a point cloud segmentation network, points with probabilities lower than a set threshold are filtered out, and the millimeter wave point cloud is denoised to obtain the denoised millimeter wave radar point cloud.
3. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 1 is characterized in that: The first feature and the second feature are obtained by extracting features of the sparse lidar point cloud and the denoised millimeter-wave radar point cloud within the range of each point cloud column, respectively. For the sparse lidar point cloud, in each point cloud column, the coordinates of the midpoint of the sparse lidar point cloud, the deviation of the point relative to the center of all the sparse lidar point cloud points in the point cloud column, and the reflection intensity of the point are extracted as the first feature; for the denoised millimeter-wave radar point cloud, in each point cloud column, the coordinates of the midpoint of the denoised millimeter-wave radar point cloud, the deviation of the point relative to the center of all the denoised millimeter-wave radar point cloud points in the point cloud column, the radar reflection cross-sectional area and the Doppler information are extracted as the second feature.
4. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 3 is characterized in that: In each point cloud column, the first feature and the second feature with the same coordinate position are fused with each other's modal features to obtain the corresponding first fused features and second fused features, including: For the first feature and the second feature with the same coordinate position, the deviation of the first feature relative to the center of all sparse lidar point cloud points in the point cloud column is interleaved and integrated into the second feature, and the deviation of the second feature relative to the center of all denoised millimeter-wave radar point cloud points in the point cloud column is interleaved and integrated into the first feature; For the first feature and the second feature with the same coordinate position, the reflection intensities of all the laser radar points in the first feature within the point cloud column where the first feature is located are averaged and added to the second feature; For the first feature and the second feature with the same coordinate position, the radar reflection cross-sectional area and Doppler information of the millimeter-wave radar in all the second features within the point cloud column where the second feature is located are averaged and added to the first feature.
5. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 3 is characterized in that: The multimodal target detection network includes: a radar feature encoding network, an image feature encoding network, an image feature enhancement network, an enhanced feature fusion module and a detection head; the radar feature encoding network uses the first fusion feature and the second fusion feature Extract radar feature F p The image feature encoding network encodes the image into a densely embedded image feature V, the image feature enhancement network uses radar features to enhance image features, the enhanced feature fusion module fuses the enhanced image features with the radar features, and then the detection head based on the region proposal network uses the fusion results to perform target detection.
6. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 5 is characterized in that: The radar feature encoding network first fuses the features and the second fusion feature Stitching to get radar modal aggregation features The first fusion feature Second fusion feature and radar modal aggregation features Convolutional features are extracted through a convolutional layer with batch normalization and ReLU activation function. After each convolution, the convolutional features of the radar modal aggregation features are processed through convolution and Sigmoid activation function to generate weights. The weights are shared with the convolutional features of the first fusion features, the second fusion features, and the radar modal aggregation features weighted extracted in the three convolutions. After multiple layers of processing, the final convolutional features of the first fusion features, the second fusion features, and the radar modal aggregation features are concatenated to obtain the radar features.
7. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 5 is characterized in that: Image feature enhancement network uses dilated convolution to enlarge radar feature F p The receptive field of the radar feature is generated by the convolution layer and the Sigmoid activation function S = Sigmoid (Conv (AtrousConv (F p ))), where Sigmoid(), Conv(), and AtrousConv() are sigmoid activation, convolution, and dilated convolution functions, respectively. After broadcasting the spatial pattern of the radar feature along the channel dimension, the spatial pattern is combined with the image feature through element-by-element multiplication to generate enhanced image features: ⊙ is element-wise multiplication.
8. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 5 is characterized in that: The enhanced feature fusion module will enhance the image features With radar signature F p Splicing, and then fusion through the convolution layer: Among them, Conv() and Concat() are convolution and concatenation functions respectively, and then based on the fusion results, the convolution layer and Softmax are used to generate image feature weights and radar weights: W V , W F =Softmax(Conv(F)), based on image feature weights W V and radar weight W F Fusion enhanced image features With radar signature F p Get the final radar signature 9. The method for supervising power engineering near-line operations based on multi-mode data target detection according to claim 8 is characterized in that: A set of 3D anchor points are predefined at each position of the final radar feature; each anchor point contains the following attributes: the coordinates of the center point of the anchor point, the three-dimensional size of the anchor point and the direction of the anchor point; the region proposal network contains: a shared convolutional layer, the final radar feature is further extracted through the shared convolutional layer and then input into the classification branch and regression branch based on the multi-layer perceptron; the classification branch outputs the probability of whether each anchor point contains the target, and uses the binary cross entropy loss for supervision; the regression branch outputs the position offset of each anchor point relative to the true detection frame, wherein the position offset includes the coordinate offset of the center point of the anchor point, the three-dimensional size offset of the anchor point and the direction offset of the anchor point; the regression branch uses the smooth L1 loss for supervision; the anchor point is adjusted according to the classification score and the regression offset to generate the candidate region, and the candidate regions with high overlap are removed using non-maximum suppression to obtain the target detection result.
10. A power engineering near-line operation monitoring device based on multi-mode data target detection, characterized in that: include: At least one processing unit, wherein the processing unit is connected to a storage unit, a camera, a laser radar and a millimeter-wave radar through a bus unit, the storage unit stores a computer program, and when the computer program is executed by the processing unit, the method for supervising near-electrical operations in power engineering based on multi-mode data target detection as described in any of claims 1-9 is implemented.