Target detection and positioning method and system for substation inspection
Through the combination of inverse perspective transformation and deep learning algorithm, two-dimensional images and three-dimensional point cloud data are processed, which solves the problem of low detection accuracy and efficiency in substation inspection, and achieves high-precision, high-efficiency and robust long-distance object detection.
Patent Information
- Application Number
- CN202510182448.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to take into account high detection accuracy, high detection efficiency and high robustness in substation inspections. Especially during long-distance detection, the two-dimensional object detection method is greatly affected by environmental priors and image quality, while the three-dimensional object detection method has problems such as point cloud sparseness and large data volume.
The inverse perspective transformation algorithm is used to process two-dimensional image data, combined with the deep learning algorithm to process three-dimensional point cloud data, and high-precision and efficient detection results are obtained through the Wasserstein distance measurement algorithm and confidence fusion technology.
It realizes high accuracy, high efficiency and high robustness detection and positioning of long-distance targets during substation inspection, and is suitable for detecting small targets and improves detection accuracy and robustness.
Smart Images

Figure CN120298652A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of substation inspection, and particularly to a method and system for target detection and positioning in substation inspection. Background Art
[0002] Detecting and positioning moving targets is the core task of substation inspection. By promptly discovering targets and determining their categories and distances, it helps the control terminal make targeted decisions to avoid risks.
[0003] Two-dimensional target detection methods are the mainstream methods for substation inspection, which locate the corresponding distances by combining geometric prior assumptions or depth information in images. However, methods based on geometric priors are greatly affected by environmental priors and image quality, resulting in low detection accuracy. Additionally, the position information obtained through data-level information fusion is overly dependent on the depth information provided by lidar. When a single sensor fails, it will not be able to work properly, and the robustness is low.
[0004] Three-dimensional target detection methods can provide high-quality depth information by using high-line-number lidar to generate point cloud data. However, the sparsity of point clouds brings the following problems: the shape of the target to be detected is incomplete and semantic information is missing; the target to be detected is easily confused with background points, resulting in false detections; the proportion of foreground points compared to the scene point cloud is small. These problems lead to low accuracy when using relevant three-dimensional target detection methods for long-distance detection, and it is difficult to apply them in substations where the inspection distance can reach hundreds of meters at most. Additionally, due to the large amount of point cloud data and irregular data structure, the detection efficiency and robustness of relevant three-dimensional target detection methods are low.
[0005] Regarding the problem that the target detection methods in the related art are difficult to balance high detection accuracy, high detection efficiency, and high robustness, no effective solution has been proposed yet. Summary of the Invention
[0006] A method and system for target detection and positioning in substation inspection provided by an embodiment of the present invention at least solve the problem that the target detection methods in the related art are difficult to balance high detection accuracy, high detection efficiency, and high robustness.
[0007] A method for target detection and positioning in substation inspection provided by an embodiment of the present invention includes: collecting two-dimensional image data and three-dimensional point cloud data of the current environment; processing the two-dimensional image data based on the inverse perspective transformation algorithm to obtain the first positioning information of the detection target; processing the three-dimensional point cloud data based on the deep learning algorithm to obtain the second positioning information of the detection target; processing the first positioning information and the second positioning information based on the Wasserstein distance metric algorithm to obtain matching information; and fusing the matching information based on the confidence level to obtain the detection result.
[0008] The target detection and positioning method for substation inspection provided by the embodiment of the present invention creates a patent, processes two-dimensional image data based on the inverse perspective transformation algorithm, and obtains the first positioning information of the detection target, including: detecting the two-dimensional image data based on a single-stage target detection algorithm to obtain a two-dimensional detection frame; determining the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error; processing the two-dimensional detection frame based on the inverse perspective matrix to obtain the three-dimensional positioning coordinates of the detection target; wherein, the first positioning information includes the two-dimensional detection frame and the three-dimensional positioning coordinates.
[0009] The target detection and positioning method for substation inspection provided by the embodiment of the present invention creates a patent, detects two-dimensional image data based on a single-stage target detection algorithm to obtain a two-dimensional detection frame, including: extracting the first semantic features of the two-dimensional image data based on multiple extraction modules of the backbone network; splicing the first semantic features together according to the channel dimension based on the downsampling module of the first detection head network to obtain spliced semantic features; reducing the number of channels of the spliced semantic features based on the downsampling module to obtain a two-dimensional detection frame; wherein, the single-stage target detection algorithm includes a backbone network and a first detection head network, the extraction module includes a convolutional layer, a batch normalization layer, and a scale scaling layer, and the convolutional layer of one of the multiple extraction modules adopts a full-dimensional dynamic convolutional network, and the full-dimensional dynamic convolutional network includes a position-level multiplication operation along the spatial dimension, a channel-level multiplication operation along the input channel dimension, a filter-level multiplication operation along the output channel dimension, and a kernel-level multiplication operation along the convolutional kernel space.
[0010] The target detection and positioning method for substation inspection provided by the embodiment of the present invention creates a patent, determines the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error, including: collecting multiple non-collinear control points on the two-dimensional image based on the multi-point averaging method and marking the pixel values of the control points; determining the corresponding coordinates of the control points in the lidar image; solving the inverse perspective transformation matrix based on the pixel values of the control points and the corresponding coordinates; solving the reprojection error and the inverse perspective matrix through the Levenberg-Marquardt algorithm based on the inverse perspective transformation matrix; in the case where the reprojection error is greater than or equal to the preset threshold, re-collecting the control points to obtain new pixel values of the control points and new corresponding coordinates, and solving the new inverse perspective matrix and the new reprojection error until the finally obtained reprojection error is less than the preset threshold.
[0011] The object detection and positioning method for substation inspection provided by the embodiment of the present invention creates uses the PointPillars algorithm as the deep learning algorithm. The three-dimensional point cloud data is processed based on the PointPillars algorithm to obtain the second positioning information, including: converting the three-dimensional point cloud data into a two-dimensional pseudo-image based on the Pillar Feature Network; performing downsampling, upsampling, and splicing operations on the pseudo-image based on the backbone network to obtain the second semantic feature; processing the second semantic feature based on the second detection head network to obtain a two-dimensional detection result from the bird's-eye view perspective; estimating the height of the three-dimensional detection box based on the three-dimensional point cloud data by the second detection head network; obtaining the three-dimensional detection box based on the two-dimensional detection result from the bird's-eye view perspective and the estimated value of the height of the three-dimensional detection box. Among them, the PointPillars algorithm includes a Pillar Feature Network, a backbone network, and a second detection head network, and the second positioning information includes a three-dimensional detection box and corresponding positioning coordinates.
[0012] The object detection and positioning method for substation inspection provided by the embodiment of the present invention creates processes the first positioning information and the second positioning information based on the Wasserstein distance metric algorithm to obtain matching information, including: projecting the three-dimensional detection box of the second positioning information onto the two-dimensional image plane to obtain a three-dimensional projected detection box; calculating the similarity between the three-dimensional projected detection box and the two-dimensional detection box of the first positioning information based on the Wasserstein distance metric algorithm according to the extended cost matrix, where the extended cost matrix includes the coordinates of the three-dimensional projected detection box, the coordinates of the two-dimensional detection box, and the coordinates of the virtual target; matching the two-dimensional detection box and the three-dimensional projected detection box based on the similarity using the Hungarian algorithm to obtain matching information, and the matching information is the positioning coordinates in the first positioning information and the positioning coordinates in the second positioning information.
[0013] The object detection and positioning method for substation inspection provided by the embodiment of the present invention creates projects the three-dimensional detection box of the second positioning information onto the two-dimensional image plane to obtain a three-dimensional projected detection box, including: obtaining the coordinates of the corner points of the three-dimensional detection box in the lidar coordinate system; obtaining the projected point coordinates of the corner points on the two-dimensional image plane based on the camera internal parameter matrix, the rotation matrix, and the translation matrix from the lidar coordinate system to the camera coordinate system; calculating the convex hull based on the projected point coordinates; and taking the minimum circumscribed rectangle of the convex hull as the three-dimensional projected detection box.
[0014] The object detection and positioning method for substation inspection provided by the embodiment of the present invention creates fuses the matching information based on the confidence level to obtain the detection result, including: ; where represents the center point position of the fused detection target and is used to determine the detection result; represents the positioning coordinates in the first positioning information, represents the positioning coordinates in the second positioning information, Represents the normalized Wasserstein distance metric between the two-dimensional detection box and the three-dimensional projected detection box; is the confidence level weight coefficient of, and the confidence level is confidence level of; Information mismatch means that the two-dimensional detection box does not match the three-dimensional projected detection box.
[0015] The object detection and positioning method for substation inspection provided by the embodiment of the present invention further includes: determining the confidence level through the following formula : ; In the formula, 1(·) is an indicator function, which is 1 when none of the sides of the object detection box intersects the image edge, and 0 when at least one side of the object detection box intersects the image edge; x and y represent the coordinates of the detection target in the lidar coordinate system; a and b represent the coordinates (a, b) of the camera optical center in the lidar coordinate system; represents the confidence level of the two-dimensional detection box.
[0016] An object detection and positioning system for substation inspection provided by the embodiment of the present invention executes any of the above methods.
[0017] The object detection and positioning method and system for substation inspection provided by the embodiment of the present invention detect two-dimensional image data based on the inverse perspective transformation algorithm, and can obtain first positioning information with lower false detection rate and missed detection rate; detect based on three-dimensional point cloud data, and can obtain second positioning information with higher ranging accuracy, and the efficiency of processing three-dimensional point cloud data based on the point-column algorithm is higher. On this basis, by fusing the first positioning information and the second positioning information, a detection and positioning result with high detection accuracy, high detection efficiency and high robustness can be obtained, and it is applicable to detecting small targets at a long distance. To solve the problem that the object detection method in the related art is difficult to balance high detection accuracy, high detection efficiency and high robustness. Description of the Drawings
[0018] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other embodiments according to these drawings without creative efforts.
[0019] Figure 1 is the step flow chart of an object detection and positioning method for substation inspection in the embodiment of the present invention.
[0020] Figure 2It is a schematic diagram of the position-level multiplication operation for adding an attention mechanism along the spatial dimension in the full-dimensional dynamic convolution network in the embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of the channel-level multiplication operation for adding an attention mechanism along the input channel dimension in the full-dimensional dynamic convolution network in the embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the filter-level multiplication operation for adding an attention mechanism along the output channel dimension in the full-dimensional dynamic convolution network in the embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of the kernel-level multiplication operation for adding an attention mechanism along the convolution kernel space in the full-dimensional dynamic convolution network in the embodiment of the present invention.
[0024] Figure 6 It is a schematic diagram of the full-dimensional dynamic convolution network in the embodiment of the present invention.
[0025] Figure 7 It is a schematic diagram of the downsampling for retaining fine-grained features in the embodiment of the present invention.
[0026] Figure 8 It is a schematic diagram of the control point positions in the embodiment of the present invention.
[0027] Figure 9 It is a schematic diagram of the algorithm framework for two-dimensional object detection and ranging positioning in the embodiment of the present invention.
[0028] Figure 10 It is a schematic diagram of the test image for converting a three-dimensional detection box into a three-dimensional projection detection box in the embodiment of the present invention.
[0029] Figure 11 It is a schematic diagram of the structure of the electronic device of the present invention. Detailed implementation manners
[0030] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.
[0031] In substation inspection tours, two-dimensional object detection methods are the mainstream. They locate the corresponding distance by means of image geometric prior assumptions or depth information. However, methods based on geometric priors are greatly affected by environmental priors and image quality, with low detection accuracy. Moreover, data-level information fusion strongly depends on lidar depth information, and it cannot work when a single sensor fails, resulting in poor robustness. Three-dimensional object detection methods use high-line-number lidar to generate point cloud data, which can provide high-quality depth information. However, there are sparsity problems in point clouds, such as incomplete shapes of objects to be detected, missing semantic information, being easily confused with background points leading to false detections, and a small proportion of foreground points, etc. This makes the detection accuracy at long distances low and it is difficult to be used in substations where the inspection distance can reach hundreds of meters. At the same time, due to its large amount of data and irregular structure, the detection efficiency and robustness are also relatively low.
[0032] Therefore, please refer to Figure 1 As shown, an object detection and positioning method for substation inspection tours provided by an embodiment of the present invention includes: Step S101, collecting two-dimensional image data and three-dimensional point cloud data of the current environment.
[0033] Step S102, processing the two-dimensional image data based on the inverse perspective transformation algorithm to obtain the first positioning information of the detection object.
[0034] Step S103, processing the three-dimensional point cloud data based on a deep learning algorithm to obtain the second positioning information of the detection object.
[0035] Step S104, processing the first positioning information and the second positioning information based on the Wasserstein distance metric algorithm to obtain matching information.
[0036] Step S105, fusing the matching information based on confidence to obtain a detection result.
[0037] It can be understood that Step S102 and Step S103 can be carried out simultaneously.
[0038] The main body executing Steps S101 to S105 can be an inspection robot or a fixed detector. The embodiment of the present invention uses a more flexible inspection robot for illustration. The inspection robot executing the above method can perform automated detection and positioning to obtain the category and distance information of the detection object. The substation control terminal can make targeted decisions based on the above category and distance information to avoid related risks.
[0039] The current environment includes but is not limited to the in-station environment of the substation and the adjacent area outside the station.
[0040] The detection objects include but are not limited to objects invading the substation and in-station activity objects, such as pedestrians and vehicles. Among them, pedestrians belong to small objects and vehicles belong to large objects. The difficulty of detecting small objects is greater than that of detecting large objects.
[0041] The inverse perspective transformation algorithm can be divided into two parts: two-dimensional object detection and inverse perspective transformation. In the two-dimensional object detection part, based on a single-stage object detection algorithm, two-dimensional image data is detected to obtain two-dimensional detection boxes for object detection. In the inverse perspective transformation part, based on the inverse perspective matrix, the two-dimensional detection boxes are processed to obtain the three-dimensional positioning coordinates of the detection targets for positioning and ranging. The single-stage object detection algorithm provided by the embodiments of the present invention is the YOLO series algorithm, which will be specifically described later.
[0042] Simply put, the above inverse perspective transformation algorithm is a two-dimensional object detection algorithm that can obtain three-dimensional positioning coordinates.
[0043] The first positioning information includes but is not limited to two-dimensional detection boxes and corresponding positioning coordinates.
[0044] The deep learning algorithm can be any one of the Pointpillar algorithm, the PointNet algorithm, and the VoteNet algorithm. The embodiments of the present invention use the Pointpillar algorithm for illustration. The Pointpillar algorithm can avoid complex three-dimensional convolution calculations by converting three-dimensional point cloud data into two-dimensional pseudo-images and performing object detection on the pseudo-images, thereby effectively improving the efficiency of the three-dimensional object detection algorithm.
[0045] Simply put, the above Pointpillar algorithm is an efficient three-dimensional object detection algorithm.
[0046] The second positioning information includes but is not limited to three-dimensional detection boxes and corresponding positioning coordinates.
[0047] The Wasserstein distance metric algorithm, also known as the earth mover's distance algorithm, is used to measure the distance between two probability distributions. It can be imagined that there are two piles of soil (representing two probability distributions), and the soil needs to be moved (re-distribute the probability mass) to make one pile of soil become the other pile. The Wasserstein distance is the minimum amount of work required to measure this movement.
[0048] The embodiments of the present invention calculate the similarity between the above two-dimensional detection boxes and three-dimensional projection detection boxes based on the Wasserstein distance metric algorithm to facilitate the matching of the positioning coordinates in the first positioning information and the positioning coordinates in the second positioning information. Among them, the three-dimensional projection detection box is obtained by projecting the above three-dimensional detection box.
[0049] The detection results include but are not limited to being presented in the form of detection boxes and positioning coordinates, and can be converted into the category and distance information of the detection targets.
[0050] The above method provided by the embodiments of the present invention creates and implements detection on two-dimensional image data based on the inverse perspective transformation algorithm, and can obtain first positioning information with a low false detection rate and a low missed detection rate; detection based on three-dimensional point cloud data can obtain second positioning information with a high ranging accuracy, and the efficiency of processing three-dimensional point cloud data based on the point pillar algorithm is relatively high. On this basis, by fusing the first positioning information and the second positioning information, a detection and positioning result with high detection accuracy, high detection efficiency, and high robustness can be obtained, and it is applicable to detecting small targets at a long distance.
[0051] Exemplarily, in step S101, RGB image data of the current environment is collected based on a camera, and three-dimensional point cloud data of the current environment is collected based on a lidar.
[0052] It can be understood that both the camera and the lidar are configured on the inspection robot.
[0053] Preferably, in step S102, the two-dimensional image data is processed based on the inverse perspective transformation algorithm to obtain first positioning information of the detection target, including: step S1021, the two-dimensional image data is detected based on a single-stage object detection algorithm to obtain a two-dimensional detection box.
[0054] Step S1022, determining the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error.
[0055] Step S1023, processing the two-dimensional detection box based on the inverse perspective matrix to obtain the three-dimensional positioning coordinates of the detection target.
[0056] Among them, the first positioning information includes a two-dimensional detection box and three-dimensional positioning coordinates.
[0057] Exemplarily, the above single-stage object detection algorithm is a YOLO series two-dimensional object detection algorithm, such as YOLOv7 or YOLOX.
[0058] Taking YOLOv7 as an example, YOLOv7 extends the efficient long-range attention network E-ELAN, which can improve the learning ability of the network without destroying the original gradient path; it can reparameterize the convolution in a planned manner to generate models of different scales while maintaining the characteristics of the model in the initial design to meet the requirements of different inference speeds. Therefore, YOLOv7 can maintain a very high accuracy while improving the detection speed.
[0059] Furthermore, considering the volume limitation of the on-vehicle perception kit of the inspection robot, the above single-stage object detection algorithm can only run on edge computing devices. Therefore, it is necessary to balance lightweight and detection accuracy. The above single-stage object detection algorithm provided by the embodiments of the present invention is preferably a lightweight model algorithm of the YOLO series, such as YOLOv7-tiny or YOLOX-Nano.
[0060] It can be understood that YOLOv7-tiny is a lightweight design based on YOLOv7, and YOLOX-Nano is a lightweight design based on YOLOX.
[0061] Next, the lightweight model YOLOv7-tiny will be used as a baseline to further illustrate the above method provided by the embodiments of the present invention.
[0062] Although YOLOv7-tiny can provide high-quality two-dimensional detection results, a sensing distance of more than 100 meters is still a huge challenge to its sensing ability. Therefore, the embodiments of the present invention have improved both its backbone network and detection head network.
[0063] Preferably, in step S1021, the two-dimensional image data is detected based on the single-stage object detection algorithm to obtain a two-dimensional detection box, including: extracting the first semantic features of the two-dimensional image data based on multiple extraction modules of the backbone network.
[0064] Based on the downsampling module of the first detection head network, the first semantic features are concatenated together according to the channel dimension to obtain concatenated semantic features.
[0065] Based on the downsampling module, the channel number of the concatenated semantic features is reduced to obtain a two-dimensional detection box.
[0066] Among them, the single-stage object detection algorithm includes a backbone network and a first detection head network. The extraction module includes a convolutional layer, a batch normalization layer, and a scale scaling layer. The convolutional layer of one of the multiple extraction modules adopts a full-dimensional dynamic convolutional network, and the full-dimensional dynamic convolutional network includes a position-level multiplication operation along the spatial dimension, a channel-level multiplication operation along the input channel dimension, a filter-level multiplication operation along the output channel dimension, and a kernel-level multiplication operation along the convolutional kernel space.
[0067] It can be understood that the above "first" detection head network refers to the detection head network of the single-stage object detection algorithm. The description of "first" is to distinguish it from the "second" detection head network of the point-column algorithm, and it has no other specific meaning. The same applies to the "first" semantic features and the "second" semantic features.
[0068] During the feature extraction process, the above-mentioned first semantic features are usually extracted through pooling or strided convolution, resulting in the loss of a large number of fine-grained features. Since distant objects usually occupy only dozens or even a few pixels in the image, the information loss during the feature extraction process has a great impact on distant object detection.
[0069] Dynamic convolution is a method that can better adapt to the features of input data by dynamically adjusting the size of the convolution kernel, so as to improve the adaptability of the convolution network for feature extraction. Therefore, adding a dynamic convolution layer in the early stage of feature extraction can avoid the loss of fine-grained features related to distant objects.
[0070] Different from the conventional dynamic convolution that only focuses on the dynamics of one dimension, the above-mentioned Omni-Dimentional Dynamic Convolution (ODConv) simultaneously focuses on the dynamics in dimensions such as the spatial domain, input channels, and output channels. Its ability to consider multi-dimensional feature learning enables ODConv with only one kernel to achieve comparable or even better performance than the dynamic convolution method with multiple kernels, which can reduce additional parameters and is suitable for lightweight models.
[0071] The above-mentioned omni-dimensional dynamic convolution network can be described by the following formula: ; In the formula, x represents the input feature, y represents the output feature, represents the convolution kernel 's attention scalar, respectively represent the newly introduced attentions along the spatial dimension, input channel dimension, and output channel dimension.
[0072] Taking the k*k feature vector as an example, the position-level multiplication operation along the spatial dimension is as Figure 2 shown, the channel-level multiplication operation along the input channel dimension is as Figure 3 shown, the filter-level multiplication operation along the output channel dimension is as Figure 4 shown, and the kernel-level multiplication operation along the convolution kernel space is as Figure 5 shown. Among them, and represent the corresponding lengths of the feature vectors, and n represents the number of feature vectors.
[0073] See Figure 6 shown, the specific implementation process of the above-mentioned omni-dimensional dynamic convolution network is as follows: Through the global average pooling module GAP, fully connected layer FC, and two activation functions ReLU and Sigmoid, the input feature x is shrunk into a feature vector with a length of , and then four different types of attention mechanism values are generated through four detection heads.
[0074] Exemplarily, four extraction modules (CBS modules) are provided in the backbone network of the lightweight model YOLOv7-tiny. One of the static convolutional layers of a CBS module is replaced with the above-mentioned full-dimensional dynamic convolutional network, so as to help the backbone network adaptively extract global information at the initial stage of feature extraction, and guide the model to retain more fine-grained features, especially the fine-grained features of distant targets.
[0075] Specifically, in the object detection algorithm, adopting multiple detection heads to output prediction results on feature maps of different scales is convenient for improving the model's perception ability of targets of different sizes. Usually, a large number of fine-grained features can be retained in the low-level semantic layer. Therefore, the low-level semantic layer is suitable for the regression of small targets. However, the features of the low-level semantic layer lack context information, resulting in its performance in detecting small targets being worse than that in detecting targets of other scales. In addition, the high-level semantic layer can provide global context semantics, but a large amount of local details will be lost during the process of strided convolution and pooling, so it is not sensitive to small targets and its performance in detecting small targets is also poor.
[0076] In other words, how to expand the receptive field without losing local details is of great significance for improving the performance of small target detection. For this reason, the embodiment of the present invention proposes a downsampling method for retaining fine-grained features, as Figure 7 shown. Taking the feature vector of s*s as an example, represents the length of the feature vector, and (i, j) represents the elements in the feature vector. Different from directly discarding features, the above-mentioned downsampling method concatenates the features together according to the channel dimension, and reduces the number of channels through a filter, and retains a large number of fine-grained features by means of space-for-depth, so it is convenient to learn local details.
[0077] Exemplarily, two downsampling modules (MP-2 modules) are provided in the detection head network of the lightweight model YOLOv7-tiny. The max pooling layer MaxPool in the MP-2 module is replaced with the SPD-Conv module. The SPD-Conv module executes the above-mentioned downsampling method, concatenates the first semantic features together according to the channel dimension to obtain the concatenated semantic features, and then reduces the number of channels of the concatenated semantic features, so as to retain the features related to small targets on the high-level semantic feature map, and improve the small target detection performance of the model through multi-scale feature fusion.
[0078] Preferably, in step S1022, determining the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error includes: collecting multiple non-collinear control points on the two-dimensional image based on the multi-point averaging method, and marking the pixel values of the control points.
[0079] Determine the corresponding coordinates of the control points in the lidar image.
[0080] Solve the inverse perspective transformation matrix based on the control point pixel values and the corresponding coordinates.
[0081] Based on the inverse perspective transformation matrix, solve the reprojection error and the inverse perspective matrix through the Levenberg-Marquardt algorithm.
[0082] When the reprojection error is greater than or equal to the preset threshold, re-collect the control points to obtain new control point pixel values and new corresponding coordinates, and solve the new inverse perspective matrix and the new reprojection error until the finally obtained reprojection error is less than the preset threshold.
[0083] Specifically, the ground of the substation is relatively flat, and it can be assumed that it is an ideal plane with little change in the pitch angle. Therefore, the camera fixed on the head of the inspection robot can calculate the projection coordinates of any ground target in the field of view on the ground plane in the three-dimensional world through inverse perspective mapping (IPM for short).
[0084] Due to the weak target features caused by the open field of view of the substation, markers are usually added and the three-dimensional coordinates are manually obtained in combination with a high-precision depth sensor to obtain the two-dimensional to three-dimensional point pair relationship. However, manually extracting coordinates will introduce large errors. For this reason, the embodiment of the present invention determines the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error, so as to process the two-dimensional detection frame to obtain the three-dimensional positioning coordinates.
[0085] Exemplarily, the test environment includes: parking the inspection robot on a straight road, and the left and right sides of the road meet at the vanishing point E. Taking the coordinate system of the lidar firmly installed in the middle of the inspection robot as the reference coordinate system, ideally, the vanishing point is constrained to the x-axis of the reference coordinate system. Four moving personnel not standing on a straight line are used as test targets.
[0086] Based on the multi-point averaging method, four non-collinear control points A, B, C, and D are collected on the two-dimensional image of the above test environment. See Figure 8 shown, where the center point of the contact part of the feet of the moving personnel with the ground is used as the above control point. The pixel values of the four control points are marked as , , , , and the road vanishing point is marked as
[0087] It can be understood that collecting control points based on the multi-point averaging method can reduce errors.
[0088] Determine the corresponding coordinates of the above four control points in the lidar image, denoted as , , , 。
[0089] It can be understood that the above lidar image is the lidar image of the above experimental environment, and the above lidar image and the two-dimensional image of the above test environment are acquired by the inspection robot at the same time. To reduce errors, the corresponding coordinates are also determined based on the multi-point averaging method.
[0090] Taking the control point A as an example, according to the perspective transformation relationship, we get: ; In the formula, represents the inverse perspective transformation matrix.
[0091] Substitute and , and we get: ; ; Perform the same processing on the control points B, C, and D, and solve the inverse perspective transformation matrix by the least squares method.
[0092] Based on the inverse perspective transformation matrix , solve the reprojection error and the inverse perspective matrix through the Levenberg-Marquardt algorithm. Among them, the Levenberg-Marquardt (LM for short) algorithm is an improvement and extension of the least squares method, which is used to solve the non-linear least squares problem. In the embodiment of the present invention, the cvGetPerspectiveTransform() function in the OpenCV library (Open Source Computer Vision Libarary) is used to solve the inverse perspective matrix.
[0093] It can be understood that in the case where the reprojection error is greater than or equal to the preset threshold, re-collect the control points A, B, C, and D to obtain the new control point pixel values and the new corresponding coordinates, and solve the new inverse perspective matrix and the new reprojection error until the finally obtained reprojection error is less than the preset threshold. The preset threshold can be set by those skilled in the art according to the prior value and the actual situation.
[0094] It can be understood that the embodiment of the present invention processes the two-dimensional image data based on the inverse perspective transformation algorithm to obtain the first positioning information of the detection target, and can realize two-dimensional target detection and ranging positioning. For the corresponding algorithm framework, see Figure 9 as shown.
[0095] Exemplarily, Figure 9 the two-dimensional image in
[0096] Preferably, the deep learning algorithm in step S103 is the PointPillars algorithm.
[0097] Processing the 3D point cloud data based on the PointPillars algorithm to obtain the second positioning information, including: converting the 3D point cloud data into a 2D pseudo-image based on a pillar feature network.
[0098] Performing downsampling, upsampling, and splicing operations on the pseudo-image based on a backbone network to obtain the second semantic feature.
[0099] Processing the second semantic feature based on a second detection head network to obtain a 2D detection result from a bird's-eye view perspective.
[0100] Estimating the height of the 3D detection box based on the second detection head network according to the 3D point cloud data.
[0101] Obtaining the 3D detection box based on the 2D detection result from a bird's-eye view perspective and the estimated height value of the 3D detection box.
[0102] Among them, the PointPillars algorithm includes a pillar feature network, a backbone network, and a second detection head network, and the second positioning information includes a 3D detection box and corresponding positioning coordinates.
[0103] Specifically, the PointPillars algorithm constructs an end-to-end 3D object detection network. The 3D point cloud data is obtained by lidar scanning and is represented in a Cartesian coordinate system.
[0104] In the pillar feature network, grids are divided according to the X-axis and Y-axis of the Cartesian coordinate system, and all the point cloud data falling into the same grid are considered to form a pillar.
[0105] A point cloud is a set of points, and point cloud data is the storage and description of the point cloud. Each point in the point cloud is represented by a 9-dimensional vector where, represents the true coordinate information of the point, represents the reflection intensity of the point, represents the geometric center of all points in the pillar where the point is located, reflects the relative position of the point to the geometric center.
[0106] Sample N pieces of point cloud data in each pillar. If the number of point cloud data in each pillar exceeds N, randomly sample to N; if the number of point cloud data in each pillar is less than N, the corresponding part is filled with 0. Represent the data of P pillars as (D, P, N), that is, stacked pillars, where D represents the corresponding vector.
[0107] In the cylindrical feature network, a simplified version of PointNet is used to process the stacked cylinders, obtaining a tensor of size (C, P, N). Max Pooling is performed according to the dimension where the stacked cylinders are located, and a feature map of dimension (C, P) can be obtained.
[0108] The first dimension of the feature map is transformed into (H, W) to obtain a two-dimensional pseudo-image, so that the point cloud can be processed by two-dimensional convolution, greatly improving the efficiency of the algorithm operation.
[0109] In the backbone network, a feature map with more advanced semantic features is obtained through a top-down network downsampling, and then the feature maps are upsampled to the same size and concatenated, thereby expanding the receptive field of the feature map, which is beneficial for the network to detect smaller objects.
[0110] Exemplarily, in the second detection head network, a Single Shot Multibox Detector (SSD detection head for short) is used to perform 3D object detection, obtaining a 2D detection result from the bird's-eye view. At the same time, the height estimation of the 3D detection box is taken as an additional regression task according to the 3D point cloud data.
[0111] It can be understood that the additional regression task means measuring the difference between the predicted height (estimated height) and the true height through the corresponding loss function, and combining the true height information contained in the 3D point cloud data to accurately predict the height of the 3D detection box during the process of updating the network.
[0112] Exemplarily, the above 3D detection box is defined by where, represents the center position of the 3D detection box, represents the size of the 3D detection box, represents the rotation angle of the 3D detection box. The regression residuals between the predicted value and the true value of the 3D detection box are defined as follows: ; ; ; In the formula, represents the center coordinates of the true value box; represents the center coordinates of the prior box (anchor box); represents the size of the true value box; represents the size of the prior box; represents the rotation angle of the true value box; represents the rotation angle of the prior box; Represents the normalization parameter, which is used to normalize the position residual to facilitate network learning.
[0113] The loss function of the above point-column algorithm is: ; In the formula, represents the number of anchors, represents the localization loss, represents the object classification loss, represents the angle localization loss, = 2, = 1, = 0.2.
[0114] Localization loss Specifically: ; In the formula, is the smooth L1 loss function.
[0115] Since the angle localization loss cannot distinguish flipped boxes, the softmax classification loss function is used to calculate the angle localization loss .
[0116] Object classification loss Specifically: ; In the formula, is the balance parameter, which is used to balance positive and negative samples, represents the probability that the model prediction point belongs to the foreground; is the influence parameter, which is used to reduce the loss contribution of easily distinguishable samples.
[0117] Preferably, in step S104, based on the Wasserstein distance metric algorithm, the first localization information and the second localization information are processed to obtain matching information, including: step S1041, projecting the three-dimensional detection box based on the second localization information onto the two-dimensional image plane to obtain a three-dimensional projected detection box.
[0118] Step S1042, based on the Wasserstein distance metric algorithm, calculate the similarity between the three-dimensional projected detection box and the two-dimensional detection box of the first localization information according to the extended cost matrix, where the extended cost matrix includes the coordinates of the three-dimensional projected detection box, the coordinates of the two-dimensional detection box, and the coordinates of the virtual target.
[0119] Step S1043, based on the Hungarian algorithm, match the two-dimensional detection box and the three-dimensional projected detection box according to the similarity to obtain matching information, and the above matching information is the localization coordinates in the first localization information and the localization coordinates in the second localization information.
[0120] Preferably, in step S1041, the three-dimensional detection box based on the second positioning information is projected onto the two-dimensional image plane, and obtaining the three-dimensional projection detection box includes: obtaining the coordinates of the corner points of the three-dimensional detection box in the lidar coordinate system.
[0121] Based on the camera intrinsic matrix, the rotation matrix and the translation matrix from the lidar coordinate system to the camera coordinate system, the coordinates of the projection points of the corner points on the two-dimensional image plane are obtained.
[0122] Calculate the convex hull based on the projection point coordinates.
[0123] Take the minimum bounding rectangle of the convex hull as the three-dimensional projection detection box.
[0124] Specifically, considering that the projection of the three-dimensional detection box on the two-dimensional image plane should fit closely with its corresponding two-dimensional detection box, after projecting the three-dimensional detection box onto the two-dimensional image plane, take the minimum bounding matrix as the corresponding projection detection box, and this projection detection box is two-dimensional.
[0125] Obtain the coordinates of the corner points of the three-dimensional detection box in the lidar coordinate system , and the projection point of this point in the two-dimensional image can be obtained through the coordinate transformation formula , and the calculation is as follows: ; In the formula, K represents the camera intrinsic matrix, R represents the rotation matrix from the lidar coordinate system to the camera coordinate system, and T represents the translation matrix from the lidar coordinate system to the camera coordinate system.
[0126] Exemplarily, after obtaining the projection point coordinates of each corner point, the embodiment of the present invention uses the convexHull function in the OpenCV library to calculate the convex hull.
[0127] The embodiment of the present invention uses the minAreaRect function in the OpenCV library to calculate the minimum bounding rectangle of the above convex hull, so as to determine the above three-dimensional projection detection box.
[0128] Exemplarily, the embodiment of the present invention provides a test image for converting a three-dimensional detection box into a three-dimensional projection detection box, see Figure 10 as shown.
[0129] Specifically, in step S1042, based on the Wasserstein distance metric algorithm, calculate the similarity between the three-dimensional projection detection box and the two-dimensional detection box of the first positioning information according to the extended cost matrix.
[0130] It is considered that the two-dimensional object detection algorithm provided by the embodiment of the present invention can detect distant objects, but the obtained position information is not accurate enough; due to the sparse corresponding point cloud and the lack of semantic features of distant objects, the false detection rate and missed detection rate of the three-dimensional object detection algorithm provided by the embodiment of the present invention are relatively high.
[0131] In other words, in actual situations, restricted by the characteristics of the data source, the detection output results of the two-dimensional object detection algorithm and the three-dimensional object detection algorithm provided by the embodiment of the present invention cannot correspond one by one.
[0132] Therefore, the embodiment of the present invention expands the cost matrix and adds virtual objects to the object set with fewer detected objects. At the same time, in order to prevent virtual objects from affecting the matching result, a relatively large matching cost can be assigned to each virtual object.
[0133] It is considered that due to calibration errors or detection errors, the three-dimensional projection detection frame of small objects may not coincide with the two-dimensional detection frame. Therefore, the embodiment of the present invention uses the Wasserstein distance metric algorithm to measure the similarity between the three-dimensional projection detection frame and the two-dimensional detection frame.
[0134] Preferably, the embodiment of the present invention uses the Normalized Wasserstein Distance Loss (NWD Loss for short) algorithm to be applicable to the similarity measurement of tiny detection objects. Its normalization processing method can reduce the sensitivity of the model to the object size, making the performance of the model in detecting objects of different sizes more balanced.
[0135] Specifically, the horizontal object detection frame is modeled as a two-dimensional Gaussian distribution N, where represents the center point coordinates of the object detection frame, and w and h respectively represent the width and height of the object detection frame.
[0136] ; = , = 。
[0137] Exemplarily, the normalized Wasserstein distance metric loss algorithm measures the similarity between the three-dimensional projection detection frame and the two-dimensional detection frame through the Gaussian distribution modeling method, specifically: the three-dimensional projection detection frame , and the two-dimensional detection frame are respectively modeled as two-dimensional Gaussian distributions and , and the Wasserstein distance between the three-dimensional projection detection frame and the two-dimensional detection frame : ; In the formula, the superscript T represents the transpose, indicating the square of the Euclidean norm.
[0138] Perform exponential normalization on the Wasserstein distance to obtain the normalized Wasserstein distance , and the normalized Wasserstein distance loss , specifically: ; ; ; In the formula, C is an influence parameter used to characterize the data density.
[0139] and The greater the difference between the two, the lower the similarity, and the greater the value of the normalized Wasserstein distance .
[0140] Specifically, in step S1043, based on the Hungarian algorithm, match the first positioning information and the second positioning information according to the similarity to obtain the matching information.
[0141] The Hungarian algorithm can be used to solve the maximum matching problem of a bipartite graph and is simple and efficient.
[0142] Exemplarily, based on the detection result set obtained from the above two-dimensional object detection algorithm, and the detection result set of the above three-dimensional object detection algorithm, set up a matching list and initialize the cost matrix . The size of the initialized cost matrix is , where the function is used to obtain the object length or the number of elements; calculate the matching costs of each element in and respectively.
[0143] In the case of , add virtual rows to until becomes a square matrix.
[0144] In the case of , add virtual columns to until becomes a square matrix.
[0145] In the case of In the case of a square matrix, a relatively high matching cost is set for all virtual rows / columns.
[0146] After using the Hungarian algorithm to find the optimal match, remove the items in the match list that match the virtual rows / columns. Among them, the embodiments of the present invention call the function in the scipy library to solve the Hungarian algorithm.
[0147] The embodiments of the present invention calculate the similarity between the three-dimensional projection detection frame and the two-dimensional detection frame through the above Wasserstein distance metric algorithm, and perform corresponding matching through the above Hungarian algorithm, so as to realize the data association between the three-dimensional target detection data and the two-dimensional target detection data.
[0148] Specifically, in step S105, the matching information is fused based on the confidence level to obtain the detection result.
[0149] Considering that the above two-dimensional target detection algorithm provided by the embodiments of the present invention may have the following errors in the process of two-dimensional target detection and ranging (positioning): Camera distortion introduction error: In the inverse perspective transformation, the distortion of the target closer to the optical axis is relatively small, while the distortion of the target closer to the image edge is relatively large; at the same time, the inverse perspective transformation ranging error of the target closer to the camera is relatively small, and the inverse perspective transformation ranging error of the target farther from the camera is relatively large.
[0150] Detection accuracy introduction error: The confidence level represents the degree of belief of the detection algorithm in the detected target category and the degree of fit of the detection frame. The higher the confidence level, the more accurate the detection result. In the embodiments of the present invention, the midpoint of the bottom edge of the two-dimensional detection frame is used as the intersection point of the detected target and the ground plane in the image coordinate system. If the target detection accuracy decreases, the positioning accuracy will decrease accordingly.
[0151] Target occlusion introduction error: In two-dimensional target detection, the detected target may be occluded. For example, for the target at the edge of part of the image, its intersection point with the ground plane is outside the field of view, which will bring a large error to obtaining the correct positioning information of the detected target on the ground plane in the image coordinate system.
[0152] In order to reduce the influence of the above errors, improve the robustness of the algorithm, and achieve quantitative evaluation, the embodiments of the present invention calculate the confidence level of the positioning coordinates of the first positioning information through the following formula : ; In the formula, 1(·) is an indicator function, which is 1 when none of the sides of the object detection box intersects with the image edge, and 0 when at least one side of the object detection box intersects with the image edge; x and y represent the coordinates of the detected object in the lidar coordinate system; a and b represent the coordinates (a, b) of the camera optical center in the lidar coordinate system. represents the confidence of the two-dimensional detection box in the first positioning information.
[0153] It can be understood that the above object detection box refers to a two-dimensional detection box. The positioning coordinates of the first positioning information are three-dimensional positioning coordinates obtained through inverse perspective transformation, but the height coordinate is set to 0, so only the coordinates x and y need to be considered, which are used to characterize the ranging (positioning) result. The confidence of the two-dimensional detection box in the first positioning information The larger it is, the more accurately the detected object can be detected, but it does not characterize the ranging accuracy. The ranging accuracy of the above two-dimensional object detection algorithm is quantitatively characterized by the confidence as follows.
[0154] The matching information is fused through the following formula: ; In the formula, represents the center point position of the fused detected object, which is used to determine the final detection result; represents the positioning coordinates in the first positioning information, represents the positioning coordinates in the second positioning information, represents the normalized Wasserstein distance metric between the two-dimensional detection box and the three-dimensional projection detection box; is the confidence of the weight coefficient; Information mismatch means that the two-dimensional detection box does not match the three-dimensional projection detection box.
[0155] Exemplarily, the weight coefficient is 0.2.
[0156] It can be understood that, compared with the data-level and feature-level multi-sensor fusion methods, the above fusion method provided by the embodiment of the present invention belongs to the decision level, and it is less affected by time and space alignment and upstream modules, is convenient for adding sensors, and the system design is relatively simple.
[0157] The above method provided by the embodiment of the present invention makes full use of two-dimensional image data and three-dimensional point cloud data, combines the respective advantages of the improved two-dimensional object detection algorithm and the three-dimensional object detection algorithm, and thus outputs a more robust detection result.
[0158] An embodiment of the present invention also provides a target detection and positioning system for substation inspection, and this system executes the above method provided by the embodiment of the present invention.
[0159] An embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program is used to cause a computer to execute the method of the embodiment of the present invention when executed by a processor of the computer.
[0160] An embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The above memory stores a computer program that can be executed by the at least one processor, and the computer program is used to cause the electronic device to execute the method of the embodiment of the present invention when executed by the at least one processor.
[0161] Reference Figure 11 , the structural block diagram of the electronic device that can be used as the server or client of the embodiment of the present invention will now be described. It is an example of the hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0162] As Figure 11 shown, the electronic device includes a computing unit 1101, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1102 or the computer program loaded from the storage unit 1108 into the random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the electronic device can also be stored. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other through a bus 1104. The input / output (I / O) interface 1105 is also connected to the bus 1104.
[0163] Multiple components in the electronic device are connected to the I / O interface 1105, including: an input unit 1106, an output unit 1107, a storage unit 1108, and a communication unit 1109. The input unit 1106 can be any type of device capable of inputting information into the electronic device. The input unit 1106 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 1107 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1108 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 1109 allows the electronic device to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0164] The computing unit 1101 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly embodied in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 1102 and / or the communication unit 1109. In some embodiments, the computing unit 1101 can be configured to execute the above methods in any other suitable manner (e.g., by means of firmware).
[0165] The computer program for implementing the method of the embodiments of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0166] In the context of embodiments of the present inventive concept, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0167] It should be noted that the term "including" and its variations used in the embodiments of the present inventive concept are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present inventive concept are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more". The descriptions of terms "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features.
[0168] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present inventive concept are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or reject.
[0169] The various steps recorded in the method embodiments provided by the embodiments of the present inventive concept may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The protection scope of the present inventive concept is not limited in this regard.
[0170] As used herein, the term "embodiment" means that the specific features, structures, or characteristics described in connection with an embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean that it is independent or alternative to other embodiments and mutually exclusive. The embodiments in this specification are all described in a related manner, and the same or similar parts between the embodiments are cross-referenced. In particular, for embodiments of devices, equipment, and systems, since they are basically similar to embodiments of methods, the description is relatively simple, and the relevant parts refer to the partial description of the embodiments of methods.
[0171] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the appended claims.
Claims
1. A target detection and positioning method for substation inspection, characterized in that Including: Collecting two-dimensional image data and three-dimensional point cloud data of the current environment; Processing the two-dimensional image data based on the inverse perspective transformation algorithm to obtain the first positioning information of the detection target; Processing the three-dimensional point cloud data based on the deep learning algorithm to obtain the second positioning information of the detection target; Processing the first positioning information and the second positioning information based on the Wasserstein distance metric algorithm to obtain matching information; Fusing the matching information based on the confidence level to obtain the detection result.
2. The method according to claim 1, characterized in that, Processing the two-dimensional image data based on the inverse perspective transformation algorithm to obtain the first positioning information of the detection target, including: Detecting the two-dimensional image data based on the single-stage object detection algorithm to obtain a two-dimensional detection box; Determining the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error; Processing the two-dimensional detection box based on the inverse perspective matrix to obtain the three-dimensional positioning coordinates of the detection target; Wherein, the first positioning information includes the two-dimensional detection box and the three-dimensional positioning coordinates.
3. The method according to claim 2, wherein Detecting the two-dimensional image data based on the single-stage object detection algorithm to obtain a two-dimensional detection box, including: Extracting the first semantic features of the two-dimensional image data based on multiple extraction modules of the backbone network; Splicing the first semantic features together according to the channel dimension based on the downsampling module of the first detection head network to obtain the spliced semantic features; Reducing the channel number of the spliced semantic features based on the downsampling module to obtain the two-dimensional detection box; Wherein, the single-stage object detection algorithm includes the backbone network and the first detection head network, the extraction module includes a convolutional layer, a batch normalization layer and a scale scaling layer, and the convolutional layer of one of the multiple extraction modules adopts a full-dimensional dynamic convolutional network, and the full-dimensional dynamic convolutional network includes a position-level multiplication operation along the spatial dimension, a channel-level multiplication operation along the input channel dimension, a filter-level multiplication operation along the output channel dimension, and a kernel-level multiplication operation along the convolutional kernel space.
4. The method according to claim 2, wherein Determining the inverse perspective matrix of the inverse perspective transformation algorithm based on minimizing the reprojection error, including: Collecting multiple non-collinear control points on the two-dimensional image based on the multi-point averaging method and marking the pixel values of the control points; Determining the corresponding coordinates of the control points in the lidar image; Solving the inverse perspective transformation matrix based on the pixel values of the control points and the corresponding coordinates; Based on the inverse perspective transformation matrix, solving the reprojection error and the inverse perspective matrix through the Levenberg-Marquardt algorithm; In the case where the reprojection error is greater than or equal to the preset threshold, re-collecting the control points to obtain new control point pixel values and new corresponding coordinates, and solving the new inverse perspective matrix and the new reprojection error until the finally obtained reprojection error is less than the preset threshold.
5. The method according to claim 1, characterized in that, The deep learning algorithm is the PointPillars algorithm; Processing the three-dimensional point cloud data based on the PointPillars algorithm to obtain the second positioning information, including: Converting the three-dimensional point cloud data into a two-dimensional pseudo-image based on the cylindrical feature network; Performing downsampling, upsampling, and splicing operations on the pseudo-image based on the backbone network to obtain the second semantic features; Process the second semantic feature based on the second detection head network to obtain a two-dimensional detection result from a bird's-eye view perspective; Estimate the height of the three-dimensional detection box based on the three-dimensional point cloud data by the second detection head network; Obtain the three-dimensional detection box based on the two-dimensional detection result from the bird's-eye view perspective and the estimated value of the height of the three-dimensional detection box; Wherein, the point pillar algorithm includes the columnar feature network, the backbone network and the second detection head network, and the second positioning information includes the three-dimensional detection box and the corresponding positioning coordinates.
6. The method according to claim 1, characterized in that Process the first positioning information and the second positioning information based on the Wasserstein distance metric algorithm to obtain matching information, including: Project the three-dimensional detection box of the second positioning information onto the two-dimensional image plane to obtain a three-dimensional projected detection box; Calculate the similarity between the three-dimensional projected detection box and the two-dimensional detection box of the first positioning information based on the Wasserstein distance metric algorithm according to the extended cost matrix, wherein the extended cost matrix includes the coordinates of the three-dimensional projected detection box, the coordinates of the two-dimensional detection box, and the coordinates of the virtual target; Match the two-dimensional detection box and the three-dimensional projected detection box based on the similarity according to the Hungarian algorithm to obtain the matching information, and the matching information is the positioning coordinates in the first positioning information and the positioning coordinates in the second positioning information.
7. The method according to claim 6, wherein Project the three-dimensional detection box of the second positioning information onto the two-dimensional image plane to obtain a three-dimensional projected detection box, including: Obtain the coordinates of the corner points of the three-dimensional detection box in the lidar coordinate system; Based on the camera internal parameter matrix, the rotation matrix and the translation matrix from the lidar coordinate system to the camera coordinate system, obtain the projection point coordinates of the corner points on the two-dimensional image plane; Calculate the convex hull based on the projection point coordinates; Take the minimum circumscribed rectangle of the convex hull as the three-dimensional projected detection box.
8. The method according to claim 6, wherein Fuse the matching information based on the confidence level to obtain a detection result, including: ; In the formula, represents the center point position of the fused detection target and is used to determine the detection result; represents the positioning coordinates in the first positioning information, represents the positioning coordinates in the second positioning information, represents the normalized Wasserstein distance metric between the two-dimensional detection frame and the three-dimensional projection detection frame; is the confidence level is the weight coefficient of, and the confidence level is for the confidence level; Information mismatch means that the two-dimensional detection frame does not match the three-dimensional projection detection frame.
9. The method according to claim 8, characterized in that The method further includes: The confidence level is determined by the following formula :[[]]END]] ; where 1(·) is an indicator function, which is 1 when none of the sides of the object detection box intersects with the image edge, and 0 when at least one side of the object detection box intersects with the image edge; x and y represent the coordinates of the detected object in the lidar coordinate system; a and b represent the coordinates (a, b) of the camera optical center in the lidar coordinate system; represents the confidence of the two-dimensional detection box.
10. A target detection and positioning system for substation inspection, characterized in that, The system executes the method according to any one of claims 1 to 9.
Citation Information
Cited By
Road image processing method based on inverse perspective transformation and electronic equipment
CN122492431A