Method for identifying rod-shaped object in plateau mountain complex environment
Through fast 3D point nearest neighbor search, spherical projection and improved 3D-MiniNet network, combined with K nearest neighbor and normal vector angle constraint filtering technology, the precise positioning problem of rod-shaped objects in complex environments on plateaus and mountains is solved, and safe and efficient weeding of unmanned vehicles is achieved.
Patent Information
- Application Number
- CN202410024278.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the complex environment of plateau and mountainous areas, it is difficult for the existing technology to accurately locate rod-like objects, resulting in the unmanned vehicles being unable to accurately determine the location and range of weeding during the weeding process, which may cause damage to the equipment.
Fast 3D point nearest neighbor search combined with spherical projection and sliding window method, the improved 3D-MiniNet network performs semantic segmentation, and the K nearest neighbor method is used to assign semantic labels to each 3D point, combining normal vector angle constraints and RANSAC filtering technology for member segmentation.
It realizes accurate positioning and accurate identification of rod-shaped objects in complex plateau and mountainous environments, ensuring that unmanned vehicles can weed safely and efficiently and avoid damage to equipment.
Smart Images

Figure CN120279523A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and more specifically, to a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains. Background Art
[0002] Weed growth in solar photovoltaic power plants is a common problem. The height of weeds often affects the normal operation of solar panels, thereby affecting the power generation of photovoltaic power generation. As weeds grow, they will block solar panels, reducing the light they receive and resulting in a decrease in power generation efficiency. In addition, weeds can also damage the equipment of the photovoltaic power station, such as damaging solar panels and blocking ventilation openings, increasing the maintenance cost and difficulty. To solve this problem, effective measures need to be taken for weeding. Traditional weeding methods usually require manual operation, which is not only inefficient but also easily damages the equipment. Therefore, it is necessary to use unmanned vehicles for intelligent weeding. Unmanned vehicles have the advantages of high efficiency, precision, and no labor cost, and can automatically identify weeds and perform weeding to avoid damaging the equipment.
[0003] However, when an unmanned vehicle operates in a complex terrain (such as high plateaus and mountains), it is not only necessary to detect three-dimensional targets in the photovoltaic power station in real time, including the height of solar panels, other equipment, and weeds, but also to accurately output these three-dimensional position information, which can ensure that the unmanned vehicle can accurately judge the weeding position and weeding range and avoid unnecessary damage to the equipment. Due to the serious occlusion in the operation environment, GPS cannot be used for accurate positioning, so there is an urgent need to locate the information of rod-shaped objects. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains, which can accurately locate the information of rod-shaped objects, ensure that the unmanned vehicle can accurately judge the weeding position and weeding range, and avoid unnecessary damage to the equipment.
[0005] The technical solution of the present invention is as follows: A method for identifying rod-shaped objects in a complex environment of high plateaus and mountains, including the steps of:
[0006] S1. Fast 3D point nearest neighbor search. The original three-dimensional point cloud is projected through a spherical surface, and the point nearest neighbor search is performed in the spherical projection space using a sliding window method;
[0007] S2. Improve the DDRNet segmentation network in the 2D segmentation network MiniNetV2, perform semantic segmentation on the projected point cloud data, and retain the projection learning module;
[0008] S3. After semantic segmentation, use the K-nearest neighbor method for search to assign a semantic label to each 3D point;
[0009] S4. Segment the point cloud data of the rod members, and then perform filtering to obtain the point cloud data of each rod member.
[0010] Further, in S1, it specifically includes: projecting the input three-dimensional point cloud data by means of spherical projection to obtain the mapping of each 3D point and 2D point coordinates, setting the size of the sliding window to k×k, performing nearest neighbor search by sliding the window in the spherical projection space to generate P groups, each group containing N points, grouping each point in the spherical projection space, and each point is represented as an 11-dimensional enhanced feature, and the features are respectively 5 features of x, y, z, depth, and remission, the relative values of these 5 features, and the 3D Euclidean distance from each point to the mean point.
[0011] Further, in S2, the projection learning module converts the original 3D points into a 2D representation available for segmentation by extracting local features, context features, and spatial features.
[0012] Further, the layer structure of local feature extraction includes four layers, and each layer is respectively composed of a Linear linear layer, a BatchNorm batch normalization layer, and a LeakyRelu activation function layer.
[0013] Further, in S2, it specifically includes the following steps:
[0014] S21. Extract local features from the input point cloud data, and use the output of the second layer of local feature extraction as the input of the context feature;
[0015] S22. Perform max pooling operation to convert the features of the point cloud data into a fixed-size feature representation;
[0016] S23. Perform a fast nearest neighbor search on the pooled context features to obtain point groups and find the nearest neighbors most relevant to the current point;
[0017] S24. Perform convolution operations on each point group using 3×3 sliding windows with different dilation rates, and the dilation rates are 1, 2, and 3 respectively to capture context information at different scales;
[0018] S25. Use the padding operation with a padding value of 0 and a stride of 1;
[0019] S26. Perform an operation of a linear layer, a BatchNorm layer, and a LeakyRelu layer once, connect the local features and the context features together, and perform a max pooling operation.
[0020] Further, S2 also includes: concatenating the extracted spatial features and context features, and the concatenation method can be element-wise addition or element-wise multiplication; reorganizing the concatenated features to form a new tensor; applying a self-attention mechanism to the new tensor; and finally performing 1×1 convolution, batch normalization operation, and using an activation function to perform non-linear transformation on the processed features. Among them, the spatial features are extracted using a convolution with a kernel function of 1×N, where N is the size of the convolution kernel.
[0021] Further, the 3D-MiniNet network is improved as follows:
[0022] Step 1: Use the DDRNet deep dual-resolution network to improve the MiniNetV2 segmentation network. The DDRNet segmentation network includes: a low-resolution branch part with a convolution kernel of 3×3 and a stride of 2; a high-resolution branch part that uses 3×3 convolution to extract features; a DAPPM deep aggregation pyramid pooling module, which inputs low-resolution feature maps and combines feature aggregation and pyramid pooling to obtain context information at different scales, helping the DDRNet segmentation network better understand the local features of images or point cloud data.
[0023] Step 2: Fuse the high-resolution features and low-resolution features, and then perform bilinear interpolation upsampling and depthwise separable convolution operations on the fused features.
[0024] Step 3: Perform batch normalization on the features after depthwise separable convolution.
[0025] Step 4: Apply the LeakyRelu activation function to perform non-linear transformation on the features after batch normalization, and then fuse the features with the features extracted from spherical projection.
[0026] Step 5: Execute Step 2 to Step 4 again to obtain a segmentation map through a 3×3 convolution.
[0027] Further, in S3, the specific steps include:
[0028] S31: Set a search range threshold, which sets the maximum allowable distance, and only the points within this distance will be considered neighboring points.
[0029] S32: Use the improved MiniNetV2 segmentation network to predict the input 3D data to obtain a preliminary segmentation result.
[0030] S33: Identify the points that are not projected onto the two-dimensional sphere. For each point, use the K-nearest neighbor algorithm to search for the k closest points on the two-dimensional sphere according to its relative depth value.
[0031] S34. For each point, based on the k nearest neighbors found, count the number of occurrences of each semantic category among the k nearest neighbors, and then perform a consensus vote based on these counts;
[0032] S35. According to the result of the consensus vote, assign a semantic label to each 3D point, and the category with the majority vote is the label of this point.
[0033] Furthermore, both the K-nearest neighbor algorithm and the consensus vote are accelerated using a GPU.
[0034] Furthermore, in S4, the specific steps include:
[0035] S41. Use the Euclidean clustering point cloud segmentation algorithm based on the normal vector angle constraint to separate the adhered rods;
[0036] S42. Measure the similarity between adjacent points by calculating the angle between their normal vectors. The smaller the angle, the closer the normal vectors of the two points are, and they may belong to the same cluster;
[0037] S43. Use the RANSAC (Random Sample Consensus) method for line fitting and filtering, and retain the points belonging to the line to obtain accurate rod point cloud data.
[0038] For the present invention according to the above solution, its beneficial effects are as follows:
[0039] (1) A method for identifying rod-shaped objects in a complex environment of high plateaus and mountains provided by the present invention performs point nearest neighbor search in the projection space through fast 3D point nearest neighbor search, using the spherical projection and sliding window methods; this method can quickly locate and find associated points, thereby providing an accurate data basis for subsequent steps.
[0040] (2) A method for identifying rod-shaped objects in a complex environment of high plateaus and mountains provided by the present invention performs semantic segmentation on point cloud data in an improved 3D-MiniNet network, combining the 2D segmentation network of MiniNetV2 and the segmentation network of DDRNet, enhancing the functionality and generalization ability of the network. Secondly, retaining the projection learning module ensures the retention and accuracy of the original point cloud data during the processing.
[0041] (3) A method for identifying rod-shaped objects in a complex environment of high plateaus and mountains provided by the present invention, after semantic segmentation, uses the K-nearest neighbor method for search, assigns a semantic label to each 3D point, and can accurately identify the category of each point, improving the accuracy of subsequent processing.
[0042] (4) A method for identifying rod-shaped objects in a complex environment of high plateaus and mountains provided by the present invention performs segmentation processing on the point cloud data of the rods, and then performs filtering to obtain the point cloud data of each rod. This step ensures that each rod can be accurately identified and segmented, is particularly suitable for complex environments such as high plateaus and mountains, and can effectively identify rod-shaped objects in these environments, providing strong technical support for performing related tasks (such as monitoring, analysis, design, etc.) in complex environments such as high plateaus and mountains. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0044] Figure 1 It is a flowchart of a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains in an embodiment of the present invention;
[0045] Figure 2 It is a system block diagram of a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The following further describes the embodiments of the present invention in detail in conjunction with the drawings and the embodiments. The following detailed description and the drawings are used to exemplarily illustrate the principle of the present invention, but cannot be used to limit the scope of the present invention, that is, the present invention is not limited to the described embodiments.
[0047] In order to better understand the present invention, the following further describes the present invention in conjunction with the drawings and the embodiments:
[0048] See Figures 1 to 2 As shown, an embodiment of the present invention provides a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains, including the steps:
[0049] S1. Fast 3D point nearest neighbor search. The original three-dimensional point cloud is projected through a spherical projection, and the point nearest neighbor search is performed in the spherical projection space using a sliding window method;
[0050] In this embodiment, the S1 step specifically includes: projecting the input three-dimensional point cloud data by means of spherical projection to obtain the mapping of each 3D point and 2D point coordinates, setting the size of the sliding window to k×k, performing nearest neighbor search by sliding the window in the spherical projection space, generating P groups, each group containing N points, grouping each point in the spherical projection space, and each point is represented as an 11-dimensional enhanced feature, and the features are respectively 5 features of x, y, z, depth, and remission, the relative values of these 5 features, and the 3D Euclidean distance from each point to the mean point.
[0051] Specifically, a method for identifying rod-shaped objects in a complex environment of high plateaus and mountains provided by an embodiment of the present invention projects three-dimensional point cloud data into a two-dimensional space through spherical projection, reducing the dimension of the data, which can reduce the complexity of subsequent calculations while retaining sufficient information for object recognition. Secondly, using a sliding window to perform nearest neighbor search in the spherical projection space can quickly locate the neighboring points of each point. This search method is more efficient than the traditional nearest neighbor search in three-dimensional space, improving the processing speed; through nearest neighbor search and grouping, the point cloud data is divided into P groups, each group containing N points. This structured grouping helps to better utilize the local characteristics of the data in subsequent processing. Moreover, each point is grouped in the spherical projection space, and each point is represented as an 11-dimensional enhanced feature. These features cover information such as the spatial position (x, y, z), depth (depth), reflection intensity (remission) of the point, and the distance from the point to the mean point within the group. This rich feature representation helps to improve the accuracy of object recognition. Since this method is based on a sliding window and grouping, it can be easily extended to process larger-scale data sets. At the same time, the size of the sliding window and the method of feature extraction can also be adjusted according to needs to adapt to different requirements and application scenarios, and it is especially suitable for complex environments such as high plateaus and mountains, which provides strong technical support for related tasks (such as monitoring, analysis, design, etc.) in complex environments such as high plateaus and mountains.
[0052] S2. Improve the DDRNet segmentation network in the 2D segmentation network MiniNetV2, perform semantic segmentation on the projected point cloud data, and retain the projection learning module;
[0053] Specifically, by combining the 2D segmentation network MiniNetV2 and the DDRNet segmentation network, the improved 3D-MiniNet network can better learn and understand the semantic information of the point cloud data, so as to more accurately segment different objects. Secondly, MiniNetV2 is a lightweight 2D segmentation network. After combining with the DDRNet segmentation network, it can reduce the amount of calculation while maintaining high accuracy, improving the segmentation speed.
[0054] In this embodiment, the projection learning module converts the original 3D points into a 2D representation for segmentation by extracting local features, context features, and spatial features. The layer structure for local feature extraction consists of four layers, each of which is composed of a Linear layer, a BatchNorm layer, and a LeakyRelu activation function layer.
[0055] Specifically, in the projection learning module, the convolutional layer is used to extract local features from the input 3D point cloud. Through convolutional operations, the network can learn the features and patterns in the local regions of the point cloud. These features include the shape, texture, and spatial distribution of the points. The BatchNorm (batch normalization) layer is used to normalize the output of the convolutional layer to accelerate network training and improve the generalization ability of the model. Through the BatchNorm layer, the network can better handle the distribution changes of the input data, reduce internal covariate shift, and thus make the network training more stable. In the projection learning module, the LeakyRelu layer helps the network learn and model complex non-linear relationships, thus better understanding and representing the local features of the point cloud data. LeakyRelu is an activation function used to introduce non-linearity into the neural network. Compared with the traditional ReLU activation function, LeakyReLU has a non-zero slope in the negative part, which helps to avoid the problem of neuron "death". In the projection learning module, the linear layer is used to integrate and transform the features extracted by the previous layers for further processing or output. Through the linear layer, the network can flexibly adjust the representation of the features to meet the requirements of subsequent tasks.
[0056] Specifically, in step S2, it specifically includes the following steps:
[0057] S21. Extract local features from the input point cloud data, and use the output of the second layer of local feature extraction as the input of the context features; this step aims to extract local features from the input point cloud data. Local features refer to the characteristics of the points adjacent to each point, and these characteristics can provide useful information about the shape, surface texture, and orientation of the objects in the point cloud data.
[0058] S22. Perform max pooling operation to convert the features of the point cloud data into a fixed-size feature representation, which can reduce the dimension of the feature map and the number of parameters, while retaining the most important information. By converting the features of the point cloud data into a fixed-size feature representation, the generalization ability of the model can be improved and the computational amount can be reduced.
[0059] S23. Perform a fast nearest neighbor search on the pooled context features to obtain point groups and find the nearest neighbors most relevant to the current point. Specifically, perform a fast nearest neighbor search on the context features after pooling, aiming to find the nearest neighbors most relevant to the current point. The nearest neighbor search helps capture the local structural information in the point cloud data and is beneficial for improving the network's understanding of the shape and position of the object.
[0060] S24. Perform a convolution operation on each point group using 3×3 sliding windows with different dilation rates, which are 1, 2, and 3 respectively, to capture context information at different scales. Among them, the dilation rate is used to control the size and coverage range of the convolution kernel. By using different dilation rates, context information at different scales can be captured. This multi-scale convolution operation helps the network better understand the structure and details of the point cloud data.
[0061] S25. Use the padding operation with a padding value of 0 and a stride of 1. Specifically, the padding operation is used to add additional data to the edges of the input data to increase the receptive field of the network and prevent boundary effects. In step S25, using a padding value of 0 means not adding any additional data to the edges and keeping the size of the input data unchanged. A stride of 1 means that the distance moved each time during the convolution process is 1 unit.
[0062] S26. Perform an operation of a linear layer, a BatchNorm layer, and a LeakyRelu layer once, connect the local features and the context features together, and perform a max pooling operation. Specifically, the linear layer is used to integrate and transform the features, the BatchNorm layer is used for normalization processing, and the activation function of the LeakyRelu layer is LeakyReLU. The LeakyReLU activation function has a small slope (non-zero) when the input value is negative, which helps avoid the problem of neuron "death". In this embodiment, the LeakyReLU activation function is used to introduce non-linearity. This combination enables the network to learn more complex and diverse feature representations, and the max pooling operation further reduces the dimension of the output.
[0063] It is worth mentioning that in step S2 of this embodiment, it also includes: connecting the extracted spatial features and the context features together, and the connection method can be element-wise addition or element-wise multiplication; reorganizing the connected features to form a new tensor; applying a self-attention mechanism to the new tensor; finally, performing a 1×1 convolution, batch normalization operation, and using an activation function to perform a non-linear transformation on the processed features. Among them, the spatial feature extraction uses a convolution with a kernel function of 1×N to extract the spatial features, where N is the size of the convolution kernel.
[0064] Specifically, spatial features describe the position, orientation, or geometric properties of points in the point cloud, while context features capture the proximity relationships or spatial distributions of points. Connecting these two types of features can provide more comprehensive information, which helps to better understand the structure and patterns in the point cloud data. The self-attention mechanism is a method to enhance the attention to specific regions, which works by calculating the correlation scores between different positions in the input sequence. In point cloud data processing, the self-attention mechanism can help the model better understand the structure and patterns of the point cloud data and emphasize the neighboring points related to the current point. The 1×1 convolutional operation is a special convolution with a kernel size of 1×1, which is used to independently transform the features of each channel. This operation can enhance the representation ability of the model and allow the network to learn more complex and abstract feature representations. The batch normalization operation helps to accelerate training and improve the stability of the model by normalizing each batch of data to reduce the internal covariate shift, thereby helping the network better learn the distribution of the data.
[0065] Preferably, the 3D-MiniNet network is improved as follows:
[0066] Step 1: Use the DDRNet deep dual-resolution network to improve the MiniNetV2 segmentation network. The DDRNet segmentation network includes: a low-resolution branch part with a 3×3 convolutional kernel and a stride of 2; a high-resolution branch part that uses 3×3 convolutions to extract features; a DAPPM deep aggregation pyramid pooling module that inputs low-resolution feature maps and combines feature aggregation with pyramid pooling to obtain context information at different scales, which helps the DDRNet segmentation network better understand the local features of images or point cloud data;
[0067] Step 2: Fuse the high-resolution features and the low-resolution features, and then perform bilinear interpolation upsampling and depthwise separable convolution operations on the fused features;
[0068] Step 3: Perform batch normalization on the features after depthwise separable convolution;
[0069] Step 4: Apply the LeakyRelu activation function to perform a non-linear transformation on the features after batch normalization, and then fuse the features with the features extracted from the spherical projection;
[0070] Step 5: Execute Steps 2 to 4 again to obtain a segmentation map through a 3×3 convolution.
[0071] S3. After semantic segmentation, use the K-Nearest Neighbor (KNN) method for searching to assign a semantic label to each 3D point;
[0072] In this embodiment, the S3 step specifically includes:
[0073] S31, setting a search range threshold, which sets a maximum allowed distance, within which points are considered to be neighboring points;
[0074] S32. Use the improved MiniNetV2 segmentation network to predict the input 3D data and obtain preliminary segmentation results;
[0075] S33, identifying points that are not projected onto the two-dimensional spherical surface, and for each point, searching for the closest k points on the two-dimensional spherical surface using a K nearest neighbor algorithm according to its relative depth value;
[0076] S34, for each point, based on the searched k neighboring points, count the number of occurrences of each semantic category in the k neighboring points, and then perform consensus voting based on these counts;
[0077] S35. According to the result of consensus voting, a semantic label is assigned to each 3D point, and the category of the majority vote is the label of the point.
[0078] Furthermore, the K-nearest neighbor algorithm and consensus voting are both accelerated by GPU.
[0079] S4. Segment the point cloud data of the rods, and then filter them to obtain the point cloud data of each rod.
[0080] In this embodiment, step S4 specifically includes:
[0081] S41. Separate the adhered rods using the Euclidean clustering point cloud segmentation algorithm based on the normal vector angle constraint;
[0082] S42. The similarity between adjacent points is measured by calculating the angle between their normal vectors. The smaller the angle, the closer the normal vectors of the two points are, and they may belong to the same cluster.
[0083] S43. Use the RANSAC random sampling consensus method to perform straight line fitting and filtering, retain the points belonging to the straight line, and obtain accurate rod point cloud data.
[0084] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.
[0085] The above is an exemplary description of the present invention in conjunction with the accompanying drawings. It is obvious that the implementation of the present invention is not limited to the above-mentioned method. As long as various improvements are made by adopting the method concept and technical solution of the present invention, or the concept and technical solution of the present invention are directly applied to other occasions without improvement, they are all within the protection scope of the present invention.
Claims
1. A method for identifying rod-shaped objects in a complex plateau mountain environment, characterized in that, Including the steps: S1. Fast 3D point nearest neighbor search. The original three-dimensional point cloud is projected through spherical projection, and the point nearest neighbor search is performed in the spherical projection space using a sliding window method; S2. Improve the DDRNet segmentation network in the 2D segmentation network MiniNetV2, perform semantic segmentation on the projected point cloud data, and retain the projection learning module; S3. After semantic segmentation, use the K-nearest neighbor method for search to assign a semantic label to each 3D point; S4. Segment the point cloud data of the rod, and then perform filtering to obtain the point cloud data of each rod.
2. The method for identifying rod-shaped objects in a complex plateau mountain environment according to claim 1, wherein: In S1, it specifically includes: Project the input three-dimensional point cloud data through spherical projection to obtain the mapping of each 3D point and 2D point coordinates. Set the size of the sliding window to k×k. Perform nearest neighbor search by sliding the window in the spherical projection space to generate P groups, each group containing N points. Group each point in the spherical projection space, and each point is represented as an 11-dimensional enhanced feature, including 5 features of x, y, z, depth, remission, the relative values of these 5 features, and the 3D Euclidean distance from each point to the mean point.
3. The rod-shaped object recognition method in a complex environment of high plateaus and mountains according to claim 1, wherein: In S2, the projection learning module converts the original 3D points into a 2D representation available for segmentation by extracting local features, context features, and spatial features.
4. The method for identifying rod-shaped objects in a complex plateau mountain environment according to claim 3, wherein: The layer structure for local feature extraction includes four layers, each layer is composed of a Linear linear layer, a BatchNorm batch normalization layer, and a LeakyRelu activation function layer.
5. The method for identifying rod-shaped objects in a complex plateau mountain environment according to claim 3, wherein: In S2, it specifically includes the following steps: S21. Extract local features from the input point cloud data, and use the output of the second layer of local feature extraction as the input of the context feature; S22. Perform max pooling operation to convert the features of the point cloud data into a fixed-size feature representation; S23. Perform a fast nearest neighbor search on the pooled context features to obtain a point group and find the nearest neighbor points most relevant to the current point; S24. Perform convolution operations on each point group using 3×3 sliding windows with different dilation rates, which are 1, 2, and 3 respectively, to capture context information at different scales; S25. Use the padding operation with a padding value of 0 and a stride of 1; S26. Perform operations of a linear layer, a BatchNorm layer, and a LeakyRelu layer once, connect the local features and context features together, and perform a max pooling operation.
6. The rod-shaped object recognition method in a complex plateau mountain environment according to claim 5, wherein: S2 also includes: Connect the extracted spatial features and context features together, and the connection method can be element-wise addition or element-wise multiplication; Recombine the connected features to form a new tensor; Apply the self-attention mechanism to the new tensor; Finally, perform 1×1 convolution, batch normalization operation, and use an activation function to perform non-linear transformation on the processed features. Among them, the spatial feature extraction uses a convolution with a kernel function of 1×N to extract spatial features, where N is the size of the convolution kernel.
7. The rod-shaped object recognition method in a complex plateau mountain environment according to claim 1, wherein: The 3D-MiniNet network is improved as follows: Step 1: Improve the MiniNetV2 segmentation network using the DDRNet deep dual-resolution network. The DDRNet segmentation network includes: a low-resolution branch part with a 3×3 convolutional kernel and a stride of 2; a high-resolution branch part that uses 3×3 convolutions to extract features; and a DAPPM deep aggregation pyramid pooling module that inputs low-resolution feature maps and combines feature aggregation with pyramid pooling to obtain context information at different scales, which helps the DDRNet segmentation network better understand the local features of image or point cloud data. Step 2: Fuse the high-resolution features and the low-resolution features, and then perform bilinear interpolation upsampling and depthwise separable convolution operations on the fused features; Step 3: Perform batch normalization on the features after depthwise separable convolution; Step 4: Apply the LeakyRelu activation function to perform a non-linear transformation on the batch-normalized features, and then fuse the features with the features extracted from the spherical projection; Step 5: Execute Steps 2 to 4 again to obtain a segmentation map through a 3×3 convolution.
8. The method for identifying rod-shaped objects in a complex plateau mountain environment according to claim 1, characterized in that: In S3, the specific steps include: S31: Set a search range threshold that sets the maximum allowed distance within which points are considered neighboring points; S32: Use the improved MiniNetV2 segmentation network to predict the input 3D data to obtain a preliminary segmentation result; S33: Identify the points that are not projected onto the two-dimensional sphere. For each point, use the k-nearest neighbor algorithm to search for the k closest points on the two-dimensional sphere according to its relative depth value; S34: For each point, based on the k neighboring points found, count the number of occurrences of each semantic category among the k neighboring points, and then perform a consensus vote based on these counts; S35: According to the result of the consensus vote, assign a semantic label to each 3D point, and the category with the majority vote is the label of the point.
9. The method for identifying rod-shaped objects in a complex highland mountain environment according to claim 8, wherein: Both the k-nearest neighbor algorithm and the consensus vote are accelerated using a GPU.
10. The method for identifying rod-shaped objects in a complex highland and mountain environment according to claim 8, characterized in that: In S4, the specific steps include: S41: Use the Euclidean clustering point cloud segmentation algorithm based on normal vector angle constraint to separate the adhered rods; S42: Measure the similarity between neighboring points by calculating the angle between their normal vectors. The smaller the angle, the closer the normal vectors of the two points are and they may belong to the same cluster; S43: Use the RANSAC random sample consensus method for line fitting and filtering, and retain the points belonging to the line to obtain accurate rod point cloud data.