Construction method of cylinder type full-sparse lightweight three-dimensional target detector based on sparse ConvNeXt

By adopting the construction method of a cylinder-type fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt in the autonomous driving system, the problem of waste of computing resources and low detection efficiency caused by the sparseness of point cloud data is solved, and efficient, accurate and real-time object detection is achieved.

CN120071328APending Publication Date: 2025-05-30CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510233629.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing point cloud data, it is difficult for the prior art to effectively solve the problem of sparsity, resulting in waste of computing resources and low detection efficiency.

Method used

Using the construction method of a cylinder-type fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt, an efficient detector architecture that can adapt to point cloud data characteristics is designed through sparse convolution technology and the advantages of ConvNeXt.

Benefits of technology

It realizes efficient processing of point cloud data, significantly reduces computing resource consumption, improves the accuracy and real-timeness of target detection, and meets the needs of high precision and low latency of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071328A_ABST
    Figure CN120071328A_ABST
Patent Text Reader

Abstract

The invention provides a method for constructing a cylinder type full-sparse lightweight three-dimensional target detector based on sparse ConvNeXt. The method comprises the following steps: S1, acquiring point cloud data; s2, performing feature coding on the point cloud data; s3, inputting the data into an SCN-Pill network architecture module, wherein the SCN-Pill network architecture module comprises a down-sampling module, a sub-manifold sparse feature extraction module and a sparse ConvNeXt feature extraction module; and S4, outputting a detection result. According to the obtained point cloud data, the accuracy and the real-time performance of target detection can be realized by using the method, and the calculation amount is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method for constructing a columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt. Background Art

[0002] In the context of the rapid development of autonomous driving technology, multi-line lidar, as the core device for vehicles to perceive the surrounding environment, continuously generates a large amount of point cloud data. These point cloud data not only contain rich environmental information, but also pose severe challenges to three-dimensional object detection.

[0003] Point cloud data has significant high sparsity and unstructured characteristics. Among them, high sparsity is manifested as the extremely low proportion of valid data points representing target objects (such as vehicles, pedestrians, and obstacles, etc.) in three-dimensional space, and a large number of spatial regions are empty data. This results in the inevitable operation on a large number of invalid regions when traditional methods for processing dense data are used to process point cloud data, thus wasting a large amount of computing resources. The unstructured characteristic is reflected in the irregular distribution of data points in space, lacking the regular arrangement form of grid data, which further increases the difficulty of extracting effective features from the original point cloud.

[0004] To cope with these complex characteristics of point cloud data, existing technologies mainly convert unstructured point cloud data into a structured format to achieve efficient feature extraction. Among them, the method of constructing point clouds into voxels or pillars has become the mainstream research direction. In the voxel method, the Voxel R-CNN technology proposed by some people divides the original point cloud in the horizontal and vertical dimensions and divides it into multiple tiny grids, and each grid is a voxel. This method realizes the structured processing of point cloud data to a certain extent, but there are also significant limitations. First, voxel division is carried out comprehensively in three-dimensional space, resulting in an extremely large number of voxels. In the subsequent feature extraction stage, when three-dimensional convolution is used to process these voxel features, the computational complexity increases exponentially, and the demand for computing resources is extremely high, which is a great challenge for autonomous driving systems with limited hardware resources.

[0005] In contrast, the pillar method shows unique advantages. PointPillar proposed by some people is a typical representative. This method only divides the point cloud data in the horizontal dimension and divides it into several pillars. This method significantly reduces the number of data structures, thus greatly reducing the computational complexity, and is particularly suitable for scenarios with limited resources of edge devices, providing the possibility for autonomous driving systems to achieve real-time and efficient object detection.

[0006] However, despite the significant progress made by the pillar method in optimizing computational complexity, its sparsity problem has not been fundamentally solved. When traditional convolution operations process pillar data, they still include the entire input feature map with a large number of invalid data regions in the operation scope. This processing method is similar to screening areas without sand during the sand sifting process, and a large amount of computing resources are wasted on processing empty data regions, severely restricting the efficiency and real-time performance of object detection.

[0007] To address the above problems, sparse convolution technology has emerged. Its core advantage lies in being able to precisely focus the operations on valid data points (i.e., non-zero value points), and by skipping the operations in invalid regions, it achieves efficient utilization of computing resources and a significant improvement in computing efficiency, providing an innovative and practical solution for processing sparse data. At the same time, ConvNeXt, as a breakthrough in the field of convolutional neural networks, with its excellent performance and efficient design architecture, demonstrates strong feature extraction and representation capabilities in numerous image processing tasks.

[0008] In summary, how to effectively combine the advantages of sparse convolution technology and ConvNeXt to design a three-dimensional object detector architecture that can fully adapt to the characteristics of point cloud data and has high performance has become a key technical problem in the field of autonomous driving. A breakthrough in this innovative direction will have a profound impact on promoting autonomous driving technology to a higher level. Summary of the Invention

[0009] The present invention aims to at least solve the technical problems existing in the prior art, and particularly innovatively proposes a construction method for a pillar-style fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt.

[0010] To achieve the above object of the present invention, the present invention provides a construction method for a pillar-style fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt, including the following steps:

[0011] S1, obtaining point cloud data;

[0012] S2, performing feature encoding on the point cloud data;

[0013] S3, inputting the data into the SCN-Pillar network architecture;

[0014] S4, outputting the detection result.

[0015] In a preferred embodiment of the present invention, in step S3, the SCN-Pillar network architecture includes six fully sparse feature extraction modules, namely the 1st fully sparse feature extraction module, the 2nd fully sparse feature extraction module, the 3rd fully sparse feature extraction module, the 4th fully sparse feature extraction module, the 5th fully sparse feature extraction module, and the 6th fully sparse feature extraction module;

[0016] The data output end of the first full sparse feature extraction module is connected to the data input end of the second full sparse feature extraction module, the data output end of the second full sparse feature extraction module is connected to the data input end of the third full sparse feature extraction module, the data output end of the third full sparse feature extraction module is connected to the data input end of the fourth full sparse feature extraction module, the data output end of the fourth full sparse feature extraction module is connected to the data input end of the fifth full sparse feature extraction module, and the data output end of the fifth full sparse feature extraction module is connected to the data input end of the sixth full sparse feature extraction module;

[0017] Data is input from the data input end of the first full sparse feature extraction module;

[0018] The data output from the data output ends of the fourth full sparse feature extraction module, the fifth full sparse feature extraction module, and the sixth full sparse feature extraction module are merged and then output.

[0019] In a preferred embodiment of the present invention, the j-th full sparse feature extraction module includes Q sub-manifold sparse convolutional feature extraction modules and q feature extraction modules based on sparse next-generation convolutions; j = 1, 2, 3, 4, 5, 6;

[0020] The Q sub-manifold sparse convolutional feature extraction modules are respectively the first sub-manifold sparse convolutional feature extraction module, the second sub-manifold sparse convolutional feature extraction module, the third sub-manifold sparse convolutional feature extraction module,..., the Q-th sub-manifold sparse convolutional feature extraction module;

[0021] The q feature extraction modules based on sparse next-generation convolutions are respectively the first feature extraction module based on sparse next-generation convolutions, the second feature extraction module based on sparse next-generation convolutions, the third feature extraction module based on sparse next-generation convolutions,..., the q-th feature extraction module based on sparse next-generation convolutions;

[0022] The data output end of the first sub-manifold sparse convolutional feature extraction module is connected to the data input end of the second sub-manifold sparse convolutional feature extraction module, the data output end of the second sub-manifold sparse convolutional feature extraction module is connected to the data input end of the third sub-manifold sparse convolutional feature extraction module, the data output end of the third sub-manifold sparse convolutional feature extraction module is connected to the data input end of the fourth sub-manifold sparse convolutional feature extraction module,..., the data output end of the Q - 1 sub-manifold sparse convolutional feature extraction module is connected to the data input end of the Q-th sub-manifold sparse convolutional feature extraction module;

[0023] The data output end of the Q-th sub-manifold sparse convolutional feature extraction module is connected to the data input end of the first feature extraction module based on sparse next-generation convolutions;

[0024] The data output end of the first sparse-based next-generation convolution feature extraction module is connected to the data input end of the second sparse-based next-generation convolution feature extraction module. The data output end of the second sparse-based next-generation convolution feature extraction module is connected to the data input end of the third sparse-based next-generation convolution feature extraction module. The data output end of the third sparse-based next-generation convolution feature extraction module is connected to the data input end of the fourth sparse-based next-generation convolution feature extraction module, ……, the data output end of the (q - 1)th sparse-based next-generation convolution feature extraction module is connected to the data input end of the qth sparse-based next-generation convolution feature extraction module;

[0025] Data is input from the data input end of the first submanifold sparse convolution feature extraction module;

[0026] Data is output from the data output end of the qth sparse-based next-generation convolution feature extraction module.

[0027] In a preferred embodiment of the present invention, it further includes a downsampling module. The data output end of the downsampling module is connected to the data input end of the first submanifold sparse convolution feature extraction module;

[0028] Data is input from the data input end of the downsampling module.

[0029] In a preferred embodiment of the present invention, when q = 2;

[0030] The data output end of the downsampling module is connected to the data input end of the first submanifold sparse convolution feature extraction module;

[0031] The data output end of the first submanifold sparse convolution feature extraction module is connected to the data input end of the second submanifold sparse convolution feature extraction module. The data output end of the second submanifold sparse convolution feature extraction module is connected to the data input end of the third submanifold sparse convolution feature extraction module. The data output end of the third submanifold sparse convolution feature extraction module is connected to the data input end of the fourth submanifold sparse convolution feature extraction module, ……, the data output end of the (Q - 1)th submanifold sparse convolution feature extraction module is connected to the data input end of the Qth submanifold sparse convolution feature extraction module;

[0032] The data output end of the Qth submanifold sparse convolution feature extraction module is connected to the data input end of the first sparse-based next-generation convolution feature extraction module;

[0033] The data output end of the first sparse-based next-generation convolution feature extraction module is connected to the data input end of the second sparse-based next-generation convolution feature extraction module;

[0034] Input data at the data input end of the downsampling module;

[0035] Output data at the data output end of the second sparse-based next-generation convolution feature extraction module.

[0036] In a preferred embodiment of the present invention, the Jth sparse-based next-generation convolution feature extraction module includes a submanifold sparse convolution layer, a normalization layer, a dual linear layer, and a GELU activation function layer; J = 1, 2, 3,..., q; q represents the number of sparse-based next-generation convolution feature extraction modules;

[0037] The dual linear layer includes a first linear layer and a second linear layer;

[0038] The data output end of the submanifold sparse convolution layer is connected to the data input end of the normalization layer, the data output end of the normalization layer is connected to the data input end of the first linear layer, the data output end of the first linear layer is connected to the data input end of the GELU activation function layer, and the data output end of the GELU activation function layer is connected to the data input end of the second linear layer;

[0039] Input data at the data input end of the submanifold sparse convolution layer;

[0040] Merge the data input to the submanifold sparse convolution layer and the data output at the data output end of the second linear layer, and then output the data.

[0041] In a preferred embodiment of the present invention, the τth submanifold sparse convolution feature extraction module includes a dual submanifold sparse convolution layer, a dual batch normalization layer, and a dual ReLU activation function layer; τ = 1, 2, 3,..., Q; Q represents the number of submanifold sparse convolution feature extraction modules;

[0042] The dual submanifold sparse convolution layer includes a first submanifold sparse convolution layer and a second submanifold sparse convolution layer;

[0043] The dual batch normalization layer includes a first batch normalization layer and a second batch normalization layer;

[0044] The dual ReLU activation function layer includes a first ReLU activation function layer and a second ReLU activation function layer;

[0045] The data output end of the first submanifold sparse convolution layer is connected to the data input end of the first batch normalization layer, the data output end of the first batch normalization layer is connected to the data input end of the first ReLU activation function layer, the data output end of the first ReLU activation function layer is connected to the data input end of the second submanifold sparse convolution layer, and the data output end of the second submanifold sparse convolution layer is connected to the data input end of the second batch normalization layer;

[0046] Input data at the data input end of the first sub-manifold sparse convolution layer;

[0047] Merge the data input to the first sub-manifold sparse convolution layer and the data output at the data output end of the second batch normalization layer, and then input it to the data input end of the second ReLU activation function layer, and output the data from the data output end of the second batch normalization layer.

[0048] The present invention also discloses a computer system, including:

[0049] A processor;

[0050] A memory for storing processor-executable instructions;

[0051] Wherein, the processor is configured to implement the construction method of the columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt when executing the executable instructions.

[0052] The present invention also discloses a computer-readable storage medium, including:

[0053] A memory with a computer program stored thereon;

[0054] A processor for executing the program in the memory to implement the construction method of the columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt.

[0055] In summary, due to the adoption of the above technical solution, the present invention can utilize the present method to achieve the accuracy, real-time performance of object detection and reduce the computational amount according to the acquired point cloud data.

[0056] Additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] The above and / or additional aspects and advantages of the present invention will become apparent and easy to understand from the description of the embodiments in conjunction with the following drawings, wherein:

[0058] Figure 1 is a schematic flowchart of the present invention.

[0059] Figure 2 is a schematic connection diagram of the sparse ConvNeXt module of the present invention

[0060] Figure 3 is a schematic connection diagram of the FSFE-Module structure of the present invention.

[0061] Figure 4 is a schematic diagram of the SCN-Pillar network architecture of the present invention.

[0062] Figure 5 It is a schematic diagram of the connection of the submanifold sparse convolution module of the present invention. Specific implementation manner

[0063] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0064] The present invention provides a construction method of a columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt, as Figure 1 shown, including the following steps:

[0065] S1, obtaining point cloud data;

[0066] S2, performing feature encoding on the point cloud data;

[0067] S3, inputting the data into the SCN-Pillar network architecture;

[0068] S4, outputting detection results.

[0069] In a preferred implementation manner of the present invention, in step S3, the SCN-Pillar network architecture includes six fully sparse feature extraction modules, namely the 1st fully sparse feature extraction module, the 2nd fully sparse feature extraction module, the 3rd fully sparse feature extraction module, the 4th fully sparse feature extraction module, the 5th fully sparse feature extraction module, and the 6th fully sparse feature extraction module; as Figure 4 shown.

[0070] The data output end of the 1st fully sparse feature extraction module is connected to the data input end of the 2nd fully sparse feature extraction module, the data output end of the 2nd fully sparse feature extraction module is connected to the data input end of the 3rd fully sparse feature extraction module, the data output end of the 3rd fully sparse feature extraction module is connected to the data input end of the 4th fully sparse feature extraction module, the data output end of the 4th fully sparse feature extraction module is connected to the data input end of the 5th fully sparse feature extraction module, and the data output end of the 5th fully sparse feature extraction module is connected to the data input end of the 6th fully sparse feature extraction module;

[0071] Input data from the data input end of the 1st fully sparse feature extraction module;

[0072] Merge the data output from the data output ends of the 4th fully sparse feature extraction module, the 5th fully sparse feature extraction module, and the 6th fully sparse feature extraction module and then output the data.

[0073] In a preferred embodiment of the present invention, the j-th full sparse feature extraction module includes Q sub-manifold sparse convolutional feature extraction modules and q sparse-based next-generation convolutional feature extraction modules; j = 1, 2, 3, 4, 5, 6;

[0074] The Q sub-manifold sparse convolutional feature extraction modules are respectively the first sub-manifold sparse convolutional feature extraction module, the second sub-manifold sparse convolutional feature extraction module, the third sub-manifold sparse convolutional feature extraction module,..., the Q-th sub-manifold sparse convolutional feature extraction module;

[0075] The q sparse-based next-generation convolutional feature extraction modules are respectively the first sparse-based next-generation convolutional feature extraction module, the second sparse-based next-generation convolutional feature extraction module, the third sparse-based next-generation convolutional feature extraction module,..., the q-th sparse-based next-generation convolutional feature extraction module;

[0076] The data output end of the first sub-manifold sparse convolutional feature extraction module is connected to the data input end of the second sub-manifold sparse convolutional feature extraction module, the data output end of the second sub-manifold sparse convolutional feature extraction module is connected to the data input end of the third sub-manifold sparse convolutional feature extraction module, the data output end of the third sub-manifold sparse convolutional feature extraction module is connected to the data input end of the fourth sub-manifold sparse convolutional feature extraction module,..., the data output end of the Q - 1 sub-manifold sparse convolutional feature extraction module is connected to the data input end of the Q-th sub-manifold sparse convolutional feature extraction module;

[0077] The data output end of the Q-th sub-manifold sparse convolutional feature extraction module is connected to the data input end of the first sparse-based next-generation convolutional feature extraction module;

[0078] The data output end of the first sparse-based next-generation convolutional feature extraction module is connected to the data input end of the second sparse-based next-generation convolutional feature extraction module, the data output end of the second sparse-based next-generation convolutional feature extraction module is connected to the data input end of the third sparse-based next-generation convolutional feature extraction module, the data output end of the third sparse-based next-generation convolutional feature extraction module is connected to the data input end of the fourth sparse-based next-generation convolutional feature extraction module,..., the data output end of the q - 1 sparse-based next-generation convolutional feature extraction module is connected to the data input end of the q-th sparse-based next-generation convolutional feature extraction module;

[0079] Data is input from the data input end of the first sub-manifold sparse convolutional feature extraction module;

[0080] Data is output from the data output end of the q-th sparse-based next-generation convolutional feature extraction module.

[0081] In a preferred embodiment of the present invention, it further includes a downsampling module, and the data output end of the downsampling module is connected to the data input end of the first submanifold sparse convolution feature extraction module;

[0082] Data is input from the data input end of the downsampling module.

[0083] In a preferred embodiment of the present invention, when q is 2;

[0084] The data output end of the downsampling module is connected to the data input end of the first submanifold sparse convolution feature extraction module;

[0085] The data output end of the first submanifold sparse convolution feature extraction module is connected to the data input end of the second submanifold sparse convolution feature extraction module, the data output end of the second submanifold sparse convolution feature extraction module is connected to the data input end of the third submanifold sparse convolution feature extraction module, the data output end of the third submanifold sparse convolution feature extraction module is connected to the data input end of the fourth submanifold sparse convolution feature extraction module, ……, the data output end of the Q-1th submanifold sparse convolution feature extraction module is connected to the data input end of the Qth submanifold sparse convolution feature extraction module; as Figure 3 shown.

[0086] The data output end of the Qth submanifold sparse convolution feature extraction module is connected to the data input end of the first sparse-based next-generation convolution feature extraction module;

[0087] The data output end of the first sparse-based next-generation convolution feature extraction module is connected to the data input end of the second sparse-based next-generation convolution feature extraction module;

[0088] Data is input from the data input end of the downsampling module;

[0089] Data is output from the data output end of the second sparse-based next-generation convolution feature extraction module.

[0090] In a preferred embodiment of the present invention, the Jth sparse-based next-generation convolution feature extraction module includes a submanifold sparse convolution layer, a normalization layer, a double linear layer, and a GELU activation function layer; J = 1, 2, 3, ……, q; q represents the number of sparse-based next-generation convolution feature extraction modules;

[0091] The double linear layer includes a first linear layer and a second linear layer;

[0092] The data output end of the submanifold sparse convolution layer is connected to the data input end of the normalization layer, the data output end of the normalization layer is connected to the data input end of the first linear layer, the data output end of the first linear layer is connected to the data input end of the GELU activation function layer, and the data output end of the GELU activation function layer is connected to the data input end of the second linear layer; as Figure 2 shown.

[0093] Data is input from the data input end of the submanifold sparse convolution layer;

[0094] The data input to the submanifold sparse convolution layer and the data output from the data output end of the second linear layer are merged and then the data is output.

[0095] In a preferred embodiment of the present invention, the τ-th submanifold sparse convolution feature extraction module includes a double submanifold sparse convolution layer, a double batch normalization layer, and a double ReLU activation function layer; τ = 1, 2, 3,..., Q; Q represents the number of submanifold sparse convolution feature extraction modules;

[0096] The double submanifold sparse convolution layer includes a first submanifold sparse convolution layer and a second submanifold sparse convolution layer;

[0097] The double batch normalization layer includes a first batch normalization layer and a second batch normalization layer;

[0098] The double ReLU activation function layer includes a first ReLU activation function layer and a second ReLU activation function layer;

[0099] The data output end of the first submanifold sparse convolution layer is connected to the data input end of the first batch normalization layer, the data output end of the first batch normalization layer is connected to the data input end of the first ReLU activation function layer, the data output end of the first ReLU activation function layer is connected to the data input end of the second submanifold sparse convolution layer, and the data output end of the second submanifold sparse convolution layer is connected to the data input end of the second batch normalization layer; as Figure 5 shown.

[0100] Data is input from the data input end of the first submanifold sparse convolution layer;

[0101] The data input to the first submanifold sparse convolution layer and the data output from the data output end of the second batch normalization layer are merged and then input to the data input end of the second ReLU activation function layer, and the data is output from the data output end of the second batch normalization layer.

[0102] The present invention also discloses a computer system, including:

[0103] a processor;

[0104] a memory for storing processor-executable instructions;

[0105] Among them, when the processor is configured to execute the executable instructions, the construction method of the columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt is implemented.

[0106] The present invention also discloses a computer-readable storage medium, including:

[0107] A memory with a computer program stored thereon;

[0108] A processor for executing the program in the memory to implement the construction method of the columnar fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt.

[0109] In the context of the rapid development of autonomous driving technology, three-dimensional object detection, as the core link for an autonomous driving system to perceive the environment, directly affects the safety and reliability of vehicle driving. The present invention proposes a newly designed and efficient three-dimensional object detector architecture, aiming to fully adapt to the complex scene requirements of autonomous driving. Through a series of technological innovations, it achieves a perfect balance between high-precision, high real-time performance, and low computational resource consumption in point cloud data object detection.

[0110] 1) Improve the accuracy of object detection

[0111] The primary goal of the present invention is to significantly improve the accuracy of three-dimensional object detection. During the driving process, an autonomous driving vehicle needs to accurately identify key information such as the spatial position, size, shape, and posture of various target objects. Any deviation may lead to potential safety hazards. However, traditional object detectors are restricted by the sparsity and unstructured characteristics of point cloud data, prone to missed detections or false detections, and difficult to meet the requirements of complex traffic scenarios. The present invention introduces an innovative sparse ConvNeXt module and combines it with a unique fully sparse feature extraction architecture, which can deeply mine the potential information in point cloud data. Even in complex traffic environments or adverse weather conditions, this object detector can still accurately detect targets, significantly reducing the probability of missed detections and false detections, achieving high-precision identification of multiple targets such as vehicles, pedestrians, and cyclists, and providing a solid and reliable foundation for autonomous driving decision-making.

[0112] 2) Achieve real-time detection

[0113] Real-time performance is another key performance indicator for 3D object detection. In the context of autonomous driving, the environmental information changes rapidly, and every judgment and decision at each moment depends on the real-time performance of object detection. Existing object detectors often cause delays due to high computational complexity, making it difficult to meet the immediate response requirements of autonomous driving systems. By optimizing the feature extraction process, this invention makes full use of sparse convolution technology to reduce the burden of invalid calculations, combined with efficient voxel feature encoding and a streamlined network architecture design, significantly improving the data processing efficiency. This object detector can output accurate detection results in an extremely short time, helping autonomous vehicles perceive environmental changes in real time, make decisions in a timely manner, thus effectively avoiding potential risks and ensuring the smooth driving and safety performance of the vehicle.

[0114] 3) Reduce computational resource consumption

[0115] Autonomous vehicles have limited computational resources. Traditional object detectors consume a large amount of hardware resources due to a large number of redundant calculations, severely restricting the scalability and operating efficiency of the system. This invention deeply implements the concept of full sparse computing, and realizes the efficient utilization of computational resources by avoiding invalid calculations in the data processing flow. While maintaining high-precision and high-real-time detection capabilities, this invention significantly reduces the dependence on hardware resources, lightens the load on the vehicle's computing devices, extends the service life of the devices, and provides the possibility to integrate more intelligent functions into the system. In addition, this resource-saving design will create more conditions for the wide application of autonomous driving technology and its popularization in the intelligent transportation ecosystem, contributing core technical support to the future intelligent transportation development.

[0116] In summary, through the innovative design of the sparse ConvNeXt module, the full sparse feature extraction architecture, and the optimized voxel encoding technology, this invention provides a complete solution for the key requirements of 3D object detection. Its breakthroughs in accuracy, real-time performance, and resource efficiency can not only meet the core needs of complex autonomous driving scenarios but also provide the industry with a cutting-edge technical path with both theoretical value and application potential, driving the autonomous driving technology towards a higher level.

[0117] I. Voxel Feature Encoding

[0118] This invention proposes a voxel-based feature encoding method for efficiently processing point cloud data and converting it into a structured format suitable for subsequent calculations. The specific steps are as follows:

[0119] Grid division: According to the set voxel size, the original point cloud data is evenly divided into a regular grid structure in the horizontal dimension, and each grid contains the corresponding point cloud data set.

[0120] Feature expansion: Using a multi-layer perceptron (MLP) to expand the feature dimension of each point cloud data, mining its potential feature information, and laying the foundation for subsequent calculations.

[0121] Normalization and activation: Batch Normalization (BatchNorm) is performed on the point cloud data to alleviate the impact of data distribution differences on training. At the same time, the ReLU activation function is used to introduce non - linear factors, enhancing the network's fitting ability for complex data patterns.

[0122] Feature aggregation: By extracting the maximum value of the point cloud data inside the grid along each dimension, a columnar feature encoding is generated to form a structured representation.

[0123] Compared with the traditional voxel encoding method that requires a full partition in three - dimensional space and relies on three - dimensional convolution for feature extraction, the columnar feature encoding method of the present invention only requires two - dimensional convolution to complete feature extraction, significantly reducing the computational amount and model complexity, providing technical support for constructing an efficient and real - time three - dimensional object detector.

[0124] II. Construction of the Sparse ConvNeXt Module

[0125] The Sparse ConvNeXt module (the J - th feature extraction module based on sparse next - generation convolution) is as Figure 2 shown, and it is the core innovative part of the present invention, aiming to efficiently extract key features in point cloud data while minimizing the consumption of computing resources. Its design includes the following key technical points:

[0126] Sub - manifold sparse convolutional layer (SubMConv2D): Sub - manifold sparse convolution limits the convolution operation to only non - empty data positions, maintaining the original sparse characteristics of the input data, avoiding the deterioration of sparsity caused by convolution, and effectively reducing redundant calculations.

[0127] Optimization of the normalization layer: Layer Normalization (LayerNorm) is used to replace the traditional Batch Normalization (BatchNorm) to overcome the drawback of unstable performance under small - batch data. LayerNorm does not depend on the batch size, can accurately normalize the data distribution, and at the same time amplify subtle features, improving the model's feature expression and generalization ability.

[0128] The structure of two Multi-Layer Perceptrons (MLPs) is improved by referring to the inverted bottleneck structure of MobileNetV2. The two MLP layers are designed as follows: the first layer (the first linear layer) expands the feature dimension to four times the input dimension to enhance the feature processing ability; the second layer (the second linear layer) contracts the dimension back to the original scale to improve the information compression and transmission efficiency. This architecture significantly enhances the network's ability to capture complex point cloud features while reducing the computational overhead. Technical advantages: High efficiency, sparse convolution and optimized network design significantly reduce the computational resource requirements; Precision, innovative MLPs and normalization methods enhance the ability to capture point cloud details; Real-time performance: Optimized design for autonomous driving scenarios to ensure that the module can quickly respond to environmental changes.

[0129] III. Design of the Fully Sparse Feature Extraction Module (FSFE-Module) (the j-th fully sparse feature extraction module)

[0130] The Fully Sparse Feature Extraction Module (FSFE-Module) plays a crucial role in connecting the upper and lower parts in the target detector architecture. Its design aims to efficiently extract deep features from point cloud data while strictly controlling the consumption of computational resources. Figure 3 As shown in the figure, it plays a crucial role in connecting the upper and lower parts in the target detector architecture. Its design aims to efficiently extract deep features from point cloud data while strictly controlling the consumption of computational resources.

[0131] Module structure and key technologies:

[0132] 1) Downsampling module

[0133] Function: Reduce the size of the input feature map through spatial sparse convolution.

[0134] Characteristics: Only perform convolution calculations and feature aggregation on non-empty data regions, avoiding resource waste caused by global calculations. While reducing the resolution of the feature map, it effectively retains key feature information, laying a foundation for subsequent processing.

[0135] 2) Submanifold Sparse Convolution Module (the τ-th submanifold sparse convolution feature extraction module)

[0136] Core components: Dual submanifold sparse convolution layer, Batch Normalization (BatchNorm) layer, ReLU activation function, and residual structure.

[0137] Processing flow:

[0138] Convolution operation: Two submanifold sparse convolutions are used to mine deep features from the input data.

[0139] Normalization and activation: Batch normalization is performed after each convolution to stabilize the data distribution, and then a non-linear transformation is introduced through ReLU activation to further enhance the model's ability to fit complex data patterns.

[0140] Residual structure: Skip connections allow gradients to propagate backward more smoothly, alleviate the vanishing gradient problem in deep training, improve feature extraction efficiency, and maintain data sparsity to reduce computational resource consumption.

[0141] Sparse ConvNeXt module: Two sparse ConvNeXt modules are integrated at the backend of the module as deep miners for high-level semantic information and subtle feature differences. With an innovative structure and powerful feature processing capabilities, it provides a reliable feature expression foundation for subsequent object detection.

[0142] Overall module collaboration: Components collaborate closely to achieve layer-by-layer extraction and processing of sparse point cloud features. During the feature extraction process, it not only ensures the efficient expression of multi-scale depth features but also avoids resource waste caused by redundant calculations.

[0143] Technical advantages

[0144] Efficiency: Only calculate sparse data regions, significantly reducing computational resource overhead.

[0145] Accuracy: Through dual convolution and residual structures, fully explore the potential of features and enhance the expression ability for sparse point cloud data.

[0146] Real-time performance: The lightweight design of the module meets the high requirements for real-time environmental perception in autonomous driving scenarios.

[0147] Application value

[0148] The FSFE-Module provides core technical support for building lightweight and high-performance 3D object detectors, significantly improving the object detection ability of the model in scenarios such as autonomous driving.

[0149] IV. SCN-Pillar network architecture construction

[0150] The SCN-Pillar network architecture is as Figure 4 shown, which is the core integration solution of the technology of the present invention. Its design fully considers the characteristics of point cloud data and the requirements of object detection, and realizes an end-to-end object detection process through a modular and efficient structure.

[0151] 1) Data preprocessing and pillar feature encoding

[0152] After the network receives the original point cloud data, it first divides the data into multiple uniform grid structures according to a predetermined pillar size.

[0153] Pillar feature encoding: Using the pillar feature encoding method, convert the point cloud data within each grid into a pillar representation form with high-dimensional feature expression ability. This process completes the preliminary structured processing of the point cloud data, significantly enhancing the feature expression ability and providing a solid foundation for subsequent feature extraction.

[0154] 2) Multi-scale Feature Extraction: The fully sparse feature extraction module (FSFE-Module) is cascaded

[0155] The main body of the network is composed of six cascaded fully sparse feature extraction modules (FSFE-Module), constructing a multi-scale feature pyramid structure.

[0156] Convolution Kernel and Feature Pyramid Structure: To meet the real-time requirements of the autonomous driving scenario, the original large convolution kernel design of ConvNeXt is abandoned, and multi-scale and progressive feature extraction is realized through the feature pyramid structure. The feature pyramid structure reduces the computational complexity while ensuring the accurate extraction of large-scale features, thus greatly improving the detection speed and real-time performance.

[0157] Hierarchical Processing Flow: Each FSFE-Module sequentially performs downsampling and deep feature extraction operations on the input column data. As the hierarchy progresses, the resolution of the feature map gradually decreases, while the feature abstraction degree and semantic richness continuously increase, forming a multi-scale feature representation from low-level details to high-level semantics.

[0158] Sparsity Optimization: Each FSFE-Module strictly maintains the sparse characteristics of the data during the processing, avoiding redundant calculations and greatly improving the computational efficiency.

[0159] 3) Multi-scale Feature Fusion

[0160] To further enhance the target representation ability, the network performs multi-scale feature fusion operations based on the output feature maps of the last three FSFE-Modules.

[0161] Fusion Advantage: The combination of feature maps of different scales enables the generated final multi-scale feature map to have both rich details and high-level semantic information. By integrating the feature advantages of different levels, the adaptability to target diversity and complex environments is significantly improved.

[0162] 4) Detection Head Module and Target Output

[0163] Based on the fused multi-scale sparse feature map, the detection head module completes the target detection task.

[0164] Classification and Regression: The classification algorithm identifies the category of the target (such as vehicles, pedestrians, cyclists, etc.). The regression algorithm accurately predicts the spatial position, size, and pose of the target.

[0165] Output Information: The final network output includes complete information such as target category, position, size, and pose, meeting the high-precision requirements of the autonomous driving scenario for environmental perception.

[0166] Technical Advantages

[0167] High efficiency: The module structure designed based on the sparse characteristics significantly reduces the computational cost and meets the requirements of real-time detection.

[0168] Multi-scale representation ability: The feature pyramid and multi-scale feature fusion mechanism enable the network to accurately capture target features at different scales.

[0169] End-to-end process: From the input of the original point cloud to the output of the detection result, the architecture achieves seamless connection and is suitable for diverse application scenarios.

[0170] Application value

[0171] The SCN-Pillar network architecture provides efficient and accurate point cloud object detection capabilities for autonomous driving systems, has extremely strong environmental perception performance in complex dynamic scenarios, and provides reliable support for autonomous driving decision-making and planning.

[0172] V. Experimental tests

[0173] 5.1. Dataset and experimental settings

[0174] 1) Selection of dataset

[0175] This patent application selects the Waymo Open Dataset as the experimental basis. This dataset covers a variety of target types and scenarios and has high diversity and generalization ability.

[0176] Target annotation: Includes common target types such as vehicles, pedestrians, cyclists, etc., and the annotation is accurate and the scenarios are rich.

[0177] Data scale:

[0178] Training set: 798 training sequences, about 160,000 frames of data;

[0179] Validation set: 202 validation sequences, about 40,000 frames of data.

[0180] Difficulty level: The targets are divided into two difficulty levels according to the number of inliers in the point cloud to reflect the different detection performances for sparse and dense targets.

[0181] Evaluation metrics: The following two criteria are used to evaluate the model performance:

[0182] Average Precision (AP): Measures the overall accuracy of object detection.

[0183] Average Precision with Weighting (APH): Further emphasizes the prediction accuracy of the target direction and comprehensively considers the target localization and pose estimation capabilities.

[0184] 2) Training and inference configuration

[0185] To ensure the fairness and reproducibility of the experiments, the training and inference configurations follow standardized principles, and a consistent experimental environment is constructed:

[0186] Infrastructure: Based on the VoxelNeXt architecture, the code implementation of VoxelNeXt-2D is reproduced and adjusted to ensure the fairness of comparison between methods.

[0187] Hardware environment: The training and inference of the model are completed using a single NVIDIA A100 GPU.

[0188] Data allocation:

[0189] Training stage: The complete training set is used to fully utilize the large-scale data to enhance the learning ability of the model;

[0190] Validation stage: The model is evaluated on the complete validation set to ensure the objectivity and comprehensiveness of the performance test.

[0191] Target detector settings:

[0192] The parameters and configurations of the target detector completely follow the default settings of VoxelNeXt-2D to ensure benchmark consistency and lay a reliable foundation for subsequent performance comparison.

[0193] 3) Technical advantages and significance

[0194] Multi-source collected data: The Waymo dataset is collected by multiple sensors, covering rich scenarios and diverse targets, which helps to train a model with strong generalization ability. Direction-sensitive evaluation metric: As a weighted precision metric, APH enhances the consideration of the target direction prediction accuracy, further meeting the actual needs of autonomous driving. Consistent experimental environment: Standardized configurations and reproduction strategies ensure the reliability of the experimental results and the fairness of comparative research.

[0195] 5.2. Performance testing

[0196] 1) Vehicle detection task

[0197] The results are shown in Table 1. In the vehicle detection task, the SCN-Pillar target detector of the present invention performs excellently, second only to PillarNet-34, but significantly better than other comparison methods. Especially compared with voxel methods such as PV-RCNN that rely on 3D convolution, SCN-Pillar shows obvious advantages.

[0198] Table 1 Vehicle detection results

[0199]

[0200] Performance improvement: In different difficulty levels, the average precision (AP) and weighted average precision (APH) of SCN-Pillar are 0.3 higher than those of VoxelNet-2D, demonstrating its advantages in vehicle feature extraction and efficient processing.

[0201] Advantage analysis: SCN-Pillar can accurately capture the position, contour, and pose of vehicles, reducing misjudgments and missed detections in complex traffic scenarios and providing a safe path planning basis for autonomous driving systems.

[0202] 2) Pedestrian detection task

[0203] SCN-Pillar also demonstrates superior performance in pedestrian detection tasks compared to other methods. Compared with methods based on complex transformer architectures such as SWFormer, SCN-Pillar has a simple and efficient structure and outstanding performance. LEVEL 1 task: The average AP of SCN-Pillar exceeds that of SWFormer by 0.1, and the APH advantage reaches 0.5; LEVEL 2 task: The AP increases by 1.9, showing an obvious advantage. Compared with VoxelNeXt-2D: The AP of SCN-Pillar increases by more than 0.7 at each level, and the APH increases by more than 1.2, demonstrating its strong performance in crowded pedestrian scenarios.

[0204] Advantage analysis: SCN-Pillar can accurately identify pedestrians in different poses and occlusion levels, ensuring that autonomous vehicles can respond to pedestrian dynamics in a timely manner and guaranteeing pedestrian safety.

[0205] 3) Bicycle detection task

[0206] In bicycle detection tasks, SCN-Pillar also performs excellently, outperforming various voxel methods. LEVEL 1 task: Compared with PillarNet34, the AP and APH of SCN-Pillar increase by 2.3 respectively; LEVEL 2 task: The two indicators increase by 2.0. Compared with VoxelNeXt-2D: The AP and APH increase by more than 0.68 at each level.

[0207] Advantage analysis: SCN-Pillar can accurately extract bicycle features, quickly locate its trajectory, effectively reduce the collision risk in a mixed traffic environment, and improve the safety and reliability of autonomous driving.

[0208] 4) Comprehensive evaluation

[0209] Considering the mean Average Precision (mAP) and mean Average Precision with weighted Human (mAPH) metrics for all target types, SCN-Pillar comprehensively outperforms other compared object detectors. Compared with PillarNet34: SCN-Pillar exceeds by 0.5 and 0.6 in mAP and mAPH respectively; compared with VoxelNeXt-2D: the mAP and mAPH of SCN-Pillar are improved by 0.6 and 0.8 respectively.

[0210] Conclusion: These results prove that SCN-Pillar has excellent overall performance in various object detection tasks and can provide high-precision and stable object detection capabilities for autonomous driving systems.

[0211] 5.3. Ablation Study Conclusions

[0212] 1) Influence of the Sparse ConvNeXt Module

[0213] After integrating the Sparse ConvNeXt module into the VoxelNet-2D framework, the overall detection accuracy has been significantly improved, especially in pedestrian detection tasks.

[0214] Performance improvement: Using AP and APH as evaluation metrics, the detection accuracy has been improved by more than 0.7, and the improvement of APH compared to AP is more than 0.5.

[0215] Advantage analysis: The Sparse ConvNeXt module can better capture the characteristics of pedestrian pose changes by precisely focusing on the effective data area. Especially in complex pedestrian scenarios, it reduces the misjudgment of directions and improves the detection accuracy.

[0216] 2) Influence of the Convolution Kernel Size

[0217] When exploring the influence of the sparse convolution kernel size on performance, it is found that different tasks have different requirements for the convolution kernel:

[0218] Vehicle and bicycle detection: Although the 5×5 convolution kernel has some improvement, the improvement amplitude is less than 0.1. The reason is that these two types of targets occupy a large area and have regular structures in the point cloud. The large-size convolution kernel can cover some features, but the effect improvement is limited.

[0219] Pedestrian detection: The 3×3 convolution kernel has better results because the features of pedestrian targets in the point cloud are more scattered, and the smaller convolution kernel can capture local features more finely and improve the detection accuracy.

[0220] Performance analysis: When using the 3×3 convolution kernel, the detection speed of SCN-Pillar reaches 18.28 FPS, which is 2 FPS higher than 16.26 FPS of the 5×5 convolution kernel, meeting the real-time requirements of autonomous driving.

[0221] 5.4) Conclusion

[0222] The SCN-Pillar target detector proposed in the present invention has significant advantages in three-dimensional target detection in the field of autonomous driving, and comprehensively improves the performance and driving safety of the autonomous driving system.

[0223] 1) Breakthrough in detection accuracy

[0224] In vehicle detection tasks, SCN-Pillar can accurately identify subtle features of vehicles, whether it is the body contour, model style, or position and posture in complex traffic flow. Its innovative feature extraction mechanism and optimized architecture design effectively reduce the risk of false detection and missed detection in scenarios with multiple vehicles running in parallel and frequent occlusion, providing reliable path planning support for the autonomous driving system and ensuring accurate decision-making of vehicles in complex environments.

[0225] 2) Pedestrian Detection Advantages

[0226] In terms of pedestrian detection, SCN-Pillar shows more outstanding advantages. Facing pedestrians of various postures, various clothing and partial occlusion, SCN-Pillar can accurately capture the key feature points of pedestrians and accurately determine the pedestrian's walking direction, speed and behavioral intention. Especially in crowded environments such as busy city streets, campuses and commercial areas, SCN-Pillar provides powerful real-time dynamic tracking capabilities, significantly improving the accuracy of pedestrian direction prediction, ensuring that the autonomous driving system can respond to pedestrian behavior changes in a timely manner, thereby effectively preventing collisions and greatly enhancing pedestrian protection capabilities.

[0227] 3) Bicycle detection performance

[0228] SCN-Pillar also performs well in bicycle detection tasks. Bicycles have flexible structures and changeable motion states. SCN-Pillar can accurately identify the characteristics of bicycles under different riding postures, speeds, and environmental conditions. In mixed traffic environments, whether running parallel to motor vehicles or weaving between pedestrians, SCN-Pillar can quickly and accurately identify bicycle targets, providing a reasonable driving strategy for the autonomous driving system and enhancing the system's adaptability and safety in a complex traffic ecosystem.

[0229] 4) Detection efficiency and real-time performance

[0230] The SCN-Pillar optimizes the utilization of computing resources through a unique sparse ConvNeXt module and a full-sparse feature extraction architecture, significantly enhancing the processing efficiency. By focusing on the effective data regions and eliminating redundant calculations in traditional methods, the SCN-Pillar achieves a high-speed detection capability of 18.28 FPS, meeting the real-time perception requirements of autonomous driving systems. Especially during high-speed driving, the SCN-Pillar can quickly identify potential dangers and promptly make decisions such as avoidance, braking, or speed adjustment, thus ensuring the smoothness and safety of driving.

[0231] 5) Computational resource consumption and lightweight design

[0232] The SCN-Pillar features an efficient lightweight design. Its full-sparse computing architecture significantly reduces the amount of ineffective computations, enabling this object detector to operate efficiently on resource-constrained autonomous driving hardware platforms. The lower computational resource requirements not only relieve the load on in-vehicle computing devices but also extend the service life of the devices, reduce energy consumption and heat dissipation pressure. At the same time, this design creates room for integrating more intelligent functions into the system, such as real-time map updates, vehicle-to-vehicle communication and cooperation, and intelligent road condition prediction, comprehensively enhancing the intelligence level and user experience of autonomous driving systems.

[0233] 6) Promoting the development of autonomous driving technology

[0234] The innovative design of the SCN-Pillar object detector not only solves the core problems of object detection accuracy and efficiency in autonomous driving systems but also provides strong support for the safety, intelligence, and reliability of future transportation systems. This technological breakthrough provides a solid foundation for the advancement of autonomous driving from theoretical research to large-scale practical applications and promotes the construction of an efficient, safe, and intelligent future transportation system.

[0235] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the claims and their equivalents.

Claims

1. A method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt, characterized in that: The following steps are involved: S1, obtain point cloud data; S2, feature encoding of point cloud data; S3, inputs data into the SCN-Pillar network architecture; S4, output the detection results.

2. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: In step S3, the SCN-Pillar network architecture includes six fully sparse feature extraction modules, namely, a first fully sparse feature extraction module, a second fully sparse feature extraction module, a third fully sparse feature extraction module, a fourth fully sparse feature extraction module, a fifth fully sparse feature extraction module and a sixth fully sparse feature extraction module; The data output end of the first fully sparse feature extraction module is connected to the data input end of the second fully sparse feature extraction module, the data output end of the second fully sparse feature extraction module is connected to the data input end of the third fully sparse feature extraction module, the data output end of the third fully sparse feature extraction module is connected to the data input end of the fourth fully sparse feature extraction module, the data output end of the fourth fully sparse feature extraction module is connected to the data input end of the fifth fully sparse feature extraction module, and the data output end of the fifth fully sparse feature extraction module is connected to the data input end of the sixth fully sparse feature extraction module; Input data from the data input terminal of the first fully sparse feature extraction module; The data outputted from the data output terminal of the 4th fully sparse feature extraction module, the data output terminal of the 5th fully sparse feature extraction module and the data output terminal of the 6th fully sparse feature extraction module are combined and then outputted.

3. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: The jth fully sparse feature extraction module includes Q sub-manifold sparse convolution feature extraction modules and q sparse next-generation convolution-based feature extraction modules; j = 1, 2, 3, 4, 5, 6; The Q sub-manifold sparse convolution feature extraction modules are respectively a first sub-manifold sparse convolution feature extraction module, a second sub-manifold sparse convolution feature extraction module, a third sub-manifold sparse convolution feature extraction module, ..., a Qth sub-manifold sparse convolution feature extraction module; The q sparse next generation convolution based feature extraction modules are respectively a first sparse next generation convolution based feature extraction module, a second sparse next generation convolution based feature extraction module, a third sparse next generation convolution based feature extraction module, ..., a q sparse next generation convolution based feature extraction module; The data output end of the first sub-manifold sparse convolution feature extraction module is connected to the data input end of the second sub-manifold sparse convolution feature extraction module, the data output end of the second sub-manifold sparse convolution feature extraction module is connected to the data input end of the third sub-manifold sparse convolution feature extraction module, the data output end of the third sub-manifold sparse convolution feature extraction module is connected to the data input end of the fourth sub-manifold sparse convolution feature extraction module, ..., the data output end of the Q-1th sub-manifold sparse convolution feature extraction module is connected to the data input end of the Qth sub-manifold sparse convolution feature extraction module; The data output terminal of the Q-th sub-manifold sparse convolution feature extraction module is connected to the data input terminal of the first sparse next generation convolution based feature extraction module; The data output end of the first next generation convolution-based feature extraction module is connected to the data input end of the second next generation convolution-based feature extraction module, the data output end of the second next generation convolution-based feature extraction module is connected to the data input end of the third next generation convolution-based feature extraction module, the data output end of the third next generation convolution-based feature extraction module is connected to the data input end of the fourth next generation convolution-based feature extraction module, ..., the data output end of the q-1th next generation convolution-based feature extraction module is connected to the data input end of the qth next generation convolution-based feature extraction module; Input data from the data input terminal of the first sub-manifold sparse convolution feature extraction module; The data is outputted by the data output terminal of the qth sparse next generation convolution based feature extraction module.

4. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: It also includes a down-sampling module, wherein a data output terminal of the down-sampling module is connected to a data input terminal of the first sub-manifold sparse convolution feature extraction module; The data is input from the data input terminal of the down-sampling module.

5. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: When q is 2; The data output terminal of the downsampling module is connected to the data input terminal of the first sub-manifold sparse convolution feature extraction module; The data output end of the first sub-manifold sparse convolution feature extraction module is connected to the data input end of the second sub-manifold sparse convolution feature extraction module, the data output end of the second sub-manifold sparse convolution feature extraction module is connected to the data input end of the third sub-manifold sparse convolution feature extraction module, the data output end of the third sub-manifold sparse convolution feature extraction module is connected to the data input end of the fourth sub-manifold sparse convolution feature extraction module, ..., the data output end of the Q-1th sub-manifold sparse convolution feature extraction module is connected to the data input end of the Qth sub-manifold sparse convolution feature extraction module; The data output terminal of the Q-th sub-manifold sparse convolution feature extraction module is connected to the data input terminal of the first sparse next generation convolution based feature extraction module; The data output terminal of the first sparse next generation convolution based feature extraction module is connected to the data input terminal of the second sparse next generation convolution based feature extraction module; Input data from the data input terminal of the downsampling module; The data is output from the data output terminal of the second sparse next generation convolution based feature extraction module.

6. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: The Jth sparse next generation convolution based feature extraction module includes a submanifold sparse convolution layer, a normalization layer, a bilinear layer and a GELU activation function layer; J = 1, 2, 3, ..., q; q represents the number of sparse next generation convolution based feature extraction modules; The dual linear layer includes the first linear layer and the second linear layer; The data output end of the submanifold sparse convolution layer is connected to the data input end of the normalization layer, the data output end of the normalization layer is connected to the data input end of the first linear layer, the data output end of the first linear layer is connected to the data input end of the GELU activation function layer, and the data output end of the GELU activation function layer is connected to the data input end of the second linear layer; Input data from the data input terminal of the submanifold sparse convolution layer; The data input to the submanifold sparse convolution layer and the data output from the data output end of the second linear layer are combined and then output.

7. The method for constructing a cylindrical fully sparse lightweight three-dimensional object detector based on sparse ConvNeXt according to claim 1, characterized in that: The τ-th submanifold sparse convolution feature extraction module includes a dual submanifold sparse convolution layer, a dual batch normalization layer, and a dual ReLU activation function layer; τ = 1, 2, 3, ..., Q; Q represents the number of submanifold sparse convolution feature extraction modules; The dual submanifold sparse convolution layer includes the first submanifold sparse convolution layer and the second submanifold sparse convolution layer; The double batch normalization layer includes the first batch normalization layer and the second batch normalization layer; The double ReLU activation function layer includes the first ReLU activation function layer and the second ReLU activation function layer; The data output end of the first submanifold sparse convolution layer is connected to the data input end of the first batch normalization layer, the data output end of the first batch normalization layer is connected to the data input end of the first ReLU activation function layer, the data output end of the first ReLU activation function layer is connected to the data input end of the second submanifold sparse convolution layer, and the data output end of the second submanifold sparse convolution layer is connected to the data input end of the second batch normalization layer; Input data from the data input terminal of the first sub-manifold sparse convolutional layer; The data input to the first submanifold sparse convolution layer and the data output from the data output end of the second batch normalization layer are merged and input to the data input end of the second ReLU activation function layer, and the data output from the data output end of the second batch normalization layer is output.

8. A computer system, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method for constructing a cylindrical fully sparse lightweight three-dimensional target detector based on sparse ConvNeXt as described in one of claims 1 to 7 when executing the executable instructions.

9. A computer-readable storage medium, characterized in that: include: a memory having a computer program stored thereon; A processor is used to execute the program in the memory to implement the method for constructing a cylindrical fully sparse lightweight three-dimensional target detector based on sparse ConvNeXt as described in one of claims 1 to 7.