Method and system for detecting obstacle with nonstandard scale in channel, and readable storage medium
Through the sparse convolution and bird's-eye view feature pyramid structure combined with 2D multi-scale fusion convolution module and anchor point detection method, the accuracy of multi-scale objects in the detection of water surface obstacles in the waterway is solved, the detection accuracy and speed are improved, and the safety of autonomous ships is ensured.
Patent Information
- Application Number
- CN202510462829.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-15
AI Technical Summary
In the detection of water surface obstacles in the waterway, the detection effect of objects of the same category but with large size differences is poor, and it is prone to detect misalignment of detection size and position deviation, and missed detection and missed detection. Especially when large and small bridges, freighters and yachts coexist, it is difficult for 3D point cloud detection algorithms to accurately identify multi-scale obstacles.
The feature pyramid structure is constructed using sparse convolution and bird's-eye view, combined with the 2D multi-scale fusion convolution module and anchor point-based detection method, and the detection accuracy and speed of multi-scale obstacles are improved through the intersection and weighted non-maximum value suppression of variable distances.
It improves the accuracy of the autonomous driving perception module, ensures the safety of autonomous driving ships in the inland waterway, and realizes simultaneous detection of large and small-size targets, improving detection speed and accuracy.
Smart Images

Figure CN120496020A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and more specifically, to a method and system for detecting obstacles of irregular scale in a waterway, and a readable storage medium. Background Art
[0002] When applied to waterway surface obstacle detection, existing mainstream technologies perform poorly for objects of the same category but with significantly different sizes. Detection errors often occur, with detected sizes and positions significantly offset from their actual locations. Missed detections and false positives are also common. Furthermore, the shipping environment often contains objects of the same category but with significantly different scales (for example, bridges vary greatly in size, with some large and some small), as well as the coexistence of large and small objects (large cargo ships and small yachts). Furthermore, the point cloud of targets farther from the radar viewpoint is sparse, creating challenges for current 3D point cloud detection algorithms. Summary of the Invention
[0003] This application aims to solve or improve the above technical problems.
[0004] To this end, the first purpose of this application is to provide a method for detecting obstacles of irregular scale in a waterway.
[0005] The second purpose of this application is to provide a system for detecting obstacles of irregular scale in waterways.
[0006] The third purpose of this application is to provide a waterway irregular-scale obstacle detection system.
[0007] The fourth object of this application is to provide a readable storage medium.
[0008] To achieve the first purpose of the present application, the technical solution of the first aspect of the present application provides a method for detecting irregular-scale obstacles in waterways, including: obtaining point cloud data of irregular-scale obstacles in waterways; preprocessing the point cloud data; constructing a feature pyramid structure based on sparse convolution and a bird's-eye view; inputting the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; inputting the first extracted data into a 2D (2-dimensional) multi-scale fusion convolution module to obtain second extracted data; and detecting the second extracted data using an anchor-based method to obtain detection data.
[0009] According to the method for detecting irregular-scale obstacles in waterways provided in this application, point cloud data of irregular-scale obstacles in waterways is first acquired. The point cloud data is then preprocessed. A feature pyramid structure is then constructed based on sparse convolution and a bird's-eye view. The preprocessed point cloud data is input into the feature pyramid structure for feature extraction, yielding first extracted data. The first extracted data is then input into a 2D multi-scale fusion convolution module, yielding second extracted data. Finally, the second extracted data is tested using an anchor-based method to yield detection data. By replacing traditional 3D convolution with sparse convolution and using an anchor-based approach to inherit the output features of sparse convolution, the anchor-based approach allows for flexible configuration of prior bounding boxes of different sizes, solving the current problem of difficulty in accurately detecting objects at multiple scales. Furthermore, prior bounding boxes of multiple sizes ensure simultaneous detection of both large and small objects in the scene, improving the problem of missed detection of objects of different scales in the same scene. Sparse convolution also ensures the detection speed of the algorithm during deployment, enabling it to maintain real-time performance even when running on current autonomous driving chips. Furthermore, after obtaining 3D features through sparse convolution, deep 2D feature extraction is performed to further improve the detection accuracy of large objects and ensure the accuracy of the detection bounding box in the Bird's Eye View (BEV) perspective, thereby further improving the accuracy of 3D detection and the accuracy of the autonomous driving perception module. The application of sparse convolution significantly increases detection speed, facilitating the practical deployment of the algorithm and effectively ensuring the safety of autonomous ships in inland waterways.
[0010] In some technical solutions, the point cloud data is optionally preprocessed, including: dividing the point cloud data into voxels of equal size according to preset rules; and calculating the features of each voxel to obtain preprocessed data.
[0011] In this technical solution, the point cloud data is preprocessed. Specifically, each frame of the point cloud space is divided into voxels of equal size according to empirical rules, non-empty voxels are selected and the features of each voxel are calculated, so that the voxel data with features and its position index are sent to the feature pyramid structure to extract deep features.
[0012] In some technical solutions, optionally, the first extracted data is input into a 2D multi-scale fusion convolution module to obtain second extracted data, including: inputting the first extracted data into a 6-layer convolution module to obtain first processed data; inputting the first processed data into a 6-layer convolution module and halving the feature map size and doubling the number of feature channels to obtain second processed data; fusing and splicing the first processed data with the second processed data to obtain second extracted data.
[0013] In this technical solution, the first extracted data is input into the 2D multi-scale fusion convolution module to obtain the second extracted data. Specifically, the data passed down from the upstream first passes through a 6-layer convolution module, and then a separate copy of the processed data is retained. The processed data continues to pass through the 6-layer convolution module downward and the feature map size is halved, and the number of feature channels is doubled. This processed data is fused with the data that has just been retained separately, so that large-scale target features and small-scale feature targets can be fully learned.
[0014] In some technical solutions, optionally, the second extracted data is detected by an anchor point-based method to obtain detection data, including: detecting by variable distance intersection-over-union weighted non-maximum suppression to obtain detection data.
[0015] In this technical solution, the second extracted data is detected using an anchor-based method to obtain detection data. Specifically, this method uses variable-distance Intersection-in-Union (IoU) weighted non-maximum suppression to obtain detection data. This variable-distance IoU weighted non-maximum suppression can improve the detection rate of sparse objects in point clouds far from the viewpoint.
[0016] In some technical solutions, optionally, detection is performed by weighted non-maximum suppression of intersection-over-union (IoU) with variable distance to obtain detection data, including: calculating the Euclidean distances of multiple prediction boxes and their corresponding bounding boxes; performing weighted correction on the confidence; finding the prediction box with the highest confidence; calculating the IoU similarity between the prediction boxes; calculating the Gaussian weighted average between all prediction boxes; and adding the prediction boxes that meet the similarity threshold to the output list.
[0017] In this technical solution, detection is performed by weighted non-maximum suppression of the intersection-and-union ratio with variable distance to obtain detection data. Specifically, the Euclidean distance between multiple prediction boxes and their corresponding bounding boxes is first calculated. The confidence is then weighted and corrected. The prediction box with the highest confidence is then found. The intersection-and-union similarity between the prediction boxes is calculated. The Gaussian weighted average between all prediction boxes is calculated. Finally, the prediction boxes that meet the similarity threshold are added to the output list. Prediction boxes at close range are usually more accurate because the point cloud is very dense, while prediction boxes at long distances are poor, especially in the direction of the pre-model measurement box, which is prone to left and right swings. The distance between the center of the prediction box and the viewpoint is taken into account in the non-maximum suppression. The prediction box with the highest confidence is first selected. For prediction boxes at close range, those with a larger intersection-and-union ratio with it will be assigned higher weights. On the contrary, for prediction boxes at long distances, the prediction boxes around it will be assigned relatively uniform weights to obtain a smoother effect.
[0018] In some technical solutions, optionally, detection is performed through weighted non-maximum suppression of the intersection-over-union ratio of variable distance to obtain detection data, which also includes: judging whether all prediction boxes are popped up; if not, re-searching the prediction box with the highest confidence; if so, outputting the detection data.
[0019] In this technical solution, detection is performed through weighted non-maximum suppression of the intersection-over-union ratio with variable distance to obtain detection data. It also includes judging whether all prediction boxes pop up. If not, the prediction box with the highest confidence is searched again. If all pop up, the detection data is output.
[0020] In some technical solutions, optionally, before inputting the first extracted data into the 2D multi-scale fusion convolution module, the method further includes: splicing all voxel features at the same position.
[0021] In this technical solution, before the first extracted data is input into the 2D multi-scale fusion convolution module, all voxel features at the same position are also spliced. It can be understood that before adding the 2D multi-scale feature fusion convolution module, the 3D features extracted by the upstream sparse convolution need to be converted into 2D features. Densifying the sparse data means filling the empty voxel positions with 0. This will not increase the amount of calculation, and the features of different heights in the densified 3D features are spliced, thereby maximizing the retention of the spatial height z-axis information extracted by the sparse convolution, better retaining the spatial features of the object, and avoiding the object's spatial features being averaged even when there are large differences.
[0022] To achieve the second purpose of the present application, the technical solution of the second aspect of the present application provides a waterway irregular-scale obstacle detection system, including: an acquisition module for acquiring point cloud data of waterway irregular-scale obstacles; a preprocessing module for preprocessing the point cloud data; a model building module for constructing a feature pyramid structure based on sparse convolution and a bird's-eye view; a first extraction module for inputting the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; a second extraction module for inputting the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data; and a detection module for detecting the second extracted data through an anchor-based method to obtain detection data.
[0023] The waterway irregular scale obstacle detection system provided by the present application includes an acquisition module, a preprocessing module, a model building module, a first extraction module, a second extraction module and a detection module. Among them, the acquisition module is used to obtain point cloud data of waterway irregular scale obstacles. The preprocessing module is used to preprocess the point cloud data. The model building module is used to construct a feature pyramid structure based on sparse convolution and a bird's-eye view. The first extraction module is used to input the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data. The second extraction module is used to input the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data. The detection module is used to detect the second extracted data using an anchor-based method to obtain detection data. By replacing traditional 3D convolution with sparse convolution and using an anchor-based approach to inherit the output features of sparse convolution, the anchor-based approach allows for flexible configuration of prior bounding boxes of different sizes, solving the current problem of difficulty in accurately detecting objects at multiple scales. Furthermore, prior bounding boxes of multiple sizes ensure simultaneous detection of both large and small objects in the scene, improving the problem of missed detection of objects of different scales in the same scene. Sparse convolution also ensures the detection speed of the algorithm during deployment, enabling it to maintain real-time performance even when running on current autonomous driving chips. Furthermore, after obtaining 3D features through sparse convolution, deep 2D feature extraction is performed to further improve the detection accuracy of large objects and ensure the accuracy of the detection bounding box in the Bird's Eye View (BEV) perspective, thereby further improving the accuracy of 3D detection and the accuracy of the autonomous driving perception module. The application of sparse convolution significantly increases detection speed, facilitating the practical deployment of the algorithm and effectively ensuring the safety of autonomous ships in inland waterways.
[0024] In order to achieve the third purpose of this application, the technical solution of the third aspect of this application provides a waterway irregular scale obstacle detection system, including: a memory and a processor, wherein the memory stores a program or instruction that can be run on the processor, and when the processor executes the program or instruction, it implements the waterway irregular scale obstacle detection method of any one of the technical solutions of the first aspect, and therefore has the technical effect of any of the technical solutions of the first aspect mentioned above, which will not be repeated here.
[0025] In order to achieve the fourth purpose of this application, the technical solution of the fourth aspect of this application provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method for detecting irregular-scale obstacles in the waterway are implemented in any one of the technical solutions of the first aspect. Therefore, it has the technical effect of any of the technical solutions of the first aspect mentioned above, and will not be repeated here.
[0026] Additional aspects and advantages of the present application will become apparent in the following description or may be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which: Figure 1 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 2 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 3 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 4 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 5 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 6 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 7 This is a schematic flow chart of the steps of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 8 This is a schematic block diagram of the structure of a waterway irregular-scale obstacle detection system according to one embodiment of the present application; Figure 9 This is a schematic block diagram of the structure of a waterway obstacle detection system with irregular dimensions according to another embodiment of the present application; Figure 10 This is a schematic diagram of the location of adding a 2D multi-scale feature fusion convolution module to the method for detecting obstacles of irregular scale in waterways according to one embodiment of the present application; Figure 11 This is a schematic diagram of a specific construction process of a 2D multi-scale feature fusion convolution module of a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application; Figure 12 This is a flow chart of the variable distance IOU weighted NMS of the method for detecting obstacles with irregular scales in a waterway according to an embodiment of the present application.
[0028] in, Figure 8 and Figure 9 The corresponding relationship between the reference numerals and component names is as follows: 10: Waterway irregular scale obstacle detection system; 110: Acquisition module; 120: Preprocessing module; 130: Model building module; 140: First extraction module; 150: Second extraction module; 160: Detection module; 20: Waterway irregular scale obstacle detection system; 300: Memory; 400: Processor. DETAILED DESCRIPTION
[0029] In order to more clearly understand the above-mentioned objects, features and advantages of the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features therein can be combined with each other in the absence of conflict.
[0030] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present application is not limited to the specific embodiments disclosed below.
[0031] Refer to the following Figures 1 to 12 Describes methods and systems for detecting obstacles of irregular scale in waterways and readable storage media according to some embodiments of the present application.
[0032] like Figure 1 As shown, the embodiment of the first aspect of the present application provides a method for detecting obstacles of irregular scale in a waterway, comprising the following steps: Step S102: Acquire point cloud data of obstacles with irregular scales in the waterway; Step S104: pre-processing the point cloud data; Step S106: constructing a feature pyramid structure based on sparse convolution and bird's-eye view; Step S108: inputting the pre-processed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; Step S110: inputting the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data; Step S112: Detect the second extracted data using an anchor point-based method to obtain detection data.
[0033] According to the method for detecting irregular-scale obstacles in waterways provided in this embodiment, point cloud data of irregular-scale obstacles in waterways is first acquired. The point cloud data is then preprocessed. A feature pyramid structure is then constructed based on sparse convolution and a bird's-eye view. The preprocessed point cloud data is input into the feature pyramid structure for feature extraction, generating first extracted data. The first extracted data is then input into a 2D multi-scale fusion convolution module, generating second extracted data. Finally, the second extracted data is tested using an anchor-based method to generate detection data. By replacing traditional 3D convolution with sparse convolution and using an anchor-based approach to inherit the output features of sparse convolution, the anchor-based approach allows for flexible configuration of prior bounding boxes of different sizes, solving the current problem of difficulty in accurately detecting objects at multiple scales. Furthermore, prior bounding boxes of multiple sizes ensure simultaneous detection of both large and small objects in the scene, improving the problem of missed detection of objects of different scales in the same scene. Sparse convolution also ensures the detection speed of the algorithm during deployment, enabling it to maintain real-time performance even when running on current autonomous driving chips. Furthermore, after obtaining 3D features through sparse convolution, deep 2D feature extraction is performed to further improve the detection accuracy of large objects and ensure the accuracy of the detection bounding box in the Bird's Eye View (BEV) perspective, thereby further improving the accuracy of 3D detection and the accuracy of the autonomous driving perception module. The application of sparse convolution significantly increases detection speed, facilitating the practical deployment of the algorithm and effectively ensuring the safety of autonomous ships in inland waterways.
[0034] like Figure 2 As shown, according to a method for detecting obstacles of irregular scale in a waterway according to an embodiment of the present application, point cloud data is preprocessed, specifically comprising the following steps: Step S202: dividing the point cloud data into voxels of equal size according to a preset rule; Step S204: Calculate the features of each voxel to obtain pre-processed data.
[0035] In this embodiment, the point cloud data is preprocessed, specifically, each frame of the point cloud space is divided into voxels of equal size according to empirical rules, non-empty voxels are selected and the features of each voxel are calculated, so that the voxel data with features and its position index are sent to the feature pyramid structure to extract deep features.
[0036] like Figure 3 As shown, according to an embodiment of the present application, a method for detecting obstacles of irregular scales in a waterway is provided, wherein the first extracted data is input into a 2D multi-scale fusion convolution module to obtain the second extracted data, specifically comprising the following steps: Step S302: input the first extracted data into a 6-layer convolution module to obtain first processed data; Step S304: input the first processed data into a 6-layer convolution module and reduce the feature map size by half and double the number of feature channels to obtain second processed data; Step S306: fusing and splicing the first processed data with the second processed data to obtain second extracted data.
[0037] In this embodiment, the first extracted data is input into the 2D multi-scale fusion convolution module to obtain the second extracted data. Specifically, the data passed down from the upstream first passes through a 6-layer convolution module, and then a separate copy of the processed data is retained. The processed data continues to pass through the 6-layer convolution module downward and the feature map size is halved and the number of feature channels is doubled. This processed data is fused with the data that has just been retained separately, so that large-scale target features and small-scale feature targets can be fully learned.
[0038] like Figure 4 As shown, according to an embodiment of the present application, a method for detecting obstacles of irregular scale in a waterway is provided, wherein the second extracted data is detected by an anchor point-based method to obtain detection data, specifically comprising the following steps: Step S402: Detection is performed by weighted non-maximum suppression of intersection-over-union ratio with variable distance to obtain detection data.
[0039] In this embodiment, the second extracted data is detected using an anchor-based method to obtain detection data. Specifically, the detection data is obtained by performing weighted non-maximum suppression of the intersection-over-union ratio with variable distance. Using weighted non-maximum suppression with the intersection-over-union ratio with variable distance can improve the detection rate of sparse objects in point clouds far from the viewpoint.
[0040] like Figure 5 As shown, according to an embodiment of the present application, a method for detecting obstacles of irregular scale in a waterway is proposed, which detects obstacles by weighted non-maximum suppression of the intersection-over-union ratio with variable distances to obtain detection data, specifically comprising the following steps: Step S502: Calculate the Euclidean distances between the multiple prediction boxes and their corresponding bounding boxes; Step S504: performing weighted correction on the confidence level; Step S506: Find the prediction box with the highest confidence; Step S508: Calculate the intersection-over-union similarity between the predicted frames; Step S510: Calculate the Gaussian weighted average between all prediction frames; Step S512: adding the predicted boxes that meet the similarity threshold to the output list.
[0041] In this embodiment, detection is performed by weighted non-maximum suppression of the intersection-and-union ratio with variable distance to obtain detection data. Specifically, the Euclidean distances between multiple prediction boxes and their corresponding bounding boxes are first calculated. The confidence is then weighted and corrected. The prediction box with the highest confidence is then found. The intersection-and-union similarity between the prediction boxes is calculated. The Gaussian weighted average between all prediction boxes is calculated. Finally, the prediction boxes that meet the similarity threshold are added to the output list. Prediction boxes at close distances are usually more accurate because the point cloud is very dense, while prediction boxes at long distances are poor, especially in the direction of the pre-model measurement box, which is prone to left and right swaying. The distance between the center of the prediction box and the viewpoint is taken into account in the non-maximum suppression. The prediction box with the highest confidence is first selected. For prediction boxes at close distances, those with a larger intersection-and-union ratio will be assigned higher weights. On the contrary, for prediction boxes at long distances, the prediction boxes around them will be assigned relatively uniform weights to obtain a smoother effect.
[0042] like Figure 6 As shown, according to an embodiment of the present application, a method for detecting obstacles of irregular scale in a waterway is provided, wherein detection is performed by weighted non-maximum suppression of the intersection-over-union ratio of variable distances to obtain detection data, and the method further includes the following steps: Step S602: Determine whether all prediction boxes are popped up. If so, execute S606; if not, execute S604. Step S604: re-searching the prediction box with the highest confidence; Step S606: Output detection data.
[0043] In this embodiment, detection is performed through weighted non-maximum suppression of intersection-over-union ratio with variable distance to obtain detection data, and it also includes judging whether all prediction boxes pop up. If not, the prediction box with the highest confidence is searched again. If all pop up, the detection data is output.
[0044] like Figure 7 As shown, according to an embodiment of the present application, the method for detecting obstacles of irregular scales in a waterway further includes the following steps before inputting the first extracted data into the 2D multi-scale fusion convolution module: Step S702: All voxel features at the same position are spliced together.
[0045] In this embodiment, before the first extracted data is input into the 2D multi-scale fusion convolution module, all voxel features at the same position are also spliced. It can be understood that before adding the 2D multi-scale feature fusion convolution module, the 3D features extracted by the upstream sparse convolution need to be converted into 2D features. Densifying the sparse data, that is, filling the empty voxel positions with 0, will not increase the amount of calculation, and the features of different heights in the densified 3D features will be spliced, thereby maximizing the retention of the spatial height z-axis information extracted by the sparse convolution, better retaining the spatial features of the object, and avoiding the object's spatial features being averaged even when there are large differences.
[0046] like Figure 8 As shown, an embodiment of the second aspect of the present application provides a waterway irregular-scale obstacle detection system 10, including: an acquisition module 110, used to acquire point cloud data of waterway irregular-scale obstacles; a preprocessing module 120, used to preprocess the point cloud data; a model construction module 130, used to construct a feature pyramid structure based on sparse convolution and a bird's-eye view; a first extraction module 140, used to input the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; a second extraction module 150, used to input the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data; and a detection module 160, used to detect the second extracted data using an anchor-based method to obtain detection data.
[0047] The waterway irregular scale obstacle detection system 10 provided in this embodiment includes an acquisition module 110, a preprocessing module 120, a model building module 130, a first extraction module 140, a second extraction module 150 and a detection module 160. Among them, the acquisition module 110 is used to obtain point cloud data of waterway irregular scale obstacles. The preprocessing module 120 is used to preprocess the point cloud data. The model building module 130 is used to construct a feature pyramid structure based on sparse convolution and a bird's-eye view. The first extraction module 140 is used to input the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data. The second extraction module 150 is used to input the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data. The detection module 160 is used to detect the second extracted data using an anchor-based method to obtain detection data. By replacing traditional 3D convolution with sparse convolution and using an anchor-based approach to inherit the output features of sparse convolution, the anchor-based approach allows for flexible configuration of prior bounding boxes of different sizes, solving the current problem of difficulty in accurately detecting objects at multiple scales. Furthermore, prior bounding boxes of multiple sizes ensure simultaneous detection of both large and small objects in the scene, improving the problem of missed detection of objects of different scales in the same scene. Sparse convolution also ensures the detection speed of the algorithm during deployment, enabling it to maintain real-time performance even when running on current autonomous driving chips. Furthermore, after obtaining 3D features through sparse convolution, deep 2D feature extraction is performed to further improve the detection accuracy of large objects and ensure the accuracy of the detection bounding box in the Bird's Eye View (BEV) perspective, thereby further improving the accuracy of 3D detection and the accuracy of the autonomous driving perception module. The application of sparse convolution significantly increases detection speed, facilitating the practical deployment of the algorithm and effectively ensuring the safety of autonomous ships in inland waterways.
[0048] like Figure 9 As shown, an embodiment of the third aspect of the present application provides a waterway irregular-scale obstacle detection system 20, comprising: a memory 300 and a processor 400, wherein the memory 300 stores a program or instruction that can be run on the processor 400, and when the processor 400 executes the program or instruction, it implements the steps of the waterway irregular-scale obstacle detection method of any one of the embodiments of the first aspect, and thus has the technical effects of any one of the embodiments of the first aspect mentioned above, which will not be repeated here.
[0049] The embodiment of the fourth aspect of the present application provides a readable storage medium on which a program or instruction is stored. When the program or instruction is executed by the processor, the steps of the method for detecting irregular-scale obstacles in a waterway of any one of the embodiments of the first aspect are implemented, and thus the technical effects of any one of the embodiments of the first aspect are achieved, which will not be repeated here.
[0050] like Figures 10 to 12 As shown, according to a specific embodiment of the present application, a method for detecting obstacles of irregular scales in waterways replaces traditional 3D convolution with sparse convolution, and uses an anchor-base method to inherit the output features of the sparse convolution. Because the anchor-base method can freely and flexibly configure prior boxes of different sizes, it solves the problem that current mainstream methods have difficulty accurately detecting multi-scale phenomena of the same object. At the same time, the prior boxes of multiple sizes ensure the simultaneous detection of large and small targets in the scene, improving the problem of missed detection of targets of different scales in the same scene. Sparse convolution also ensures the detection speed of the algorithm during deployment, allowing it to maintain real-time performance even when running on current autonomous driving chips. In addition, to further improve the detection accuracy of large-scale targets, this embodiment performs deep extraction of 2D features after obtaining 3D features through sparse convolution to ensure the accuracy of the detection box from a BEV (bird's eye view) perspective, thereby further improving the accuracy of 3D detection. Notably, this embodiment modifies the BEV implementation used in mainstream detection methods, replacing the traditional cumulative summation of all voxel features at the same location with a concatenation of all voxel features at the same location. This better preserves the spatial characteristics of objects and avoids averaging even when there are significant differences in their spatial characteristics. Finally, this embodiment modifies the traditional NMS (Non-Maximum Suppression) method by taking into account the distance between the detection box and the radar origin, further improving the detection accuracy of objects far from the radar viewpoint where the point cloud is sparse.
[0051] 3D convolution is replaced with sparse convolution, and multiple layers of sparse convolution are designed as a feature pyramid structure to learn the feature information of objects of different scales. Combining sparse convolution with the anchor-base method improves detection accuracy while ensuring real-time computation.
[0052] Adding a Bird's Eye View (BEV) structure and a 2D multi-scale fusion convolution module after the sparse convolution of the Feature Pyramid Networks (FPN) structure for deep feature extraction can better regress the position and size of objects in the 2D plane world.
[0053] The variable distance IOU (Intersection over Union) weighted NMS improves the detection rate of sparse objects in point clouds with distant viewpoints.
[0054] Specifically, anchors are a set of preset bounding boxes used to learn the offset between the actual bounding box position and the preset bounding boxes during training. In simple terms, the approximate locations of the target are pre-set, and then bounding boxes (anchors) of different sizes are designed at these preset locations. With each training iteration, the model continuously updates and learns the offset between these anchors and the actual location and size.
[0055] Specifically, there are two main design aspects: one is the distribution of positive and negative samples during training, and the other is the number and size of model output heads.
[0056] During training, the positive and negative samples are allocated, and the anchor labels and detection frame size and position are allocated using the anchor-base method. It is worth mentioning that the distribution of the true value of the detection frame direction is divided separately in the present invention, that is, it is not placed together with the size and position of the detection frame. The purpose of this is not to predict the sine and cosine values of the angle, but to directly predict the angle and use a separate loss function to predict the direction of the detection frame (cross entropy loss function) to improve the accuracy of direction detection. The direction value is changed from the original sine and cosine values of the angle value to directly predict the angle value. The purpose of this is to reduce the content of network learning. In order to avoid the problem of forward and reverse prediction errors, a two-category unique hot encoding is performed during training, that is, 0°~90° and 270°~360° are 0, and 90°~270° is 1. The model outputs the number and size, and the detection head is established according to the category, size position, and direction. The size of its feature map is designed according to the point cloud range and voxel size.
[0057] Add 2D multi-scale feature fusion convolution module: Before adding the 2D multi-scale feature fusion convolution module, the 3D features extracted by the upstream sparse convolution need to be converted to 2D features. The existing method is to accumulate and sum all the data on each grid of the feature_map, but this method is no longer used in this embodiment. Instead, the sparse data is densified (specifically, empty voxel positions are filled with 0, which does not increase the computational complexity) and the features of different heights in the densified 3D features are spliced together. This is done to maximize the preservation of the spatial height z-axis information extracted by the sparse convolution.
[0058] In order to improve the feature extraction capability of the model, this embodiment adds a multi-scale feature map fusion module based on the 2D convolution layer after the sparse convolution (and after BEV). The specific location is as follows: Figure 10As shown. The specific process is S802: input point cloud data; S804: divide the point cloud into voxels of equal size according to the rules; S806: calculate the features of each voxel; S808: send it to the sparse convolution of the Feature Pyramid Networks to extract features; S810: 2D multi-scale feature fusion convolution module; S812: use the anchor-base method to regress the detection frame; S814: output the detection frame. The specific implementation method is that the data passed down from the upstream first passes through a 6-layer convolution module (each module consists of a convolution layer + a normalization layer BN (Batch Normalization) + a ReLU (Rectified Linear Unit) activation function layer), and then retains a separate copy of the processed data. The processed data continues to pass through a 6-layer convolution module and the feature map size is halved and the number of feature channels is doubled. This processed data is fused with the data that has just been retained separately. The purpose of this is to fully learn large-scale target features and small-scale feature targets. The construction of the network module is as follows Figure 11 The specific process is as follows: S902: convolution module; S904: convolution module; S906: convolution module; S908: convolution module; S910: convolution module; S912: convolution module; S914: convolution module, halving the feature map scale and doubling the number of channels; S916: convolution module; S918: convolution module; S920: convolution module; S922: convolution module; S924: convolution module; S926: fusion and splicing.
[0059] IOU-weighted NMS with variable distance: like Figure 12 As shown, the specific process is S1002: calculate the Euclidean distance between N predicted Bboxes and their corresponding anchors; S1004: perform weighted correction on the confidence; S1006: find the predicted box with the highest confidence; S1008: calculate the IOU similarity between the predicted boxes; S1010: calculate the Gaussian weighted average between all Bboxes; S1012: add the Bboxes that meet the similarity threshold to the output list; S1014: determine whether all Bboxes are popped up. If not, return to S1006. If so, output. Specific implementation steps: Calculate the Euclidean distance dist between N predicted Bboxes and their corresponding anchors with respect to the (x, y) coordinates.
[0060] Correct the confidence again, that is, weight the corrected confidence S again, and the weight is 1-softmax(dist).
[0061] Find the prediction box Bbox with the highest confidence in C, denoted as c.
[0062] Multiply and accumulate the predicted IOU values corresponding to the remaining prediction boxes except c and their IOU values with c, and record the data as cnt (that is, considering c as a "pseudo-GT", then this multiplication and accumulation is equivalent to calculating the inner product to compare the similarity. The higher the similarity value, the more accurate the prediction of c).
[0063] For these Bboxes, calculate the Gaussian weighted average.
[0064] If cnt is greater than the set threshold, then the Gaussian weighted prediction box is added to the output list.
[0065] Delete the Bbox filtered out in f from C, and search for the predicted box Bbox with the highest confidence in C again until all Bboxes in C are judged.
[0066] Reason for addition: The model is generally more accurate for predicted boxes at close range because point clouds are dense; however, predicted boxes at long distances are less accurate, especially due to the tendency for the prediction box's orientation to sway. The distance between the prediction box center and the viewpoint (the ship) is factored into the NMS. The prediction box with the highest confidence is selected first. For predicted boxes at close range, those with a large IOU are given a higher weight. Conversely, for predicted boxes at long distances, the surrounding prediction boxes are given a relatively even weight, resulting in a smoother result.
[0067] In summary, the beneficial effects of the embodiments of the present application are as follows: the detection rate of objects of the same category at multiple sizes is significantly increased, the size and position of the detection box of multi-scale targets in complex scenes are also greatly improved, the accuracy of the autonomous driving perception module is improved, and the application of sparse convolution significantly improves the detection speed, which provides the possibility for the actual deployment of the algorithm and effectively ensures the safety of autonomous driving ships in inland waterways.
[0068] In this application, the terms "first," "second," and "third" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance. The term "plurality" refers to two or more, unless expressly limited otherwise. Terms such as "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, "connected" can mean a fixed connection, a detachable connection, or an integral connection; "connected" can mean a direct connection or an indirect connection through an intermediary. Those skilled in the art can understand the specific meanings of the above terms in this application based on the specific circumstances.
[0069] In the description of this application, it should be understood that the terms "upper", "lower", "front", "back", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, rather than indicating or implying that the device or module referred to must have a specific direction, be constructed and operated in a specific direction. Therefore, they should not be understood as limitations on this application.
[0070] Throughout this specification, terms such as "one embodiment," "some embodiments," and "specific embodiments" mean that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present application. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0071] The above are merely preferred embodiments of the present application and are not intended to limit the present application. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for detecting obstacles of irregular scale in a waterway, characterized in that: include: Obtain point cloud data of irregular-scale obstacles in waterways; Preprocessing the point cloud data; Construct a feature pyramid structure based on sparse convolution and bird's-eye view; Inputting the preprocessed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; Inputting the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data; The second extracted data is detected using an anchor point-based method to obtain detection data.
2. The method for detecting obstacles of irregular scale in waterways according to claim 1, characterized in that: The preprocessing of the point cloud data includes: Dividing the point cloud data into voxels of equal size according to preset rules; Calculate the features of each voxel to obtain preprocessed data.
3. The method for detecting obstacles of irregular scale in waterways according to claim 1, characterized in that: The step of inputting the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data includes: Inputting the first extracted data into a 6-layer convolution module to obtain first processed data; Input the first processed data into a 6-layer convolution module and reduce the feature map size by half and double the number of feature channels to obtain second processed data; The first processed data and the second processed data are fused and spliced to obtain second extracted data.
4. The method for detecting obstacles of irregular scale in waterways according to any one of claims 1 to 3, characterized in that: The detecting the second extracted data by using the anchor point-based method to obtain the detection data includes: Detection is performed through weighted non-maximum suppression of intersection-over-union ratio with variable distance to obtain detection data.
5. The method for detecting obstacles of irregular scale in waterways according to claim 4, characterized in that: The detection is performed by weighted non-maximum suppression of the intersection-over-union ratio of variable distances to obtain detection data, including: Calculate the Euclidean distance between multiple prediction boxes and their corresponding bounding boxes; Make weighted corrections to confidence levels; Find the prediction box with the highest confidence; Calculating the intersection-over-union similarity between the prediction frames; Calculate the Gaussian weighted mean between all the predicted boxes; The predicted boxes that meet the similarity threshold are added to the output list.
6. The method for detecting obstacles of irregular scale in waterways according to claim 5, characterized in that: The detection is performed by weighted non-maximum suppression of the intersection-over-union ratio of variable distances to obtain detection data, and further includes: Determine whether all prediction boxes pop up; If not, then re-search for the prediction box with the highest confidence; If so, the detection data is output.
7. The method for detecting obstacles of irregular scale in waterways according to any one of claims 1 to 3, characterized in that: Before inputting the first extracted data into the 2D multi-scale fusion convolution module, the method further includes: All voxel features at the same position are stitched together.
8. A waterway irregular-scale obstacle detection system, characterized in that: include: An acquisition module (110) is used to acquire point cloud data of obstacles of irregular scale in the waterway; A preprocessing module (120), configured to preprocess the point cloud data; A model building module (130) for building a feature pyramid structure based on sparse convolution and bird's-eye view; A first extraction module (140) is used to input the pre-processed point cloud data into the feature pyramid structure for feature extraction to obtain first extracted data; A second extraction module (150) is used to input the first extracted data into a 2D multi-scale fusion convolution module to obtain second extracted data; The detection module (160) is used to detect the second extracted data using an anchor point-based method to obtain detection data.
9. A waterway irregular-scale obstacle detection system, characterized in that: include: A memory (300) and a processor (400), wherein the memory (300) stores a program or instruction that can be run on the processor (400), and when the processor (400) executes the program or the instruction, the steps of the method for detecting irregular-scale obstacles in a waterway as described in any one of claims 1 to 7 are implemented.
10. A readable storage medium having a program or instruction stored thereon, characterized in that: When the program or the instruction is executed by a processor, the steps of the method for detecting obstacles of irregular scale in a waterway are implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Object detection method, device and equipment in automatic driving and storage medium
CN122218646A