Underwater point cloud segmentation method, device and program product
Through density adaptive clustering and sparse feature screening, combined with multi-scale feature construction and time series features, the effect of underwater point cloud segmentation under density unevenness and noise interference is solved, and efficient, stable and high-precision underwater point cloud segmentation is achieved.
Patent Information
- Application Number
- CN202510091916.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
The existing underwater point cloud segmentation technology has limited segmentation effects under uneven density and sparse data, and it is difficult to maintain high accuracy and stability under noise interference.
Density adaptive clustering and sparse feature screening strategies are adopted to divide point cloud data into dense and sparse areas according to density thresholds. Through multi-scale feature construction and layer-by-layer feature fusion, combining time series features and adaptive weights, the final feature representation is generated and input into a multi-scale convolutional network for segmentation.
The segmentation efficiency and accuracy are improved, the model's adaptability to complex underwater targets is enhanced, the stability and consistency of segmentation results are ensured, and the problem of uneven density of underwater point clouds is effectively solved.
Smart Images

Figure CN120107284A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an underwater point cloud segmentation method, device and program product. Background Art
[0002] In the field of modern navigation and underwater detection, accurate identification and classification of underwater targets are crucial, which is of great significance to ensuring navigation safety, developing marine resources and protecting the ecological environment. As an important carrier for recording the geometric shape, position and spatial distribution of underwater objects, underwater point cloud data is widely used in fields such as shipwreck detection and coral reef monitoring. However, the underwater environment is special, with problems such as insufficient lighting, complex clutter, sparse data and dynamic changes, which bring many challenges to the accurate segmentation of underwater targets. Existing underwater point cloud segmentation technology mainly relies on fixed feature extraction and segmentation strategies. The segmentation effect is limited in the case of uneven density and sparse data. In addition, underwater noise is frequent, and traditional algorithms are difficult to maintain high accuracy and stability under noise interference. With the increasing demand for efficient underwater detection and identification in the field of navigation, the development of point cloud semantic segmentation methods that adapt to the complexity of the underwater environment, high segmentation accuracy and stability has become a technical development trend. Summary of the invention
[0003] The embodiments of the present invention provide an underwater point cloud segmentation method, device and program product to improve segmentation efficiency and accuracy and enhance the adaptability of the model to complex underwater targets.
[0004] In order to achieve the above object, on the one hand, a method for underwater point cloud segmentation is provided, the method comprising:
[0005] S1, acquiring point cloud data and 2D image data of a selected area, wherein a plurality of continuous 2D image data constitute a time series of image data, and preprocessing the point cloud data, wherein the point cloud data is a set of three-dimensional coordinates;
[0006] S2, clustering the preprocessed point cloud data using a predetermined clustering algorithm to obtain a dense area and a sparse area, wherein the dense area and the sparse area contain a plurality of clusters;
[0007] S3, defining a first neighborhood of each point in each cluster, and calculating a first local feature of each point in the corresponding first neighborhood, wherein the first local feature includes: density and normal vector; wherein the definition formula of the first neighborhood is:
[0008] N(p i )={p j |||p i -p j ||≤d}
[0009] N(p i) represents the first neighborhood, p j and p i represents a point in the point cloud data, and d represents a predetermined neighborhood radius;
[0010] The density is calculated as follows:
[0011]
[0012] ρ i represents the density, |N(p i )| represents the number of points in the first neighborhood, V(N(p i )) represents the volume of the first neighborhood;
[0013] S4, screening out the sparse feature points whose density is less than a predetermined threshold from the sparse area to obtain a sparse feature point set, wherein each sparse feature point in the sparse feature point set contains information of the normal vector corresponding to the first neighborhood;
[0014] S5, extracting features from the sparse feature point set and all point cloud data using convolution kernels of different scales to obtain a multi-scale feature set;
[0015] S6, extracting semantic information from the 2D image of the time series by a predetermined semantic segmentation algorithm, projecting the semantic information onto the three-dimensional point cloud coordinates of the point cloud data to generate a corresponding 2D feature representation, and obtaining a 2D feature set; wherein the projection formula is:
[0016]
[0017] f i proj represents the 2D feature representation, T represents the projection matrix, Represents the semantic information at time t;
[0018] Combining the multi-scale feature set with the 2D feature set to obtain a fused feature set;
[0019] S7, for each point in the first neighborhood, weightedly fuse the fused feature corresponding to the point in the fused feature set with the fused features corresponding to other points in the first neighborhood except the point, to generate a second local feature; wherein the second local feature is:
[0020]
[0021] f i loc Represents point p i The second local feature of , Z represents the normalization factor, f jRepresents the fusion features of other points in the first neighborhood except the point;
[0022] S8, performing weighted aggregation on the second local feature in a predetermined second neighborhood to obtain a mesoscale feature; wherein the second neighborhood is larger than the first neighborhood;
[0023] S9, aggregating the mesoscale features in a predetermined third neighborhood by weighted averaging to obtain global features; wherein the third neighborhood is larger than the second neighborhood;
[0024] S10, performing weighted fusion on the second local feature, the mesoscale feature and the global feature to obtain a final feature representation;
[0025] S11, inputting the final feature representation into a predetermined multi-scale convolutional network to obtain a category label for each point in the point cloud data, and mapping the category label to an actual object category.
[0026] Preferably, the underwater point cloud segmentation method further comprises one or more of the following:
[0027] In step S1, preprocessing the point cloud data includes: applying three-dimensional Gaussian filtering to the point cloud data;
[0028] In step S2, the preprocessed point cloud data is clustered by a predetermined clustering algorithm to obtain a dense area and a sparse area, wherein the dense area and the sparse area contain a plurality of clusters, wherein:
[0029] The clusters in the dense area are:
[0030] C k ={p i |ρ(p i )≥ρ min}
[0031] The clusters in the sparse region are:
[0032] C k ={p i |ρ(p i )<ρ min}
[0033] Among them, C k represents the cluster, p i represents a point, ρ(p i ) represents point p i The density of min represents the density threshold of the predetermined cluster;
[0034] In step S11, conditional random field smoothing is applied to the category label, wherein the conditional random field smoothing is:
[0035]
[0036] L i Represents a single point p i The category label, L j Represents a single point p i The category labels of the neighboring points, ψ u (L i ) represents the category label consistency term, ψ p (L i , L j ) is the category label smoothing term of adjacent points, and E(L) represents the energy function.
[0037] Preferably, the underwater point cloud segmentation method, after generating the global features in step S9 and before outputting the final feature representation in step S10, further comprises:
[0038] S301, fusing each frame image in the time series, the 2D feature set, the second local feature and / or the global feature to obtain a time series feature of each point in each frame image;
[0039] The time series features of each point in each frame are concatenated and aggregated according to a predetermined index for each point in the point cloud data to obtain a time series feature set;
[0040] Performing cosine similarity calculation on each of the time series features corresponding to the continuous frames of the time series to obtain feature similarity;
[0041] S302, assigning an adaptive weight to each point in the point cloud data according to the density calculated in the current frame and the feature similarity.
[0042] Preferably, in the underwater point cloud segmentation method, step S302 comprises:
[0043] When the density of a point in the point cloud data in the current frame is greater than or equal to a first predetermined threshold or the feature similarity with the previous frame is greater than or equal to a second predetermined threshold, increasing the adaptive weight of the point in the local feature or the global feature;
[0044] When the density of a point in the point cloud data in the current frame is less than a third predetermined threshold or the feature similarity with the previous frame is less than a fourth predetermined threshold, the adaptive weight of the point in the local feature or the global feature is reduced.
[0045] Preferably, in the underwater point cloud segmentation method, the adaptive weight is adjusted according to the density by the following formula:
[0046]
[0047] Represents point p in the current frame i The adaptive weights, represents the point p in the previous frame i The adaptive weights, Represents point p in the current frame i The density, ρ baseline Indicates the predetermined initial density reference value.
[0048] Preferably, the underwater point cloud segmentation method further comprises: calculating the importance of the multi-dimensional features of each point in the point cloud data by a predetermined calculation method, and adjusting the feature channel weight of each dimensional feature in the multi-dimensional features according to the importance;
[0049] The calculation method includes: calculation based on statistics, gradient or attention mechanism;
[0050] The multi-dimensional feature includes: the first local feature, the second local feature, the mesoscale feature, the global feature, the 2D feature representation, the fusion feature and / or the time series feature;
[0051] The importance includes: discrimination, stability and / or relevance.
[0052] Preferably, the underwater point cloud segmentation method further comprises:
[0053] S701, recording and aggregating feature changes of each point in the cluster in each frame of the time series to obtain a time average feature;
[0054] S702, the time average feature is weightedly calculated by the density to obtain a priori feature; wherein the priori feature is calculated by the following formula:
[0055]
[0056] B k represents the prior feature, ρ i Represents point p i The density, f i avg represents said time-averaged characteristic;
[0057] S703, assigning the same category label to each point in each cluster to obtain a preliminary segmentation area;
[0058] The time series features of all points in the current frame are compared with the prior features by cosine similarity, and then the category label is adjusted by the following formula:
[0059]
[0060] represents the adjusted category label, CosSim represents the cosine similarity function, and f i (t) Representing the time series feature of a point in the current frame;
[0061] S704 , performing feature comparison on the continuous frames of the time series to obtain a quantization parameter of label consistency, and when the value of the quantization parameter is less than a preset quantization threshold, recalculating the category label of the preliminary segmented region.
[0062] On the other hand, an embodiment of the present invention provides a device for underwater point cloud segmentation, which includes a memory and a processor, the memory stores at least one program, and the at least one program is executed by the processor to implement any of the underwater point cloud segmentation methods described above.
[0063] On the other hand, an embodiment of the present invention provides a computer program product, including a computer program, wherein when the computer program is executed by a processor, it implements any of the underwater point cloud segmentation methods described above.
[0064] The above technical solution has the following technical effects:
[0065] The embodiments of the present invention adopt density adaptive clustering and sparse feature screening strategies to divide underwater point cloud data into dense and sparse areas according to density thresholds, thereby achieving effective extraction of key features and compression of redundant information, improving segmentation efficiency and accuracy; among them, sparse area feature screening reduces redundant information processing and retains points with significant geometric information, ensuring the stability and consistency of segmentation results and improving segmentation efficiency; when faced with differences in the sparsity and density of underwater data, the stability and consistency of segmentation results are enhanced, thereby effectively solving the problem of uneven density of underwater point clouds. Secondly, multi-scale feature construction and layer-by-layer feature fusion improve the adaptability of the model to complex underwater targets, extracting and fusing features layer by layer at local, mesoscale and global scales, making the segmentation algorithm more flexible in capturing subtle edges and complex shapes;
[0066] In a further embodiment, by calculating the feature similarity between consecutive frames and adaptively updating the weights, when the feature similarity is lower than a threshold, the weight update is triggered to ensure consistency and stability in the time series, thereby reducing the segmentation jitter caused by changes in the target morphology and making the segmentation result more stable, thereby being able to cope with the segmentation instability problem caused by changes in the target position or morphology;
[0067] In a further embodiment, dynamic prior information is generated by weighted aggregation of multi-frame data features, providing consistent guidance for the target area in subsequent segmentation, achieving segmentation guidance optimization, ensuring that the segmentation model maintains consistent marking of the same target area on multi-frame data, and showing superior adaptability and accuracy in complex scene changes;
[0068] In a further embodiment, a timing consistency check mechanism is added to verify the segmentation results by measuring the label consistency between adjacent frames. This can suppress label instability caused by noise or environmental disturbances, help the segmentation results remain consistent in a dynamic environment, and provide stable support for target recognition in underwater continuous monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 A schematic diagram of a flow chart of an underwater point cloud segmentation method according to an embodiment of the present invention;
[0070] Figure 2 Schematic diagram of the structure of an underwater point cloud segmentation device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0071] To further illustrate the various embodiments, the present invention provides drawings. These drawings are part of the disclosure of the present invention, which are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, a person of ordinary skill in the art should be able to understand other possible implementations and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0072] The present invention will now be further described with reference to the accompanying drawings and specific implementation methods.
[0073] Embodiment 1:
[0074] In order to improve segmentation efficiency and accuracy and enhance the adaptability of the model to complex underwater targets, an embodiment of the present invention provides an underwater point cloud segmentation method. Figure 1 FIG. 1 is a flow chart of underwater point cloud segmentation according to an embodiment of the present invention. Figure 1 As shown, the method includes:
[0075] S1, obtaining point cloud data and 2D image data of the selected area, wherein a plurality of continuous 2D image data constitute a time series of image data, and preprocessing the point cloud data, wherein the point cloud data is a set of three-dimensional coordinates;
[0076] S2, clustering the preprocessed point cloud data using a predetermined clustering algorithm to obtain dense areas and sparse areas, where the dense areas and sparse areas contain multiple clusters;
[0077] S3, defining the first neighborhood of each point in each cluster, and calculating the first local feature of each point in the corresponding first neighborhood, the first local feature includes: density and normal vector; wherein the definition formula of the first neighborhood is:
[0078] N(p i )={p j |||p i -p j ||≤d}
[0079] N(p i ) represents the first neighbor, p j and p i represents a point in the point cloud data, and d represents a predetermined neighborhood radius;
[0080] The density is calculated as:
[0081]
[0082] ρ i represents density, |N(p i )| represents the number of points in the first neighborhood, V(N(p i )) represents the volume of the first neighborhood;
[0083] S4, screening out sparse feature points whose density is less than a predetermined threshold from the sparse area to obtain a sparse feature point set, wherein each sparse feature point in the sparse feature point set contains information of a normal vector corresponding to the first neighborhood;
[0084] S5, extract features from the sparse feature point set and all point cloud data through convolution kernels of different scales to obtain a multi-scale feature set;
[0085] S6, extracting semantic information from the 2D image of the time series through a predetermined semantic segmentation algorithm, projecting the semantic information onto the three-dimensional point cloud coordinates of the point cloud data to generate a corresponding 2D feature representation, and obtaining a 2D feature set; wherein the projection formula is:
[0086]
[0087] f i projrepresents 2D feature representation, T represents the projection matrix, Represents the semantic information at time t;
[0088] Combine the multi-scale feature set and the 2D feature set to obtain a fused feature set;
[0089] S7, for each point in the first neighborhood, weightedly fuse the fused feature corresponding to the point in the fused feature set and the fused features corresponding to other points in the first neighborhood except the point, to generate a second local feature; wherein the second local feature is:
[0090]
[0091] f i loc Represents point p i The second local feature of , Z represents the normalization factor, f j Indicates the fusion features of other points in the first neighborhood except this point;
[0092] S8, performing weighted aggregation on the second local feature in a predetermined second neighborhood to obtain a mesoscale feature; wherein the second neighborhood is larger than the first neighborhood;
[0093] S9, aggregating the mesoscale features in a predetermined third neighborhood by weighted averaging to obtain global features; wherein the third neighborhood is larger than the second neighborhood;
[0094] S10, performing weighted fusion on the second local feature, the mesoscale feature and the global feature to obtain a final feature representation;
[0095] S11, input the final feature representation into a predetermined multi-scale convolutional network to obtain the category label of each point in the point cloud data, and map the category label to the actual object category.
[0096] Embodiment 2:
[0097] The method of this embodiment of the present invention includes:
[0098] 1. Obtain the point cloud data and 2D image data of the selected area. Multiple consecutive 2D image data constitute the time series of image data, and pre-process the point cloud data;
[0099] Preferably, the point cloud data is acquired by an underwater sonar or lidar sensor.
[0100] Preferably, the point cloud data is a three-dimensional coordinate set, which is expressed as:
[0101] P = {p i |p i =(x i ,y i, z i ), i=1, 2, ..., N}
[0102] P represents a three-dimensional coordinate set, p i Represents any point in the point cloud data, (x i ,y i , z i ) represents the three-dimensional coordinates of any point in the point cloud data, and N represents the number of points in the point cloud data.
[0103] Preferably, preprocessing the point cloud data includes:
[0104] (1) Normalize the coordinates of each point in the point cloud data and standardize each point to a predetermined range to ensure the consistency of data at different scales; specifically, the predetermined range is: [-1, 1].
[0105] (2) Apply three-dimensional Gaussian filtering to the point cloud data to reduce the noise in the underwater environment. Specifically, the filtering formula is:
[0106]
[0107] N 1 (p i ) represents point p i The neighborhood is determined by the Gaussian distribution, and σ represents the smoothing parameter.
[0108] 2. The pre-processed point cloud data is clustered by a predetermined clustering algorithm to obtain dense areas and sparse areas, which contain multiple clusters;
[0109] Preferably, the predetermined clustering algorithm is a density-based clustering algorithm (DBSCAN); specifically, first find the core points that meet the dense requirements (i.e., the number of surrounding neighborhood points is large enough and the density is high enough) according to the predetermined clustering density threshold, and then determine the points with insufficient surrounding density as sparse areas or outliers; in terms of implementation ideas, the predetermined clustering density threshold can be used to find those clusters that are "tightly clustered" in space (i.e., dense areas), and can also distinguish sparse areas that are "dispersed" or "fragmented" in space, so as to meet the needs of underwater environments for separate processing of sparse and dense point clouds.
[0110] In a specific embodiment, the predetermined clustering algorithm divides the point cloud data into multiple clusters by analyzing the density structure of the point cloud. Density clustering assigns each point to a corresponding cluster; specifically, when the density of a point in the point cloud data is greater than or equal to a predetermined cluster density threshold, the point is regarded as a "core point" or a "core cluster" member of a dense area, wherein the clusters in the dense area are represented as:
[0111] C k={p i |ρ(p i )≥ρ min}
[0112] When the density of a point in the point cloud data is less than the predetermined cluster density threshold, it is regarded as a sparse area or "outlier point", where the cluster in the sparse area is represented as:
[0113] C k ={p i |ρ(p i )<ρ min}
[0114] Among them, C k represents the cluster, ρ(p i ) represents point p i The density of min Indicates the density threshold of the predetermined cluster, which is used to distinguish dense areas from sparse areas.
[0115] 3. Define the first neighborhood of each point in each cluster, and calculate the first local feature of each point in the corresponding first neighborhood, the first local feature includes: density and normal vector;
[0116] Preferably, the first neighborhood consists of all points whose distance from the selected point is less than or equal to a predetermined neighborhood radius, and the definition formula of the first neighborhood is:
[0117] N(p i )={p j |||p i -p j ||≤d}
[0118] N(p i ) represents the first neighbor, p j and p i represents a point in the point cloud data, and d represents a predetermined neighborhood radius, which is used to capture local structure information.
[0119] Preferably, the first local feature of each point in the corresponding first neighborhood is calculated to provide finer-grained local geometric information for the subsequent segmentation step. The normal vector represents the average direction of the point in the first neighborhood; the density is the ratio of the number of points in the first neighborhood to the volume of the first neighborhood. Specifically, the density is calculated as:
[0120]
[0121] ρ i represents density, |N(p i )| represents the number of points in the first neighborhood, V(N(p i )) represents the volume of the first neighborhood.
[0122] 4. Filter out sparse feature points whose density is less than a predetermined threshold from the sparse area to obtain a sparse feature point set, wherein each sparse feature point in the sparse feature point set contains information of a corresponding normal vector in the first neighborhood;
[0123] The sparse feature point set is selected to ensure that points with key geometric structures are retained while optimizing computational efficiency for use in subsequent multi-scale feature construction. Preferably, sparse feature points with lower density (i.e., sparse areas) but containing significant geometric features are selected based on density and normal vectors to obtain a sparse feature point set, which is defined as a set of points whose density is lower than a specified threshold.
[0124] In a specific embodiment, the sparse feature point set is defined as:
[0125] S={s i |s i ∈C k ,ρ(s i )<ρ threshold}
[0126] S represents a sparse feature point set, s i represents sparse feature points, ρ threshold Indicates the density threshold for sparse feature point screening, which is used to ensure that the selected points are distributed in a sparser area. In the process of screening sparse feature points, the geometric information such as the normal vectors of these sparse feature points is retained, so that in subsequent steps (such as multi-scale feature extraction, time series consistency check, etc.), key geometric structure information (such as edges, corners, surface mutations, etc.) can be extracted from the sparse area to avoid the loss of potential key information in the sparse area, which can effectively improve the feature retention and segmentation accuracy of the sparse area. The generation of sparse feature point sets not only reduces the amount of data and improves the computational efficiency, but also retains a detailed description of the sparse area.
[0127] 5. Use convolution kernels of different scales to extract features from sparse feature point sets and all point cloud data to obtain a multi-scale feature set.
[0128] 6. Extract semantic information from the 2D images of the time series through a predetermined semantic segmentation algorithm, project the semantic information onto the three-dimensional point cloud coordinates of the point cloud data to generate the corresponding 2D feature representation, and obtain a 2D feature set; then combine the multi-scale feature set with the 2D feature set to obtain a fused feature set;
[0129] Since underwater scenes have more uncertain factors such as noise and lighting changes, semantic segmentation models based on convolutional neural networks (CNN) are often used to improve robustness and accuracy.
[0130] The predetermined semantic segmentation algorithm is not limited to a specific network, but refers to a general 2D convolutional neural network or its extended structure. The predetermined semantic segmentation algorithm can effectively extract pixel-level features from 2D images and output semantic segmentation results. Since the embodiments of the present invention are directed to 2D images of time series rather than single-frame images, a timing module (such as ConvLSTM or 3D convolution) can be further added to the above-mentioned CNN framework or a multi-frame cascade method can be used to capture the correlation information of the previous and next frames, thereby improving the recognition effect of targets in underwater dynamic environments.
[0131] Preferably, the predetermined semantic segmentation algorithms include: FCN (Fully Convolutional Network), U-Net, PSPNet (Pyramid Scene Parsing Network), DeepLab and other 2D semantic segmentation frameworks.
[0132] Preferably, the projection formula is:
[0133]
[0134] f i proj represents 2D feature representation, T represents the projection matrix, Represents the semantic information at time t.
[0135] Preferably, the multi-scale feature set and the 2D feature set are combined and expressed as: Among them, F represents the fusion feature set, F s represents the multi-scale feature set, F proj represents a 2D feature set, Represents a feature cascade operation.
[0136] Preferably, the fused feature set is normalized so that the mean of the fused feature value in each dimension is 0 and the variance is 1, thereby ensuring the consistency of the fused feature in subsequent processing.
[0137] 7. For each point in the first neighborhood, perform weighted fusion of the fusion feature corresponding to the point in the fusion feature set and the fusion features corresponding to other points in the first neighborhood except the point to generate a second local feature;
[0138] Preferably, the second local feature is:
[0139]
[0140] f i loc Represents point p i The second local feature of , Z represents the normalization factor, which is used to ensure the balance of the weighted results, f jRepresents the fused features of all points in the first neighborhood except this point.
[0141] 8. Perform weighted aggregation on the second local features within a predetermined second neighborhood to obtain mesoscale features to describe a larger regional structure; wherein the second neighborhood is larger than the first neighborhood.
[0142] 9. Aggregate the mesoscale features in a predetermined third neighborhood by weighted averaging to obtain global features to capture the overall structure of the scene; wherein the third neighborhood is larger than the second neighborhood.
[0143] 10. Perform weighted fusion on the second local features, mid-scale features and global features to obtain the final feature representation.
[0144] 11. Input the final feature representation into a predetermined multi-scale convolutional network to obtain a preliminary category label for each point in the point cloud data;
[0145] Preferably, the predetermined multi-scale convolutional network includes multiple convolutional layers, each layer performs feature extraction at different scales and generates refined segmentation results, gradually improving the fineness and accuracy of the segmentation, and each layer of feature mapping gradually enhances the recognition ability of the target.
[0146] In a specific embodiment, conditional random field smoothing is applied to the preliminary category labels to eliminate misclassification of small areas and ensure the consistency of category labels between adjacent points;
[0147] Preferably, the conditional random field is smoothed as:
[0148]
[0149] L i Represents a single point p i The preliminary category label, L j Represents a single point p i The preliminary category labels of the neighboring points, ψ u (L i ) represents the preliminary category label consistency term, ensuring that the category label of each point is consistent with its features, ψ p (L i , L j ) is the preliminary category label smoothing item of the adjacent points, which is used to reduce the misclassification and noise interference of the edge, and E(L) represents the energy function.
[0150] 12. Fusing each 2D image frame, 2D feature set, second local feature and / or global feature in the time series to obtain the time series feature of each point in each frame of the image;
[0151] The time series features of each point in each frame are spliced and aggregated according to the predetermined index for each point in the point cloud data to obtain a time series feature set; the time series feature set is used to reflect the shape or position changes of the target in different frames.
[0152] 13. Calculate the cosine similarity of each time series feature corresponding to the continuous frames of the time series to obtain the feature similarity;
[0153] Preferably, feature similarity is obtained by cosine similarity calculation to ensure consistency in time series:
[0154]
[0155] f i (t) and f i (t-1) are the feature representations of the current frame and the previous frame respectively.
[0156] 14. Assign an adaptive weight to each point in the point cloud data according to the density and feature similarity calculated in the current frame, and adjust the proportion of local features and global features in dense and sparse areas; wherein the density calculated in the current frame includes: local density or global density;
[0157] In a specific embodiment, when the density of a point in the point cloud data in the current frame is greater than or equal to a first predetermined threshold or the feature similarity with the previous frame is greater than or equal to a second predetermined threshold, the adaptive weight of the point in the local feature or the global feature is increased; wherein, when the density of a point in the current frame is high, it means that the point is located in a dense area; when the feature similarity of a point in the current frame is high with the feature similarity of the previous frame, it means that the consistency in the time series is good;
[0158] When the density of a point in the point cloud data in the current frame is less than the third predetermined threshold or the feature similarity with the previous frame is less than the fourth predetermined threshold, the adaptive weight of the point in the local feature or the global feature is reduced; wherein, when the density of a point in the current frame is extremely low, it means that the point is located in a sparse / discrete area; when the feature similarity of a point in the current frame is lower than the feature similarity of the previous frame, it means that the time series is discontinuous or unstable.
[0159] In high-density areas (or areas that remain stable and consistent with the previous frame), more emphasis is placed on local features; in low-density areas or areas with large temporal fluctuations, the proportion of global features or other stability factors may be gradually increased.
[0160] Adaptive weights are also assigned in combination with specific functions such as thresholds and weighting coefficients. Usually, a function is fitted first (for example, according to Gaussian decay or linear gain), and adaptive weights are adaptively calculated based on specific values such as density and feature similarity calculated in the current frame.
[0161] In a specific embodiment, when it is detected that the feature similarity (CosSim) of adjacent frames is too low, the weight of the feature inheritance of the corresponding point in the previous frame of the point in the current frame is reduced, and the dependence of the point on global features or other prior information is increased. In other words, if the current frame is significantly different from the previous frame, it means that the original time series inheritance is unreliable, and more reliance should be placed on stable information such as global or dense areas. Specifically, a scaling factor (such as λ < 1) is used to attenuate the adaptive weight of the time frame before the current frame, and the density change of the point is combined to determine whether to further increase / decrease the proportion of local or global features, so that when the feature similarity is low, the adaptive weight of the point in the current frame is reallocated to enhance the stability and consistency of the segmentation result.
[0162] In a specific embodiment, the weight is adaptively adjusted by the following formula:
[0163]
[0164] Represents point p in the current frame i The adaptive weight of represents the point p in the previous frame i The adaptive weight of Represents point p in the current frame i The density of baseline Indicates a predetermined initial density reference value, used to balance density changes in different frames.
[0165] 15. Calculate the importance of the multi-dimensional features of each point in the point cloud data by a predetermined calculation method, and adjust the feature channel weight of each dimensional feature in the multi-dimensional features according to the importance;
[0166] Preferably, the predetermined calculation method includes: calculation based on statistics, gradient or attention mechanism; specifically, calculation based on statistics includes: calculation based on consistency between feature channels and segmentation accuracy or target labels; calculation based on gradient includes: during deep network training, analysis of the contribution of feature channels to the decrease in loss function is performed for calculation; calculation based on attention mechanism includes: calculation by weighted learning of feature channels through attention module.
[0167] Preferably, the multi-dimensional features include: first local features, second local features, mesoscale features, global features, 2D feature representations, fused features and / or time series features; specifically, the features ultimately carried by each point in the point cloud data actually include multiple sources, such as semantic information obtained from the local / mesoscale / global three-layer aggregation, from 2D image projection, and dynamic features accumulated or updated in time series, etc. Therefore, multi-dimensional features refer to each dimension (or each type) in the point cloud feature vector, including both the continuously updated dynamic components in the time series features and other features fused from local geometric information, global scene information, 2D semantic projection, etc.
[0168] Preferably, the importance includes: discrimination, stability and / or relevance. Specifically, discrimination refers to whether a certain dimensional feature can significantly improve the discrimination of the target geometry / semantics;
[0169] Stability refers to whether the numerical fluctuation of the dimension feature is too large in a multi-frame time series, or whether it can remain stable in a sparse / noisy environment;
[0170] The relevance is that if a certain dimension feature has strong semantic consistency with the prior label or the current frame / previous frame, its feature channel weight can be increased.
[0171] 16. Record and aggregate the feature changes of each point in the cluster in each frame of the time series to obtain the time average feature;
[0172] Specifically, in the feature aggregation process between consecutive frames, the feature change trajectory T of each point in each frame is recorded i ={f i (t) |t=1,2,...,T}, and the time average feature is formed by aggregating these feature changes.
[0173] Preferably, the time-averaged feature is expressed as:
[0174]
[0175] f i avg Represents the time average feature, which is used to represent the stable feature of the point in the entire time series, and T represents the total number of time frames.
[0176] 17. The time average feature is weighted by density to obtain the prior feature;
[0177] Specifically, the density-weighted time average features are calculated for all points in each cluster to obtain the prior features of the cluster, which can be used to guide subsequent segmentation and ensure consistency in multiple frames of data.
[0178] Preferably, the prior feature is calculated by the following formula:
[0179]
[0180] B k represents the prior feature, ρ i Represents point p i The density of is used to smooth the noise effect in the environment changes.
[0181] 18. Assign the same category label to each point in each cluster to obtain the preliminary segmentation area;
[0182] Specifically, the points in each cluster are assigned the same category label L i = k, L i Represents the category label of the point, and k represents a certain category label assigned; thereby ensuring that the category labels of all points in the cluster are the same, and the category label preliminarily divides the area to maintain consistency in the same area.
[0183] 19. Compare the time series features of all points in the current frame with the prior features by cosine similarity, and then dynamically adjust the preliminary category labels using the following formula to ensure consistency with the prior information:
[0184]
[0185] Represents the adjusted category label, and CosSim represents the cosine similarity function, which is used to calculate the similarity between the current frame features and the prior features.
[0186] 20. Perform feature comparison on the continuous frames of the time series to obtain the quantization parameter of label consistency. When the value of the quantization parameter is less than the preset quantization threshold, recalculate the category label of the preliminary segmented area;
[0187] In a specific embodiment, in the time series, semantic consistency is checked by comparing the features of consecutive frames to ensure that the same object keeps the same label in consecutive frames. Preferably, the formula for measuring the consistency of labels of adjacent frames is:
[0188]
[0189] C i Represents the label consistency measurement parameter, δ represents the indicator function; if the label consistency measurement parameter is lower than the set threshold, the category label of the preliminary segmented area is recalculated to ensure consistency in consecutive frames.
[0190] 21. The final category label after optimization in the above steps is the semantic segmentation result of the underwater target, and the final category label of each point is mapped to the actual object category;
[0191] Preferably, the class labels are based on class definitions in the training data.
[0192] Preferably, the categories of actual objects include: shipwrecks, coral reefs, marine life, etc.
[0193] Embodiment three:
[0194] The present invention also provides a device for underwater point cloud segmentation, such as Figure 2 As shown, the device includes a processor 201, a memory 202, a bus 203, and a computer program stored in the memory 202 and executable on the processor 201. The processor 201 includes one or more processing cores. The memory 202 is connected to the processor 201 via the bus 203. The memory 202 is used to store program instructions. When the processor 201 executes the computer program, the steps in the above method embodiment of the first embodiment of the present invention are implemented.
[0195] Further, as an executable solution, the underwater point cloud segmentation device may be a computer unit, which may be a computing device such as a desktop computer, a notebook, a PDA, and a cloud server. The computer unit may include, but is not limited to, a processor and a memory. Those skilled in the art may understand that the composition structure of the above-mentioned computer unit is only an example of a computer unit and does not constitute a limitation on the computer unit, and may include more or less components than the above-mentioned, or a combination of certain components, or different components. For example, the computer unit may also include input and output devices, network access devices, buses, etc., which are not limited in the embodiments of the present invention.
[0196] Further, as an executable solution, the processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the computer unit, and various interfaces and lines are used to connect the various parts of the entire computer unit.
[0197] The memory can be used to store the computer program and / or module, and the processor realizes various functions of the computer unit by running or executing the computer program and / or module stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system and at least one application required for a function; the data storage area can store data created according to the use of the mobile phone, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0198] Embodiment 4:
[0199] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the steps of the method described above are implemented.
[0200] Although the present invention has been specifically shown and described in conjunction with the preferred embodiments, it should be understood by those skilled in the art that various changes may be made to the present invention in form and details without departing from the spirit and scope of the present invention as defined by the appended claims, all of which are within the scope of protection of the present invention.
Claims
1. An underwater point cloud segmentation method, characterized in that: include: S1, acquiring point cloud data and 2D image data of a selected area, wherein a plurality of continuous 2D image data constitute a time series of image data, and preprocessing the point cloud data, wherein the point cloud data is a set of three-dimensional coordinates; S2, clustering the preprocessed point cloud data using a predetermined clustering algorithm to obtain a dense area and a sparse area, wherein the dense area and the sparse area contain a plurality of clusters; S3, defining a first neighborhood of each point in each cluster, and calculating a first local feature of each point in the corresponding first neighborhood, wherein the first local feature includes: density and normal vector; wherein the definition formula of the first neighborhood is: N(p i )={p j |||p i -p j ||≤d} N(p i ) represents the first neighborhood, p j and p i represents a point in the point cloud data, and d represents a predetermined neighborhood radius; The density is calculated as follows: ρ i represents the density, |N(p i )| represents the number of points in the first neighborhood, V(N(p i )) represents the volume of the first neighborhood; S4, screening out the sparse feature points whose density is less than a predetermined threshold from the sparse area to obtain a sparse feature point set, wherein each sparse feature point in the sparse feature point set contains information of the normal vector corresponding to the first neighborhood; S5, extracting features from the sparse feature point set and all point cloud data using convolution kernels of different scales to obtain a multi-scale feature set; S6, extracting semantic information from the 2D image of the time series by a predetermined semantic segmentation algorithm, projecting the semantic information onto the three-dimensional point cloud coordinates of the point cloud data to generate a corresponding 2D feature representation, and obtaining a 2D feature set; wherein the projection formula is: f i proj represents the 2D feature representation, T represents the projection matrix, Represents the semantic information at time t; Combining the multi-scale feature set with the 2D feature set to obtain a fused feature set; S7, for each point in the first neighborhood, weightedly fuse the fused feature corresponding to the point in the fused feature set with the fused features corresponding to other points in the first neighborhood except the point, to generate a second local feature; wherein the second local feature is: f i loc Represents point p i The second local feature of , Z represents the normalization factor, f j Represents the fusion features of other points in the first neighborhood except the point; S8, performing weighted aggregation on the second local feature in a predetermined second neighborhood to obtain a mesoscale feature; wherein the second neighborhood is larger than the first neighborhood; S9, aggregating the mesoscale features in a predetermined third neighborhood by weighted averaging to obtain global features; wherein the third neighborhood is larger than the second neighborhood; S10, performing weighted fusion on the second local feature, the mesoscale feature and the global feature to obtain a final feature representation; S11, inputting the final feature representation into a predetermined multi-scale convolutional network to obtain a category label for each point in the point cloud data, and mapping the category label to an actual object category.
2. The underwater point cloud segmentation method according to claim 1, characterized in that: Also includes one or more of the following: In step S1, preprocessing the point cloud data includes: applying three-dimensional Gaussian filtering to the point cloud data; In step S2, the preprocessed point cloud data is clustered by a predetermined clustering algorithm to obtain a dense area and a sparse area, wherein the dense area and the sparse area contain a plurality of clusters, wherein: The clusters in the dense area are: C k ={p i Iρ(p i )≥ρ min } The clusters in the sparse region are: C k (p i |ρ(p i )<ρ min }} Among them, C k represents the cluster, p i represents a point, ρ(p i ) represents point p i The density of min represents the density threshold of the predetermined cluster; In step S11, conditional random field smoothing is applied to the category label, wherein the conditional random field smoothing is: L i Represents a single point p i The category label, L j Represents a single point p i The category labels of the neighboring points, ψ u (L i ) represents the category label consistency term, ψp ( L i , L j ) is the category label smoothing term of adjacent points, and E(L) represents the energy function.
3. The underwater point cloud segmentation method according to claim 1, characterized in that: After generating the global features in step S9 and before mapping the category labels to the actual object categories in step S11, the method further includes: S301, fusing each frame image in the time series, the 2D feature set, the second local feature and / or the global feature to obtain a time series feature of each point in each frame image; The time series features of each point in each frame are concatenated and aggregated according to a predetermined index for each point in the point cloud data to obtain a time series feature set; Performing cosine similarity calculation on each of the time series features corresponding to the continuous frames of the time series to obtain feature similarity; S302, assigning an adaptive weight to each point in the point cloud data according to the density calculated in the current frame and the feature similarity.
4. The underwater point cloud segmentation method according to claim 3, characterized in that: The step S302 includes: When the density of a point in the point cloud data in the current frame is greater than or equal to a first predetermined threshold or the feature similarity with the previous frame is greater than or equal to a second predetermined threshold, increasing the adaptive weight of the point in the local feature or the global feature; When the density of a point in the point cloud data in the current frame is less than a third predetermined threshold or the feature similarity with the previous frame is less than a fourth predetermined threshold, the adaptive weight of the point in the local feature or the global feature is reduced.
5. The underwater point cloud segmentation method according to claim 3, characterized in that: The adaptive weight is adjusted according to the density by the following formula: Represents point p in the current frame i The adaptive weights, represents the point p in the previous frame i The adaptive weights, Represents point p in the current frame i The density, ρ baseline Indicates the predetermined initial density reference value.
6. The underwater point cloud segmentation method according to claim 3, characterized in that: Also includes: Calculating the importance of the multi-dimensional features of each point in the point cloud data by a predetermined calculation method, and adjusting the feature channel weight of each dimensional feature in the multi-dimensional features according to the importance; The calculation method includes: calculation based on statistics, gradient or attention mechanism; The multi-dimensional feature includes: the first local feature, the second local feature, the mesoscale feature, the global feature, the 2D feature representation, the fusion feature and / or the time series feature; The importance includes: discrimination, stability and / or relevance.
7. The underwater point cloud segmentation method according to claim 3, characterized in that: Also includes: S701, recording and aggregating feature changes of each point in the cluster in each frame of the time series to obtain a time average feature; S702, the time average feature is weightedly calculated by the density to obtain a priori feature; wherein the priori feature is calculated by the following formula: B k represents the prior feature, ρ i Represents point p i The density, f i avg represents said time-averaged characteristic; S703, assigning the same category label to each point in each cluster to obtain a preliminary segmentation area; The time series features of all points in the current frame are compared with the prior features by cosine similarity, and then the category label is adjusted by the following formula: represents the adjusted category label, CosSim represents the cosine similarity function, and f i (t) Representing the time series feature of a point in the current frame; S704 , performing feature comparison on the continuous frames of the time series to obtain a quantization parameter of label consistency, and when the value of the quantization parameter is less than a preset quantization threshold, recalculating the category label of the preliminary segmented region.
8. A device for underwater point cloud segmentation, characterized in that: It comprises a memory and a processor, wherein the memory stores at least one program, and the at least one program is executed by the processor to implement the underwater point cloud segmentation method according to any one of claims 1 to 7.
9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the underwater point cloud segmentation method as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Semantic segmentation method and system based on 3D point cloud
CN121353677A
Semantic segmentation method and system based on 3D point cloud
CN121353677B