Three-dimensional point cloud unsupervised change detection method
Through an unsupervised 3D point cloud change detection method, mask consistency loss and self-supervisory signals are used to decouple features and segmentation, which solves the problems of high labeling cost and weak generalization ability of existing methods and achieves efficient and accurate urban change detection.
Patent Information
- Application Number
- CN202510754741.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-12
AI Technical Summary
Existing 3D urban change detection methods rely on handcrafted features, with poor versatility and portability. Methods based on supervised or weakly supervised learning have high labeling costs and limited generalization capabilities, making them unable to effectively detect changes in semantic information.
An unsupervised 3D point cloud change detection method is adopted. By training a change segmenter and feature extractor, mask consistency loss and self-supervisory signals are used to decouple features and segmentation, thereby achieving accurate identification of point cloud change areas.
It significantly improves the accuracy and robustness of detection results, reduces the cost of annotation, enhances the model's prediction ability in unseen scenarios, and improves the generalization and computational efficiency of the detection process.
Smart Images

Figure CN120635882A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and artificial intelligence technology, and specifically relates to a three-dimensional point cloud unsupervised change detection method. Background Art
[0002] Digital city building change detection is a process that uses multi-source spatiotemporal data such as remote sensing images, aerial photography, laser radar (LiDAR), satellite data, and urban three-dimensional models, combined with computer vision, deep learning, and geographic information system (GIS) technologies, to automatically or semi-automatically identify, locate, and analyze information such as the increase, decrease, reconstruction, and height changes of buildings in urban areas over time. It plays a core role in the construction of smart cities and plays a vital role.
[0003] Existing 3D urban change detection methods can be roughly categorized into two main categories: traditional methods and deep learning. Traditional methods primarily include algebraic-based methods, classification-based methods (random forests, support vector machines), and others (Markov random fields, octrees). Algebraic-based methods determine whether point clouds have changed by calculating the height, distance, or texture differences between multiple time points and comparing them with a threshold. The most widely used method is to calculate the spatial Euclidean distance between point clouds and detect changes by determining whether the distance exceeds a threshold. Other methods convert 3D point clouds into other data formats to facilitate the application of existing change detection models. Deep learning-based methods typically build a twin architecture from two feature extraction backbone networks, extracting feature differences between point clouds at different times to detect changes. These methods exhibit outstanding learning and generalization capabilities. Excellent feature difference calculation methods, including direct difference methods, multi-scale difference methods, and feature fusion methods, are an important foundation for efficient 3D change detection. These methods facilitate the extraction of differences between point clouds by subtracting corresponding features or fusing multiple temporal features.
[0004] Traditional methods mostly rely on hand-crafted features, which cannot fully represent point clouds and have poor versatility and portability, resulting in reduced accuracy. Deep learning-based methods are divided into volumetric grid methods and point-based methods. However, when converting point clouds into structured data, the rasterization process of volumetric grid-based methods inevitably results in significant information loss. Point-based methods can easily obtain the differences between point clouds by subtracting corresponding features or fusing multiple temporal features, but suffer from insufficient change features and insufficient attention to change. In addition, the success of all existing methods is based on fully supervised or weakly supervised learning, which requires a large amount of point-level labels. Specifically, these methods use labels or clustered pseudo-labels to train conjoined point cloud segmentation networks to complete the point cloud change detection task. In addition, due to supervised training on specific labeled datasets, these networks have limited generalization capabilities and cannot effectively detect scenes involving changes in semantic information.
[0005] Therefore, there is an urgent need to provide a change detection method under three-dimensional buildings in digital cities to improve the defects of existing technologies. Summary of the Invention
[0006] In order to solve the above problems existing in the prior art, the present invention provides a method for unsupervised change detection of three-dimensional point clouds. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for unsupervised change detection of a three-dimensional point cloud, comprising:
[0008] For the same scene, first point cloud data and second point cloud data are acquired at different times, and the first point cloud data and the second point cloud data are preprocessed to obtain processed first point cloud data and processed second point cloud data;
[0009] Using the trained variation segmenter to process the processed first point cloud data and the processed second point cloud data, respectively obtaining a first probability and a second probability;
[0010] According to the first probability, the change points of the first point cloud data are segmented, and according to the second probability, the change points of the second point cloud data are segmented.
[0011] Beneficial effects of the present invention:
[0012] The present invention provides an unsupervised change detection method for three-dimensional point clouds, which significantly improves the accuracy and robustness of detection results by decoupling features and segmentation guidance, and accurately identifies changed and unchanged areas in the point cloud.
[0013] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flow chart of a method for unsupervised change detection of three-dimensional point clouds provided by an embodiment of the present invention;
[0015] Figure 2 is a schematic diagram of a variable splitter portion provided by an embodiment of the present invention;
[0016] Figure 3 is a schematic diagram of a model provided by an embodiment of the present invention;
[0017] Figure 4 This is a schematic diagram comparing the training data set before and after processing provided by an embodiment of the present invention;
[0018] Figure 5 is a schematic diagram of a mask consistency module provided by an embodiment of the present invention;
[0019] Figure 6 This is a schematic diagram of feature initialization provided by an embodiment of the present invention;
[0020] Figure 7 This is a schematic diagram of a training model provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The present invention will be further described in detail below with reference to specific examples, but the embodiments of the present invention are not limited thereto.
[0022] Point clouds, a commonly used data format in 3D tasks, can capture the geometric and visual characteristics of objects and scenes in detail and are robust to perspective distortion and lighting effects. These characteristics make point cloud data advantageous in solving complex change detection problems in 2D images and can more accurately describe the spatial relationships between objects and their surroundings. Consequently, point cloud data is becoming increasingly important in 3D city change detection (3DCD) tasks, and is also becoming increasingly crucial for analyzing changes in multi-temporal point cloud data.
[0023] See Figure 1 , Figure 1 3D point cloud unsupervised change detection method provided by the embodiment of the present invention is a flow chart, the present invention provides a 3D point cloud unsupervised change detection method, comprising:
[0024] S101. For the same scene, respectively acquire first point cloud data and second point cloud data at different times, and pre-process the first point cloud data and the second point cloud data to obtain processed first point cloud data and processed second point cloud data.
[0025] Specifically, in this embodiment, preprocessing the first point cloud data and the second point cloud data includes:
[0026] The first point cloud data and the second point cloud data are segmented into corresponding prism shapes.
[0027] S102: Use the trained variation segmenter to process the processed first point cloud data and the processed second point cloud data to obtain a first probability and a second probability, respectively.
[0028] Specifically, in this embodiment, the trained change segmenter includes a twin point cloud network, a nearest neighbor fusion module, a first multilayer perceptron, a second multilayer perceptron, a first activation function, and a second activation function; the trained change segmenter is used to process the processed first point cloud data and the processed second point cloud data to obtain a first probability and a second probability, respectively, including:
[0029] The processed first point cloud data is processed using a twin point cloud network to obtain a first segmentation feature, and the processed second point cloud data is processed using a twin point cloud network to obtain a second segmentation feature;
[0030] Taking the first segmentation feature as a benchmark, the nearest neighbor fusion module is used to perform a nearest neighbor operation to obtain a first fusion feature. Taking the second segmentation feature as a benchmark, the nearest neighbor fusion module is used to perform a nearest neighbor operation to obtain a second fusion feature.
[0031] The first fusion feature is processed by a first multilayer perceptron to obtain a first feature, and the first feature is processed by a first activation function to obtain a first probability; the second fusion feature is processed by a second multilayer perceptron to obtain a second feature, and the second feature is processed by a second activation function to obtain a second probability.
[0032] It should be noted that the parameters of the trained variation segmenter have been optimized and determined during the training process.
[0033] In this embodiment, in current research, many advanced segmentation models and techniques are used for change detection, such as feature embedding and attention mechanism. However, the focus of this invention is to propose a feasible unsupervised point cloud change detection training method. In order to reduce the complexity of the network, only the most basic structure is used as the segmenter. For the change segmenter, this embodiment selects the basic Unet network structure PointNet++. Since the change detection task requires the fusion of features extracted from the two branches of the Siamese network, a simple module NFF (Nearest Feature Fusion) based on the nearest neighbor is added to the last layer, such as Figure 2 As shown, Figure 2This is a schematic diagram of the change segmenter part provided by an embodiment of the present invention; for each point in the time point cloud, the nearest neighbor point is selected from other time point clouds for feature fusion. For the fused feature, Sigmoid is used as the final activation function to ensure that the output range is between 0 and 1.
[0034] In this embodiment, the expression of the first probability is:
[0035] M x =SIGMOD1(φ1(NN(PN(P x ),PN(P y )));
[0036] in, Represents the first point cloud data after processing, represents the second point cloud data after processing, PN(·) represents the twin point cloud network processing, which is the PointNet++ network, NN(·) represents the nearest neighbor operation, φ1(·) represents the multi-layer perceptron processing, mapping the features to one-dimensional space, and SIGMOD1(·) represents the activation function processing;
[0037] The expression of the second probability is:
[0038] M y =SIGMOD2(φ2(NN(PN(P x ),PN(P y )));
[0039] Among them, φ2(·) represents the multi-layer perceptron processing, and SIGMOD2(·) represents the activation function processing.
[0040] S103 : Segment the change points of the first point cloud data according to the first probability, and segment the change points of the second point cloud data according to the second probability.
[0041] In an optional embodiment of the present invention, in this embodiment, as Figure 3 , Figure 3 This is a schematic diagram of the model provided by an embodiment of the present invention. First, a model is constructed. The model mainly includes a change segmenter, a feature extractor, a mask consistency module, and a feature initialization module (a hybrid self-reconstruction task and a predicted point comparison task module). The core of three-dimensional point cloud change detection is to learn discriminative and robust features to capture and distinguish shape information in multiple point clouds. In order to achieve this goal in an unsupervised manner, the present invention proposes to train the change segmenter through a bidirectional optimization problem of a shared weight feature extractor and a change segmenter, thereby segmenting the change point cloud.
[0042] Among them, the training process of the trained change segmenter includes:
[0043] S1. Obtain a training data set, preprocess the training data set, obtain a processed training data set, and obtain a ground point cloud set; the training data set includes multiple training samples, and the training samples are point cloud data 1 and point cloud data 2 obtained for the same scene at different times.
[0044] In this embodiment, the preprocessing includes segmenting large-scale urban point cloud data, extracting ground point cloud sets, etc.
[0045] The present invention aims to quickly detect areas where changes have occurred in large urban point clouds, which means that it is not possible to directly process datasets of millions of points, because ordinary networks cannot withstand such large datasets and have limited computing power. Therefore, this embodiment extracts a cylindrical area near the center point of the changing scene in the street-level point cloud (Change3D), similar to extracting a patch in a SAR image. This training dataset marks deleted objects in the time-1 point cloud and added objects in the time-2 point cloud, and the shape of the deleted objects can be determined. The changed objects marked in the training dataset include pedestrians, vehicles, benches, signs, billboards, street lights, bicycles and other categories. In addition, the voxelized coordinates of the urban building point cloud (Urb3DCD) are used to segment it into prisms to facilitate network processing. The training dataset processing process is as follows: Figure 4 As shown, Figure 4 This is a schematic diagram showing the comparison of a training data set before and after processing provided by an embodiment of the present invention.
[0046] In this embodiment, predicting point contrast loss requires operations on the ground point cloud, making extraction crucial. For simplicity, this embodiment only operates on coordinates. After normalizing the point cloud data, a threshold is set. When the point cloud coordinate z falls below this threshold, the point is considered a ground point cloud, thus constructing the ground point cloud.
[0047] It should be noted that this embodiment adopts an unsupervised training method and does not require labeling of the training data set.
[0048] S2. Initialize the trainable parameters α of the feature extractor to be trained, and initialize the trainable parameters β of the variation segmentor to be trained.
[0049] S3. In the jth iteration, the encoder of the feature extractor to be trained is used to process the training sample to obtain feature 1 and feature 2 respectively.
[0050] S4. Use the decoder of the feature extractor to be trained to reconstruct feature 1 and feature 2 to obtain reconstructed training samples; during the reconstruction process, calculate the self-reconstruction loss and the predicted point comparison loss to optimize the trainable parameter α.
[0051] S5. Use the change segmenter to be trained to process the training samples to obtain probability 1 and probability 2 respectively. The probabilities are the change probabilities, and the change points are obtained during inference.
[0052] S6. Use the first mask consistency module to process feature one, feature two and probability one, and calculate the consistency loss of the masked point cloud data one; use the second mask consistency module to process feature one, feature two and probability two, and calculate the consistency loss of the masked point cloud data two; and introduce the norm constraint loss to optimize the trainable parameters α and β.
[0053] S7. Use the updated trainable parameters α and β as the trainable parameters of the feature extractor and the trainable parameters of the change segmentor in the j+1th iteration to obtain the j+1th feature extractor to be trained and the change segmentor to be trained; iterate in this way until the number of training times or the degree of convergence meets the preset conditions, and obtain the trained feature extractor and the trained change segmentor.
[0054] In this embodiment, the basic assumption of 3D change detection is that for the invariant scenes in multi-temporal point clouds, they have a certain consistency in geometric space. For the changed scenes, their features will be very different. Based on this assumption, the feature relationship of the invariant points is used as a free and rich supervisory signal to train the segmenter for point cloud change detection. This assumption can be similarly expressed as: for an invariant scene in one point cloud, adjacent regions with similar features can be found in another point cloud, while the changed scenes cannot be found or the features of adjacent regions are very different. Therefore, the goal of masked consistency loss (MCL) is to mine the semantic knowledge shared by the invariant scenes in multiple point clouds, such as Figure 5 As shown, Figure 5 This is a schematic diagram of a mask consistency module provided by an embodiment of the present invention.
[0055] Since the invariant points and the change points always appear as a region, it is meaningless to consider the features of only one point. In addition, unlike the pixel pairs in the image pair, the point pairs in the paired point cloud are not one-to-one corresponding. x ={p x,1 ,p x,2 ,…,p x,N Any query point in}, from point cloud data P y ={p y,1 ,p y,2 ,…,p y,N Then, for each point, the consistency is measured by the average value of the features of the searched KNN points, and then multiplied by the change probability to calculate the masked consistency loss.
[0056] Let Fx ={f x,1 ,f x,2 ,…,f x,N}、F y ={f y,1 ,f y,2 ,…,f y,N}、 P x and P y Point level features. For P x Any point P in x,i , first in P y Search its KNN points in F and then get these points in F y The features in the point cloud data are averaged and P y The i-th point in the point cloud data P x The average feature of the k nearest neighbors is expressed as:
[0057]
[0058] in, Represents point p x,i The corresponding P y The features of the adjacent regions in the point cloud data are shown in Figure 2. y Search point cloud p x,i k nearest neighbors, f y,k It represents the point features searched in the k-nearest neighbor set. The symbol k represents the number of neighboring points to be searched, which can be set according to the actual processing process.
[0059] A direct way to optimize the feature relationship is to minimize f x,i and The absolute difference between However, this goal may not be the best choice, as it imposes a linear penalty on the error of each feature dimension. In addition, in a high-dimensional feature space, the presence of a large number of feature dimensions may lead to the influence of noise and redundant features. Therefore, this embodiment chooses to supervise the relative quality of features and the quality of segmentation prediction through an unsupervised metric learning task. Specifically, for P x The feature f of each point in x,i , forcing them to approach P y The features of the adjacent regions in the image are filtered out using the change probability obtained by the segmentation algorithm to filter out the influence of the change points on the feature consistency learning. The masked consistency loss is expressed as:
[0060]
[0061] Among them, m x,i Indicates the change probability of the i-th point cloud in point cloud data 1, m y,irepresents the change probability of the i-th point cloud in point cloud data 2, f x,i Represents point cloud data - P x The feature of the i-th point in , Represents point cloud data - P x The i-th point in the point cloud data P y The average feature of the k nearest neighbors, f y,i Represents point cloud data 2P y The feature of the i-th point in , Represents point cloud data 2P y The i-th point in the point cloud data P x The average characteristics of the k nearest neighbors.
[0062] By minimizing this loss, the segmenter is forced to maximize the probability of changing points as much as possible and thus segment them by the threshold.
[0063] It should be noted that the output of the feature extractor is normalized before calculating the similarity, and the dot product is used to obtain the feature similarity.
[0064] When m x,i When both are 1, the objective function will be 0. The norm constraint is applied to the segmentation map to avoid completely changing the output, which is expressed as:
[0065]
[0066] Among them, M x represents the probability of one, M y represents the probability of two, express norm.
[0067] In summary, we get the mask consistency loss L mcon , whose expression is:
[0068] L mcon =L mxcon +L mycon +λL sum ;
[0069] Among them, L mxcon represents the consistency loss of the masked point cloud data, L mycon represents the consistency loss of the masked point cloud data, L sum represents the norm constraint loss, and λ represents the weight, which is used to balance the impact of the 1-norm norm on the optimization results.
[0070] It should be noted that mask consistency and 1-norm together can achieve the sought optimization of the segmenter, while also optimizing the quality of the feature extractor.
[0071] In this example, since discovering useful change detection knowledge from unlabeled data is often quite difficult, the masked consistency loss may not necessarily lead to useful optimization. Intuitively, the quality of the feature extractor is crucial, as the masked consistency loss only supervises the segmenter to obtain points with similar features. In other words, if the feature extractor is initialized well, it will provide good supervision for the segmenter, creating a virtuous cycle for the learning of the segmenter and feature extractor.
[0072] Conversely, due to poor initialization of the feature extractor, the learning process can lead to unpredictable results, a fact also pointed out by many unsupervised tasks. To avoid this problem, this paper proposes an auxiliary initialization task to supervise the network and jointly learn useful knowledge. Specifically, two simple tasks, including self-reconstruction and predicted point contrast loss, are used as two self-supervisory signals.
[0073] Self-reconstruction or self-encoding is a widely used unsupervised point cloud learning technique. To perform self-reconstruction, an encoder E and a decoder D based on an autoencoder are used to reconstruct the 3D coordinates of the point cloud on multiple time point clouds. The self-reconstruction loss L rec Defined as the chamfer distance, the expression is:
[0074]
[0075] Among them, P x Represents point cloud data, P y Represents point cloud data 2, E(·) represents a hierarchical point cloud feature learning network, D(·) represents a point cloud reconstruction decoder, and ChamferDist(·) represents the chamfer distance. In this embodiment, features are extracted from different levels of abstraction, and the following three nearest neighbor interpolation hierarchical upsampling is used to obtain point-level features.
[0076] In this embodiment, unlike image CD, in 3DCD, whether some points have been changed can be known in advance. For example, the grounding point in the same scene and the non-grounding point in different scenes. By comparing the loss of these predicted points, the feature extractor can be initialized well. The grounding points in different time periods of the current scene are used as positive samples, and the non-grounding points in different scenes are used as negative samples, such as Figure 6 As shown, Figure 6 This is a schematic diagram of feature initialization provided by an embodiment of the present invention. fpc The predicted point can be expressed as:
[0077]
[0078] Among them, G x represents the ground point cloud set, represents the indicator function, Represents point cloud data - P x The feature of the i-th point in , Represents point cloud data - P x The i-th point in the point cloud data P y The average feature of the k nearest neighbors, b j Indicates that different scene sets are not point cloud data. x Corresponding scene, b represents point cloud data-P x Corresponding scene, N represents point cloud data - P x points.
[0079] In this embodiment, compared with previous change detection methods that only focus on the quality of change features or the effectiveness of segmentation, which usually focus on the optimization of a single method and often ignore the global interaction between different modules, this embodiment combines the mask consistency module and the initialization module of the feature extractor to design the overall target loss of the framework, which is expressed as:
[0080] L mucd =L mcon +L rec +L fpc ;
[0081] By minimizing the above total loss function, the segmenter and feature extractor can be optimized simultaneously.
[0082] In an optional embodiment of the present invention, the training of the model mainly includes model initialization, learning rate and optimizer settings, selection of loss function hyperparameters, and data set sampling strategy. Figure 7 This is a schematic diagram of a training model provided by an embodiment of the present invention.
[0083] S1. Dataset sampling. 8192 points are extracted from each processed point cloud slice for training and testing, and only the coordinate information is normalized to zero mean and unit variance as input.
[0084] S2, learning rate and optimizer settings. Use the Adam optimizer, set the batch size to 4, and set the initial learning rate to 0.001, and reduce it exponentially (decay rate 0.7).
[0085] S3. Selection of loss function hyperparameters. The k value for all region features related to feature metric learning is set to 8. The hyperparameter λ of the balanced mask consistency module is set to 0.6.
[0086] S4, network initialization and training process. In order to train the network normally, the feature initialization epoch is set to 40, and the joint training epoch is also set to 40. In each epoch of the entire training process, the target L is used. rec and L fpcTo optimize the parameters of the feature extractor to better guide the extracted features for segmentation.
[0087] When the current epoch exceeds the feature initialization epoch, calculate L mcon To optimize the parameters of the segmenter. In summary, the objective of the initialization operation in the feature extractor makes it possible to learn discriminative features, while the mask consistency loss forces the network to predict similar features in invariant regions.
[0088] In this embodiment, the verification of the model mainly includes the reasoning and evaluation of the model.
[0089] S1. Model Inference. During the inference phase, when using the proposed model to detect changes in point cloud data, only the trained change segmenter is used for inference, and a confidence threshold of 0.5 is used to segment change points. Points with a confidence threshold above 0.5 are predicted as change points.
[0090] S2. Model Evaluation. For evaluation metrics, we use the most commonly used methods in change detection, namely overallaccuracy (OA) and mean intersection over union (mIoU). mIoU is the average of the IoU values of three landmarks: unchanged, deleted, and added.
[0091] In summary, the present invention provides an innovative unsupervised change detection method for 3D point clouds. This method addresses the high annotation costs and weak generalization capabilities of traditional urban point cloud change detection methods. By decoupling feature and segmentation guidance, the method significantly improves the accuracy and robustness of detection results. By introducing an innovative mask consistency loss and a feature extractor initialization module, the method decouples features from segmentation, accurately identifying changed and unchanged regions in the point cloud. Furthermore, the method does not rely on manual annotation, reducing costs and time consumption. Furthermore, by replacing it with a self-supervisory signal, the model's predictive ability in unseen scenes is effectively enhanced, improving the generalization of the detection process. The method provided by the present invention exhibits significant advantages in multiple aspects. First, the feature extractor is initialized using a self-reconstruction task and a priori knowledge point contrast loss, which improves the quality of feature extraction and provides a better supervision signal for the change segmentor, forming a virtuous cycle. Second, the mask consistency loss, as a new loss function, utilizes the shared information of unchanged regions in point clouds at different times as a free supervision signal to train the change segmentor, effectively mining the semantic knowledge of unchanged scenes. In addition, the method provided by the present invention uses a simple network structure, making the model both efficient and easy to promote.
[0092] In terms of practical applications, the method proposed in the present invention significantly improves the computational efficiency of the model while maintaining detection performance, reduces dependence on computing resources, and is suitable for actual deployment and application. Compared with existing supervised and weakly supervised methods, the method proposed in the present invention has demonstrated competitiveness in multiple evaluation indicators, and even surpasses these methods in some cases. In addition, the method proposed in the present invention has advantages in the number of parameters and training time, making it more efficient in practical applications. Through migration experiments on different data sets, it is proved that the present method has good migration capabilities and can perform effective change detection on a variety of different invisible data. These beneficial effects show that the method of this study provides a new and effective unsupervised learning solution for the field of point cloud change detection, which is applicable to a variety of complex urban and street scenes, such as urban development planning, environmental monitoring or disaster assessment, and significantly expands the application scope of change detection technology.
[0093] Based on the same inventive concept, the present invention further provides a three-dimensional point cloud unsupervised change detection device for implementing the three-dimensional point cloud unsupervised change detection method provided in the above embodiment of the present invention. The embodiment of the method is referred to above and will not be described in detail here. The device includes:
[0094] a data acquisition module, configured to acquire first point cloud data and second point cloud data at different times for the same scene, and preprocess the first point cloud data and the second point cloud data to obtain processed first point cloud data and processed second point cloud data;
[0095] a data processing module, configured to process the processed first point cloud data and the processed second point cloud data using a trained change segmenter to obtain a first change probability and a second change probability, respectively;
[0096] The result acquisition module is used to segment the change points of the first point cloud data according to the first probability, and segment the change points of the second point cloud data according to the second probability.
[0097] In an optional embodiment of the present invention, the effect of the unsupervised change detection method of three-dimensional point clouds provided by the above embodiment is verified through simulation experiments, specifically:
[0098] This simulation experiment verifies the effectiveness of the proposed method on the SLPCCD dataset. 80 training cycles were performed on the dataset, including 40 cycles of feature initialization and 40 cycles of joint training. The dataset includes 621 pairs of training samples, consisting of 398 pairs, 95 pairs, and 128 pairs of training, validation, and test datasets, respectively. The pytorch tool is used for sampling, normalization, and extraction of ground point subscripts, and the obtained subscripts are saved as npy files for later reading. The method of using dataset training in the present invention is relatively simple, and only requires accurately setting the path of the input npy and the path for saving the final result.
[0099] After all data processing is complete, we construct a dataset reading method. We use the FromNpy method in the toolkit to read the npy file and the transform method in the torch package to convert the read point sequence into a tensor format and normalize it. Furthermore, for each point cloud, we use the KNN method to find the nearest neighbor of each point at another time and save it for subsequent model use.
[0100] Once all data has been read, the model training process can begin. After completing the model initialization, the point cloud is first encoded using the encoder in the autoencoder, mapping it from coordinate space to latent space. The acquired features are then subjected to the nearest neighbor predicted point comparison loss, and the decoded features are then subjected to the diagonal loss with the original point cloud. Secondly, after the feature initialization is completed, the variation segmenter and feature extractor are bidirectionally optimized using the mask consistency loss.
[0101] After training, the test set point cloud was fed into the trained change segmenter. The final change points were obtained using point-level change probabilities and thresholds. After inference, the generated results were evaluated by calculating the OA and mIoU between the predicted and true change points to verify the effectiveness of the proposed model. The results are shown in Table 1.
[0102] Table 1 Evaluation indicators of the model proposed in this invention
[0103] OA mIoU 94.75 60.12
[0104] It should be noted that, in this document, relational terms such as first and second are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Furthermore, the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not explicitly listed. Without further limitation, an element defined by the phrase "comprising a..." does not preclude the presence of additional identical elements in the article or device comprising the element. Terms such as "connected" or "connected" are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. References to orientations or positional relationships, such as "upper," "lower," "left," and "right," are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate description and simplify the description of the present invention. They do not indicate or imply that the device or element referred to must have, be constructed, or operate in a specific orientation, and are therefore not to be construed as limiting the present invention.
[0105] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification.
[0106] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. For those skilled in the art of the present invention, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should be considered to fall within the scope of protection of the present invention.
Claims
1. A method for unsupervised change detection of three-dimensional point clouds, characterized in that: include: For the same scene, first point cloud data and second point cloud data are acquired at different times, and the first point cloud data and the second point cloud data are preprocessed to obtain processed first point cloud data and processed second point cloud data; Processing the processed first point cloud data and the processed second point cloud data using a trained variation segmenter to obtain a first probability and a second probability, respectively; According to the first probability, the change points of the first point cloud data are segmented, and according to the second probability, the change points of the second point cloud data are segmented.
2. The unsupervised change detection method for three-dimensional point clouds according to claim 1, characterized in that: The preprocessing of the first point cloud data and the second point cloud data includes: The first point cloud data and the second point cloud data are segmented into corresponding prism shapes respectively.
3. The unsupervised change detection method for three-dimensional point clouds according to claim 1, characterized in that: The trained change segmenter includes a twin point cloud network, a nearest neighbor fusion module, a first multi-layer perceptron, a second multi-layer perceptron, a first activation function, and a second activation function; the trained change segmenter is used to process the processed first point cloud data and the processed second point cloud data to obtain a first probability and a second probability, respectively, including: Using the twin point cloud network to process the processed first point cloud data to obtain a first segmentation feature, and using the twin point cloud network to process the processed second point cloud data to obtain a second segmentation feature; Taking the first segmentation feature as a reference, the nearest neighbor fusion module is used to perform a nearest neighbor operation to obtain a first fused feature; taking the second segmentation feature as a reference, the nearest neighbor fusion module is used to perform a nearest neighbor operation to obtain a second fused feature; The first fused feature is processed by the first multilayer perceptron to obtain a first feature, and the first feature is processed by the first activation function to obtain the first probability; the second fused feature is processed by the second multilayer perceptron to obtain a second feature, and the second feature is processed by the second activation function to obtain the second probability.
4. The unsupervised change detection method for three-dimensional point clouds according to claim 3, characterized in that: The expression of the first probability is: M x =SIGMOD1(φ1(NN(PN(P x ),PN(P y ))); Among them, P x Represents the first point cloud data after processing, P y represents the second point cloud data after processing, PN(·) represents the twin point cloud network processing, NN(·) represents the nearest neighbor operation, φ1(·) represents the multi-layer perceptron processing, and SIGMOD1(·) represents the activation function processing; The expression of the second probability is: M y =SIGMOD2(φ2(NN(PN(P x ),PN(P y ))); Among them, φ2(·) represents the multi-layer perceptron processing, and SIGMOD2(·) represents the activation function processing.
5. The unsupervised change detection method for three-dimensional point clouds according to claim 1, characterized in that: The training process of the trained variation segmenter includes: Acquire a training data set, preprocess the training data set to obtain a processed training data set, and obtain a ground point cloud set; the training data set includes a plurality of training samples, and the training samples are point cloud data 1 and point cloud data 2 acquired at different times for the same scene; Initialize the trainable parameters α of the feature extractor to be trained, and initialize the trainable parameters β of the variation segmentor to be trained; In the jth iteration, the encoder of the feature extractor to be trained is used to process the training sample to obtain feature 1 and feature 2 respectively; Reconstructing the first feature and the second feature using a decoder of the feature extractor to be trained to obtain a reconstructed training sample; during the reconstruction process, calculating a self-reconstruction loss and a predicted point contrast loss to optimize a trainable parameter α; The training samples are processed using the variation segmenter to be trained to obtain probability one and probability two respectively; A first mask consistency module is used to process the feature 1, the feature 2, and the probability 1, and calculate the consistency loss of the masked point cloud data 1; a second mask consistency module is used to process the feature 1, the feature 2, and the probability 2, and calculate the consistency loss of the masked point cloud data 2; and a norm constraint loss is introduced to optimize the trainable parameters α and β; The updated trainable parameters α and β are used as the trainable parameters of the feature extractor and the trainable parameters of the variation segmentor in the j+1th iteration, and the j+1th feature extractor to be trained and the variation segmentor to be trained are obtained; this iteration is repeated until the number of training times or the degree of convergence meets the preset conditions, and a trained feature extractor and a trained variation segmentor are obtained.
6. The unsupervised change detection method for three-dimensional point clouds according to claim 5, characterized in that: The self-reconstruction loss L rec The expression is: Among them, P x Represents point cloud data, P y represents point cloud data 2, E(·) represents the hierarchical point cloud feature learning network, D(·) represents the point cloud reconstruction decoder, and ChamferDist(·) represents the chamfer distance.
7. The unsupervised change detection method for three-dimensional point clouds according to claim 5, characterized in that: The predicted point contrast loss L fpc The expression is: Among them, G x represents the ground point cloud set, represents the indicator function, Represents point cloud data - P x The feature of the i-th point in , Represents point cloud data - P x The i-th point in the point cloud data P y The average feature of the k nearest neighbors, b j Indicates that different scene sets are not point cloud data. x Corresponding scene, b represents point cloud data-P x Corresponding scene, N represents point cloud data - P x points.
8. The unsupervised change detection method for three-dimensional point clouds according to claim 5, characterized in that: Mask consistency loss L mcon The expression is: THE mcon =L mxcon +L mycon +λL sum ; L sum =l1(M x +M y ); Among them, L mxcon represents the consistency loss of the masked point cloud data, L mycon represents the consistency loss of the masked point cloud data, L sum represents the norm constraint loss, λ represents the weight, which is used to balance the influence of the norm on the optimization result, m x,i Indicates the change probability of the i-th point in the point cloud data, m y,i represents the change probability of the i-th point in point cloud data 2, f xi Represents point cloud data - P x The feature of the i-th point in , Represents point cloud data - P x The i-th point in the point cloud data P y The average feature of the k nearest neighbors, f y,i Represents point cloud data 2P y The feature of the i-th point in , Represents point cloud data 2P y The i-th point in the point cloud data P x The average feature of the k nearest neighbors, M x represents the probability of one, M y represents the probability of 2, and l1 represents the l1 norm.
9. The unsupervised change detection method for three-dimensional point clouds according to claim 8, characterized in that: The point cloud data P y The i-th point in the point cloud data P x The average characteristics of the k nearest neighbors The expression is: Among them, k represents the number of searched neighboring points, and KNN represents the number of points in the point cloud data P y Search point cloud p x,i k nearest neighbors, f y,k It is represented as the point features searched in the k-nearest neighbor set.