A method for training elastically decoupled point cloud models for reflectivity disturbances

By decoupling reflectivity information and geometric information in the point cloud model and introducing discriminators for dynamic adjustment, the vulnerability problem in point cloud data processing is solved, and the robustness of the point cloud graph processing algorithm is achieved while maintaining high performance.

CN118036703BActive Publication Date: 2025-05-06ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311754550.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-05-06
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

The prior art is vulnerable to point cloud data processing and is vulnerable to noise, incomplete data and adversarial attacks, resulting in loss of point cloud map information and difficulty in processing.

Method used

A flexible decoupled point cloud model training method for reflectivity perturbation is proposed. By decoupling reflectivity information and geometric information, and introducing a discriminator to dynamically adjust the dependence of the model on reflectivity during the training process, the balance between performance and robustness is achieved.

Benefits of technology

Through this method, the model can more flexibly balance performance and robustness in the face of reflectivity perturbation and adversarial attacks, reduce the dependence on reflectivity information, thereby improving the stability and accuracy of point cloud map processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118036703B_ABST
    Figure CN118036703B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of autonomous driving, and discloses a method for training a point cloud model with elastic decoupling for reflectivity disturbance, comprising the following steps: Step 1: Prove the importance of reflectivity and the fragility of the model; Step 1.1: Calculate the amount of information of reflectivity; Step 1.2: Prove the fragility of reflectivity disturbance; Step 1.3: Prove the fragility of reflectivity to adversarial attacks; Step 2: Training based on decoupling; First, decouple reflectivity information and geometric information in the model, and then train the model. The present invention explores and demonstrates the importance of reflectivity information and the fragility of deep models facing reflectivity disturbances and adversarial attacks. The present invention proposes a decoupled training method to elastically adjust the model's dependence on reflectivity information in an adversarial mode. The method of the present invention can achieve a balance between performance and fragility in a flexible manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and in particular relates to an elastic decoupling point cloud model training method for reflectivity disturbance. Background Art

[0002] With the rapid development and popularization of autonomous driving technology, autonomous driving perception has become a key technical field. As a perception system for autonomous driving vehicles, LiDAR collects information such as reflectivity and point cloud geometry of point clouds, and uses perception algorithms to accurately perceive and understand roads, obstacles, pedestrians, and traffic signs. Although deep learning perception algorithms based on point clouds have achieved remarkable results, point cloud models have vulnerabilities to information such as reflectivity and point cloud geometry, which pose challenges to point cloud processing and application. First, point cloud images are easily affected by noise and incomplete data. This model-independent error will cause point cloud image information loss, making subsequent point cloud image processing and analysis difficult. In addition, point cloud images are also susceptible to adversarial attacks. Adversarial attacks refer to intentionally designing and modifying point cloud data to deceive point cloud models based on the gradient, objective function and other structures of the algorithm model, resulting in incorrect output of the point cloud processing algorithm.

[0003] Although many works have been proposed to solve the sensitivity problem in deep learning, a large amount of work is still focused on the robustness and fragility of images. Compared with image data, the research on point cloud data is still in its infancy. Existing solutions have made some progress in addressing the fragility of point cloud images, but there are still some challenges and shortcomings. Some technologies propose various point cloud denoising algorithms, using the geometric structure and local consistency of point clouds to remove noise points and abnormal points and improve the quality of point cloud images. However, point cloud denoising algorithms have some limitations in processing complex noise models and large-scale point cloud data; some researchers have proposed robust point cloud processing algorithms, which globally model point cloud images and combine statistical learning methods to improve the robustness of point cloud images to noise and abnormal data. However, due to the complexity and diversity of point cloud images, it is difficult to design a general robust point cloud processing algorithm; some technologies propose adversarial attack defense and detection methods based on adversarial attacks, input preprocessing, etc. to solve the impact of adversarial attacks on point cloud images. However, it is difficult for defense methods to cope with the evolving adversarial attack methods;

[0004] Although some progress has been made, the current solutions for point cloud vulnerability still have some shortcomings: 1) The lack of unified standards and evaluation indicators makes it difficult to compare and verify different research results, which makes some results in point cloud vulnerability research difficult to reproduce or incomparable; 2) The balance between robustness and performance. Improving the robustness of point cloud processing algorithms often leads to performance degradation or increased computational and storage complexity. In practical applications, how to improve the robustness of point cloud processing algorithms while maintaining high performance remains a challenge; 3) The continuous evolution of adversarial attacks. Adversarial attack methods are constantly evolving, and attackers may use new techniques and strategies to bypass current defense methods. Therefore, defenses against adversarial attacks need to be continuously updated and improved to cope with new attack methods. Summary of the invention

[0005] The purpose of the present invention is to provide a method for training an elastic decoupled point cloud model for reflectivity disturbance to solve the above-mentioned technical problems.

[0006] In order to solve the above technical problems, the specific technical solution of the elastic decoupling point cloud model training method for reflectivity disturbance of the present invention is as follows:

[0007] A method for training an elastic decoupled point cloud model for reflectivity disturbances comprises the following steps:

[0008] Step 1: Demonstrate the importance of reflectivity and model fragility;

[0009] Step 1.1: Calculate the information content of reflectivity;

[0010] Step 1.2: Demonstrate the vulnerability of reflectivity perturbations;

[0011] Step 1.3: Demonstrate the vulnerability of reflectivity to adversarial attacks;

[0012] Step 2: Decoupling-based training;

[0013] First, the reflectivity information and geometric information are decoupled in the model, and then the model is trained.

[0014] Furthermore, the step 1 adopts the projection-based point cloud segmentation technology, and each sample point in the data set is a point cloud frame F = {P1, P2, ..., P n}, the subscript n represents the number of points in frame F, P represents a five-dimensional vector (x, y, z, r, l), where x, y, z represent the geometric coordinates of the point cloud, l represents the true label, and r represents the reflectivity of point p. The n points in frame F are transformed by a mapping from 3D Euclidean space coordinates (x, y, z) to normalized discretized image coordinates (u, v), and the points are aligned by image coordinates. For each point, x, y, z, r are stored in the matrix respectively, creating a 5×h×w tensor T=[X,Y,Z,D,R] as the input of the point cloud model. Furthermore, the step 1 uses IoU as the model performance indicator. For the model M composed of the feature extractor G and the segmentation module S PD, given the tensor T converted from the frame F, pay attention to the performance degradation PD:

[0015]

[0016]

[0017] L is the label matrix, O(.) represents the perturbation or attack operation, PD shows the degree of performance deterioration of the model when it is subjected to different types of O(.), but the geometric information of the point cloud does not change, PD shows the degree of performance deterioration of the model when it is subjected to different types of perturbations or adversarial attacks, while the geometric information does not change.

[0018] Furthermore, in step 1.1, when training the point cloud segmentation model, the 5-channel input tensor T is replaced by only the reflectivity channel R. After mapping to the image coordinates, only the reflectivity information is used to predict the category label of each point, proving that the boundaries of different categories can be distinguished only by the reflectivity channel.

[0019] Furthermore, step 1.2 multiplies the reflectivity of each point by a random factor ranging from 0.5 to 1.5 as the input of the well-trained model, that is:

[0020] O(R)=RοN,(2)

[0021] 1.5>N ij >0.5,

[0022] Where o represents the Hadamard product;

[0023] The performance comparison results show that the IoU of all categories is significantly reduced. The model is susceptible to model-independent reflectivity perturbations, and the model relies heavily on reflectivity information during inference. Further, step 1.3: adversarial attack on the reflectivity channel without any modification of geometric information. This setting is model-aware. Through the predefined segmentation loss function Ls, the matrix δ is optimized in the following way:

[0024]

[0025]

[0026] O(R)=R+δ

[0027] For each frame, a multi-step optimization is performed and the performance comparison results show that the model is highly dependent on reflectivity and is therefore susceptible to reflectivity perturbations. Model-aware adversarial attacks lead to greater performance degradation compared to examples of random perturbations that are irrelevant to the model.

[0028] Furthermore, step 2 proposes a decoupling-based adversarial training method, which enables the model to flexibly strike a balance between performance and robustness.

[0029] First, a set of point cloud frames without reflectivity information is collected, which is achieved by setting O(R) = 0. The corresponding input is recorded as After extracting the features of the point cloud frame with reflectivity removed and the normal frame, a discriminator D is introduced to distinguish them from the extracted features. and G(T normal ), using the binary cross entropy loss function L D Optimize the discriminator D:

[0030]

[0031] Among them, if the reflectivity channel R is set to 0, then 1(T) = 1, otherwise 1(T) = 0. The goal of the discriminator D is to detect the reflectivity information in the features extracted by G. Then, G is trained in an adversarial manner, using the following loss function L G :

[0032]

[0033] Finally, the model is trained in the following way:

[0034]

[0035] λ is the weight coefficient. By adjusting λ, the model can rely more flexibly on reflectivity information between performance and vulnerability. A larger λ value makes the model focus more on deceiving the discriminator, thereby weakening the influence of the reflectivity channel and reducing and G(T normal ), if λ is smaller, the model will focus more on the segmentation task and extract the geometric and reflectivity information as much as possible.

[0036] The elastic decoupling point cloud model training method for reflectivity disturbance of the present invention has the following advantages:

[0037] The present invention studies the subject of vulnerability to reflectivity perturbations of deep learning models, and for the first time demonstrates the amount of reflectivity information, as well as the sensitivity of point cloud deep models to reflectivity perturbations. The present invention conducts a series of experiments to demonstrate and analyze adversarial attacks on reflectivity perturbations that are irrelevant to random models and gradient-based perception models. Taking point cloud segmentation as an example, the present invention conducts qualitative and quantitative experiments to show that a deep model is very sensitive to random reflectivity noise. That is, although reflectivity information has been shown to drive model performance, the use of reflectivity can also lead to vulnerability issues. Therefore, the present invention considers a better balance between point cloud model performance and vulnerability to reflectivity perturbations. Finally, the present invention proposes a decoupling method to flexibly make the model more or less dependent on reflectivity information to have more accuracy and less vulnerability. An adversarial learning process is organized to control the use of reflectivity information during training. In summary, the advantages of the present invention are mainly the following three points:

[0038] First, to complement other related works, we explore and demonstrate the importance of reflectivity information and the vulnerability of deep models to reflectivity perturbations and adversarial attacks.

[0039] 2. The present invention proposes a decoupled training method to flexibly adjust the model's dependence on reflectivity information in an adversarial mode.

[0040] Third, the method of the present invention can achieve a balance between performance and vulnerability in a flexible manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a visualization of the point cloud frame.

[0042] Figure 2 is a visualization of the prediction results with and without reflectivity perturbation. The fifth row is the visualization of the true label.

[0043] Figure 3 An overview of the disentanglement-based training method.

[0044] Figure 4 It is a graph of performance changes evaluated by IoU when facing random reflectivity perturbations.

[0045] Figure 5 It is the reflectivity distribution map of some categories in the Semantic-KITTI dataset.

[0046] Figure 6 Here are some qualitative examples of how the model performs when facing gradient-based model-aware deep attacks.

[0047] Figure 7 It is a graph showing the performance change evaluated by IoU when changing λ. DETAILED DESCRIPTION

[0048] In order to better understand the purpose, structure and function of the present invention, the following is a further detailed description of an elastic decoupled point cloud model training method for reflectivity disturbance of the present invention in conjunction with the accompanying drawings.

[0049] This paper uses point cloud semantic segmentation as an example of point cloud perception tasks, and adopts projection-based point cloud segmentation technology to demonstrate the importance of reflectivity information and the vulnerability of the model to reflectivity disturbances. IoU is used as an indicator of model performance, focusing on the performance deterioration degree PD.

[0050] The present invention uses point cloud semantic segmentation as a point cloud perception task. Point cloud semantic segmentation refers to the task of assigning each point in the point cloud data to a corresponding semantic category, which is used to understand and reason about objects and scenes in a three-dimensional environment. The present invention adopts projection-based point cloud segmentation technology. Each sample point in the data set is a point cloud frame F = {P1, P2, ..., P n}, subscript n represents the number of points in frame F, P represents a five-dimensional vector (x, y, z, r, l), where x, y, z represent the geometric coordinates of the point cloud, l represents the real label, and r represents the reflectivity of point p, including distance, object surface properties, etc. The n points in frame F are transformed by a mapping from 3D Euclidean space coordinates (x, y, z) to normalized discretized image coordinates (u, v). The points are aligned by image coordinates. For each point, x, y, z, r are stored in matrices respectively, creating a 5×h×w tensor T=[X,Y,Z,D,R] as the input of the point cloud model.

[0051] First, the present invention demonstrates the importance of reflectivity information and the model's vulnerability to reflectivity perturbations and adversarial attacks. Specifically, IoU is used as the model performance metric. For the model M composed of a feature extractor G and a segmentation module S, given a tensor T converted from a frame F, we focus on the performance degradation PD.

[0052]

[0053]

[0054] L is the label matrix, and O(.) represents the perturbation or attack operation. We show how the model deteriorates when subjected to different types of O(.), but the geometric information of the point cloud (X, Y, Z, D channels) does not change. Further, we show how the model deteriorates when subjected to different types of perturbations or adversarial attacks, while the geometric information (X, Y, Z, D channels) does not change.

[0055] Finally, this paper proposes a training method with elastic decoupling for reflectivity disturbances, dynamically adjusts the model's importance to reflectivity, and strikes a balance between performance and fragility by adjusting the model's dependence on the reflectivity channel.

[0056] Specifically, the elastic decoupling point cloud model training method for reflectivity disturbance of the present invention comprises the following steps:

[0057] Step 1: Demonstrate the importance of reflectivity and model fragility.

[0058] Reflectivity in LiDAR refers to the intensity of reflected laser light, ranging from 0 to 1, and contains rich information, such as the distance between the object and the sensor and the surface characteristics of the object, and is widely used in applications such as point cloud segmentation. Here, the present invention further demonstrates the importance of reflectivity information and the vulnerability of the model to reflectivity changes.

[0059] Step 1.1: Calculate the information content of reflectivity.

[0060] When training the point cloud segmentation model, the present invention replaces the 5-channel input tensor T with only the reflectivity channel R. After mapping to image coordinates, we only use the reflectivity information to predict the category label of each point. Figure 1 The results are visualized in . Figure 1 The first row is the distance map of the point cloud frame, which indicates the distance from the object to the sensor. The second row is the reflectivity map of the point cloud frame, where some boundaries of different categories can be roughly distinguished. The third and fourth rows are the prediction results of the SalsaNext model, but the fourth row only uses the reflectivity channel. This results in worse results than the former, however, the points can still be reasonably segmented and different categories can be roughly identified, which shows that reflectivity contains rich information and can be used for point cloud perception. The fifth row is the visualized true label. Using only reflectivity leads to worse results compared to the prediction of the original setting, but objects of different categories can still be roughly distinguished. We are also working on Figure 1 The reflectivity is visualized in Figure 3, and the results show that the boundaries of different classes can be distinguished using the reflectivity channel alone.

[0061] Step 1.2: Demonstrate vulnerability to reflectivity perturbations.

[0062] The present invention implements a specific example to demonstrate the impact of reflectivity disturbance. Considering that the reflectivity information is randomly disturbed, the reflectivity of each point is multiplied by a random factor (ranging from 0.5 to 1.5) as the input of a well-trained model, that is:

[0063] O(R)=RοN,(2)

[0064] 1.5>Nij >0.5,

[0065] Where o represents the Hadamard product. Figure 4 The performance comparison is shown in Figure 2, where we can see that the IoU of all categories is significantly reduced. Figure 2 The perturbed reflectivity channel is visualized in , which is still similar to the original channel. Figure 2 The first row is the original reflectivity map. The second row is the reflectivity map with random perturbations, which is not very noticeable. The third and fourth rows are the predicted results, but the reflectivity channel in the fourth row is perturbed. The random perturbations mislead the model, so many categories are misclassified as some similar categories (for example, cyclists are mistaken for people and bicycles, cars are mistaken for trucks, and roads are mistaken for sidewalks). However, the third and fourth rows show that these perturbations may cause the model to confuse some similar categories, such as roads and sidewalks, cyclists and pedestrians, cars and trucks. This result not only reveals that the model may be easily affected by reflectivity perturbations that are irrelevant to the model, but also shows that the model relies heavily on reflectivity information during inference.

[0066] Step 1.3: Demonstrate the vulnerability of reflectivity to adversarial attacks.

[0067] We conduct adversarial attacks on the reflectivity channel without any modification of the geometric information. This setting is model-aware because we need the gradient information of the model during the attack, and is therefore more difficult to implement in practice. We conduct this experiment only to further demonstrate the effect and importance of reflectivity. With the predefined segmentation loss function Ls (cross entropy loss, Lovasz loss, etc.), we optimize the matrix δ in the following way:

[0068]

[0069]

[0070] O(R)=R+δ

[0071] For each frame, we perform multi-step optimization and the performance comparison results are shown in Figure 4 middle. Figure 4 The performance suffers a significant degradation (down more than 15%) in the CNN model. The results show that the model is highly dependent on reflectivity and is therefore susceptible to reflectivity perturbations. Compared to examples of random perturbations that are independent of the model, model-aware adversarial attacks lead to larger performance degradation, further highlighting the impact and importance of reflectivity on the model.

[0072] Step 2: Decoupling-based training.

[0073] There is a dilemma between performance and robustness. Completely removing the reflectivity channel from training and inference can completely prevent any perturbations or attacks on the reflectivity, but it will also lead to performance degradation. Utilizing reflectivity information can improve performance in normal scenarios, but it may also increase the risk of being affected or attacked by reflectivity perturbations. To address this issue, we propose a decoupled adversarial training method that allows the model to elastically strike a balance between performance and robustness. An overview of our method is shown in Figure 3 shown.

[0074] Inspired by ElasticNet, this paper proposes a method to elastically adjust the importance of reflectivity during training and inference. First, the reflectivity information and geometric information are decoupled in the model.

[0075] The present invention proposes to achieve this goal in an adversarial way. Specifically, we first collect a set of point cloud frames that do not contain reflectivity information, which is achieved by setting O(R) = 0. The corresponding input is recorded as After extracting the features of the point cloud frame with reflectivity removed and the normal frame, we introduce a discriminator D to distinguish them from the extracted features and G(T normal ). We use the binary cross entropy loss function L D Optimize the discriminator D:

[0076]

[0077] Where 1(T) = 1 if the reflectivity channel R is set to 0, otherwise 1(T) = 0. The goal of the discriminator D is to detect the reflectivity information in the features extracted by G. Then, G is trained in an adversarial manner using the following loss function L G :

[0078]

[0079] Finally, the model is trained in the following way:

[0080]

[0081] λ is the weight coefficient. By adjusting λ, the model can rely more flexibly on reflectivity information between performance and vulnerability. A larger λ value will make the model focus more on deceiving the discriminator, thereby weakening the influence of the reflectivity channel and reducing and G(T normal ). On the other hand, if λ is smaller, the model will focus more on the segmentation task and extract as much geometric and reflectivity information as possible.

[0082] Example:

[0083] A. Dataset and benchmark settings:

[0084] Datasets and evaluation metrics:

[0085] This paper uses the large-scale Semantic-KITTI dataset, which was collected in different cities in Germany and contains a total of 22 time series and 43,551 point cloud frames. The model is trained between sequences 00 and 10 and validated with sequence 08. The accuracy and intersection-over-union (IoU) are used as evaluation indicators. S(G(T)) c is the set of points predicted to be of category c, L c is the set of actual data categories c, and |.| represents the cardinality.

[0086] Benchmark settings:

[0087] The present invention uses the projection-based point cloud segmentation process as the benchmark setting because it has practical inference time and memory cost advantages. Specifically, the present invention uses the SalsaNext model based on the projection process as the benchmark model.

[0088] Trained on four NVIDIA Titan RTX graphics cards, with a batch size of 24 and epochs of 150, the parameters of the last epoch were selected as the general representation of the results, rather than the parameters of the epoch with the best IoU. This paper focuses on the model depth and skips the post-processing KNN module, because post-processing does not affect the training and inference process of the deep learning model, and verifies that the IoU is 53.3.

[0089] B. Informational nature of reflection intensity:

[0090] Statistics of reflection intensity information:

[0091] The present invention claims that the reflection intensity contains rich information and can be used for semantic segmentation. Figure 5 The distribution of all categories is visualized in Figure 1, and the results show that the distribution of reflection intensity varies for different categories. Some categories have peaks (such as parking lots and sidewalks), and some categories are relatively uniform (such as motorcycles and telephone poles). A considerable portion of vehicle points have weak reflection intensity (close to 0), while in the case of traffic signs, points with strong reflection intensity (close to 1) account for the largest proportion. Since different categories have different reflection intensity distributions, semantic segmentation models are able to exploit this information during training, and changes in this distribution may affect the performance of the model during inference.

[0092] Reflection strength for semantic segmentation:

[0093] The present invention further conducts qualitative experiments to demonstrate the informativeness of reflection intensity for semantic information. After projection, the visualization results of some reflection intensity channels are as follows: Figure 1 As shown in , the boundaries of different categories (such as vehicles and roads, roads and terrain, terrain and tree trunks, etc.) can be roughly distinguished. We also trained the model using only the reflection intensity channel as input and Figure 1 Some prediction results are shown. Compared to the results using the 5-channel tensor, the results using only the reflection intensity channel perform worse, especially for some similar categories such as vegetation and terrain. However, many categories such as vehicles, bicycles, trucks, roads, etc. can still be inferred with reasonable prediction results. These results also show that deep models can infer semantic information from reflection intensity for segmentation tasks.

[0094] C. Model vulnerability:

[0095] Random, model-independent perturbations:

[0096] We demonstrate the sensitivity of the model to changes in reflection intensity by performing random model-independent perturbations, such as Figure 4 Almost all categories are severely affected, with the average IoU dropping by more than 15 percentage points, indicating that the model relies heavily on reflection intensity information during inference, so simple random model-independent perturbations can significantly affect model performance.

[0097] Gradient-based Model-aware Adversarial Attacks:

[0098] Furthermore, the present invention uses adversarial attack to perturb the reflectivity intensity channel, which is a model-aware perturbation. Compared with model-independent perturbation, it requires the availability of all model parameters and calculates δ for each specific sample. Therefore, although it may cause significant degradation of the model, it is basically impractical in practical applications. The present invention only conducts these experiments to further demonstrate the vulnerability of deep models to changes in reflectivity.

[0099] For each sample, we optimize δ for a specific 5, 10, 15, 20 steps (with a step size of 0.1), as shown in Table I.

[0100]

[0101] Table I

[0102] We also list the average L-1 norm of δs, which is Compared to the normalized reflection intensity channel, δ has a mean of 0, which is relatively small, but the performance loss is large. It can be seen that after 10 steps of optimization, the performance drop is larger than in the case of model-independent perturbations (for the baseline model, the average change in reflection intensity is 0.0129). The results show that attacking reflection intensity alone can also significantly affect deep models, even if the magnitude of the perturbation is relatively small.

[0103] We are Figure 6 Some qualitative examples are shown in Figure 6 For each input frame, the gradient descent optimization is performed with a step size of 0.01 for 5, 10, 15, and 20 steps. It can be seen that as the total number of steps increases, the misclassified area also increases. In the first three columns, the very important category "road" is gradually occupied by the sidewalk or parking space, as shown in the prediction results of the baseline model, while our elastic disentanglement training method alleviates this situation. Columns 5 to 7 show some cases of cars, which are similar to the road situation. Columns 4 and 8 are some failure cases of our method. For the former, using our method, the model misclassifies the sidewalk as a parking space. However, compared with the baseline model, the important road area is protected in this case. As for column 8, the model trained using our method confuses cars with trucks, while the baseline model misleadingly classifies trucks as fences.

[0104] Effectiveness of D elastic untangle training:

[0105] Training prototype:

[0106] We use the original samples and the non-reflected samples for disentanglement training. The discriminator D has 5 residual blocks, and finally an average pooling layer and a fully connected layer. The input of D is the features of the fifth block of the baseline model (which means G is the first five blocks of the entire model). As shown in Equation (5), the disentanglement training loss L G is only computed on raw samples. This is because we mainly want to control the feature extractor G to use more or less information on the telemetry channel, so we do not need to impose constraints on the reflection-free features.

[0107] Effectiveness of the hyperparameter λ:

[0108] The present invention shows here the effect of adjusting the hyperparameter λ. Intuitively, if λ is close to 1, G will pay less attention to the semantic segmentation task, and its main focus is to deceive the discriminator. On the other hand, if λ is close to 0, it will degenerate into a kind of data augmentation (i.e., training the model with some non-reflective samples).

[0109] Changing λ and showing Figure 7The dark grey dashed line represents the average IoU when the model is inferred on uncontaminated samples, and the light grey dashed line represents the average IoU on perturbed samples (we use random irrelevant model perturbations as examples here). The solid line represents the performance of the baseline model when inferring with or without the perturbed reflection channel. It can be seen that as λ increases, the performance on uncontaminated samples decreases slightly, but the robustness to reflection perturbations increases significantly.

[0110] Furthermore, it can be observed that the performance of perturbed samples decreases unexpectedly quickly as λ approaches 1. This seems counterintuitive, as larger λ should lead to models that are more robust to reflected perturbations. However, this can likely be explained by the practice in the field of adversarial training, such as the training process of GANs, where extremely strong or weak discriminators can lead to training failures. In this case, it is found that when λ approaches 1, the loss of the discriminator is maximized, hindering the adversarial training process.

[0111] Dealing with random, irrelevant model perturbations:

[0112] Quantitative results such as Figure 7 As shown, the detailed results for each category can be seen in Table II.

[0113]

[0114] Form II

[0115] Compared to the baseline model, after disentangled training, the IoU of almost all categories increases significantly in the face of random irrelevant model perturbations. For example, when λ is set to 0.5, the performance gap at inference is reduced by about 70% (from 17.9% to 5.8%), while the performance sacrifice on uncontaminated samples is relatively small (about 2 percentage points), which shows the effectiveness of our method.

[0116] Countering Gradient-Based Model-Aware Adversarial Attacks:

[0117] We set λ to 0.5 and show the detailed results in this section. We evaluate our model in adversarial attack scenarios and report detailed IoU for each category. Even after disentanglement training, the performance still drops significantly, but the performance gap is smaller than the baseline model. The results also show that the attack algorithm needs to generate larger δ (approximately 15%, 24%, 27%, and 33% more δ than when attacking the baseline model) to attack our model, and our method consistently shows an advantage of about 5 percentage points over the baseline model in the average IoU metric.

[0118] We are still Figure 6Some qualitative examples are shown in . It can be observed that as the number of attack steps increases, the misclassified area also increases. In the first three columns, the very important category "road" is gradually encroached by sidewalks or parking spaces, as shown in the prediction results of the baseline model, while our disentanglement training method alleviates this situation. Columns 5 to 7 show the situation of some cars, which is similar to the situation of roads. Columns 4 and 8 are some failure cases of our method. For the former, the model using our method misclassifies some sidewalk points as parking spaces. However, compared with the baseline model, important road areas are protected in this case. As for column 8, the model trained using our method confuses cars with trucks, while the baseline model misleadingly classifies trucks as fences.

[0119] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A method for training an elastic decoupled point cloud model for reflectivity disturbance, characterized in that: The steps include: First, a set of point cloud frames without reflectivity information is collected, which is achieved by setting O(R) = 0. The input corresponding to the point cloud frame without reflectivity information is recorded as The input corresponding to the point cloud frame containing reflectivity information is denoted as T normal After extracting the features of the point cloud frame without reflectivity and the normal frame, a discriminator D is introduced to distinguish the features extracted by the feature extraction module G. and G(T normal ), using the binary cross entropy loss function L D Optimize the discriminator D: min L D (T;D,G)=-[1(T)*log(D(G(T)))+(1-1(T))*log(1-D(G(T)))],(4) Among them, if the reflectivity channel R is set to 0, that is, the input does not contain reflectivity information, then 1(T) = 0, otherwise, it represents the input corresponding to the point cloud frame, that is, it contains reflectivity information, then 1(T) = 1. The goal of the discriminator D is to detect the reflectivity information in the features extracted by G. Then, G is trained in an adversarial manner, using the following loss function L G : minL G (T;D,G)=-(1-1(T))*log(1-D(G(T))),(5) Finally, the model is trained in the following way: min((1-λ)*L s (T,L;S,G)+λ*L G (T;D,G)),(6) λ is the weight coefficient. By adjusting λ, the model can rely more flexibly on reflectivity information between performance and vulnerability. A larger λ value makes the model focus more on deceiving the discriminator, thereby weakening the influence of the reflectivity channel and reducing and G(T normal ), if λ is smaller, the model will focus more on the segmentation task and extract the geometric and reflectivity information as much as possible.