Robustness evaluation method and system for multimodal stereo vision data reconstruction

By generating adversarial samples through adaptive perturbation limitation and iterative optimization, the robustness of the multimodal stereo vision data reconstruction model is evaluated, which solves the security problem of the model under adversarial sample attacks and improves the security and accuracy of the model.

CN120451420BActive Publication Date: 2025-09-09SHANDONG HI SPEED CONSTRUCTION MANAGEMENT GROUP CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510934698.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-09
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

The robustness of existing multimodal stereo vision data reconstruction models in the face of adversarial sample attacks is difficult to evaluate, and the input data may be affected by interference and mispairing, affecting the security and accuracy of the model, especially posing safety risks in applications such as autonomous driving.

Method used

By generating adaptive perturbation limits and iterative optimization methods, adversarial perturbations are generated for stereo vision data and visual data, and the optimization problem of formula (1) is used to solve the perturbations, and the perturbations that exceed the range are clipped to generate the final adversarial samples to evaluate the robustness of the model.

Benefits of technology

It improves the robustness of the multimodal stereo vision data reconstruction model, provides comprehensive adversarial sample generation, enhances the security of the model, and reduces the impact of adversarial sample attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451420B_ABST
    Figure CN120451420B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for evaluating the robustness of stereo vision data reconstruction based on multimodality, which relates to the technical field of deep learning and solves the problem of difficult-to-predict model robustness. The method includes: obtaining residual point cloud data, single-view images and target point clouds, and inputting the residual point cloud data, single-view images and target point clouds into a stereo vision data reconstruction model; generating original noise for the residual point cloud data and single-view images to generate initial point cloud perturbations; calculating adaptive perturbation limits based on partial point cloud data; solving adversarial perturbations with an optimization problem, and solving the optimization problem with the constraints to constrain the perturbations to a preset range; using an iterative optimization method to update the constrained perturbations, and cropping them after exceeding the perturbation constraints; generating final adversarial samples based on the updated and cropped perturbations; and using the adversarial samples to evaluate the robustness of the stereo vision data reconstruction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of deep learning, and specifically to a method and system for evaluating the robustness of multimodal stereo vision data reconstruction. Background Art

[0002] Stereo data reconstruction is a crucial task in the field of 3D vision. It can restore complete 3D object models from incomplete stereo data (such as point cloud data), addressing the issues of missing and low resolution during stereo data acquisition. Stereo data reconstruction has become a necessary upstream process in stereo data processing, with widespread applications in autonomous driving, robotic vision, 3D reconstruction, and other fields. However, due to factors such as sensor resolution, scanning range, and object occlusion, the collected stereo data often suffers from incompleteness, sparseness, and noise, hindering its subsequent processing and application.

[0003] In recent years, deep learning methods have made significant progress in stereo vision data reconstruction tasks, but research has shown that these models are vulnerable to adversarial attacks. Adversarial attacks are a security risk for deep learning models. An attacker adds small perturbations, causing the model to produce incorrect completion results. For example, in autonomous driving scenarios, stereo vision data reconstruction models are used to perceive and understand the surrounding environment. If sensor data is attacked, the model may incorrectly complete or identify obstacles, leading to traffic accidents.

[0004] Recently, multimodal stereo data reconstruction methods have improved the completion effect by introducing image information when completing incomplete stereo data. However, little research has focused on the robustness of multimodal stereo data reconstruction models. While the introduction of image modality may reduce the ambiguity of point cloud data and improve semantic robustness, in multimodal completion scenarios, both stereo and visual data may be attacked simultaneously, making the robustness of the model difficult to predict.

[0005] Research on the robustness of multimodal stereo data reconstruction models faces challenges: First, the input of multimodal stereo data reconstruction models includes two modalities: point clouds and images, which may be simultaneously interfered with and may also be affected by mispairing. Therefore, it is necessary to design multimodal adversarial attack methods and optimize attack strategies for different scenarios. Second, multimodal stereo data reconstruction models integrate information from two modalities. How to exploit the correlation between modalities for more effective attacks is a problem worth exploring. In recent years, multimodal stereo data reconstruction methods have become a research hotspot. By introducing visual modality data, better reconstruction results have been achieved. However, the robustness of these methods has not yet been evaluated, posing hidden dangers to practical applications. Summary of the Invention

[0006] In view of the shortcomings of the related technologies mentioned above, the present application provides a method and system for robustness evaluation of stereo vision data reconstruction based on multimodality to solve the above technical problems.

[0007] In a first aspect, the present application provides a robustness evaluation method for multimodal stereo vision data reconstruction, comprising:

[0008] Obtaining residual cloud data , single-view image And the target point cloud ,The residual point cloud data, single-view image and target point cloud are input into the stereo vision data reconstruction model;

[0009] Generate original noise for residual point cloud data and generate initial point cloud perturbation , generate original noise for single-view image, generate initialization image perturbation ;

[0010] Compute adaptive perturbation limits based on partial point cloud data and ;

[0011] Solve the anti-perturbation problem through the optimization problem of formula (1);

[0012] (1);

[0013] in, is a stereo vision data reconstruction model that accepts point cloud and image modalities, d is the point cloud distance calculated by the point cloud distance function; the constraints are and , solving the optimization problem with this constraint condition will constrain the disturbance to a preset range;

[0014] Use iterative optimization method to update the constrained perturbation and , and clipping after exceeding the perturbation constraint;

[0015] Based on the updated and clipped perturbation 、 , generate the final adversarial sample ; and use adversarial examples 、 Evaluating reconstruction models for stereo vision data robustness.

[0016] In a second aspect, the present application provides a multimodal stereo vision data reconstruction robustness evaluation system, comprising:

[0017] Data input unit, used to obtain part of the point cloud data , single-view image I and target point cloud Y;

[0018] Disturbance limit calculation unit, based on part of the point cloud data Compute adaptive disturbance limits;

[0019] Stereo vision data reconstruction model unit, using neural network Process input data;

[0020] Perturbation generation unit, used to initialize and calculate and update the adversarial perturbation according to the optimization formula and , and cut the disturbances that exceed the limit range;

[0021] Adversarial example evaluation unit, used to analyze the impact of adversarial perturbations on completion results.

[0022] As described above, the method and system for dynamically assessing enterprise management risk levels provided by this application have the following beneficial effects:

[0023] The present invention proposes a robustness evaluation method for stereo vision data reconstruction based on multimodal visual data. It can simultaneously generate adversarial perturbations for stereo vision data (such as point clouds) and visual data, introduce adaptive perturbation restrictions, and provide comprehensive adversarial sample generation, filling the gap in current research and improving the security of multimodal stereo vision data reconstruction technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, serving to explain the principles of the present application. It is obvious that the drawings described below are merely some embodiments of the present application, and a person of ordinary skill in the art can derive other drawings based on these drawings without inventive effort. In the drawings:

[0025] Figure 1 1 is a flowchart of the steps of the robustness evaluation method for multimodal stereo vision data reconstruction proposed by the present invention;

[0026] Figure 2 This is a sample diagram of an example input stereo vision data reconstruction model of the present invention;

[0027] Figure 3 A schematic diagram of the attacked completion result generated by the present invention based on the updated perturbation is provided;

[0028] Figure 4 This is a schematic diagram of the process of generating adversarial samples and verifying them with a classifier according to an embodiment of the present invention;

[0029] Figure 5It is a structural diagram of a multimodal stereo vision data reconstruction robustness evaluation system proposed in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will describe the embodiments of the present application with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for the purpose of illustrating the present application and are not intended to limit the scope of protection of the present application.

[0031] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. Therefore, the illustrations only show components related to the present application and are not drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component can be changed at will, and the component layout type may also be more complicated.

[0032] In the following description, a large number of details are discussed to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.

[0033] Figure 1 This is a flowchart of the robustness evaluation method for multimodal stereo vision data reconstruction proposed by the present invention, such as Figure 1 As shown, the steps include:

[0034] S1: Obtain residual cloud data , single-view image And the target point cloud as input.

[0035] Residual point cloud data is a 3D point cloud collected from a single viewpoint, often lacking information about the backside or occluded parts of an object. The single-view image and the residual point cloud data are from the same viewpoint, providing appearance or geometric clues. The target point cloud is the complete 3D point cloud corresponding to the residual point cloud data.

[0036] Figure 2 This is a sample diagram of an example input stereo vision data reconstruction model of the present invention, such as Figure 2 As shown in FIG, there are a single-view image of the car, residual point cloud data of the car, and a complete 3D point cloud of the car, and the complete 3D point cloud of the car is used as the target point cloud.

[0037] S2: Generate original noise for residual point cloud data and generate initial point cloud perturbation , generate original noise for single-view image, generate initialization image perturbation .

[0038] 、 For smaller values, it can be randomly assigned a mean of 0 and a standard deviation of 0.001.

[0039] S3: Calculate adaptive perturbation limits based on partial point cloud data and .

[0040] S4: Solve the adversarial disturbance using the optimization problem of formula (1);

[0041] (1);

[0042] in, is a stereo vision data reconstruction model that accepts point cloud and image modalities, d is the point cloud distance calculated by the point cloud distance function, such as Chamfer or other point cloud functions; the constraints are and , solving the optimization problem with this constraint condition will constrain the disturbance to a preset range;

[0043] S5: Use iterative optimization method to update the perturbation after S4 constraints and , and clipping is performed after exceeding the perturbation constraint.

[0044] Figure 3 A schematic diagram of the attacked completion result generated by the present invention based on the updated perturbation is given. Figure 3 include Figure 1 The comparison of the incomplete defect cloud data, the attacked completion result and the correct completion result is shown. Figure 1 Residual cloud data and Figure 3 By comparing the above, we can conclude that the disturbance modification involved in the steps given in the present invention Figure 1 After the defect cloud data is added, the model's completion result deviates from the correct value. The attacked completion result causes the classifier to misclassify it as "ship" instead of Figure 1 The original incomplete cloud data shows a "car" that meets the requirements of adversarial samples.

[0045] The embodiment of the present invention also provides a specific process for executing step S3:

[0046] S31: Calculate the point using the following formula (2) The average distance to the first nearest neighbor point; the first nearest neighbor point is the nearest neighbor point of the i-th point cloud calculated by the k-nearest neighbor algorithm (k-NN); is the i-th point cloud data, is the j-th point cloud data; yes The k-nearest neighbor point set of is the number of neighbor points set by the k-nearest neighbor algorithm;

[0047] (2);

[0048] S32: Calculated by formula (3) The standard deviation of

[0049] (3);

[0050] S33: Calculate the point by formula (4) The disturbance range;

[0051] (4);

[0052] in, is the scaling factor and t is the local uniformity weight.

[0053] The embodiment of the present invention also provides the execution of step S5 for , after exceeding the perturbation constraint, the specific process of clipping is as follows:

[0054] S51: For disturbance ,like , then rescaled to ;

[0055] For example, =2, for disturbance ,like , then rescaled to ;

[0056] The above process determines that the disturbance constraint is exceeded and then scales it to ensure that the disturbance size meets the constraint. , no operation is performed.

[0057] The embodiment of the present invention also provides , execute step S5, and perform trimming after exceeding the disturbance constraint condition. Specific process:

[0058] Perturbation of I ,use Norm constraint, that is, confirmation Is it satisfied If it is not satisfied, the pixel perturbations that are out of range are cropped. If it is satisfied, no operation is performed.

[0059] The present invention provides a specific process for executing S5 to update the constrained disturbance:

[0060] S5-1: The optimization process uses an iterative gradient update strategy, and each iteration performs the following updates:

[0061] K11: Computing The gradient of the direction and update the perturbation , the process of updating the disturbance is:

[0062] (5);

[0063] is the target point cloud, Single-view image It is the difference between two point clouds and can be measured by Chamfer Distance. yes Directional gradient;

[0064] Cut to suit constraint;

[0065] K12: Computing The gradient of the direction and update the perturbation ;

[0066]

[0067] yes Gradient in the direction; clipping is performed to meet constraint.

[0068] S6: Updated and pruned perturbation based on S5 、 , generate the final adversarial sample ; and use adversarial examples 、 Evaluating reconstruction models for stereo vision data robustness.

[0069] Evaluation indicators include target reconstruction error, target normalized reconstruction error, classifier accuracy and classifier attack success rate.

[0070] Target reconstruction error (T-RE), defined as T-RE = (5), d is the difference between the two point clouds, which can be measured by the Chamfer Distance; it is used to measure the similarity between the adversarial output and the target object.

[0071] In order to take into account the inherent reconstruction error of the completed model, the target normalized reconstruction error (T-NRE) is calculated as follows:

[0072]

[0073]

[0074] in, 、 is the adversarial sample, Y is the target point cloud, is the single-view image corresponding to object Y, It is the residual cloud corresponding to Y.

[0075] Use a classifier to classify the adversarial output to evaluate the semantic effectiveness of the attack; calculate the classifier accuracy and the classifier attack success rate to quantify the effectiveness of the adversarial perturbation in misleading the classifier; when the classifier result is still the original point cloud classification, the classification is considered accurate; when the classifier result is the target point cloud classification, the classifier is considered to be successfully attacked.

[0076] Figure 4 This is a schematic diagram of the process of generating adversarial samples through the embodiment of the present invention and verifying them through the classifier. Point cloud 1 is generated by the stereo vision data reconstruction model based on the point cloud without adding disturbances. Point cloud 2 is the attack completion result point cloud generated by the stereo vision data reconstruction model after adding disturbances through the method of the present invention. Figure 4 The effectiveness of the multimodal stereo vision data reconstruction robustness evaluation method of the present invention can be intuitively confirmed, and adversarial samples are effectively generated, so that the stereo vision data reconstruction model can be effectively evaluated.

[0077] Table 1 compares T-RE and T-NRE results under different perturbation limits. This paper provides a robustness evaluation method for stereoscopic data reconstruction (completion) methods. Table 1 shows the evaluation results of the three reconstruction methods. Table 1 shows that among the three completion methods, EGII has the best robustness, while ViPC has the worst. α is the scaling factor.

[0078]

[0079] The embodiment of the present invention also proposes a robustness evaluation system for reconstruction of stereoscopic vision data based on multimodality, Figure 5 This is a schematic diagram of the structure of a robustness evaluation system for multimodal stereo vision data reconstruction proposed in an embodiment of the present invention. Figure 5, the robustness evaluation system for multimodal stereo vision data reconstruction includes:

[0080] Data input unit, used to obtain part of the point cloud data , single-view image I and target point cloud Y; specifically used to obtain residual point cloud data , single-view image And the target point cloud .

[0081] Disturbance limit calculation unit, based on part of the point cloud data Calculate adaptive perturbation limits; specifically used to calculate adaptive perturbation limits based on part of the point cloud data and .

[0082] Stereo vision data reconstruction model unit, using neural network Process input data; specifically used to input residual point cloud data, single-view images and target point clouds into the stereo vision data reconstruction model; generate original noise for the residual point cloud data and generate initial point cloud disturbance , generate original noise for single-view image, generate initialization image perturbation ;

[0083] The disturbance-limited computation unit is also used to solve the optimization problem against disturbances by using formula (1);

[0084] (1);

[0085] in, is a stereo vision data reconstruction model that accepts point cloud and image modalities, d is the point cloud distance calculated by the point cloud distance function; the constraints are and , solving the optimization problem with this constraint will constrain the disturbance to a preset range.

[0086] Perturbation generation unit, used to initialize and calculate and update the adversarial perturbation according to the optimization formula and , and cut the disturbances that exceed the limit range;

[0087] Adversarial sample evaluation unit, used to analyze the impact of adversarial perturbations on completion results; specifically used based on updated and trimmed perturbations 、 , generate the final adversarial sample ; and use adversarial examples 、 Evaluating reconstruction models for stereo vision data robustness.

[0088] In one embodiment, the disturbance limiting calculation unit is configured to disturbance Based on - Local adaptive constraints of the nearest neighbors, calculated as follows:

[0089] (1) Calculation point to its - Average distance between neighboring points:

[0090]

[0091] (2) Calculation point The local standard deviation of :

[0092]

[0093] (3) Calculation point Permitted disturbance range:

[0094]

[0095] in, is the scaling factor, is the local uniformity weight.

[0096] In one embodiment, the disturbance generating unit is for disturbance , adopt the following clipping strategy: if , then rescaled to:

[0097] To ensure that the perturbation size complies with the constraints.

[0098] In one embodiment, the disturbance generating unit is for disturbance ,use Norm constraint, that is, satisfying , and crop the pixel perturbations that are out of range.

[0099] In one embodiment, the perturbation generation unit uses an iterative gradient update strategy in the optimization process, and each iteration performs the following updates: (1) Calculate The gradient of the direction and update the perturbation:

[0100]

[0101] Cut to suit Constraints. (2) Calculation The gradient of the direction and update the perturbation:

[0102]

[0103] Cut to suit constraint.

[0104] In one embodiment, the evaluation metrics of the adversarial sample evaluation unit include target reconstruction error, target normalized reconstruction error, classifier accuracy, and classifier attack success rate. Target reconstruction error (T-RE) is defined as T-RE = , measuring the similarity between the adversarial output and the target object. In order to take into account the inherent reconstruction error of the completion model, the target normalized reconstruction error (T-NRE) is calculated as follows:

[0105]

[0106] The adversarial output is classified using a classifier to evaluate the semantic effectiveness of the attack. The classifier accuracy and attack success rate are calculated to quantify the effectiveness of the adversarial perturbation in misleading the classifier. If the classifier output still matches the original point cloud classification, the classification is considered accurate; if the classifier output matches the target point cloud classification, the classifier is considered successfully attacked.

[0107] An embodiment of the present application also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs, which, when executed by one or more processors, enables the electronic device to implement the multi-modal stereo vision data reconstruction robustness evaluation method provided in the above-mentioned embodiments.

[0108] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When executed by a computer processor, the computer program causes the computer to perform the robustness assessment method for multimodal stereoscopic data reconstruction provided in each of the above embodiments. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.

[0109] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the multimodal stereo vision data reconstruction robustness assessment method provided in each of the above embodiments.

[0110] In the embodiments of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance. Throughout the specification and claims, the terms "including" and "comprising" are open-ended terms and should be interpreted as "including but not limited to."

[0111] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, any equivalent modifications or alterations accomplished by a person of ordinary skill in the art without departing from the spirit and technical concepts disclosed in this application shall be covered by the claims of this application.

Claims

1. A robustness evaluation method for multimodal stereo vision data reconstruction, characterized in that: The method comprises: Get residual cloud data X P , single-view image I and target point cloud U, input the residual point cloud data, single-view image and target point cloud into the stereo vision data reconstruction model; Generate original noise for residual point cloud data and generate initial point cloud perturbation Generate original noise for the single-view image and generate the initialization image perturbation δ I ; Compute adaptive perturbation limits based on partial point cloud data and ∈ I ; The optimization problem is solved to combat disturbances through formula (1); Among them, f θ is a stereo vision data reconstruction model that accepts point cloud and image modalities, d is the point cloud distance calculated by the point cloud distance function; the constraints are And ||δ I || q ≤∈ I , solving the optimization problem with this constraint condition will constrain the disturbance to a preset range; Use iterative optimization method to update the constrained perturbation and δ I , and clipping after exceeding the perturbation constraint; Based on the updated and clipped perturbation δ I , generate the final adversarial sample And use adversarial examples Evaluating the reconstruction model f for stereo vision data θ robustness.

2. The method according to claim 1, characterized in that Compute adaptive perturbation limits based on partial point cloud data and ∈ I ,include: Calculate the point by the following formula (2) The average distance to the first nearest neighbor point; the first nearest neighbor point is the nearest neighbor point of the i-th point cloud calculated by the k-nearest neighbor algorithm; is the i-th point cloud data, is the j-th point cloud data; yes The k-nearest neighbor point set, k is the number of neighbor points set by the k-nearest neighbor algorithm; Calculated by formula (3) The standard deviation of Calculate the point by formula (4) The disturbance range; where α is the scaling factor and t is the local uniformity weight.

3. The method according to claim 1, characterized in that Clipping is performed after perturbation constraints are exceeded, including: Targeting X p disturbance like Then rescale to 4. The method according to claim 1, wherein Clipping is performed after perturbation constraints are exceeded, including: Perturbation δ for I I , using L ∞ Norm constraint, confirm δ I Does it satisfy ||δ I || ∞ ≤∈ I , if it is not satisfied, the pixel perturbations that exceed the range are cropped.

5. The method according to claim 1, wherein Use iterative optimization method to update the constrained perturbation and δ I ,include: The optimization process adopts an iterative gradient update strategy, and each iteration performs the following updates: Calculate X P The gradient of the direction and update the perturbation The process of updating the perturbation is: Calculate the gradient in the I direction and update the perturbation δ I , the process of updating the disturbance is:

6. The method according to claim 1, characterized in that Using adversarial examples Evaluating the reconstruction model f for stereo vision data θ Robustness, including: Evaluation indicators include target reconstruction error, target normalized reconstruction error, classifier accuracy, and classifier attack success rate; The target reconstruction error is defined as T-RE = d1(f θ (X′ P ), Y)(5); Calculate the target normalized reconstruction error (T-NRE) as follows: in, is the adversarial sample, Y is the target point cloud, I Y is the single-view image corresponding to object Y, Y P It is the residual cloud corresponding to Y.

7. A robustness evaluation system for multimodal stereo vision data reconstruction, characterized by: The method for robustness evaluation of multimodal stereo vision data reconstruction according to any one of claims 1 to 6 is applied, comprising: Data input unit, used to obtain part of the point cloud data X P , single-view image I and target point cloud Y; The disturbance limit calculation unit is based on the partial point cloud data X P Compute adaptive disturbance limits; Stereo vision data reconstruction model unit, using neural network f θ Process input data; Perturbation generation unit, used to initialize and calculate and update the adversarial perturbation according to the optimization formula and δ I , and cut the disturbances that exceed the limit range; Adversarial example evaluation unit, used to analyze the impact of adversarial perturbations on completion results.

Citation Information

Patent Citations

  • Adversarial patch generation method and device for image-point cloud fusion perception model

    CN120182772A

  • Security systems for machine learning models

    US20240400079A1