Multi-modal based three-dimensional point cloud adversarial sample intelligent supervision method and system
By constructing a continuous time factor graph and using iterative optimization techniques, the problem of online monitoring against disturbances and natural degradation in 3D point cloud perception scenarios was solved. This enabled real-time correction of multi-sensor drift and time delay, as well as adaptive adjustment of environmental quality, reducing false alarm and missed alarm rates and improving the robustness and stability of monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIMEI UNIV
- Filing Date
- 2026-04-30
- Publication Date
- 2026-06-02
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies struggle to provide stable online monitoring against disturbances and natural degradation in 3D point cloud sensing scenarios, resulting in high false alarm and false negative rates. Furthermore, multi-sensor time offsets and extrinsic parameter drifts are difficult to correct in real time, leading to the accumulation of cross-modal matching errors and a lack of adaptive mechanisms for environmental quality detection.
A continuous time factor map is constructed, multimodal data is collected for initial alignment, environmental quality observations are introduced and switchable constraints are set, external parameters and time offsets are estimated through iterative optimization, cross-modal consistency is verified by combining hybrid robust loss, and adversarial risk scores are generated.
It enables online estimation and dynamic correction of drift and time delay asynchrony of multi-sensor extrinsic parameters, reduces cross-modal registration error, suppresses false alarms due to natural degradation, improves robustness against disturbances, and enhances regulatory stability.
Smart Images

Figure CN122133081A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of multi-sensor fusion perception security and 3D point cloud processing, and in particular to an edge computing-based, multimodal 3D point cloud adversarial sample intelligent monitoring method and system. Background Technology
[0002] 3D LiDAR point clouds, due to their ability to directly characterize spatial geometry, have been widely applied in perception scenarios such as autonomous driving, robotic inspection, and UAV mapping. To improve the integrity and reliability of perception, engineering typically integrates point clouds with auxiliary modes such as cameras, millimeter-wave radar, and inertial measurement units. Through time synchronization, coordinate extrinsic parameter calibration, and cross-modal data association, functions such as target detection, tracking, and localization are achieved. Meanwhile, research on the security of point cloud perception systems continues to advance, leading to the development of adversarial perturbation and anomaly detection techniques for point clouds. Examples include anomaly detection based on model confidence or output consistency, detection based on statistical distribution offset, and cross-modal consistency verification based on reprojection or reconstruction errors. In terms of calibration, schemes utilizing graph optimization or factor graphs for extrinsic parameter estimation and self-calibration have also been developed.
[0003] However, existing technologies still have the following shortcomings: 1. Multi-sensor time offset and extrinsic parameter drift have continuous and slow-changing characteristics in actual operation. Existing methods mostly use offline calibration or discrete time correction, which makes it difficult to stably track the slow changes in time delay and extrinsic parameters in online scenarios. This leads to the accumulation of cross-modal matching errors and affects the reliability of anomaly detection.
[0004] 2. Natural degradation such as rain, fog, dust, occlusion, sparse point clouds, and reflection can significantly reduce the quality of cross-modal association. Existing consistency or residual detection usually uses fixed thresholds or fixed weights, lacking explicit modeling and adaptive adjustment mechanisms for environmental quality, which can easily misjudge natural degradation as an attack or miss real attacks.
[0005] 3. Some anomaly detection relies on a single modality or a single model output, lacking robust constraints and multimodal closed-loop verification mechanisms under conditions of spatiotemporal alignment errors, environmental degradation, and adversarial disturbances. This results in high false alarm and false negative rates, making it difficult to form a stable and implementable intelligent monitoring solution.
[0006] Therefore, there is a need for a method and system for intelligent monitoring of adversarial examples in 3D point clouds that can address the shortcomings of existing technologies. Summary of the Invention
[0007] One objective of this invention is to propose a method for intelligent surveillance of adversarial examples in 3D point clouds based on multimodal approaches. Addressing the challenges of existing technologies in point cloud sensing scenarios, where adversarial disturbances are often subtle and variable, compounded by natural degradation factors such as rain, fog, dust, occlusion, sparse point clouds, and reflections, and further complicated by time asynchrony between multiple sensors, inconsistencies in coordinate systems, and drift in extrinsic parameters, single-modal or fixed-threshold consistency detection struggles to distinguish between real-world environmental changes and adversarial attacks, resulting in high false positives and false negatives, and hindering stable online surveillance, this invention proposes a method that collects point cloud and auxiliary modal data within a preset time window and extracts environmental quality observations. It constructs a continuous time factor graph containing slowly varying extrinsic parameters, time offset parameters, and latent environmental quality variables. Switchable constraints are introduced to generate matching reliability weights, which are adaptively adjusted by the latent environmental quality variables. Iterative joint optimization is performed using a hybrid robust loss to update time and coordinate alignment. Finally, cross-modal cyclic consistency checks are performed on the updated aligned data to generate an adversarial risk score and output the surveillance results. This invention offers the technical advantages of online estimation of extrinsic parameters and time delays, adaptation to natural degradation, robustness to outlier matching and adversarial disturbances, reduced false positives and false negatives, and improved surveillance stability.
[0008] This invention provides a multimodal 3D point cloud adversarial example intelligent surveillance method, including: S1. Collect multimodal time series data of the monitored object, including 3D point cloud data and at least one auxiliary modal data. Associate time information and sensor identifiers between the 3D point cloud data and the auxiliary modal data. Perform initial time alignment and initial coordinate alignment based on initial extrinsic parameters within a preset time window to obtain initial aligned data and an environmental quality observation sequence characterizing the degree of environmental degradation. S2. Construct a continuous time factor map based on the initial aligned data. Set state variables, including time-varying extrinsic parameters, time offset parameters, and environmental quality latent variables. Introduce the environmental quality observation sequence as an observation constraint into the continuous time factor map and apply time-varying prior constraints to the extrinsic parameters and time offset parameters. S3. Perform cross-modal matching between the 3D point cloud data and the auxiliary modal data. To address the matching error, a switchable constraint is set in the continuous time factor graph to introduce a matching reliability weight for the matching error, which is adaptively adjusted according to the environmental quality latent variable. A robust loss is applied to the matching error to suppress outlier errors, resulting in a weighted robust factor graph. S4. The weighted robust factor graph is iteratively optimized by jointly estimating extrinsic parameters, time offset parameters, environmental quality latent variables, and matching reliability weights. Based on the estimation results, time alignment and coordinate alignment updates are performed on the multimodal time series data to obtain updated aligned data. S5. Based on the updated aligned data and the estimation results, cross-modal cyclic consistency verification is performed to obtain cyclic consistency residuals. Combining the cyclic consistency residuals, environmental quality latent variables, and matching reliability weights, an adversarial risk score is generated, and the regulatory results are output according to preset rules.
[0009] Optionally, S1 includes: For each frame of the three-dimensional point cloud data, a first timestamp and a first sensor identifier are recorded, and for each frame of the auxiliary modal data, a second timestamp and a second sensor identifier are recorded, forming identified multimodal data using the first timestamp, the first sensor identifier, the second timestamp, and the second sensor identifier. Within a preset sliding time window, based on the first timestamp of each frame of point cloud data, at least one frame of auxiliary modal data whose second timestamp falls within the preset sliding time window is selected from the auxiliary modal data, and the selected auxiliary modal data is aligned with the corresponding point cloud data in the first time. The auxiliary modal data that has completed the first time alignment is transformed according to the preset initial extrinsic parameter matrix to convert the auxiliary modal data to the coordinate system of the three-dimensional point cloud data, and the first alignment data is generated by the converted auxiliary modal data and the corresponding point cloud data. An environmental quality observation sequence is calculated based on the first aligned data, wherein the point cloud density index is obtained by counting the number of points within a preset spatial range and dividing by the volume of the preset spatial range, the point cloud intensity distribution index is obtained by statistically analyzing the point cloud intensity values, and the auxiliary mode signal-to-noise ratio index is obtained by statistically analyzing the signal strength and noise strength of the auxiliary mode data and calculating the ratio.
[0010] Optionally, S2 includes: Based on the first alignment data, a time series within a preset sliding time window is determined, and state variable nodes are set for each moment corresponding to the time series in the continuous time factor diagram. The state variable nodes include extrinsic parameter matrix nodes, time offset nodes, and environmental quality latent variable nodes. The extrinsic parameter matrix is used to represent the rigid body transformation from the auxiliary modal data coordinate system to the three-dimensional point cloud data coordinate system, and the time offset is used to represent the time deviation between the three-dimensional point cloud data and the auxiliary modal data. The extrinsic parameter matrix and the time offset are parameterized by an interpolation function based on multiple control times. This is achieved by setting multiple control times for the extrinsic parameter matrix control and the time offset control, and interpolating the extrinsic parameter matrix and the time offset at any time according to the control times, so that the extrinsic parameter matrix and the time offset change slowly with time. The interpolation function is any one of B-spline interpolation function, piecewise linear interpolation function, or polynomial interpolation function; In the continuous time factor diagram, a priori constraint factors are set for the extrinsic parameter matrix control quantity and time offset control quantity corresponding to adjacent control moments, so as to constrain the rate of change of the extrinsic parameter matrix and the time offset to meet the preset change range. In the continuous time factor graph, observation factors corresponding to the environmental quality observation sequence are set for the environmental quality latent variable nodes, so that the environmental quality latent variables and the environmental quality observation sequence are constrained to generate an initial graph; Furthermore, in the continuous time factor diagram, state transition constraint factors are set for the environmental quality latent variable nodes at adjacent times so that the environmental quality latent variables satisfy the preset time continuity, wherein the state transition constraint factor is a random walk constraint factor or a first-order Gauss-Markov constraint factor. Furthermore, the cross-modal cyclic consistency verification includes multi-loop cyclic consistency verification, wherein when there are at least two auxiliary modal data, at least one cyclic mapping link containing the three-dimensional point cloud data, the first auxiliary modal data and the second auxiliary modal data is constructed, and the cyclic consistency residual corresponding to each cyclic mapping link is calculated respectively, and the cyclic consistency residuals are fused to generate the adversarial risk score.
[0011] Optionally, S3 includes: In the initial diagram, a matching relationship is constructed between each pair of first-time aligned 3D point cloud data and auxiliary modal data in the first aligned data, and the matching error is calculated based on the matching relationship. The matching relationship is obtained by transforming the point cloud points in the 3D point cloud data to the observation space of the auxiliary modal data based on the current estimated value of the extrinsic parameter matrix and the preset sensor model to generate predicted observations. In the auxiliary modal data, candidate actual observations that satisfy a preset correlation threshold with the predicted observations are identified, and a correspondence is established between the actual observations that minimize the preset cost function selected from the candidate actual observations and the predicted observations. The matching error is the difference vector or distance metric between the predicted observation and the actual observation; The initial value of the current estimate of the extrinsic parameter matrix is a preset initial extrinsic parameter matrix; when the auxiliary modal data is image data, the predicted observation includes at least one of pixel coordinates obtained by projecting point cloud points through camera intrinsic parameters and predicted depth; the actual observation includes at least one of image edge features, semantic segmentation results, depth map and optical flow results; the matching error includes at least one of the following: distance residual from pixel coordinates to edge features, semantic category inconsistency residual, depth residual between predicted depth and depth map, and optical flow residual between predicted optical flow and optical flow results. When the auxiliary modal data is millimeter-wave radar data, the predicted observations include at least one of the distance, azimuth, elevation, and radial velocity of the point cloud points in the millimeter-wave radar coordinate system, and the actual observations include at least one of the distance, azimuth, elevation, and radial velocity of the target detected by the millimeter-wave radar. The matching error is the difference vector between the predicted observations and the actual observations in the measurement space. When the auxiliary modal data is inertial measurement data, inertial pre-integration is performed based on the inertial measurement data to obtain the expected pose increment, and point cloud registration is performed based on the three-dimensional point cloud data to obtain the observed pose increment. The matching error is the pose increment residual between the expected pose increment and the observed pose increment. A switchable constraint factor is set in the initial map for the matching error, and a matching reliability weight is introduced for each matching error by the switchable constraint factor, so that the matching reliability weight is used to adjust the constraint strength corresponding to the matching error. The switchable constraint factor includes a switch variable node, and the switch variable is a continuous variable or a binary variable with a value range between 0 and 1, and a priori constraint factor is set for the switch variable. The matching reliability weight is set to be adaptively adjusted as the environmental quality latent variable changes, wherein the environmental quality latent variable represents the environmental quality level corresponding to the environmental quality observation sequence, and the adaptive adjustment includes: reducing the matching reliability weight when the environmental quality level represented by the environmental quality latent variable decreases, and increasing the matching reliability weight when the environmental quality level represented by the environmental quality latent variable increases. A hybrid robust loss function is set for the matching error, and the hybrid robust loss function is combined with the matching reliability weight to weight the matching error, so as to suppress the impact of outlier matching error on the initial graph and generate a weighted robust graph. The hybrid robust loss function is a weighted combination of Huber loss and Cauchy loss, or a weighted combination of Gaussian kernel loss and Laplace kernel loss.
[0012] Optionally, S4 includes: An optimization objective function is constructed based on the weighted robust graph. The optimization objective function is the total cost of each matching error in the weighted robust graph after being processed by the hybrid robust loss function and weighted by the matching reliability weight. The optimization objective function is minimized using an iterative optimization method. In each iteration, the increment corresponding to the optimization objective function is calculated based on the extrinsic parameter matrix, time offset, environmental quality latent variable, and matching reliability weight of the current iteration. The extrinsic parameter matrix, time offset, environmental quality latent variable, and matching reliability weight are then updated based on the increment until the preset convergence condition is met, and the optimization result is obtained. Based on the optimization results, a second time alignment is performed on the labeled multimodal data, including: correcting the timestamp of the auxiliary modal data according to the time offset, and re-determining the auxiliary modal data corresponding to each 3D point cloud data within the preset sliding time window; Based on the optimization results, a second coordinate alignment is performed on the labeled multimodal data, including: transforming the auxiliary modal data aligned by the second time to the coordinate system of the three-dimensional point cloud data according to the extrinsic parameter matrix, and generating the second aligned multimodal data; Furthermore, the iterative optimization method includes sliding window edge-out processing, in which state variable nodes that exceed the preset sliding time window are edge-out to generate prior factors, and the prior factors are introduced into the weighted robust graph within the retention window to achieve online real-time optimization.
[0013] Optionally, S5 includes: A mapping relationship between three-dimensional point cloud data and auxiliary modal data is established based on a preset sensor model. The coordinate transformation of the three-dimensional point cloud data in the second aligned multimodal data is performed based on the extrinsic parameter matrix so that the three-dimensional point cloud data is mapped to the auxiliary modal coordinate system to generate the first intermediate data. Based on the preset sensor model, an inverse mapping relationship from the auxiliary modal coordinate system to the point cloud coordinate system is established, and the first intermediate data is inversely mapped back to the point cloud coordinate system to generate the second intermediate data. The second intermediate data and the three-dimensional point cloud data are compared to generate a cyclically consistent residual. The difference calculation includes statistically analyzing the point deviations between the second intermediate data and the three-dimensional point cloud data within a preset spatial range to obtain a residual statistic of the cyclically consistent residual. This residual statistic includes at least one of the median, truncated mean, root mean square error, and quantile of the point deviations. An adversarial risk score is generated based on the residual statistic, the environmental quality latent variable, and the matching reliability weight. The environmental compensation coefficient is determined using the environmental quality latent variable, and the residual statistic is normalized to increase the residual tolerance when environmental quality decreases and decrease it when environmental quality improves. The normalized residual statistic is then weighted by the matching reliability weight to obtain the adversarial risk score. The adversarial risk score is mapped to a regulatory result according to a preset risk grading rule, and the regulatory result is output. The preset risk grading rule includes one of the following: comparing the adversarial risk score with at least two preset thresholds to output low risk, medium risk, or high risk, and outputting an adversarial sample alarm information when the risk is high. Based on the adversarial risk score, a cumulative statistic is constructed, and when the cumulative statistic exceeds a preset alarm threshold, the adversarial sample alarm information is output.
[0014] On the other hand, the present invention also provides a multimodal 3D point cloud adversarial example intelligent monitoring system, comprising: The acquisition and alignment module is used to acquire three-dimensional point cloud data and at least one auxiliary modal data, and to perform initial time alignment and initial coordinate alignment based on initial extrinsic parameters within a preset time window to obtain initial alignment data and environmental quality observation sequence. The factor graph construction module is used to construct a continuous-time factor graph based on the initial aligned data, set extrinsic parameters, time offset parameters and environmental quality latent variables, and introduce observation constraints of the environmental quality observation sequence as well as time-varying prior constraints of extrinsic parameters and time offset parameters. The weighted robust modeling module is used to set switchable constraints for cross-modal matching errors to introduce matching reliability weights, and to make the matching reliability weights adaptively adjust with environmental quality latent variables. At the same time, a robust loss is applied to the matching error to obtain a weighted robust factor map. The joint optimization and update module is used to iteratively optimize the weighted robust factor graph, jointly estimate the extrinsic parameters, time offset parameters, environmental quality latent variables and matching reliability weights, and update the time alignment and coordinate alignment of the multimodal data accordingly to obtain updated aligned data. The consistency verification output module is used to perform cross-modal cyclic consistency verification based on the updated alignment data to obtain cyclic consistency residuals, and combine the cyclic consistency residuals, environmental quality latent variables and matching reliability weights to generate adversarial risk scores, and output regulatory results according to preset rules.
[0015] The beneficial effects of this invention are: 1. By constructing a continuous time factor graph containing slowly varying extrinsic parameters and time offset parameters, and iteratively optimizing it within a sliding time window, online estimation and dynamic correction of multi-sensor extrinsic parameter drift and time delay asynchrony can be achieved, reducing the accumulation of cross-modal registration errors and enabling regulatory judgments to be based on more accurate spatiotemporal alignment.
[0016] 2. By introducing latent environmental quality variables and constraining them with observations such as point cloud density, intensity distribution, and auxiliary mode signal-to-noise ratio, the matching reliability weight and residual tolerance are adaptively adjusted according to the degree of environmental degradation. This effectively suppresses false alarms in natural degradation scenarios such as rain, fog, dust, occlusion, and sparseness, and improves the sensitivity to subtle adversarial disturbances and reduces false alarms when the environment is good.
[0017] 3. For cross-modal matching errors, a weighted robust model is adopted using switchable constraints and hybrid robust loss. Cross-modal cyclic consistency verification is performed on the updated aligned data to generate an adversarial risk score. This can suppress the impact of outlier matching and abnormal perturbations on estimation and judgment, and improve the stable monitoring capability of adversarial samples with varied and covert attack patterns. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows the intelligent surveillance method for adversarial examples based on multimodal 3D point clouds proposed in this invention. Figure 2 This is a flowchart of step S3 of the present invention. Detailed Implementation
[0019] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0020] refer to Figure 1-2 A multimodal 3D point cloud adversarial example intelligent surveillance method, including: S1. Collect multimodal time series data of the monitored object, including 3D point cloud data and at least one auxiliary modal data. Associate time information and sensor identifiers between the 3D point cloud data and the auxiliary modal data. Perform initial time alignment and initial coordinate alignment based on initial extrinsic parameters within a preset time window to obtain initial aligned data and an environmental quality observation sequence characterizing the degree of environmental degradation. S2. Construct a continuous time factor map based on the initial aligned data. Set state variables, including time-varying extrinsic parameters, time offset parameters, and environmental quality latent variables. Introduce the environmental quality observation sequence as an observation constraint into the continuous time factor map and apply time-varying prior constraints to the extrinsic parameters and time offset parameters. S3. Perform cross-modal matching between the 3D point cloud data and the auxiliary modal data. To address the matching error, a switchable constraint is set in the continuous time factor graph to introduce a matching reliability weight for the matching error, which is adaptively adjusted according to the environmental quality latent variable. A robust loss is applied to the matching error to suppress outlier errors, resulting in a weighted robust factor graph. S4. The weighted robust factor graph is iteratively optimized by jointly estimating extrinsic parameters, time offset parameters, environmental quality latent variables, and matching reliability weights. Based on the estimation results, time alignment and coordinate alignment updates are performed on the multimodal time series data to obtain updated aligned data. S5. Based on the updated aligned data and the estimation results, cross-modal cyclic consistency verification is performed to obtain cyclic consistency residuals. Combining the cyclic consistency residuals, environmental quality latent variables, and matching reliability weights, an adversarial risk score is generated, and the regulatory results are output according to preset rules.
[0021] In this specific embodiment, S1 includes: The monitoring object is equipped with a three-dimensional lidar as the first sensor and a millimeter-wave radar as the second sensor. The system receives time-series data at a fixed sampling frequency and writes time information and sensor identification into each frame of data. No. Frame 3D point cloud data is denoted as Indicates the first The first frame The three-dimensional coordinate vector of a point in the point cloud coordinate system This represents the point cloud intensity value at that point. Indicates the first Count the points in the point cloud of a frame and record the first timestamp for that frame. With the first sensor identifier ; No. Frame millimeter-wave radar data is denoted as: ; in For distance measurement, For azimuth measurement, For pitch angle measurement, For radial velocity measurement, For signal strength measurement, For noise intensity measurement, Detect the number of targets in this frame and record a second timestamp for this frame. With the second sensor identifier This results in a labeled multimodal data stream; The system is set to preset the sliding time window length. And set the window update step size Each time the window is updated, the millimeter-wave radar frames are aligned in real time based on the point cloud frames falling into the current window. That is, for each frame of point cloud... Determine the set of candidate radar frames When the candidate set contains multiple frames, select the appropriate frame. The smallest unique radar frame With point cloud frames Pairing: When the candidate set is empty, the point cloud frame is marked as "no corresponding auxiliary frame" and removed from the subsequent processing chain of this window to avoid introducing uncertain associations; After completing the initial alignment, the system will, based on the preset sensor model, Each test Converting to three-dimensional points in millimeter-wave radar coordinate system using deterministic trigonometric mapping from polar coordinates to rectangular coordinates. The angles are measured in radians and the distances in meters. The two-dimensional point is then linked to its corresponding... and Spatial observation as an auxiliary mode; The system presets the initial extrinsic parameter matrix. This is used to transform spatial observations in the millimeter-wave radar coordinate system to the point cloud coordinate system, where By rotation matrix With translation vector The system is configured and the calibration results are written during the system deployment phase. The system performs calibration on each... Perform a rigid body coordinate transformation to obtain the corresponding points in the point cloud coordinate system. and will and Together they form the first alignment data; The system then sets a preset spatial range for quality assessment in the point cloud coordinate system. Align the cuboid region with axis And its volume is calculated in meters. And calculate the environmental quality observation sequence based on the first aligned data. Among them, environmental quality observation Point cloud density index Point cloud intensity distribution index With auxiliary mode signal-to-noise ratio Composition, point cloud density index according to falling The ratio of the number of points to the volume is obtained and recorded. The number of points inside is Point cloud intensity distribution index Press Fall In Set of point strengths within Construct a fixed number of boxes The normalized histogram vector and the intensity range is truncated to To ensure cross-frame comparability, the auxiliary modal signal-to-noise ratio (SNR) is calculated as the ratio of the sum of signal strength to the sum of noise strength detected within a radar frame. The calculation of these two ratios is uniformly expressed by the following formula: ; in For point cloud frame indexing, For radar detection index, To be with point cloud frames Complete the radar frame index mapping with first-time alignment. Used for characterization Inner point cloud density Used for characterization Statistical distribution pattern of reflection intensity of internal point clouds, The signal-to-noise level of millimeter-wave radar observations was used to characterize the data, ultimately yielding the first aligned data and environmental quality observation sequence for subsequent factor graph modeling.
[0022] In this specific embodiment, S2 includes: The system is in the current preset sliding time window Internal reading of the first aligned data and environmental quality observation sequence ,in This is the start time of the preset sliding time window. The time length of the preset sliding time window. In order to be with the first Environmental quality observations corresponding to frame point clouds Point cloud density index Let be the normalized histogram vector of the point cloud intensity distribution index and the number of bins. The signal-to-noise ratio (SNR) of millimeter-wave radar. This represents the number of point cloud frames involved in mapping within that time window. The system uses point cloud frame timestamp sequences within this time window. As a set of observation times for continuous-time modeling, For the first The first timestamp of the frame point cloud; The system sets the auxiliary modes to the first auxiliary mode millimeter-wave radar data and the second auxiliary mode camera data, and displays them in the continuous time factor graph. The middle section sets extrinsic parameters and time offset parameters for two auxiliary modes, and sets implicit environmental quality variables for environmental degradation modeling, specifically in the control time set. Establish external parameter control nodes and and time offset control node and And collect at the point cloud observation time Establish environmental quality hidden variable nodes ,in and Let represent the extrinsic Lie algebra vectors of the millimeter-wave radar and camera relative to the point cloud coordinate system, respectively. The first three dimensions of the Lie algebra vector are rotational vector components in rad, and the last three dimensions are translational components in meters. and These represent the time offsets between point cloud time and millimeter-wave radar time and camera time, respectively, with units of 1. Indicates the level of environmental quality and The smaller the value, the more severe the environmental degradation. For control node index; System settings control time interval , and according to Generate a set of control times and let The extrinsic control node of the millimeter-wave radar is initialized to the initial extrinsic matrix in S1. The Lie algebra vector obtained through logarithmic mapping is used to initialize the camera's extrinsic control nodes with the environmental quality observation scalars calculated during the deployment phase. And let The environmental quality observation scalar mentioned therein Through the By upper limit points Perform linear normalization and truncate to ,right Take the natural logarithm and then the upper limit Normalize and truncate to ,right Calculate Shannon entropy and according to After normalization, the inverse is taken to obtain three sub-ratings, which are then weighted by a fixed method. Linear fusion is thus guaranteed Comparable across different frames; To achieve gradual changes in extrinsic parameters and time offset parameters over time, the system uses a factor plot. Piecewise linear interpolation is used to continuously parameterize the extrinsic parameters and time migration at any point cloud observation time. A change a priori constraint factor is introduced between adjacent control times to constrain the rate of change of the extrinsic parameters and time migration to meet a preset range. The single-axis standard deviation of the extrinsic parameter change a priori for the rotation component is set as... Let the uniaxial standard deviation of the translation component be set as Let the standard deviation of the time offset change be a priori. This allows external participation delays to drift slowly during online operation and prevents cross-modal matching constraints from being dominated by transient anomalies. The system incorporates environmental quality observation sequences as observation constraints into the factor graph. Specifically, this involves establishing environmental observation factors for each frame of point cloud and using the residuals... Construct observation constraints and set the standard deviation of observation noise as This allows the latent variables of environmental quality to be influenced by both observed data and prior constraints during optimization; To ensure the temporal continuity of environmental quality latent variables, the system sets a first-order Gaussian Markov state transition constraint factor between environmental quality latent variable nodes at adjacent time points, and uses the following determination form to uniformly define continuous-time interpolation and environmental state transition: , , , ; in aux,cam Indicates the auxiliary mode type, For a moment The extrinsic Lie algebra vector at that point, For a moment The time offset at that point. To meet Control time index, The interpolation scaling factor and The environmental state transition attenuation coefficient. The time interval between adjacent point cloud frames. It is a time constant. The noise in the environmental state transition process follows a zero-mean Gaussian distribution. and take To limit the magnitude of abrupt changes in environmental quality between adjacent frames; After the factor graph is constructed, the system writes the cyclic mapping link configuration for the multi-loop cyclic consistency verification in the subsequent step S5. The cyclic mapping link includes the first closed-loop link of point cloud-camera-point cloud, the second closed-loop link of point cloud-millimeter-wave radar-point cloud, and the third closed-loop link of point cloud-camera-millimeter-wave radar-point cloud. The sensor model parameters required for each link are bound to the current continuous-time extrinsic parameterization interface in the configuration, resulting in the initial continuous-time factor graph containing slowly varying extrinsic parameters, time offset parameters, and environmental quality latent variables, and possessing multi-loop closed-loop verification topology information.
[0023] In this specific embodiment, S3 includes: The system in the initial continuous-time factor plot Above, for each frame of point cloud The millimeter-wave radar frame that was first aligned with S1 and camera frames Establish cross-modal matching relationships and calculate matching errors, where Indexing point cloud frames The corresponding millimeter-wave radar frame index mapping, Indexing point cloud frames The corresponding camera frame index mapping, For resolution An RGB image with the origin of the pixel coordinate system located at the top left corner, the horizontal axis to the right, and the vertical axis downward; The system reads time points from the continuous-time model. The extrinsic Lie algebra vector at the location and And obtain the homogeneous transformation matrix through exponential mapping respectively. and ,in This represents the rigid body transformation of millimeter-wave radar to the point cloud coordinate system. This represents the rigid body transformation from the camera to the point cloud coordinate system; For millimeter-wave radar matching, the system uses point cloud points Input and select only those falling within the preset space range Points within the range participate in the matching, pass The inverse transformation is used to convert the coordinates to the millimeter-wave radar coordinate system to obtain the predicted three-dimensional points, and then the predicted measurements are calculated according to the deterministic radar measurement geometric model. ,in To predict distance, To predict azimuth, To predict the pitch angle; The system in actual radar frames Candidate filtering is performed using an association threshold, which is set as follows: ; in For the first In the first frame of radar The range, azimuth, pitch, and radial velocity of each detected target. The radial velocity of this point is predicted by the displacement difference between adjacent frames of the point cloud. When the candidate set is not empty, the system uses a cost function to select the unique corresponding actual observation. The cost function is defined as a weighted sum of squares over each measurement dimension, with the weights set as distance weights. Angle weight Speed weight And select the detection target with the lowest cost. With point Establish a matching relationship; When the candidate set is empty, the system does not generate the millimeter-wave radar matching factor for that point to avoid introducing invalid constraints. Based on this, the system constructs the millimeter-wave radar matching error vector: ; The superscript aux indicates the millimeter-wave radar mode. Match the error index of the millimeter-wave radar generated within this frame with the point index. One-to-one correspondence; For camera matching, the system sets up a pinhole camera model and a camera intrinsic parameter matrix. Fixed as , ,in and Focal length parameter per pixel and Using the main point coordinates, the system does not use a distortion model and fixes the distortion coefficient to zero to ensure the determination of the projection model; The system will point cloud points pass The inverse transform is used to transform the data to the camera coordinate system to obtain the predicted 3D points, and then... Complete perspective projection to obtain predicted pixel coordinates ,when When the system falls within the image boundary, Binary edge map obtained by running Canny edge extraction. The high threshold is set to 150 and the low threshold is set to 50. The system uses... Centered on the radius Within a pixel search window, a set of candidate edge pixels is collected, and the edge pixel with the smallest Euclidean distance is selected. and Establish a matching relationship; when the candidate set is empty, the system does not generate a camera matching factor for that point. The system constructs the camera matching error vector accordingly: ; The superscript cam indicates the camera mode; To suppress occlusion, incorrect associations, and combat outlier matching errors induced by perturbations, the system addresses each matching error... In factor graph The text introduces switchable constraint factors and adds switch variable nodes. ,in For modal identification, For a continuous variable whose value ranges from 0 to 1 and is initially set to 1, the system assigns each... Set a prior constraint factor with a mean of 1 and a standard deviation of 0.2 to constrain its deviation, and update it after each iteration of linearization. Perform interval projection to ensure that its value falls within the range ; The system sets the matching reliability weights to adaptively adjust based on the environmental quality latent variables, where the environmental quality latent variable nodes... From step S2 and with point cloud frame index Correspondingly, the system will A smaller value is interpreted as worse environmental quality, and the system accordingly lowers the matching reliability weight. A larger value indicates better environmental quality, and the matching reliability weight is increased accordingly. The system applies a hybrid robust loss function to each matching error simultaneously. The hybrid robust loss function consists of Huber loss and Cauchy loss with fixed coefficients. The linear combination is used, where the threshold parameter of the Huber loss is fixed at 1.5, the scaling parameter of the Cauchy loss is fixed at 2.5, and the input scalar is the norm of the matching error. The hybrid robust loss function is combined with the matching reliability weight to form a weighted robust factor, the cost of which is defined as: ; in For modality In frame Upper The cost of each matching factor This is the matching reliability weight of the matching factor. Let the hybrid robust loss function be... For a norm 2 operator, This is the corresponding matching error vector. For switch variables, As a latent variable for environmental quality, As the lowest reliability weight, With the highest reliability weight, the system will assign all The corresponding matching factor and its on / off prior factor are added to the factor graph. This generates a weighted robust factor graph.
[0024] In this specific embodiment, S4 includes: The system supports weighted robust factor graphs An optimization objective is constructed and online iterative optimization with a sliding window is performed to jointly estimate extrinsic parameters, time offset parameters, environmental quality latent variables, and matching reliability weights. The system will display the current sliding time window. The parameter vector consists of all the states to be estimated. ,in Includes millimeter-wave radar extrinsic control node Camera extrinsic control node Millimeter-wave radar time offset control node Camera time offset control node Environmental quality hidden variable nodes and all switch variable nodes ,in For control time index and control time satisfies To control the time interval, For point cloud frame indexing, The number of point cloud frames within the window. For modal identification, For matching error index; The system defines the factor set as follows: It includes cross-modal matching factors, prior factors for extrinsic parameter changes, prior factors for time offset changes, environmental quality observation factors, environmental quality state transition factors, and prior factors for switch variables, and constructs the objective function as the sum of the costs of all factors: ; in To optimize the objective function, Let be the vector of parameters to be estimated. For the factor set, For factor index, As a factor In parameters The cost at the location, and the cross-modal matching factor The weighted robust form is used and already includes matching error. Switch variables Latent variables of environmental quality and hybrid robust loss The combined effect; The system employs Levenberg-Marquardt iterative minimization. The maximum number of iterations was set to 10, and the initial damping coefficient was set to... In each iteration, all factors are estimated in the current estimate. Perform linearization at the point to form the incremental equation and solve it to obtain the parameter increment. Subsequently, the external parameter Lie algebra control nodes are updated according to the exponential mapping, that is, each Use its corresponding After left-multiplying to update to SE(3), take the logarithm to map back to the Lie algebra to maintain the rigid body transformation constraints. Update the time offset control nodes by addition. The latent variable nodes for environmental quality are updated additively and then projected onto the interval after the update. To maintain the level of environmental quality as defined Update the switch variable nodes by addition and then project them onto the interval. To maintain the definition of a continuous switch with switchable constraints, i.e. ,in For interval projection operators; The system determines to stop iteration based on two types of acceptance conditions simultaneously, namely when... Or when the relative decrease in the objective function is less than Stop and output the current estimate as the optimization result; Based on the optimization results, the system performs a second time alignment on the labeled multimodal data, and the system shifts the time offset. Interpreted as the difference between point cloud time and auxiliary mode time. And at each point cloud moment The piecewise linear interpolation of S2 is used to obtain Then, the timestamp of each frame of the auxiliary modality is corrected to obtain the equivalent time of the auxiliary modality on the point cloud time axis: ; in To assist in the original timestamp of the modality, To correct the timestamp and minimize it within a preset sliding time window. To achieve reassociation, a unique auxiliary modal frame is redefined for each point cloud frame, using the criterion of reassignment. The system further performs a second coordinate alignment based on the interpolation results of the optimized extrinsic parameters at the point cloud time point, that is, it applies a second coordinate alignment to the millimeter-wave radar. obtained by exponential mapping The second-time-aligned radar spatial observations are transformed to a point cloud coordinate system to obtain the second-aligned radar data. The camera is then subjected to... obtained by exponential mapping As a coordinate transformation interface between point cloud and camera to form a second aligned camera data index relationship, thereby generating second aligned multimodal data; To achieve online real-time optimization of sliding window edge processing, the system performs edge processing on windows starting from... Updated to and At that time, all state variable nodes with timestamps earlier than the start time of the new window and their associated factors are moved from... The variables are removed from the graph, and before removal, the joint information matrix and information vector of the removed and retained variables are constructed using the current iteration endpoint as the linearization point. The removed variables are eliminated by Schur complement to obtain prior factors that only apply to the retained variables. These prior factors are added as new prior constraints to the updated in-window factor graph to maintain the continuity of historical information constraints on the current estimate and to avoid the computational load increasing over time.
[0025] In this specific embodiment, S5 includes: The system uses second-aligned multimodal data to process each frame of point cloud. Perform cross-modal cyclic consistency checks and generate adversarial risk scores. ; The system at point cloud moment Read the camera extrinsic transformation matrix obtained by continuous-time interpolation Transformation matrix of millimeter-wave radar extrinsic parameters ,in This represents a rigid body transformation from the camera coordinate system to the point cloud coordinate system. This represents the rigid body transformation from the millimeter-wave radar coordinate system to the point cloud coordinate system, and reads the latent environmental quality variables for that frame. ,in The smaller the value, the more severe the environmental degradation. The system uses a preset spatial range Filtering point clouds Falling into point Participating in the cycle-consistency computation, among which For the first Frame number The three-dimensional coordinate vector of a point in the point cloud coordinate system Point index; The system first executes a cyclic mapping link from point cloud to camera to point cloud to generate the first cyclically consistent residual, specifically by mapping each... pass The inverse transformation is used to convert the coordinates to the camera coordinate system to obtain the predicted camera points, and the pinhole camera model and camera intrinsic parameter matrix are used. Complete the forward mapping to generate the first intermediate data ,in Fixed as and Focal length parameter per pixel and Principal point coordinate parameters, and in the camera image The edge map was extracted using the Canny operator. The high threshold is 150 and the low threshold is 50. The system uses the predicted pixel coordinates of each point as the predicted observation and sets the radius... Within the pixel search window, the edge pixel with the smallest Euclidean distance is selected as the actual observation, thus making the first intermediate data... The set is defined as "actual edge pixel coordinates and predicted depth" and each point involved in the calculation corresponds to only one edge pixel. When there is no edge pixel in the window, the point is removed and not included in the residual statistics. The system then uses the first intermediate data Perform reverse mapping to generate second intermediate data The inverse mapping relationship adopts a deterministic model of "pixel backprojection + depth recovery", that is, the coordinates of each actual edge pixel are mapped to its corresponding predicted depth. Back-projection restores the 3D points to the camera coordinate system, and then through... The reconstructed points are obtained by forward transformation back to the point cloud coordinate system. ; The system uses point deviation As a point-level cycle consistency error and The internal system performs statistical analysis on all point-level cyclic consistency errors to form the cyclic consistency residual statistics for the camera link. ,in The median of point-level cyclic consistency error is used to suppress a small number of extreme outliers; The system further executes a cyclic mapping link from point cloud to millimeter-wave radar to point cloud to generate a second cyclically consistent residual, specifically by mapping each... pass The inverse transformation is used to convert to the millimeter-wave radar coordinate system and the predicted measurement is calculated according to the radar measurement geometry model, followed by the radar frame after second time alignment. The first intermediate data is selected by using the same correlation threshold and cost function as S3 to select the unique corresponding actual radar detection measurement. The actual measurement is then restored to a three-dimensional point in the radar coordinate system using the inverse geometric model of the measurement, and then... Reconstructed points are obtained by transforming back to the point cloud coordinate system. The system also uses The median of the point-level cyclic consistency error is used to form the cyclic consistency residual statistic for the millimeter-wave radar link. ; The system aggregates and matches reliability weights from the optimization results of S3 and S4 to form a frame-level reliability weighting coefficient. and ,in The matching reliability weights for all camera matching factors in this frame. arithmetic mean and For camera matching factor index, The matching reliability weight for all radar matching factors in this frame. arithmetic mean and For the radar matching factor index (in this embodiment, each radar matching factor is constructed from a radar matching error, therefore the radar matching factor index and the radar matching error index correspond one-to-one, both using...), (Indicated), the system will and It is interpreted as a measure of "reliability of cycle-consistent residuals" and used to weight residual statistics; The system is based on environmental quality hidden variables Determine the environmental compensation coefficient The residual statistics were normalized, among which... Follow The system settings are adjusted to increase residual tolerance when environmental quality deteriorates. and To limit the scope of compensation; The system fuses the normalized residuals from the camera link and the millimeter-wave radar link to obtain an adversarial risk score. The regulatory results are output using a two-threshold hierarchical approach, and their calculation and symbol definition are shown in the following formula: ; in For the first Frame adversarial risk score, For camera link fusion weights, For millimeter-wave radar link fusion weights, For the first Frame camera matching reliability weight mean, For the first Frame millimeter-wave radar matching reliability weight mean, The median of the camera link cyclic consistency residual statistic is given in units of 1. The median of the cyclic consistency residual statistics for millimeter-wave radar links, in units of 1. This is the environmental compensation coefficient. As a latent variable for environmental quality, For compensation strength parameters; The system will With preset threshold and Compare and output regulatory results, where when Output low risk, when Output risk, when Output high-risk information and adversarial example warnings.
[0026] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0027] This invention combines continuous-time factor map online external participation delay estimation with cross-modal cyclic consistency verification: On the one hand, based on the matching error between point cloud and auxiliary mode, joint optimization is performed on the factor map, so that the external parameters and time offset parameters are continuously updated within a preset time window, reducing cross-modal alignment errors caused by asynchrony, coordinate system inconsistency, and external parameter drift from the source; on the other hand, cross-modal cyclic consistency verification is performed after the alignment result is corrected, using the cyclic consistency residual to characterize the geometric and observation closed-loop deviations between multiple modes, and combining environmental quality latent variables and matching reliability weights to form an adversarial risk score, thereby improving the distinguishability of anomaly sources and reducing false alarms and missed alarms when complex environmental degradation and adversarial disturbances coexist, achieving a stable and implementable intelligent supervision effect.
[0028] To address the technical issues in this case, this invention makes scenario-oriented improvements to the algorithm structure: First, it models the external participation delay as a continuously changing state, and through interpolation parameterization and change a priori constraints, it adapts to the dynamic characteristics of slow drift and delay changes in actual operation, thereby improving the stability of online estimation. Second, it introduces switchable constraints and superimposes hybrid robust loss in cross-modal constraints, so that outlier matching caused by erroneous associations, occlusion, and abnormal residuals induced by attacks do not dominate optimization, thereby improving robustness. Third, it introduces environmental quality latent variables constrained by observations such as point cloud density, intensity distribution, and auxiliary modal signal-to-noise ratio, which drive the adaptive adjustment of matching reliability weights and residual tolerance, so that the system reduces the probability of misjudging noise residuals when the environment degrades, and improves the ability to detect subtle adversarial disturbances when the environment is good, further enhancing the consistency and stability of the regulatory effect.
Claims
1. A multimodal 3D point cloud adversarial example intelligent surveillance method, characterized in that, include: S1. Collect multimodal time series data of the monitored object, including three-dimensional point cloud data and at least one auxiliary modal data, associate time information and sensor identifiers between the three-dimensional point cloud data and the auxiliary modal data, perform initial time alignment and initial coordinate alignment based on initial external parameters within a preset time window, and obtain initial aligned data and environmental quality observation sequence characterizing the degree of environmental degradation. S2. Construct a continuous time factor graph based on the initial aligned data, set state variables, including time-varying external parameters, time offset parameters, and environmental quality latent variables, introduce the environmental quality observation sequence as an observation constraint into the continuous time factor graph, and apply time-varying prior constraints to the external parameters and time offset parameters. S3. To address the cross-modal matching error between 3D point cloud data and auxiliary modal data, switchable constraints are set in the continuous time factor map to introduce matching reliability weights to the matching error, and these weights are adaptively adjusted with the environmental quality latent variable. Robust loss is applied to the matching error to suppress outlier errors, resulting in a weighted robust factor map. S4. Iteratively optimize the weighted robust factor graph, jointly estimate the extrinsic parameters, time offset parameters, environmental quality latent variables and matching reliability weights, and update the multimodal time series data by time alignment and coordinate alignment based on the estimation results to obtain updated aligned data. S5. Perform cross-modal cyclic consistency verification based on the updated aligned data and estimation results to obtain cyclic consistency residuals. Combine the cyclic consistency residuals, environmental quality latent variables, and matching reliability weights to generate an adversarial risk score and output regulatory results according to preset rules.
2. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 1, characterized in that, S1 includes: For each frame of the three-dimensional point cloud data, a first timestamp and a first sensor identifier are recorded, and for each frame of the auxiliary modal data, a second timestamp and a second sensor identifier are recorded, forming identified multimodal data using the first timestamp, the first sensor identifier, the second timestamp, and the second sensor identifier. Within a preset sliding time window, based on the first timestamp of each frame of point cloud data, at least one frame of auxiliary modal data whose second timestamp falls within the preset sliding time window is selected from the auxiliary modal data, and the selected auxiliary modal data is aligned with the corresponding point cloud data in the first time. The auxiliary modal data that has completed the first time alignment is transformed according to the preset initial extrinsic parameter matrix to convert the auxiliary modal data to the coordinate system of the three-dimensional point cloud data, and the first alignment data is generated by the converted auxiliary modal data and the corresponding point cloud data. An environmental quality observation sequence is calculated based on the first aligned data. The point cloud density index is obtained by counting the number of points within a preset spatial range and dividing by the volume of the preset spatial range. The point cloud intensity distribution index is obtained by counting the point cloud intensity values. The auxiliary mode signal-to-noise ratio index is obtained by counting the signal intensity and noise intensity of the auxiliary mode data and calculating the ratio.
3. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 2, characterized in that, S2 include: Based on the first alignment data, a time series within a preset sliding time window is determined, and state variable nodes are set for each moment corresponding to the time series in the continuous time factor diagram. The state variable nodes include extrinsic parameter matrix nodes, time offset nodes, and environmental quality latent variable nodes. The extrinsic parameter matrix is used to represent the rigid body transformation from the auxiliary modal data coordinate system to the three-dimensional point cloud data coordinate system, and the time offset is used to represent the time deviation between the three-dimensional point cloud data and the auxiliary modal data. The extrinsic parameter matrix and the time offset are parameterized by an interpolation function based on multiple control times. This is achieved by setting multiple control times for the extrinsic parameter matrix control and the time offset control, and interpolating the extrinsic parameter matrix and the time offset at any time according to the control times, so that the extrinsic parameter matrix and the time offset change slowly with time. The interpolation function is any one of B-spline interpolation function, piecewise linear interpolation function, or polynomial interpolation function; In the continuous time factor diagram, a priori constraint factors are set for the extrinsic parameter matrix control quantity and time offset control quantity corresponding to adjacent control moments, so as to constrain the rate of change of the extrinsic parameter matrix and the time offset to meet the preset change range. In the continuous time factor graph, observation factors corresponding to the environmental quality observation sequence are set for the environmental quality latent variable nodes, so as to establish a constraint relationship between the environmental quality latent variables and the environmental quality observation sequence, and generate an initial graph.
4. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 3, characterized in that, S3 includes: In the initial image, a matching relationship is constructed between each pair of first-time aligned 3D point cloud data and auxiliary modal data in the first alignment data, and a matching error is calculated based on the matching relationship. The matching relationship is obtained by transforming the point cloud points in the 3D point cloud data to the observation space of the auxiliary modal data based on the current estimated value of the extrinsic parameter matrix and the preset sensor model to generate predicted observations. In the auxiliary modal data, candidate actual observations that satisfy a preset correlation threshold with the predicted observations are determined, and a correspondence is established between the actual observations that minimize the preset cost function selected from the candidate actual observations and the predicted observations. The matching error is the difference vector or distance metric between the predicted observation and the actual observation; The initial value of the current estimate of the extrinsic parameter matrix is a preset initial extrinsic parameter matrix; When the auxiliary modal data is image data, the predicted observation includes at least one of pixel coordinates obtained by projecting point cloud points through camera intrinsic parameters and predicted depth. The actual observation includes at least one of image edge features, semantic segmentation results, depth map and optical flow results. The matching error includes at least one of the following: distance residual from pixel coordinates to edge features, semantic category inconsistency residual, depth residual between predicted depth and depth map, and optical flow residual between predicted optical flow and optical flow results. When the auxiliary modal data is millimeter-wave radar data, the predicted observations include at least one of the distance, azimuth, elevation, and radial velocity of the point cloud points in the millimeter-wave radar coordinate system, and the actual observations include at least one of the distance, azimuth, elevation, and radial velocity of the target detected by the millimeter-wave radar. The matching error is the difference vector between the predicted observations and the actual observations in the measurement space. When the auxiliary modal data is inertial measurement data, inertial pre-integration is performed based on the inertial measurement data to obtain the expected pose increment, and point cloud registration is performed based on the three-dimensional point cloud data to obtain the observed pose increment. The matching error is the pose increment residual between the expected pose increment and the observed pose increment. A switchable constraint factor is set in the initial graph for the matching error, and a matching reliability weight is introduced for each matching error by the switchable constraint factor, so that the matching reliability weight is used to adjust the constraint strength corresponding to the matching error; The switchable constraint factor includes a switch variable node, wherein the switch variable is a continuous variable or a binary variable whose value ranges from 0 to 1, and a prior constraint factor is set for the switch variable. The matching reliability weight is set to be adaptively adjusted as the environmental quality latent variable changes, wherein the environmental quality latent variable represents the environmental quality level corresponding to the environmental quality observation sequence, and the adaptive adjustment includes: reducing the matching reliability weight when the environmental quality level represented by the environmental quality latent variable decreases, and increasing the matching reliability weight when the environmental quality level represented by the environmental quality latent variable increases. A hybrid robust loss function is set for the matching error, and the hybrid robust loss function is combined with the matching reliability weight to weight the matching error, so as to suppress the impact of outlier matching error on the initial graph and generate a weighted robust graph. The hybrid robust loss function is a weighted combination of Huber loss and Cauchy loss, or a weighted combination of Gaussian kernel loss and Laplace kernel loss.
5. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 4, characterized in that, S4 include: An optimization objective function is constructed based on the weighted robust graph. The optimization objective function is the total cost of each matching error in the weighted robust graph after being processed by the hybrid robust loss function and weighted by the matching reliability weight. The optimization objective function is minimized using an iterative optimization method. In each iteration, the increment corresponding to the optimization objective function is calculated based on the extrinsic parameter matrix, time offset, environmental quality latent variable, and matching reliability weight of the current iteration. The extrinsic parameter matrix, time offset, environmental quality latent variable, and matching reliability weight are then updated based on the increment until the preset convergence condition is met, and the optimization result is obtained. Based on the optimization results, a second time alignment is performed on the labeled multimodal data, including: correcting the timestamp of the auxiliary modal data according to the time offset, and re-determining the auxiliary modal data corresponding to each 3D point cloud data within the preset sliding time window; Based on the optimization results, a second coordinate alignment is performed on the labeled multimodal data, including: transforming the auxiliary modal data aligned by the second time to the coordinate system of the three-dimensional point cloud data according to the extrinsic parameter matrix, and generating second aligned multimodal data.
6. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 5, characterized in that, S5 includes: A mapping relationship between three-dimensional point cloud data and auxiliary modal data is established based on a preset sensor model. The coordinate transformation of the three-dimensional point cloud data in the second aligned multimodal data is performed based on the extrinsic parameter matrix so that the three-dimensional point cloud data is mapped to the auxiliary modal coordinate system to generate the first intermediate data. Based on the preset sensor model, an inverse mapping relationship from the auxiliary modal coordinate system to the point cloud coordinate system is established, and the first intermediate data is inversely mapped back to the point cloud coordinate system to generate the second intermediate data. The difference between the second intermediate data and the three-dimensional point cloud data is calculated to generate a cyclic consistency residual. The difference calculation includes statistically analyzing the point position deviations between the second intermediate data and the three-dimensional point cloud data within a preset spatial range to obtain the residual statistics of the cyclic consistency residual. The residual statistics include at least one of the point position deviation median, truncated mean, root mean square error, and quantile. A risk mitigation score is generated based on the residual statistics, the environmental quality latent variables, and the matching reliability weights. The environmental compensation coefficient is determined by the environmental quality latent variables, and the residual statistics are normalized to increase the residual tolerance when the environmental quality decreases and decrease the residual tolerance when the environmental quality increases. The normalized residual statistics are then weighted by the matching reliability weights to obtain the risk mitigation score. The adversarial risk score is mapped to a regulatory result according to a preset risk grading rule and the regulatory result is output. The preset risk grading rule includes one of the following: comparing the adversarial risk score with at least two preset thresholds to output low risk, medium risk or high risk, and outputting the adversarial sample alarm information when the risk is high. Based on the adversarial risk score, a cumulative statistic is constructed, and when the cumulative statistic exceeds a preset alarm threshold, the adversarial sample alarm information is output.
7. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 3, characterized in that, In the continuous time factor diagram, state transition constraint factors are set for environmental quality latent variable nodes at adjacent times so that the environmental quality latent variables satisfy the preset time continuity. The state transition constraint factors are random walk constraint factors or first-order Gauss-Markov constraint factors.
8. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 5, characterized in that, The iterative optimization method includes sliding window edge-out processing, wherein state variable nodes that exceed the preset sliding time window are edge-out to generate prior factors, and the prior factors are introduced into the weighted robust graph within the retention window to achieve online real-time optimization.
9. The intelligent surveillance method for adversarial examples based on multimodal 3D point clouds according to claim 7, characterized in that, The cross-modal cyclic consistency verification includes multi-loop cyclic consistency verification. When there are at least two auxiliary modal data, at least one cyclic mapping link containing the three-dimensional point cloud data, the first auxiliary modal data, and the second auxiliary modal data is constructed. The cyclic consistency residuals corresponding to each cyclic mapping link are calculated respectively, and the cyclic consistency residuals are fused to generate the adversarial risk score.
10. A multimodal 3D point cloud adversarial example intelligent monitoring system, used to execute the multimodal 3D point cloud adversarial example intelligent monitoring method according to any one of claims 1 to 9, characterized in that, include: The acquisition and alignment module is used to acquire three-dimensional point cloud data and at least one auxiliary modal data, and to perform initial time alignment and initial coordinate alignment based on initial extrinsic parameters within a preset time window to obtain initial alignment data and environmental quality observation sequence. The factor graph construction module is used to construct a continuous-time factor graph based on the initial aligned data, set extrinsic parameters, time offset parameters and environmental quality latent variables, and introduce observation constraints of the environmental quality observation sequence as well as time-varying prior constraints of extrinsic parameters and time offset parameters. The weighted robust modeling module is used to set switchable constraints for cross-modal matching errors to introduce matching reliability weights, and to make the matching reliability weights adaptively adjust with environmental quality latent variables. At the same time, a robust loss is applied to the matching error to obtain a weighted robust factor map. The joint optimization and update module is used to iteratively optimize the weighted robust factor graph, jointly estimate the extrinsic parameters, time offset parameters, environmental quality latent variables and matching reliability weights, and update the time alignment and coordinate alignment of the multimodal data accordingly to obtain updated aligned data. The consistency verification output module is used to perform cross-modal cyclic consistency verification based on the updated alignment data to obtain cyclic consistency residuals, and combine the cyclic consistency residuals, environmental quality latent variables and matching reliability weights to generate adversarial risk scores, and output regulatory results according to preset rules.