Method and system for safety assessment of multi-modal unmanned system based on neural radiance field

By using a multimodal unmanned system security assessment method based on neural radiation fields, adversarial examples are generated and input into the unmanned system, solving the problem of inaccurate assessment results in existing technologies and achieving a more efficient and robust security assessment.

CN121482562BActive Publication Date: 2026-08-25BEIJING INST OF SPACECRAFT SYST ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511580739.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-08-25
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Existing safety assessment methods for unmanned systems suffer from problems such as one-sided assessment dimensions, isolated models, and static and unrealistic assessment scenarios when facing the complex and dynamic physical world, which leads to doubts about the validity and credibility of the assessment results.

Method used

A multimodal unmanned system safety assessment method based on neural radiation fields is adopted. By collecting multimodal datasets, constructing neural radiation fields and training adversarial neural radiation fields, adversarial examples are generated, including adversarial viewpoint images and adversarial LiDAR point cloud data, which are then input into the target unmanned system for evaluation.

Benefits of technology

It enhances the ability to transfer between different perception modalities, significantly improves evaluation efficiency, ensures the system's robustness against environmental changes, and enables accurate measurement of the robustness of the perception modules of unmanned systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482562B_ABST
    Figure CN121482562B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of safety evaluation of artificial intelligence and unmanned systems, and provides a multi-modal unmanned system safety evaluation method and system based on a neural radiation field. The method comprises the following steps: step 1: collecting perspective images and LiDAR point cloud data of a target multi-modal unmanned system, and constructing a multi-modal data set; step 2: constructing a neural radiation field and training the neural radiation field by using the multi-modal data set; step 3: inputting constraints to the trained neural radiation field, constructing a multi-modal loss function for training, and obtaining an adversarial neural radiation field; step 4: generating adversarial samples by using the adversarial neural radiation field; wherein the adversarial samples comprise adversarial perspective images and adversarial LiDAR point cloud data; and step 5: inputting the adversarial samples into the target multi-modal unmanned system to evaluate the multi-modal unmanned system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safety assessment technology for artificial intelligence and unmanned systems, and in particular to a multimodal unmanned system safety assessment method and system based on neural radiation fields. Background Technology

[0002] With the widespread application of unmanned systems in fields such as autonomous driving, intelligent inspection, and deep space exploration, conducting safety assessments of their perception systems has become a prerequisite and industry consensus for ensuring reliable deployment. Currently, deep learning-based multimodal perception models are the core of unmanned systems' understanding of their environment, and their robustness directly determines the safety boundary of the entire system.

[0003] Adversarial testing is a crucial means of evaluating model security and exposing its decision-making vulnerabilities. However, existing evaluation paradigms exhibit limitations and shortcomings when faced with the complex and dynamic physical world, leading to doubts about the validity and reliability of the evaluation results. These limitations mainly lie in the following aspects:

[0004] The one-sidedness of assessment dimensions. Mainstream security assessment methods focus on the two-dimensional image domain. Although the adversarial perturbations generated by such methods can interfere with the model at the digital level, they are difficult to reproduce in the real physical world due to the lack of three-dimensional spatial structure and physical constraints. This leads to a disconnect between the assessment scenario and the actual operating environment of the unmanned system, and fails to effectively reflect the security risks in the physical world.

[0005] The isolation of evaluation models is a concern. Unmanned systems rely on multi-sensor fusion to improve perception redundancy and security. However, most existing evaluation methods are designed for single-modal applications and lack the ability to evaluate cross-modal cooperative attacks in a unified 3D scene. This makes it impossible to examine the vulnerability of the fusion decision-making mechanism when multiple sensors are simultaneously interfered with, thus overlooking systemic security vulnerabilities.

[0006] The evaluation scenario suffers from staticity and unrealistic nature. Many 3D attack methods based on modifying point clouds or mesh textures often generate evaluation scenarios that are discrete, static, and unnatural "digital specimens." These scenarios cannot maintain attack effectiveness under continuously changing viewpoints and lighting conditions, lack spatiotemporal consistency, and therefore make it difficult to effectively evaluate the continuous security performance of unmanned systems during dynamic operation. Summary of the Invention

[0007] To address the problem that existing unmanned system safety assessment methods lack a unified multimodal information and a three-dimensional dynamic assessment base that guarantees physical authenticity, leading to inaccurate assessment results, this invention proposes a multimodal unmanned system safety assessment method and system based on neural radiation fields.

[0008] In a first aspect, the present invention provides a method for safety assessment of multimodal unmanned systems based on neural radiation fields, comprising:

[0009] Step 1: Collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset;

[0010] Step 2: Construct the neural radiation field and train it using the multimodal dataset;

[0011] Step 3: Input constraints to the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field;

[0012] Step 4: Generate adversarial examples using the adversarial neural radiation field; wherein the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data;

[0013] Step 5: Input the adversarial examples into the target multimodal unmanned system to evaluate the multimodal unmanned system.

[0014] Furthermore, in step 2, the neural radiation field is trained using a first loss function, the formula of which is as follows:

[0015]

[0016] in, Denotes the first loss function. Indicates light, The three-dimensional space representing the generation of the neural radiation field. This indicates the rendering of light in three-dimensional space generated by neural radiation fields. This indicates realistic lighting rendering.

[0017] Furthermore, the three-dimensional space generated by the neural radiation field is rendered with light colors to calculate the first loss function, wherein the color rendering formula is as follows:

[0018]

[0019] in, Indicates the range of light rays The probability of being absorbed / scattered for the first time. Indicates transmittance. Indicates location, Represents volume density. Indicates the sampling interval. Represents RGB color. This represents the number of discrete sampling points when rendering a volume along a ray.

[0020] Furthermore, in step 3, the constraint is to introduce an environmental noise and terrain disturbance model to define the range against disturbances:

[0021]

[0022] in, Represents the environmental noise term. This represents scene-independent medium and terrain interference parameters. Indicates the upper limit of the disturbance amplitude. Represents the set of original scene parameters. This represents the adversarial parameters obtained after constraints.

[0023] Furthermore, in step 3, the training uses a multimodal adversarial loss function, as shown in the following formula:

[0024]

[0025] in, Represents the multimodal loss function. The cross-entropy loss function represents the visual classifier. This represents the cross-entropy loss function of the LiDAR detector. This represents the loss function that forces visual and LiDAR features to align. Represents the smoothing loss function. , , and The weights are represented by the formula for the forced visual and LiDAR feature alignment loss function, which is as follows:

[0026]

[0027] in, Indicates the image unit index. Indicates the total number of image units. Representing image unit The corresponding weights or confidence levels, The visual input representing the adversarial example. Visual input representing a clean view. Indicates visual input Mapped to the embedding space, Indicates visual input Mapped to the embedding space;

[0028] The formula for the smoothing loss function is as follows:

[0029]

[0030] in, Represents the horizontal discrete gradient. Represents a smoothed scalar field. This represents the vertical discrete gradient.

[0031] Furthermore, in step 3, the training employs a dynamic adjustment strategy to update the neural radiation field parameters, specifically including:

[0032] Training each iteration Next, update based on the relative rate of decrease of each loss. :

[0033]

[0034] in, Represents the t-th iteration , Indicates the learning rate. Indicates the change in loss;

[0035] Based on the naturalness score of adversarial examples Adaptive adjustment:

[0036]

[0037]

[0038]

[0039] in, Indicates a fixed hyperparameter. Indicates a fixed hyperparameter. This indicates a penalty point for non-natural behavior.

[0040] Furthermore, in step 5, the perceptual robustness index PRI, multimodal consistency error MCE, and dynamic adaptability score DAS are used to evaluate the multimodal unmanned system.

[0041] Secondly, the present invention provides a multimodal unmanned system safety assessment system based on neural radiation fields, comprising:

[0042] The multimodal data collection module is used to collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset;

[0043] A neural radiation field construction and training module is used to construct a neural radiation field and train it using the multimodal dataset.

[0044] The constraint introduction module is used to input constraints onto the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field;

[0045] An adversarial example generation module is used to generate adversarial examples using the adversarial neural radiation field; wherein the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data;

[0046] A multimodal unmanned system evaluation module is used to input the adversarial examples into the target multimodal unmanned system to evaluate the multimodal unmanned system.

[0047] Thirdly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect above.

[0048] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor to perform the method described in the first aspect above.

[0049] The beneficial effects of this invention are as follows:

[0050] This invention provides a multimodal unmanned system security assessment method based on neural radiation fields. Through multimodal adversarial perturbation generation, the assessment system exhibits stronger transfer capabilities across different perception modalities. Traditional multimodal adversarial perturbation generation suffers from slow convergence speeds, and improper step size settings can lead to failed sample generation or excessive perturbation. This invention optimizes attack efficiency through dynamic step size adjustment, significantly improving assessment efficiency. Furthermore, this invention employs cross-modal adversarial perturbation assessment, and the generated adversarial samples, rendered from any viewpoint, inherently possess adversarial characteristics in both images and point clouds. This allows for precise measurement of the robustness of the unmanned system's perception modules, ensuring the system's robustness against environmental changes. Attached Figure Description

[0051] Figure 1 A flowchart illustrating a multimodal unmanned system safety assessment method based on neural radiation fields, provided in an embodiment of the present invention;

[0052] Figure 2 A schematic diagram of the framework of a multimodal unmanned system safety assessment system based on neural radiation field provided in an embodiment of the present invention;

[0053] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0054] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the embodiments of this invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0055] like Figure 1As shown in the figure, an embodiment of the present invention provides a safety assessment method for a multimodal unmanned system based on neural radiation fields, comprising:

[0056] Step 1: Collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset.

[0057] Specifically, the system collects time-synchronized and spatially calibrated viewpoint images and LiDAR point cloud data of the target multimodal unmanned system in various scenarios to construct a high-quality multimodal dataset.

[0058] Step 2: Construct the neural radiation field and train it using a multimodal dataset.

[0059] Specifically, the neural radiation field constructed in this step can fuse image appearance and LiDAR geometric information, and is trained using a multimodal dataset to obtain a high-fidelity scene data twin model.

[0060] Step 3: Input constraints to the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field.

[0061] Specifically, after pre-training in step 2, the neural radiation field is fine-tuned by introducing adversarial perturbations and constructing a multimodal adversarial loss function that depends on the gradient of the downstream task model.

[0062] Step 4: Generate adversarial examples using adversarial neural radiation fields; where adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data.

[0063] Specifically, the adversarial viewpoint image and multi-resistance LiDAR point cloud data are generated using the adversarial neural radiation field obtained in step 3.

[0064] Step 5: Input adversarial examples into the multimodal unmanned system to evaluate the multimodal unmanned system.

[0065] Specifically, adversarial examples are input into the target multimodal unmanned system to quantitatively assess the degradation of its perception performance and the safety of the system's decision-making, thereby completing the safety assessment.

[0066] The method provided in this invention can systematically and efficiently expose the problems of multimodal unmanned systems in perception and decision-making by examining the three-dimensional nature of the attack scenario, thereby enabling security assessment.

[0067] The method provided in the embodiments of the present invention will be described in detail below.

[0068] Neural radiation field modeling specifically utilizes Nerf, a neural network-based 3D scene representation method, and employs volumetric rendering techniques to represent a continuous five-dimensional function.

[0069]

[0070] in, Represents the coordinates of a point in three-dimensional space. Indicates the direction of view. Represents RGB color. This represents volume density. For each ray in the 3D scene generated by the neural radiation field... ,in Indicates the starting point of the light ray. Represents the direction vector of light rays. The depth parameter along the light ray is used to calculate color rendering as follows:

[0071]

[0072] in, Indicates the range of light rays The probability of being absorbed / scattered for the first time. Indicates transmittance. Indicates location, Represents volume density. Indicates the sampling interval. Represents RGB color. This represents the number of discrete sampling points when rendering a volume along a ray.

[0073] The optimization objective of training the neural radiation field is to minimize the rendering error. Therefore, a first loss function is used to train the neural radiation field, and the formula for the first loss function is as follows:

[0074]

[0075] in, Denotes the first loss function. Indicates light, The three-dimensional space representing the generation of the neural radiation field. This indicates the rendering of light in three-dimensional space generated by neural radiation fields. This indicates realistic lighting rendering.

[0076] Adversarial example generation. First, input constraints are introduced, including an environmental noise and terrain interference model, defining the range of adversarial perturbations:

[0077]

[0078] in, Represents the environmental noise term. Indicates the intensity of sandstorm disturbance. Indicates the upper limit of the disturbance amplitude. This represents the set of original scene parameters, which can be understood as a parameter carrier that can be disturbed but is subject to physical constraints. This represents the adversarial parameters obtained within the constraints, used to generate adversarial viewpoint images and adversarial point clouds.

[0079] Then, the visual and LiDAR perception errors are jointly optimized to ensure that the generated adversarial examples can deceive both types of sensors simultaneously. A multimodal adversarial loss function is used, as shown in the following formula:

[0080] The multimodal adversarial loss function is adopted, and the formula is as follows:

[0081]

[0082] in, The cross-entropy loss function represents the visual classifier. This represents the cross-entropy loss function of the LiDAR detector. This represents a function that forces visual alignment with LiDAR features. Represents the smoothing loss function. , , and The weights are represented by the formula for the forced alignment loss function between visual and LiDAR features, which is as follows:

[0083]

[0084] in, Represents the alignment loss function. Indicates the index of the image unit, which can be a pixel, patch, ray sampling point, or the image position corresponding to a LiDAR point after projection by extrinsic parameters; This represents the total number of image units, that is, the number of pixels / patches / ray points included in the alignment calculation in a frame. This represents the weight or confidence level corresponding to the k-th image unit, used to mitigate the effects of occlusion, low-confidence samples, or boundary regions. The visual input representing the adversarial example. This represents the visual input for a clean view, such as in cross-modal alignment, where it is a reference view / feature map obtained by projecting or rendering LiDAR points via extrinsic parameters. Indicates visual input Mapped to the embedding space, Indicates visual input Mapped to the embedding space;

[0085] The formula for the smoothing loss function is as follows:

[0086]

[0087] in, Represents the smoothing loss function. Represents the horizontal discrete gradient, the first-order difference between adjacent grid points. This represents a smoothed scalar field, with the default value being the amplitude plot against perturbations. This represents the vertical discrete gradient.

[0088] When training the neural radiation field using a multimodal adversarial loss function, the strategy is dynamically adjusted; specifically, in each iteration... Next, update based on the relative rate of decrease of each loss. :

[0089]

[0090] in, Represents the t-th iteration , Indicates the learning rate. Indicates the change in loss;

[0091] Based on the naturalness score of adversarial examples Adaptive adjustment:

[0092]

[0093] in, represents a fixed hyperparameter, the baseline scaling factor used for the naturalness adaptive mapping, when the naturalness score of the adversarial example is at the baseline level. by Scaling from the starting point; This represents a fixed hyperparameter, used as an amplification factor for naturalness adaptive mapping, in conjunction with... adjust Size, thereby controlling the strength of the adaptive update's response to "naturalness"; , The non-naturalness penalty score reflects the degree to which the adversarial sample deviates from a natural appearance or reasonable geometry. PN is a non-negative real number, and the larger the value, the more unnatural it is.

[0094] System evaluation. The generated adversarial examples are injected into the virtual scene and synchronized to the real sensor interfaces, triggering the unmanned system response. System output is collected, and multimodal unmanned system evaluation metrics are calculated. These metrics include the Perception Robustness Index (PRI), the Multimodal Consistency Error (MCE), and the Dynamic Fitness Score (DAS).

[0095] Specifically, the Perceptual Robustness Index (PRI) is used to quantify the false positive rate of a system under adversarial examples:

[0096]

[0097] Multimodal consistency error (MCE) is used to assess the difference between visual and LiDAR decision-making.

[0098]

[0099] in, and These represent the predicted probability distributions for the two types of sensors, respectively.

[0100] Dynamic fitness score (DAS) is used to measure the stability of a system under continuous disturbances.

[0101]

[0102] in, Indicates the first The increment of perceived error in each iteration. This represents the attenuation coefficient.

[0103] Furthermore, during system evaluation, the parameters of adversarial examples can be optimized based on the feedback results.

[0104] Specifically, the parameters of adversarial examples are optimized based on indicators (such as...). This process is repeated until the metric converges. The parameters are adaptively adjusted, with the loss weights dynamically adjusted based on the evaluation results.

[0105] .

[0106] In one embodiment, a multimodal unmanned system environment is constructed using 30 classes of lunar objects (such as craters and rocks) from the IM3D dataset as a baseline. Each class contains 100 3D meshes. LiDAR point clouds are simulated using the KITTI dataset with a sampling density of [missing information]. Add a lunar dust noise model The multimodal unmanned system evaluated included a visual model and a LiDAR model, with the visual model being Inception-v3 (pre-trained on ImageNet) and an input resolution of [missing information]. The LiDAR model is PointRCNN (pre-trained with KITTI), and the input is a point cloud number. .

[0107] Step 101: Collect real-world environmental data to build a training set.

[0108] Step 102: Construct and train a neural radiation field to reconstruct a 3D environment. The neural radiation field network structure is an 8-layer MLP (256 neurons per layer), with position encoding frequency... The Adam optimizer was used during training, with an initial learning rate of... Batch size The light rays are analyzed, and a first loss function is used to minimize the reconstruction loss.

[0109] Step 103: Introduce constraints and construct a multimodal loss function to train the neural radiation field, iteration number. Save intermediate parameters every 50 iterations to overcome budget constraints. Alignment loss and smoothing loss are introduced into the multimodal loss function. Alignment loss ensures that adversarial examples remain consistent across different perspectives, while smoothing loss ensures the naturalness of the adversarial examples and avoids unrealistic noise. The parameters of the multimodal loss function are initialized as follows: .

[0110] Step 104: Generate adversarial samples using the adversarial neural radiation field trained in Step 103.

[0111] Step 105: Utilize adversarial examples to assess the security of multimodal unmanned systems.

[0112] Specifically, adversarial examples were injected into the lunar rover's perception module to test the following scenarios: static obstacles: impact craters, rock piles (size...) );

[0113] Dynamic interference: Simulated lunar dust storm (visibility) );

[0114] Multitasking load: Performance degradation when performing navigation, sampling, and communication tasks simultaneously.

[0115] like Figure 2 As shown, this embodiment of the invention also provides a multimodal unmanned system safety assessment system based on neural radiation fields, comprising:

[0116] The multimodal data collection module is used to collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset;

[0117] The neural radiation field construction and training module is used to construct neural radiation fields and train them using multimodal datasets.

[0118] The constraint introduction module is used to input constraints onto the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field;

[0119] The adversarial example generation module is used to generate adversarial examples using adversarial neural radiation fields; the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data;

[0120] The multimodal unmanned system evaluation module is used to input adversarial examples into the target multimodal unmanned system to evaluate the multimodal unmanned system.

[0121] Based on the above embodiments, such as Figure 3 As shown, this embodiment also provides an electronic device, which may include: a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304. The processor 301 can call logical instructions in the memory 303 to execute the methods provided in the above embodiment, such as including:

[0122] Step 1: Collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset; Step 2: Construct a neural radiation field and train it using the multimodal dataset; Step 3: Input constraints to the trained neural radiation field and construct a multimodal loss function for training to obtain an adversarial neural radiation field; Step 4: Generate adversarial examples using the adversarial neural radiation field; the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data; Step 5: Input the adversarial examples into the multimodal unmanned system to evaluate the multimodal unmanned system.

[0123] Furthermore, when the logical instructions in the aforementioned memory 303 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0124] Based on the above embodiments, this embodiment also provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the method provided in the above embodiments, for example including:

[0125] Step 1: Collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset; Step 2: Construct a neural radiation field and train it using the multimodal dataset; Step 3: Input constraints to the trained neural radiation field and construct a multimodal loss function for training to obtain an adversarial neural radiation field; Step 4: Generate adversarial examples using the adversarial neural radiation field; the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data; Step 5: Input the adversarial examples into the multimodal unmanned system to evaluate the multimodal unmanned system.

[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A safety assessment method for multimodal unmanned systems based on neural radiation fields, characterized in that, Include: Step 1: Collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset; Step 2: Construct the neural radiation field and train it using the multimodal dataset; Step 3: Input constraints onto the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field; the training uses a multimodal adversarial loss function, the formula of which is as follows: in, Represents the multimodal loss function. The cross-entropy loss function represents the visual classifier. This represents the cross-entropy loss function of the LiDAR detector. This represents the loss function that forces visual and LiDAR features to align. Represents the smoothing loss function. , , and The weights are represented by the formula for the forced visual and LiDAR feature alignment loss function, which is as follows: in, Indicates the image unit index. Indicates the total number of image units. Representing image unit The corresponding weights or confidence levels, The visual input representing the adversarial example. Visual input representing a clean view. Indicates visual input Mapped to the embedding space, Indicates visual input Mapped to the embedding space; The formula for the smoothing loss function is as follows: in, Represents the horizontal discrete gradient. This represents a smoothed scalar field. Represents the vertical discrete gradient; The training employs a dynamic adjustment strategy to update neural radiation field parameters, specifically including: Training each iteration Next, update based on the relative rate of decrease of each loss. : in, Represents the t-th iteration , Indicates the learning rate. Indicates the change in loss; Based on the naturalness score of adversarial examples Adaptive adjustment: in, Indicates a fixed hyperparameter. Indicates a fixed hyperparameter. Indicates non-natural penalty points; Step 4: Generate adversarial examples using the adversarial neural radiation field; wherein the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data; Step 5: Input the adversarial examples into the target multimodal unmanned system to evaluate the multimodal unmanned system.

2. The method for safety assessment of a multimodal unmanned system based on neural radiation fields according to claim 1, characterized in that, In step 2, the neural radiation field is trained using a first loss function, the formula of which is as follows: in, Denotes the first loss function. Indicates light, The three-dimensional space representing the generation of the neural radiation field. This indicates the rendering of light in three-dimensional space generated by neural radiation fields. This indicates realistic lighting rendering.

3. The method for safety assessment of a multimodal unmanned system based on neural radiation fields according to claim 2, characterized in that, The three-dimensional space generated by the neural radiation field is rendered with light colors to calculate the first loss function, wherein the color rendering formula is as follows: in, Indicates the range of light rays The probability of being absorbed / scattered for the first time. Indicates transmittance. Indicates location, Represents volume density. Indicates the sampling interval. Represents RGB color. This represents the number of discrete sampling points when rendering a volume along a ray.

4. The method for safety assessment of a multimodal unmanned system based on neural radiation fields according to claim 1, characterized in that, In step 3, the constraint is to introduce an environmental noise and terrain disturbance model to define the range against disturbances: in, Represents the environmental noise term. Indicates scene-independent media and terrain interference parameters. Indicates the upper limit of the disturbance amplitude. Represents the set of original scene parameters. This represents the adversarial parameters obtained after constraints.

5. The method for safety assessment of a multimodal unmanned system based on neural radiation fields according to claim 1, characterized in that, In step 5, the perceptual robustness index PRI, multimodal consistency error MCE, and dynamic adaptability score DAS are used to evaluate the multimodal unmanned system.

6. A multimodal unmanned system safety assessment system based on neural radiation fields, characterized in that, include: The multimodal data collection module is used to collect viewpoint images and LiDAR point cloud data of the target multimodal unmanned system to construct a multimodal dataset; A neural radiation field construction and training module is used to construct a neural radiation field and train it using the multimodal dataset. The constraint introduction module is used to input constraints onto the trained neural radiation field and construct a multimodal loss function for training to obtain the adversarial neural radiation field. The training uses a multimodal adversarial loss function, the formula of which is as follows: in, Represents the multimodal loss function. The cross-entropy loss function represents the visual classifier. This represents the cross-entropy loss function of the LiDAR detector. This represents the loss function that forces visual and LiDAR features to align. Represents the smoothing loss function. , , and The weights are represented by the formula for the forced visual and LiDAR feature alignment loss function, which is as follows: in, Indicates the image unit index. Indicates the total number of image units. Representing image unit The corresponding weights or confidence levels, The visual input representing the adversarial example. Visual input representing a clean view. Indicates visual input Mapped to the embedding space, Indicates visual input Mapped to the embedding space; The formula for the smoothing loss function is as follows: in, Represents the horizontal discrete gradient. This represents a smoothed scalar field. Represents the vertical discrete gradient; The training employs a dynamic adjustment strategy to update neural radiation field parameters, specifically including: Training each iteration Next, update based on the relative rate of decrease of each loss. : in, Represents the t-th iteration , Indicates the learning rate. Indicates the change in loss; Based on the naturalness score of adversarial examples Adaptive adjustment: in, Indicates a fixed hyperparameter. Indicates a fixed hyperparameter. Indicates non-natural penalty points; An adversarial example generation module is used to generate adversarial examples using the adversarial neural radiation field; wherein the adversarial examples include adversarial viewpoint images and adversarial LiDAR point cloud data; A multimodal unmanned system evaluation module is used to input the adversarial examples into the target multimodal unmanned system to evaluate the multimodal unmanned system.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • New view angle synthesis method based on depth image and neural radiation field

    CN113706714A

  • Visual three-dimensional model construction method and device for power equipment

    CN118212356A