Automatic driving scene library dynamic generation method based on multi-modal data fusion

By combining a distributed data acquisition layer, a multimodal semantic encoding layer, and a dynamic evolution engine with a federated learning framework and a virtual-real coexistence scene generator, the problem of dynamic generation and safe sharing of autonomous driving scene libraries is solved, the coverage of extreme working conditions and data security are improved, and high-fidelity testing and verification are provided.

CN121483045BActive Publication Date: 2026-04-17BEIJING SMART CAR MZONE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SMART CAR MZONE CO LTD
Filing Date
2026-01-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing methods for building autonomous driving scenario libraries cannot dynamically capture extreme working conditions. Fragmented multi-source data leads to scenario distortion, high risk of privacy leakage, serious spatiotemporal conflicts, single evaluation model, inability to quantify system failure probability, and lack of security mechanisms for data sharing.

Method used

By establishing a distributed data acquisition layer, a multimodal semantic encoding layer, a dynamic evolution engine, and a virtual-real coexistence scenario generator, and combining a federated learning framework, we can achieve secure collaborative acquisition of multi-source data, integrate multimodal data streams, generate a high-fidelity, multi-dimensional test scenario library, and generate risk assessment reports using physical interference simulation and group interaction rules.

Benefits of technology

It enables the dynamic generation and continuous evolution of the autonomous driving scenario library, improves the coverage of extreme working conditions, ensures data security, solves the problem of spatiotemporal conflicts, and provides comprehensive and reliable testing and verification support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483045B_ABST
    Figure CN121483045B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multimodal data fusion's automatic driving scene library dynamic generation method, belong to the field of automatic driving test, solve the problem of insufficient extreme working condition coverage, data fragmentation and single evaluation of static scene library.Its technical scheme includes: establishing distributed data acquisition layer, through the federal learning framework access vehicle terminal and V2X roadside equipment data stream;Construct multi-modal semantic coding layer, fuse data extraction features and convert laser radar point cloud data into space-time constraint matrix, output scene DNA coding sequence;Deploy dynamic evolution engine, map DNA coding sequence to traffic flow state parameter and calculate entropy gradient field;Start virtual-real symbiosis scene generator, receive high-entropy value area coordinates and physical interference intensity parameters, simulate to generate controlled physical interference test scene, while based on behavior uncertainty injection group interaction rules, finally output dynamic updated scene library and risk report, for automatic driving system test evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous driving system technology. More specifically, this invention relates to a method for dynamically generating an autonomous driving scene library based on multimodal data fusion. Background Technology

[0002] The testing and verification of autonomous driving technology heavily relies on the construction of scenario libraries, but current mainstream methods have significant shortcomings. Static scenario library architectures are limited by preset rules and historical data boundaries, making it difficult to dynamically capture extreme conditions in real-world roads, such as long-tail scenarios like chain reactions from sudden accidents or extreme weather changes, resulting in insufficient safety verification coverage. The lack of a multi-source heterogeneous data collaborative processing mechanism further exacerbates scenario distortion: data streams independently collected by onboard sensing systems and roadside facilities cannot be effectively integrated due to protocol differences and temporal misalignments, leading to a fragmentation of the scenario's semantic integrity. Traditional evaluation models rely only on limited-dimensional parameters, failing to quantify the probability of system failure under cognitive uncertainty, resulting in a singular dimension of safety risk assessment.

[0003] Serious privacy vulnerabilities exist in the distributed data acquisition process. Existing federated learning frameworks lack end-to-end data flow protection mechanisms, making raw sensor data from vehicle terminals and commercial data from roadside units vulnerable to man-in-the-middle attacks during transmission. User trajectory information and confidential equipment operation data are at risk of reverse engineering. The lack of enterprise-level confidentiality isolation solutions during data sharing in commercial scenarios further hinders cross-entity data collaboration.

[0004] Multimodal data fusion faces challenges related to spatiotemporal conflicts. LiDAR point clouds and visual perception data exhibit millisecond-level timing discrepancies due to equipment acquisition delays. When intelligent traffic light phase switching coincides with changes in vehicle trajectories, traditional hard synchronization mechanisms cannot eliminate semantic contradictions within spatiotemporal units. Accuracy drift caused by long-term operation of roadside equipment leads to systematic errors in the mapping of traffic participant locations and road topology. Sudden traffic events, such as temporary traffic control or accident scenes, lack real-time topology correction capabilities, resulting in logical conflicts in scene coding.

[0005] The aforementioned shortcomings create a triple technical barrier: static architecture prevents the scene library from evolving adaptively; data fragmentation leads to the failure of cross-domain feature extraction; and the simplistic evaluation model masks systemic risks. Privacy leaks hinder large-scale data collaboration, while spatiotemporal conflicts directly reduce the credibility of scene coding. Summary of the Invention

[0006] This invention provides a method for dynamically generating autonomous driving scenario libraries based on multimodal data fusion. It can construct an autonomous driving test scenario library system that is dynamically evolving, high-fidelity, and multi-dimensionally evaluated. It achieves safe and collaborative acquisition of multi-source data through a federated learning framework, improves the coverage of extreme working conditions by utilizing multimodal semantic encoding and a dynamic evolution engine, and combines physically realistic interference simulation and group interaction rule generation technology to output accurate risk assessment reports, providing comprehensive and reliable test and verification support for high-level autonomous driving systems.

[0007] To achieve these objectives and other advantages of the present invention, a method for dynamically generating an autonomous driving scene library based on multimodal data fusion is provided, comprising:

[0008] A distributed data acquisition layer is established, which accesses the data streams of the vehicle terminal autonomous driving system and V2X roadside equipment through a federated learning framework with time buffering, and uses a streaming processing pipeline to achieve time decoupling between data acquisition and model training.

[0009] A multimodal semantic encoding layer is constructed, and multimodal data streams from vehicle terminals and V2X roadside devices are fused through a federated learning framework to extract cross-domain semantic features. At the same time, LiDAR point cloud data is converted into a spatiotemporal constraint matrix through a topology graph mapping engine. The semantic features and the spatiotemporal constraint matrix are input together into the driving strategy tree generation module to output a scene DNA encoding sequence.

[0010] A dynamic evolution engine is deployed to map DNA coding sequences to traffic flow state parameters through traffic flow field theory, and to calculate the scene entropy gradient field based on the Navier-Stokes equations; the device scheduling force vector is dynamically generated according to the gradient field strength, driving the distributed data acquisition devices to focus and migrate to high information entropy areas;

[0011] The virtual-real symbiotic scene generator is activated, receiving the high-entropy region coordinates and physical interference intensity parameters output by the dynamic evolution engine. The physical interference intensity parameters include visibility value, precipitation intensity level and radar wave attenuation coefficient. A controlled physical interference test scene of the target area is generated using 3D scene rendering technology. At the same time, the uncertainty gradient of group behavior is calculated based on the dynamic evolution engine, and group interaction rule parameters that conform to the characteristics of traffic flow field are injected through a conditional adversarial generative network.

[0012] The final output includes a dynamically updated scenario library and a risk assessment report.

[0013] Preferably, the federated learning framework incorporates a three-level privacy filtering mechanism to collaboratively process the data streams from the onboard terminal's autonomous driving system and the V2X roadside equipment, specifically as follows:

[0014] 1) Terminal-level noise injection stage: Locally at the vehicle terminal and V2X roadside equipment, adaptive Gaussian noise is injected into the original sensor data stream through the differential noise injection module to generate an encrypted feature stream;

[0015] 2) Edge-level feature desensitization stage: The edge server receives encrypted feature streams from multiple terminals, uses knowledge distillation technology to remove device identity information and aggregate cross-domain features, and outputs the desensitized cross-domain scene feature stream to the multimodal semantic coding layer;

[0016] 3) Cloud-level security update phase: The model parameters of each terminal are aggregated on the cloud platform to generate a global model. The global weight parameter stream is protected by homomorphic encryption algorithm and distributed to the vehicle terminal and V2X roadside equipment to dynamically optimize their data acquisition capabilities. The vehicle terminal includes LiDAR, camera and millimeter-wave radar; the V2X roadside equipment includes roadside unit, smart traffic light and weather station.

[0017] Preferably, the driving strategy tree generation module includes a conflict resolution mechanism, specifically:

[0018] When lidar point cloud data and smart traffic light phase data conflict in the same spatiotemporal unit, a confidence fusion algorithm based on Bayesian inference is activated to calculate the reliability weight of each data source in the spatiotemporal unit. The same spatiotemporal unit meets the following conditions: time tolerance ≤ 100ms and spatial tolerance ≤ 1m.

[0019] Simultaneously, a dynamic topology correction engine is deployed. When temporary traffic control, road construction, or sudden factors at an accident scene are identified, the engine automatically reconstructs the road topology connection relationship and updates the branch weights of the driving strategy tree. Specifically, the dynamic topology correction engine reconstructs the road topology connection relationship by: a) loading a pre-set topology template based on the type of sudden factor; b) fusing real-time multimodal data to generate a semantic raster map; c) dynamically adjusting the topology connection matrix according to the semantic raster status; and d) updating the branch weight coefficients of the driving strategy tree according to the urgency of the event.

[0020] Preferably, 3D scene rendering technology and environmental interference simulation models are used in conjunction to generate controlled physical interference scenes that conform to physical laws; the environmental interference simulation model specifically includes:

[0021] Visual interference sub-model: Based on the physical properties of rain and fog particles, calculate the degree of light attenuation and color deviation caused by them to visible and near-infrared light;

[0022] Radar jamming sub-model: Based on electromagnetic wave characteristics, predict the distortion and position offset of lidar point cloud data in rain and fog environments;

[0023] Among them, the 3D scene rendering technology combines the pixel-level degradation parameters output by the visual interference sub-model with the point cloud distortion parameters output by the radar interference sub-model to generate a physically realistic controlled physical interference scene.

[0024] Meanwhile, the conditional adversarial generative network introduces group behavior rules to adjust the interactive decision-making mechanism of traffic participants and generate group interaction rule parameters.

[0025] Preferably, the group behavior rules generate group interaction rule parameters through the following steps:

[0026] 1) Behavioral decision modeling: Extract the historical movement trajectories of traffic participants and analyze their speed changes, direction adjustments, and interactive avoidance habits;

[0027] 2) Environmental Status Perception: Real-time monitoring of traffic light status, road congestion levels, and the location of surrounding road users;

[0028] 3) Conflict resolution mechanism: When the movement paths of multiple traffic participants conflict, right-of-way is dynamically allocated according to the priority of traffic regulations;

[0029] 4) Interaction parameter generation: Integrate behavioral decision-making habits with environmental conditions to output probability parameters of interaction actions such as acceleration, deceleration or turning.

[0030] Preferably, a version compatibility verification layer is set in the dynamic evolution engine to handle compatibility issues of non-traditional traffic participant data. When importing non-traditional traffic participant data, a scene migration verification test is first run in an isolated sandbox. By comparing the degree of difference between the DNA encoding of the new and old versions of the scene, it is determined whether the topology mapping rules need to be reconstructed. If the degree of difference exceeds a preset threshold, the incremental learning module is started to fine-tune and optimize the encoder, while retaining the parallel call interface of the historical version scene library to ensure the continuous compatibility of the old version scene library.

[0031] Preferably, the distributed data acquisition layer enables a data completion mechanism in areas with missing roadside data. Specifically, based on the scene entropy distribution of the acquired areas, a spatiotemporal adversarial generative network is trained to synthesize simulated data that conforms to the traffic flow characteristics of the target area. At the same time, a data authenticity verification unit is set up to dynamically optimize the training parameters of the generative network by comparing the distribution differences between the generated data and the real data at the driving strategy decision nodes, thereby filling the data gaps in areas with weak infrastructure.

[0032] Preferably, the steps for generating the risk assessment report specifically include:

[0033] 1) Perform multi-agent reinforcement learning risk assessment based on a dynamically updated scenario library, and load the driving strategy confusion index assessment model into the test sandbox;

[0034] 2) Quantify the failure probability of autonomous driving systems under cognitive uncertainty;

[0035] 3) Generate dynamically updated risk assessment reports based on failure probability.

[0036] Preferably, the driving strategy confusion index evaluation model includes a cognitive load monitoring module, which monitors three types of decision-making anomaly indicators in real time: the frequency of strategy tree branch switching exceeds the safety threshold; the control command output fluctuates abnormally under the same input conditions; and the spatiotemporal constraint matrix satisfaction rate drops sharply.

[0037] Specifically, when the cognitive load monitoring module detects any decision-making anomaly indicator that triggers an alarm, it automatically initiates a scenario backtracking mechanism to compare and analyze the current decision chain with the historical best decision-making strategy to locate the root cause of the decision anomaly.

[0038] Preferably, the dynamically updated scenario library drives the iterative optimization of the autonomous driving system through a closed-loop feedback mechanism of testing and verification, specifically including:

[0039] The dynamically updated scenario library is input into the autonomous driving system test platform to perform stress tests and collect system response data;

[0040] Based on system response data, coordinates of blind spots in scene coverage and characteristics of strategy defects are generated.

[0041] The coordinates of the blind spots and the characteristics of strategy defects are fed back to the dynamic evolution engine;

[0042] The dynamic evolution engine drives the distributed data acquisition device to enhance data acquisition in areas corresponding to the coordinates of the coverage blind spots and the characteristics of the strategy defects, and drives the virtual and real symbiotic scene generator to generate adversarial test scenarios targeting the characteristics of the strategy defects.

[0043] Ultimately, this forms a closed-loop iterative chain encompassing scenario library generation, system testing, defect feedback, and dynamic optimization.

[0044] The present invention has at least the following beneficial effects:

[0045] First, this invention achieves the dynamic generation and continuous evolution of an autonomous driving scenario library by constructing an integrated technical framework comprising a distributed data acquisition layer, a multimodal semantic encoding layer, a dynamic evolution engine, and a virtual-real symbiotic scene generator. The federated learning framework with temporal buffering ensures temporal decoupling and secure access to multi-source data streams; the multimodal semantic encoding layer generates a DNA-encoded sequence that accurately characterizes the essence of the scene by fusing cross-domain semantic features with a spatiotemporal constraint matrix derived from LiDAR point cloud transformation; the dynamic evolution engine calculates the scene entropy gradient field based on traffic flow field theory and the Navier-Stokes equations, driving data acquisition devices to focus on high-information-entropy regions, improving coverage of extreme conditions and long-tailed scenarios; the virtual-real symbiotic scene generator combines 3D scene rendering technology, physical interference parameters, and group interaction rules to generate high-fidelity test scenarios and output multi-dimensional risk assessment reports, providing comprehensive and reliable testing and verification support for high-level autonomous driving systems.

[0046] Secondly, this invention establishes a three-tiered privacy filtering mechanism—terminal-level noise injection, edge-level feature desensitization, and cloud-level secure updates—to achieve collaborative and secure processing of the entire data stream of the in-vehicle terminal autonomous driving system and the V2X roadside equipment. The terminal-level differential noise injection module injects adaptive Gaussian noise into the original sensor data stream locally, generating an encrypted feature stream that makes the data "usable but invisible," blocking privacy leaks at the source. At the edge level, knowledge distillation technology is used to strip away device identity information and aggregate cross-domain features, outputting a desensitized cross-domain scene feature stream, ensuring privacy security during the aggregation process. At the cloud level, homomorphic encryption algorithms protect the global model weight parameter stream, ensuring zero exposure of trade secrets during distributed updates. This mechanism achieves secure sharing of multi-source data while dynamically optimizing device data acquisition capabilities through weight distribution, thus enhancing both privacy protection and model performance.

[0047] Third, this invention effectively solves the problem of spatiotemporal conflicts caused by device latency, accuracy drift, or sudden traffic events in multimodal data by deploying a conflict resolution mechanism and a dynamic topology map correction engine within the driving strategy tree generation module. Based on a Bayesian inference-based confidence fusion algorithm, within the same spatiotemporal unit with a time tolerance ≤100ms and a spatial tolerance ≤1m, it dynamically calculates the reliability weights of each data source, intelligently resolving semantic contradictions and preserving the value of conflicting data. The dynamic topology map correction engine can identify sudden factors such as temporary traffic control, road construction, or accident scenes. By loading a pre-set topology template, fusing real-time multimodal data to generate a semantic raster map, dynamically adjusting the topology connection matrix, and updating the driving strategy tree branch weights according to the urgency of the event, it automatically reconstructs the road topology connection relationships, ensuring the spatiotemporal consistency of the scene's DNA encoding under extreme disturbances, and providing highly reliable input to the autonomous driving decision-making module.

[0048] Fourth, by collaborating 3D scene rendering technology with an environmental interference simulation model (including a visual interference sub-model and a radar interference sub-model), a controlled physical interference test scenario conforming to the laws of physics and the propagation characteristics of electromagnetic waves is generated. The visual interference sub-model calculates the attenuation of visible and near-infrared light and color deviation values ​​caused by rain and fog particles based on their physical properties; the radar interference sub-model predicts the distortion and positional offset of lidar point cloud data in rain and fog environments based on electromagnetic wave characteristics; the 3D scene rendering technology combines these pixel-level degradation parameters with point cloud distortion parameters to generate a physically realistic test scenario. Simultaneously, the conditional adversarial generative network, by introducing group behavior rule parameters, adjusts the interactive decision-making mechanism of virtual traffic participants, solving the problem of sensor perception distortion and behavioral interaction distortion caused by traditional virtual scenarios deviating from physical laws, thus providing a highly realistic and consistent multi-sensor perception test environment for autonomous driving systems.

[0049] Fifth, the group behavior rules are generated through four steps: behavioral decision-making modeling, environmental state perception, conflict resolution mechanisms, and interaction parameter generation. These steps produce group interaction rule parameters that conform to human driving habits and traffic regulations. By extracting the historical movement trajectories of traffic participants to analyze their behavioral habits, monitoring environmental states in real time, and dynamically allocating right-of-way according to traffic regulations, the system ultimately outputs probability parameters for interactive actions such as acceleration, deceleration, or steering. This process endows traffic participants in the virtual scenario with the decision-making inertia, risk avoidance awareness, and regulatory compliance of human drivers, significantly improving the realism and complexity of group behavior simulation. This allows for more effective testing and verification of the decision-making and response capabilities of autonomous driving systems in real, complex interactive scenarios.

[0050] Sixth, by setting a version compatibility verification layer in the dynamic evolution engine, compatibility issues when introducing non-traditional traffic participant data are effectively addressed. When importing new data, the system first runs scene migration verification tests in an isolated sandbox. By comparing the differences in the DNA encoding of the new and old versions of the scenes, it scientifically determines whether the topology mapping rules need to be reconstructed. If the difference exceeds a preset threshold, the incremental learning module is activated to fine-tune and optimize the multimodal semantic encoder, rather than completely reconstructing it, thus maximizing system stability while adapting to the new data structure. At the same time, the parallel call interface of the historical version scene library is retained to ensure the continued compatibility of the old version test system. This mechanism breaks through the upgrade barrier caused by the solidification of the data structure of traditional scene libraries, supporting the rapid and smooth access of new traffic elements.

[0051] Seventh, the data completion mechanism activated by the distributed data acquisition layer in areas with missing roadside data scientifically fills the data gaps caused by weak infrastructure through the collaborative work of a spatiotemporal adversarial generative network and a data authenticity verification unit. This mechanism trains the generative network based on the scene entropy distribution of the collected areas to synthesize simulated data that conforms to the traffic flow characteristics of the target area. The authenticity verification unit dynamically feeds back and optimizes the training parameters of the generative network by comparing the distribution differences between the generated data and real data at key nodes of driving strategy decision-making. The simulated data generated by this method not only possesses physical constraints but also maintains statistical distribution characteristics consistent with real data, providing an effective solution to the data scarcity problem in emerging and high-entropy areas and laying a data foundation for generating high-quality test scenarios.

[0052] Eighth, the risk assessment report is dynamically generated through steps such as multi-agent reinforcement learning risk assessment and quantifying the failure probability under cognitive uncertainty. This process loads a driving strategy confusion index assessment model into the test sandbox, overcoming the limitations of traditional methods that rely solely on single indicators such as collision rate. It can deeply quantify the probability of decision failures caused by cognitive confusion in unknown or complex scenarios for autonomous driving systems. Based on a dynamically updated scenario library and failure probability-generated risk assessment reports, the reports can more accurately reflect the cognitive boundaries and safety bottlenecks of autonomous driving systems, providing multi-dimensional and in-depth data support and decision-making basis for system safety verification and performance optimization.

[0053] Ninth, the cognitive load monitoring module included in the driving strategy confusion index evaluation model provides a quantitative tool for identifying the system's decision-making state by real-time monitoring three types of decision anomaly indicators: policy tree branch switching frequency, control command output fluctuations, and spatiotemporal constraint matrix satisfaction rate. When any indicator triggers an alarm, the system automatically initiates a scenario backtracking mechanism, comparing and analyzing the current decision chain with historical optimal decision strategies. This allows for precise identification of the root cause of decision anomalies (such as specific sensor failure, algorithm module defects, or scene misunderstanding). This breaks the predicament of the "black box" of autonomous driving decision-making being difficult to trace, enabling developers to optimize policy trees or related modules in a targeted manner, thus improving the efficiency of system debugging and robustness optimization.

[0054] Tenth, this invention constructs a closed-loop feedback mechanism for testing and verification, linking a dynamically updated scenario library, autonomous driving system testing, defect discovery, and dynamic optimization into an automated iterative chain. The coordinates of blind spots and policy defect characteristics exposed by stress testing are fed back to the dynamic evolution engine in real time. The engine then drives distributed data acquisition devices to focus on and migrate towards the defect area, collecting richer real-world data, and drives a virtual-real co-creation scenario generator to generate targeted adversarial test scenarios. The new data and scenarios feed back into the dynamic updates of the scenario library and the retesting of the system. This closed-loop process achieves continuous iteration of "defect discovery - data supplementation - scenario generation - system optimization," greatly improving the efficiency and relevance of testing and verification, and continuously driving the evolution of the autonomous driving system's ability to cope with extreme working conditions.

[0055] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating the method for dynamically generating an autonomous driving scenario library based on multimodal data fusion, as described in this invention. Detailed Implementation

[0057] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.

[0058] It should be understood that terms such as “having,” “comprising,” and “including” as used herein do not exclude the presence or addition of one or more other elements or combinations thereof.

[0059] like Figure 1 As shown, this embodiment of the invention provides a method for dynamically generating an autonomous driving scenario library based on multimodal data fusion, including:

[0060] S1. Establish a distributed data acquisition layer, and access the data stream of the vehicle terminal autonomous driving system and the data stream of V2X roadside equipment through a federated learning framework with time buffer. Use a streaming processing pipeline to achieve time decoupling between data acquisition and model training.

[0061] S2. Construct a multimodal semantic encoding layer, and extract cross-domain semantic features by fusing multimodal data streams from vehicle terminals and V2X roadside devices through a federated learning framework; at the same time, convert LiDAR point cloud data into a spatiotemporal constraint matrix through a topology graph mapping engine; input the semantic features and spatiotemporal constraint matrix into the driving strategy tree generation module, and output the scene DNA encoding sequence.

[0062] S3. Deploy a dynamic evolution engine to map DNA coding sequences to traffic flow state parameters through traffic flow field theory, and calculate the scene entropy gradient field based on the Navier-Stokes equations; dynamically generate device scheduling field force vectors according to the gradient field strength, and drive distributed data acquisition devices to focus and migrate to high information entropy areas.

[0063] S4. Start the virtual-real symbiotic scene generator, receive the high-entropy region coordinates and physical interference intensity parameters output by the dynamic evolution engine. The physical interference intensity parameters include visibility value, precipitation intensity level and radar wave attenuation coefficient. Use 3D scene rendering technology to generate a controlled physical interference test scene for the target area. At the same time, calculate the uncertainty gradient of group behavior based on the dynamic evolution engine, and inject group interaction rule parameters that conform to the characteristics of traffic flow field through a conditional adversarial generative network.

[0064] S5. The final output includes a dynamically updated scenario library and a risk assessment report.

[0065] In the above embodiments, the distributed data acquisition layer accesses data streams from the vehicle-mounted terminal autonomous driving system and V2X roadside equipment through a federated learning framework with a time-series buffering mechanism. The vehicle-mounted terminal may include sensor devices such as LiDAR, cameras, and millimeter-wave radar, while the V2X roadside equipment encompasses infrastructure such as roadside units, smart traffic lights, and weather monitoring stations. The time-series buffering mechanism is primarily used to process multi-source asynchronous data streams, resolving data timing misalignment issues caused by network latency or inconsistent device sampling frequencies. Its buffer window can be dynamically adjusted according to actual network conditions, for example, between 100 milliseconds and 500 milliseconds. The federated learning framework allows raw data to remain on the local device, aggregating it only through encrypted model parameters or features, thereby achieving multi-party collaborative training while ensuring data privacy. The streaming pipeline further decouples the data acquisition and model training processes in time. The acquisition end continuously receives real-time data and performs preliminary cleaning and caching, while the training end extracts batch data as needed for model updates. The two communicate asynchronously through message queues or data buses, thereby improving the overall system throughput and response efficiency.

[0066] The multimodal semantic coding layer, relying on a federated learning framework, fuses various modal data from vehicle terminals and V2X roadside equipment, such as images, point clouds, millimeter-wave signals, traffic light status, and meteorological information, to extract feature representations with cross-domain semantic meaning, such as vehicle behavior intent, road structure semantics, and traffic event types. Simultaneously, this layer also includes a point cloud data processing module. Through a topology graph mapping engine, it converts the 3D point cloud data collected by LiDAR into a spatiotemporally constrained matrix representation. This spatiotemporally constrained matrix not only includes the spatial distribution of obstacles and road topology but also introduces a temporal dimension to express the movement trends of dynamic traffic participants, such as a 5m×5m×100ms spatiotemporally constrained matrix. The semantic features and spatiotemporal constraint matrix are then fed into the driving strategy tree generation module, which is essentially a structured decision logic generator that can infer reasonable driving strategies based on multimodal inputs and encode them into a compact and parsable scene DNA sequence. This sequence describes the core features and decision logic of the current scene in the form of binary or real number vectors. For example, "001011" represents a left-turn conflict scenario at an intersection.

[0067] It's important to note that the driving strategy tree generation module, as a core component of the multimodal semantic encoding layer, is responsible for converting the fused multimodal data into structured driving decision logic and outputting a standardized scene DNA sequence. This module first receives cross-domain semantic features from a federated learning framework, including vehicle trajectories, traffic light status, and weather data. Simultaneously, it integrates a gridded spatial model generated from laser point clouds via a topology mapping engine, for example, labeling obstacle positions and motion vectors with 5m×5m grid cells. Based on these inputs, the module constructs a hierarchical decision logic tree structure: the root node corresponds to high-level driving objectives, such as safely navigating intersections; branch nodes represent various decision conditions, such as traffic light status and obstacle distances; and leaf nodes correspond to the final action, such as braking or steering. During the decision tree's operation, the spatiotemporal constraint matrix provides physical constraint information to the nodes, such as the specific coordinates of obstacles within the grid, while semantic features inject environmental state information, such as automatically prioritizing braking operations in rainy conditions. The weights of each node in the tree can be dynamically adjusted based on real-time data to reflect the latest environmental state. When inconsistencies arise between input data, such as when the lidar detects an obstacle while the traffic light shows green, the module can activate a conflict resolution mechanism. By calculating the reliability weight of each data source under specific conditions—for example, the lidar confidence level might drop to 0.7 in rainy or foggy weather—the module dynamically selects the high-confidence data branch to adjust the decision path. Ultimately, the module encodes the entire decision path into a binary scene DNA sequence, where each bit represents the state of a key decision node in the policy tree. For example, the sequence "10110" can represent the complex decision logic of "green light, brake in rain, yield to pedestrians."

[0068] The dynamic evolution engine draws on the fundamental ideas of traffic flow field theory, mapping the aforementioned scenario's DNA-encoded sequence to traffic flow state parameters, such as vehicle density, velocity, or pressure fields. Based on this, it uses the Navier-Stokes equations or their simplified forms to calculate the entropy gradient field across the entire scenario, quantifying the information density or uncertainty level in different regions. Regions with high entropy values ​​typically correspond to traffic anomalies, sudden accidents, or complex interaction scenarios, and are key areas for testing autonomous driving systems. The entropy calculation formula is as follows: Where N is the number of observable objects in the scene, such as discrete states like vehicle trajectory, signal phase, and obstacle distribution; Let be the probability of the i-th state occurring, and let be the normalized scene entropy gradient value, with dimensions ranging from [0, 1]. The gradient direction points to the region where entropy increases the fastest, i.e., the high-entropy region, and the field strength represents the rate of change of the degree of disorder. For example, if the entropy value at a certain ramp accident point surges from 0.28 to 0.86, the gradient field shows that the field strength is highest in the northwest direction, indicating that this region is the center of sudden high disorder. The dynamic evolution engine dynamically generates a virtual "device scheduling field force vector" based on the strength and direction of the entropy gradient field. This vector can guide distributed data acquisition devices, such as mobile inspection vehicles or drones, to perform focused migration and data acquisition towards high-entropy regions, thereby achieving efficient allocation of test resources and targeted enhancement of the scene library.

[0069] The virtual-real symbiotic scene generator receives input from the dynamic evolution engine, including the spatial coordinates of high-entropy regions and a series of physical interference intensity parameters, such as visibility values, precipitation intensity levels, and radar wave attenuation coefficients. Utilizing 3D scene rendering technology combined with an environmental interference simulation model, the generator produces highly realistic controlled interference test scenarios in both visual and physical characteristics, such as road environments under extreme weather conditions like heavy rain, dense fog, and snow. Simultaneously, the generator introduces a group behavior uncertainty gradient. In the dynamic evolution engine, the quantification of this gradient is achieved through multi-dimensional behavioral entropy calculation. Specifically, it extracts dynamic features such as the rate of change of speed, acceleration standard deviation, and following distance fluctuation coefficient of traffic participants from a historical trajectory database to construct a spatiotemporal behavioral feature matrix; employs a sliding time window (default 5-second window / 1-second step size) to calculate individual behavioral entropy values ​​in real time; and uses the DBSCAN spatial clustering algorithm to identify cooperative group units with speed correlation > 0.8 and spatial distance < 10m, calculating the spatial gradient field of behavioral entropy within each unit. Finally, the gradient field is normalized to dimensionless parameters in the range [0,1]. Regions with gradient strength > 0.7 are marked as high-uncertainty interaction scenarios, such as emergency braking clusters and conflict lane-changing hotspots. The gradient vector field is output to the conditional adversarial generative network (GCN) to drive virtual traffic participants to inject group interaction rule parameters that conform to the physical characteristics of traffic flow. By injecting group interaction rule parameters that conform to the characteristics of real traffic flow through the GCN, complex interactive behaviors between traffic participants, such as vehicle following, lane changing, and yielding decision-making processes, are simulated. Ultimately, the system outputs a continuously dynamically updated scenario library and a comprehensive risk assessment report, which reflects the performance and potential failure modes of the autonomous driving system in different scenarios.

[0070] Compared to traditional static scenario library construction methods, this embodiment improves coverage of extreme operating conditions and long-tail scenarios through multimodal fusion and dynamic evolution mechanisms, avoiding testing blind spots caused by limited preset scenarios. Simultaneously, the application of a federated learning framework and various privacy enhancement technologies effectively prevents the leakage of sensitive information while supporting the collaborative use of multi-source data, solving the trust challenges inherent in data sharing within the industry. Furthermore, relying on physically realistic interference simulation and group behavior modeling, the generated test scenarios possess higher realism and challenge at both the perception and behavioral levels, more accurately reflecting the performance of autonomous driving systems in the real world, thus providing a more reliable basis for safety verification and performance optimization. Overall, this solution constructs a closed-loop system from data acquisition and scenario generation to test evaluation, supporting the continuous iteration and improvement of autonomous driving systems.

[0071] In one specific implementation, the federated learning framework establishes a three-level privacy filtering mechanism to collaboratively process the data streams of the onboard terminal's autonomous driving system and the V2X roadside equipment, specifically as follows:

[0072] 1) Terminal-level noise injection stage: Locally at the vehicle terminal and V2X roadside equipment, adaptive Gaussian noise is injected into the original sensor data stream through the differential noise injection module to generate an encrypted feature stream;

[0073] 2) Edge-level feature desensitization stage: The edge server receives encrypted feature streams from multiple terminals, uses knowledge distillation technology to remove device identity information and aggregate cross-domain features, and outputs the desensitized cross-domain scene feature stream to the multimodal semantic coding layer;

[0074] 3) Cloud-level security update phase: The model parameters of each terminal are aggregated on the cloud platform to generate a global model. The global weight parameter stream is protected by homomorphic encryption algorithm and distributed to the vehicle terminal and V2X roadside equipment to dynamically optimize their data acquisition capabilities. The vehicle terminal includes LiDAR, camera and millimeter-wave radar; the V2X roadside equipment includes roadside unit, smart traffic light and weather station.

[0075] In the above implementation, the system deploys a differential noise injection module in the local operating environment of various vehicle-mounted terminals and V2X roadside equipment. The core principle of this differential noise injection module is based on differential privacy technology. It dynamically adds adaptive noise conforming to a Gaussian distribution to the original sensor data stream, thereby encrypting and perturbing the original information. Specifically, the variance of the noise can be adaptively adjusted according to the data type and sensitivity. For example, for highly sensitive data such as precise vehicle location information, a relatively high noise variance can be selected, such as between 0.3 and 0.7; while for general road environment perception data, a lower noise intensity can be selected, such as a variance in the range of 0.1 to 0.3. In actual operation, the module receives raw data streams from sensors such as LiDAR, cameras, and millimeter-wave radar in real time, and injects noise without affecting subsequent feature extraction, thereby generating an encrypted feature stream. This encrypted data retains sufficient statistical properties for model training, but the original sensitive information cannot be reconstructed, achieving the security goal of "data usable but not visible."

[0076] In edge-level feature desensitization, the encrypted feature stream is transmitted to the edge server for further processing. The edge server employs knowledge distillation technology to aggregate and refine the encrypted features uploaded from multiple terminals, stripping away information that may contain device identification, such as device MAC addresses, serial numbers, or geographic coordinates. Simultaneously, this layer fuses data features from different domains; for example, it cross-domain correlates the point cloud features of vehicle-mounted LiDAR with meteorological information from roadside weather stations, ultimately outputting a cross-domain scene feature stream that maintains semantic integrity while completing the desensitization process. This process not only eliminates the risk of data tracing but also improves the efficiency and security of cross-device data collaboration. Specifically, after receiving the encrypted feature stream from multiple terminals, the edge server initiates a processing flow based on knowledge distillation technology. First, it strips the device identification information embedded in the data packets, such as removing elements that can directly or indirectly identify the terminal, like device MAC addresses and GPS coordinates, to completely block data tracing pathways. Building upon this foundation, the system further implements cross-domain feature aggregation, deeply fusing and semantically refining data from different sources and modalities, such as point cloud information from vehicle-mounted LiDAR and meteorological parameters collected by roadside weather stations, ultimately generating an anonymized cross-domain scene feature stream. For example, by fusing encrypted trajectory data from 10 vehicles with real-time meteorological information reported by 5 roadside weather stations, the system can output a composite feature description such as "northeast wind speed at the intersection 5 m / s + average vehicle speed 30 km / h." This feature stream fully preserves semantic validity while completely eliminating individual device identity information, thus achieving a balance between privacy protection and data utility.

[0077] During cloud-based security updates, the system aggregates model parameters from various terminal devices on the cloud platform and uses homomorphic encryption algorithms to encrypt and securely transmit the global model weight parameters. Homomorphic encryption allows model aggregation and update calculations to be performed in encrypted form, ensuring that model parameters are not leaked throughout the entire process of cloud processing and distribution back to terminal devices. Furthermore, this mechanism can dynamically optimize the data acquisition strategies of each terminal based on the performance of the global model, such as adjusting the sampling frequency or resolution of sensors, thereby continuously improving the overall perception and modeling capabilities of the system while protecting data privacy. The types of devices involved include LiDAR, cameras, and millimeter-wave radar in vehicle terminals, as well as roadside units, smart traffic lights, and weather stations in V2X roadside equipment.

[0078] This implementation method achieves full lifecycle security protection for multi-source data within the federated learning framework through a three-level collaborative privacy filtering mechanism at the terminal, edge, and cloud levels. It effectively prevents the risk of original data leakage and malicious attacks, while ensuring efficient fusion of cross-domain data and iterative optimization of models. It significantly enhances the reliability and compliance of sensitive data processing in autonomous driving systems, providing a solid security foundation for multi-entity data collaboration.

[0079] In one specific implementation, the driving strategy tree generation module includes a conflict resolution mechanism, specifically:

[0080] When lidar point cloud data and smart traffic light phase data conflict in the same spatiotemporal unit, a confidence fusion algorithm based on Bayesian inference is activated to calculate the reliability weight of each data source in the spatiotemporal unit. The same spatiotemporal unit meets the following conditions: time tolerance ≤ 100ms and spatial tolerance ≤ 1m.

[0081] Simultaneously, a dynamic topology correction engine is deployed. When temporary traffic control, road construction, or sudden factors at an accident scene are identified, the engine automatically reconstructs the road topology connection relationship and updates the branch weights of the driving strategy tree. Specifically, the dynamic topology correction engine reconstructs the road topology connection relationship by: a) loading a pre-set topology template based on the type of sudden factor; b) fusing real-time multimodal data to generate a semantic raster map; c) dynamically adjusting the topology connection matrix according to the semantic raster status; and d) updating the branch weight coefficients of the driving strategy tree according to the urgency of the event.

[0082] In the above implementation, for example, during the morning rush hour in a city, a sudden road construction operation occurs at a main road intersection due to a ruptured underground pipeline, causing dynamic changes in the traffic environment. The vehicle-mounted lidar continuously scans the surrounding environment at a frequency of 10Hz and identifies a traffic cone obstacle at coordinates X=125.7m, Y=8.3m. Simultaneously, the roadside intelligent traffic lights, due to communication delays, fail to update in time and continue to maintain a green light. Because the two signals differ by 83ms in time (less than the system's set 100ms timing tolerance) and 0.9m in spatial distance (within a 1m tolerance range), they are determined to be data conflicts within the same spatiotemporal unit.

[0083] The system then initiated a Bayesian confidence fusion algorithm to resolve conflicts: by retrieving historical operating logs of the equipment, it obtained reliability records for both the lidar and traffic lights. Data showed that the lidar had experienced three false alarms in rainy or foggy weather, causing its reliability baseline to drop from 0.88 to a real-time value of 0.75; while the traffic lights had a lower false alarm rate over the past 24 hours, maintaining a reliability of 0.92. Based on a probability model, weights were assigned: lidar data had a weight of 0.45, and traffic light data had a weight of 0.55. Based on the fusion results, the driving strategy tree generated a compromise instruction, outputting a decision branch coded as 01: "Decelerate to 20 km / h and observe."

[0084] Meanwhile, the topology map dynamic correction engine also responds to sudden road events. Roadside cameras capture "road construction" warning signs and worker gestures, triggering a four-level response mechanism: loading a preset construction scenario template, closing the right-turn lane and planning an alternative route; fusing millimeter-wave radar and camera data to generate a semantic raster map, marking the location of cone clusters and safe passage areas; dynamically adjusting the topology connection matrix, correcting the original right-turn lane vector from 60 degrees to approximately straight 5 degrees; and increasing the weight of the "lane keeping" strategy to 0.85 based on the urgency of the event, while simultaneously prohibiting automatic lane changing operations.

[0085] If a traditional synchronization scheme is used, the system will forcibly align data with a fixed timestamp, discarding conflicting data within the 83ms time difference between the lidar and the traffic light. This causes vehicles to accelerate forward based on incorrect green light instructions, only attempting to brake when they are about 10 meters away from the traffic cone, ultimately resulting in a collision due to insufficient braking distance. The static topology map scheme relies on pre-stored high-precision maps and cannot respond to sudden construction in real time. Vehicles will continue to drive along the original right-turn path, and by the time the lidar identifies the traffic cone at extremely close range, it is too late to avoid it. In contrast, this embodiment effectively utilizes the value of conflicting data through dynamic confidence fusion, avoiding the loss of key information, while simultaneously reconstructing the road topology in real time. This allows the driving strategy tree to adapt to sudden traffic conditions, ensuring the consistency of the decision chain in the spatiotemporal dimensions.

[0086] It should be noted that the conditions for the same spatiotemporal unit (time tolerance ≤ 100 ms, spatial tolerance ≤ 1 m) are the system default configuration, applicable to most urban roads and medium-speed scenarios. The system supports dynamically adjusting the tolerance threshold according to the actual scenario: in high-speed scenarios such as highways, the time tolerance can be compressed to ≤ 50 ms; in low-speed or static scenarios, the spatial tolerance can be relaxed to ≤ 2 m; the system can automatically adapt the tolerance parameters according to the sensor accuracy configuration file to ensure the accuracy and robustness of multi-source data fusion.

[0087] In one specific implementation, 3D scene rendering technology and an environmental interference simulation model collaboratively generate a controlled physical interference scene that conforms to physical laws; the environmental interference simulation model specifically includes:

[0088] Visual interference sub-model: Based on the physical properties of rain and fog particles, calculate the degree of light attenuation and color deviation caused by them to visible and near-infrared light;

[0089] Radar jamming sub-model: Based on electromagnetic wave characteristics, predict the distortion and position offset of lidar point cloud data in rain and fog environments;

[0090] Among them, the 3D scene rendering technology combines the pixel-level degradation parameters output by the visual interference sub-model with the point cloud distortion parameters output by the radar interference sub-model to generate a physically realistic controlled physical interference scene.

[0091] Meanwhile, the conditional adversarial generative network introduces group behavior rules to adjust the interactive decision-making mechanism of traffic participants and generate group interaction rule parameters.

[0092] In the above embodiments, the visual interference sub-model simulates the impact of atmospheric particles such as rain and fog on the propagation of visible and near-infrared light based on their optical and physical properties. This model calculates the scattering and absorption of specific wavelengths of light by raindrops or fog droplets of different sizes, concentrations, and shapes, deriving the attenuation coefficient and color deviation value of light after passing through the interference medium. For example, under simulated rain conditions, the model may calculate that the attenuation rate of green light near a wavelength of 550 nanometers is approximately 50%-70%, and due to the combined effects of Rayleigh scattering and Mie scattering, the surface color of objects may generally lean towards a cool tone, with a hue shift of approximately 10-20 degrees. The radar interference sub-model, based on the propagation characteristics of electromagnetic waves in rain and fog environments, predicts the scattering, reflection, and attenuation phenomena that occur when the laser beam emitted by the lidar encounters precipitation particles, thereby inferring potential problems such as density reduction, positional drift, and shape distortion in point cloud data. This model typically considers multiple parameters such as radar wavelength, precipitation intensity, and particle dielectric constant, outputting the expected offset vector and confidence reduction coefficient for each point in the point cloud. For example, under heavy rain conditions, the ranging error of a 905nm lidar may increase by tens of centimeters, and the point cloud density may decrease by about 20%-40%. During operation, these two sub-models continuously receive real-time or preset meteorological parameter inputs, such as precipitation, fog concentration, and wind speed, and output a series of physical parameters that can be used for rendering and sensor simulation, providing a data foundation for constructing highly realistic interference environments.

[0093] The 3D scene rendering engine receives pixel-level optical attenuation parameters and color correction matrices generated by the visual interference sub-model, as well as point cloud distortion parameters and position corrections provided by the radar interference sub-model, and maps these physical parameters to the corresponding elements in the virtual environment. For example, when rendering a rainstorm scene, the 3D scene rendering engine dynamically adjusts the brightness, contrast, and hue of the entire scene based on the attenuation data output by the visual model to simulate the visual effect of light passing through the rain curtain. Simultaneously, it perturbs the original point cloud data of the virtual LiDAR based on the point cloud distortion model, making it reflect the typical error characteristics of real radar in rain. To achieve sensor consistency, the rendering process must also ensure that the object outlines, positions, and motion states reflected by the visual image and the radar point cloud are physically consistent, avoiding logical contradictions where the image is optically visible but undetectable by the radar. The entire rendering process typically runs in real-time or near real-time, and the interference intensity can be dynamically adjusted during testing, such as gradually increasing the rainfall or fog concentration, to verify the perception robustness of the autonomous driving system under different interference levels.

[0094] Generative Adversarial Networks (GANs) receive the uncertainty gradient of group behavior calculated by a dynamic evolutionary engine and combine it with predefined or learned traffic rules and behavioral patterns to generate interaction parameters that regulate the decision-making logic of virtual traffic participants. These interaction parameters may include adjustment coefficients for vehicle following distance and delay times for lane-changing decisions. Their values ​​are often set based on historical real traffic data or expert rules. For example, in low visibility conditions, the following distance may be reduced by 20%-30% compared to normal weather conditions, and the probability of the driver taking emergency braking increases accordingly. Through network training and conditional control, dynamic elements such as vehicles and pedestrians in the virtual environment no longer move according to simple rules, but can exhibit complex behaviors similar to human drivers, such as risk perception, conservative decision-making, and interactive avoidance, thus creating more realistic and challenging scenarios in testing.

[0095] This implementation method generates highly realistic and challenging test scenarios at both the perception and behavior levels by closely combining physical-level sensor interference simulation with behavioral-level decision rule injection. It significantly improves the test coverage and verification depth of autonomous driving systems under extreme weather and complex interaction conditions. It provides a complete and dynamically adjustable technical path to solve problems such as perception distortion, logical inconsistency, and overly idealized behavior that often exist in virtual test environments, thereby enhancing the reliability and effectiveness of autonomous driving simulation testing.

[0096] In one specific implementation, the group behavior rules generate group interaction rule parameters through the following steps:

[0097] 1) Behavioral decision modeling: Extract the historical movement trajectories of traffic participants and analyze their speed changes, direction adjustments, and interactive avoidance habits;

[0098] 2) Environmental Status Perception: Real-time monitoring of traffic light status, road congestion levels, and the location of surrounding road users;

[0099] 3) Conflict resolution mechanism: When the movement paths of multiple traffic participants conflict, right-of-way is dynamically allocated according to the priority of traffic regulations;

[0100] 4) Interaction parameter generation: Integrate behavioral decision-making habits with environmental conditions to output probability parameters of interaction actions such as acceleration, deceleration or turning.

[0101] In the above implementation, behavioral decision modeling refers to continuously collecting and analyzing the historical movement trajectories of traffic participants (such as vehicles and pedestrians) to extract micro-features in their driving or movement behaviors, such as speed change patterns, frequency and magnitude of direction adjustments, and interactive avoidance habits in different traffic situations. The system typically records these trajectories with high spatiotemporal precision, for example, collecting raw data with a spatial resolution of 0.1-1 meters and a temporal precision of 10-100 milliseconds. It then uses machine learning methods such as clustering, Hidden Markov Models, or Long Short-Term Memory (LSTM) networks to identify typical behavior categories and their probabilities from historical data. Environmental state perception is responsible for acquiring dynamic information in the traffic scene in real time, including the current phase and remaining duration of traffic lights, road congestion levels (e.g., using vehicle density or average speed as indicators), and the precise location and movement status of surrounding traffic participants. These environmental parameters can be acquired in real-time or near real-time via onboard sensors, roadside units, or V2X communication equipment and input into the group behavior rule generation module in a unified format. Behavioral modeling and environmental perception work together: behavioral models provide prior knowledge for environmental perception. For example, after identifying that a vehicle has a habit of frequently changing lanes, the system can predict its next possible behavior and adjust the interaction strategies of surrounding vehicles in advance. Meanwhile, the real-time environmental state provides contextual constraints for behavioral decisions, ensuring that the generated behavioral rules conform to the current actual traffic conditions.

[0102] When the system detects that the predicted movement paths of multiple traffic participants intersect or overlap, indicating a potential conflict, it will initiate a conflict resolution procedure based on rules or learning algorithms. This mechanism first assigns dynamic weights to each participant according to traffic regulations and traffic priorities (such as pedestrians having the highest right-of-way and emergency vehicles having priority). Then, it calculates the optimal right-of-way allocation scheme by combining factors such as the participant's current speed, acceleration, and time to the conflict point. For example, at an intersection without traffic lights, the system may assign higher priority to vehicles approaching from the right; while at pedestrian crossings, it forces vehicles to yield to pedestrians. In actual operation, this process is typically executed cyclically at millisecond intervals, continuously updating the conflict decision results based on the latest sensor data. To achieve smooth and safe interaction, the system can also introduce randomness or fuzzy logic to simulate the uncertainties in human driving, such as allowing a certain probability of yielding or preemptive action in certain edge scenarios, thus making the generated interaction rules more realistic and diverse.

[0103] This interaction parameter generation stage integrates the aforementioned behavioral models, environmental states, and conflict resolution results, ultimately outputting a set of parameters that can be used to control the behavior of virtual traffic participants. These parameters are typically expressed in the form of probability distributions or policy functions, such as the probability values ​​of a vehicle choosing to accelerate, decelerate, steer, or maintain its original speed in different scenarios. The generation process often relies on supervised learning or reinforcement learning methods, training the generative model with a large amount of real driving data so that its output parameters not only conform to traffic rules but also reflect regional driving cultures or individual differences in habits. For example, in some traffic flows, vehicles may maintain a short following distance when following another vehicle, and this system can capture such features and represent them as a higher deceleration probability parameter. In addition, the generation module also has online adaptive capabilities, dynamically adjusting the parameter distribution based on test feedback to gradually approximate real-world group interaction patterns.

[0104] This implementation method can enhance the realism and complexity of traffic participant behavior in virtual testing scenarios. By using a data-driven approach, it can recreate the habits, risk perception, and regulatory compliance characteristics in human driving decisions, thereby providing a more reliable and effective testing environment for autonomous driving systems and enhancing their verification effectiveness and decision robustness in real and complex interactive scenarios.

[0105] In one specific implementation, the dynamic evolution engine includes a version compatibility verification layer to handle compatibility issues with non-traditional traffic participant data. When importing non-traditional traffic participant data, a scene migration verification test is first run in an isolated sandbox. By comparing the degree of difference between the DNA encoding of the new and old versions of the scene, it is determined whether the topology mapping rules need to be reconstructed. If the degree of difference exceeds a preset threshold, the incremental learning module is activated to fine-tune and optimize the encoder, while retaining the parallel call interface of the historical version scene library to ensure the continued compatibility of the old version scene library.

[0106] In the above implementation, when new types of smart road signs and other non-traditional traffic participants are introduced into the urban road network, such as interactive road signs with dynamic projection functions, traditional static scene libraries, due to their fixed data structures, are often unable to parse the data fields brought by the new devices, such as projection height or flashing frequency. The emergence of this type of new data may cause the parsing logic of the original scene library to fail, thereby affecting the stability and availability of the entire testing system.

[0107] To address this issue, this embodiment introduces a version compatibility verification layer into the dynamic evolution engine. This mechanism first performs scene migration verification tests on the new road sign data stream in an isolated sandbox. By comparing the differences in scene DNA encoding between the old and new versions, it identifies changes at the data structure level. For example, the added projection height parameter of the new road sign expands the dimension of the spatiotemporal constraint matrix. If the dimensional conflict rate exceeds a preset threshold of 30%, it is determined that the encoding rules need to be adjusted.

[0108] The system then triggers an incremental learning mechanism to perform lightweight fine-tuning on the multimodal semantic coding layer. While maintaining the original core logic of laser point cloud topology mapping, a new parsing branch for parameters such as projection height is added, enabling the encoder to adapt to new data inputs. Simultaneously, the system retains the parallel call interface for historical version scene libraries, ensuring that test systems that have not yet been upgraded can continue to use scene data in the original format, avoiding system interruptions due to upgrades.

[0109] After optimization, the reconstructed encoder has the ability to output two versions. The new version of the scene DNA coding sequence can include extended fields with new parameters, such as identifying dynamic projection regions with specific binary sequences; the old version of the code will automatically remove the new fields through the interface conversion layer to generate a simplified coding sequence that is compatible with the old system.

[0110] This implementation method enables the autonomous driving scenario library to expand its ability to analyze new data features when introducing new traffic elements such as smart road signs, while effectively ensuring the continuity and stability of the original testing system, thereby supporting the smooth evolution and technological iteration of traffic infrastructure.

[0111] In one specific implementation, the distributed data acquisition layer enables a data completion mechanism in areas with missing roadside data. Specifically, based on the scene entropy distribution of the acquired areas, a spatiotemporal adversarial generative network is trained to synthesize simulated data that conforms to the traffic flow characteristics of the target area. At the same time, a data authenticity verification unit is set up to dynamically optimize the training parameters of the generative network by comparing the distribution differences between the generated data and the real data at the driving strategy decision node, thereby filling the data gaps in areas with weak infrastructure.

[0112] In the above implementation, when the distributed data acquisition layer identifies data gaps in certain areas due to weak infrastructure or insufficient equipment coverage, the system initiates a data completion process. This process first trains a dedicated spatiotemporal adversarial generative network (GDN) based on the scene entropy distribution characteristics of the acquired areas. Scene entropy is used to quantify the information complexity or uncertainty in the traffic environment; for example, areas with sudden traffic flow changes, frequent abnormal behaviors, or unusual weather patterns often have high entropy values. The system extracts typical features from these high-entropy areas, such as vehicle trajectory patterns, speed distributions, or traffic light variation patterns, as learning targets for the generative network. The generative network typically employs an encoder-decoder structure, where the encoder learns the latent distribution characteristics of real data, and the decoder generates synthetic data based on random noise or conditional input. During training, the generative network attempts to synthesize simulated data that is statistically highly similar to real data, such as generating virtual vehicle trajectories or event sequences that conform to local traffic flow characteristics. The entire training process may last from several hours to several days, depending on the size and complexity of the data gap areas. The training objective is to minimize the distribution differences between the generated data and the real data in key features.

[0113] The data authenticity verification unit is responsible for quality assessment and feedback adjustment of the synthetic data output by the generator network. During the verification process, the system compares the performance differences between the generated data and real data at driving strategy decision nodes from multiple dimensions, such as the probability of vehicle deceleration at intersections and the timing of lane-changing decisions. These decision nodes typically correspond to key behaviors or state transition moments in traffic scenarios and can effectively reflect the authenticity and rationality of the data. In the data authenticity verification unit, a significance threshold of p-value < 0.05 is set, and the Kolmogorov-Smirnov test is used to test the difference in distribution between the generated data and real data at driving strategy decision nodes. If the test result shows a p-value below 0.05, the distribution difference is determined to be significant, and the training parameters of the generator network need to be dynamically adjusted. This feedback adjustment process is usually run online or near online, forming a closed-loop optimization of generation and verification, gradually improving the quality and reliability of the synthetic data.

[0114] In practical applications, this data completion process is automatically triggered when the system detects missing or insufficient data in a certain area. First, the system extracts relevant data from neighboring already collected areas, calculates the scene entropy distribution, and configures the initial parameters of the generation network accordingly. Then, the generation network begins synthesizing simulated data, while a realism verification unit continuously evaluates the generated results. After multiple rounds of iterative optimization, the synthesized data that meets the quality requirements is injected into the scene library to enhance the test coverage of that area. For example, in the case of a newly built intersection lacking real data, the system may generate vehicle trajectory data simulating traffic flow at different times of day, including dense traffic patterns during the morning rush hour or sparse traffic patterns at night, enabling the autonomous driving system to conduct more comprehensive and effective testing in this virtual environment. The entire completion process is automated as much as possible to minimize human intervention and ensure the efficiency and consistency of data generation.

[0115] This implementation method can improve the coverage and realism of the autonomous driving test scenario library in areas with missing data. Through intelligent generation and iterative optimization mechanisms, it can effectively make up for the data gaps caused by infrastructure limitations, providing a more comprehensive and reliable test environment for autonomous driving systems, thereby enhancing their adaptability and safety in actual deployment.

[0116] For example, a newly built road in a new industrial park creates a data collection blind spot with a radius of 800 meters. Traditional linear interpolation methods cannot accurately reflect the high-frequency turning characteristics of freight traffic, causing virtual vehicles to mechanically repeat historical trajectories at T-junctions. Therefore, this embodiment employs a data completion mechanism in the distributed data acquisition layer. The dynamic evolution engine first extracts the scene entropy distribution of neighboring already collected areas, identifying the entropy peak around the freight hub in the industrial park due to frequent lane changes by trucks. Based on this, a spatiotemporal adversarial generative network is trained: the generator learns key features such as sharp turns and temporary stops by trucks in high-entropy areas, while the discriminator compares the curvature of the real trajectory with the generated data, ultimately outputting a synthetic data stream containing turning angle events of 15° or more, consistent with the actual behavior of freight vehicles. This is followed by a dynamic verification phase. The data authenticity verification unit monitors the driving strategy decision nodes and finds discrepancies between the generated data's decision distribution at intersections without traffic lights and the actual situation; for example, the actual deceleration probability of a truck passing through is 92%, while the generated data only shows 78%. By optimizing the generated network weights in reverse, the risk perception parameters of uncontrolled intersections are improved. After iterative optimization, the probability of trucks slowing down in the generated data increases to 89%, significantly approaching the true statistical characteristics.

[0117] In one specific implementation, the risk assessment report generation step specifically includes:

[0118] 1) Perform multi-agent reinforcement learning risk assessment based on a dynamically updated scenario library, and load the driving strategy confusion index assessment model into the test sandbox;

[0119] 2) Quantify the failure probability of autonomous driving systems under cognitive uncertainty;

[0120] 3) Generate dynamically updated risk assessment reports based on failure probability.

[0121] The driving strategy confusion index evaluation model includes a cognitive load monitoring module, which monitors three types of decision-making anomalies in real time: the frequency of strategy tree branch switching exceeds the safety threshold; the output of control commands fluctuates abnormally under the same input conditions; and the satisfaction rate of the spatiotemporal constraint matrix drops sharply.

[0122] When the cognitive load monitoring module detects any abnormal decision-making indicator that triggers an alarm, it automatically starts the scenario backtracking mechanism to compare and analyze the current decision chain with the historical best decision-making strategy to locate the root cause of the decision abnormality.

[0123] In the above implementation, for example, an autonomous driving system faces extreme conditions of continuous alternation between light and dark during testing in a mountainous tunnel complex. Traditional risk assessment methods only count the number of collisions, making it difficult to quantify the cognitive confusion risk caused by the system losing GPS signals. Therefore, this embodiment implements a complete cognitive risk quantification and diagnosis process.

[0124] First, a multi-agent reinforcement learning test is initiated. A virtual adversarial environment of tunnel clusters is loaded into the dynamic scene library to simulate positioning drift caused by strong light interference at the tunnel exit and signal shielding at the entrance. Simultaneously, the driving strategy confusion index evaluation model is activated, with the cognitive load monitoring module tracking the decision chain status in real time. It should be noted that the driving strategy confusion index is a comprehensive indicator that quantifies whether an autonomous driving system is "confused" or "decisive." It is calculated by monitoring three key behaviors: whether strategy switching is too frequent, whether control commands are inconsistent, and whether the system's understanding and compliance with environmental rules drops sharply. Its threshold setting is mainly based on two aspects: first, a reasonable benchmark for human driving behavior, such as no more than 10 decisions per second; and second, the stability range demonstrated by the system in a large number of historical normal tests. When the driving strategy confusion index exceeds the threshold, it indicates that the system may face the risk of failure due to its inability to understand the current scenario, thus triggering an early warning mechanism. The driving strategy confusion index is calculated by weighted fusion of three types of decision anomaly indicators, specifically: Index = w1×f1 + w2×f2 + w3×f3, where f1 is the policy tree branch switching frequency (unit: times / hundred milliseconds), f2 is the standard deviation of the control command output under the same input (normalized), f3 is the spatiotemporal constraint matrix satisfaction rate (value range [0,1]), and w1, w2, and w3 are weight coefficients, usually w1=0.4, w2=0.3, and w3=0.3.

[0125] The system then proceeds to the cognitive anomaly monitoring and failure quantification phase. This module detected three key anomalies: the frequency of policy tree branch switching surged in the light-dark transition zone, exceeding the safety threshold of 5 times per 100 milliseconds; the steering wheel angle command fluctuation amplitude was abnormal under the same lighting conditions, with its standard deviation exceeding twice the baseline; and the spatiotemporal constraint matrix satisfaction rate plummeted by more than 30% during signal loss periods and persisted for over 500 milliseconds. Based on the simultaneous occurrence of these anomalies, the system comprehensively calculated the failure probability under cognitive uncertainty.

[0126] Based on this, a backtracking and report generation mechanism is triggered. The system compares the current decision chain with the historical best strategy to pinpoint the root cause of the failure: during the signal loss period, the positioning module continuously output outdated coordinates, causing repeated switching of path planning branches. The final risk assessment report clearly identifies "positioning drift causing spatiotemporal matrix failure" as the core issue and outputs a dynamically updated failure probability curve.

[0127] Traditional testing methods have significant drawbacks, relying solely on repeated testing in fixed scenarios and recording collision results. When a test vehicle experiences a minor collision with the tunnel wall due to positioning drift, the system simply marks it as a collision, failing to reveal the abnormal switching frequency of path planning branches in the decision chain. This leads researchers to mistakenly attribute the problem to camera malfunction.

[0128] The innovative advantages of this embodiment are reflected in three dimensions: quantifying risk in the cognitive dimension through dynamic indicators such as branch switching frequency; using a backtracking mechanism to compare historical best strategies, breaking the black box of decision-making and accurately locating the root cause; and forming a closed-loop diagnostic process, so that the risk assessment report is directly linked to the defects of specific modules, providing clear guidance for targeted optimization.

[0129] This method successfully identified the true root cause of failure in mountain tunnel testing as the positioning module failing to switch to a backup strategy in time when the signal was lost, rather than a visual perception defect. This breakthrough overcomes the technical limitation of traditional statistical reports, which cannot link cognitive processes with physical failures.

[0130] In one specific implementation, the dynamically updated scenario library drives the iterative optimization of the autonomous driving system through a closed-loop feedback mechanism of testing and verification, specifically including:

[0131] The dynamically updated scenario library is input into the autonomous driving system test platform to perform stress tests and collect system response data;

[0132] Based on system response data, coordinates of blind spots in scene coverage and characteristics of strategy defects are generated.

[0133] The coordinates of the blind spots and the characteristics of strategy defects are fed back to the dynamic evolution engine;

[0134] The dynamic evolution engine drives the distributed data acquisition device to enhance data acquisition in areas corresponding to the coordinates of the coverage blind spots and the characteristics of the strategy defects, and drives the virtual and real symbiotic scene generator to generate adversarial test scenarios targeting the characteristics of the strategy defects.

[0135] Ultimately, this forms a closed-loop iterative chain encompassing scenario library generation, system testing, defect feedback, and dynamic optimization.

[0136] In the above implementation, during the autonomous driving system verification phase, a dynamically updated scenario library is input into the test platform to perform stress tests on highway merging zones. For example, the test simulates a scenario where heavy trucks frequently merge into the main road, and system response data is collected. Analysis reveals that when a truck merges at an angle greater than 30 degrees and a speed difference exceeding 15 km / h, the tested system exhibits trajectory prediction deviations, leading to delayed braking command triggering. Based on this, the following blind zone coordinates are generated: a 300-meter range from the merging zone entrance; and the strategy defect characteristic is that the trajectory prediction confidence decreases by 40% when the angle of entry exceeds 30 degrees and the speed difference is greater than 15 km / h.

[0137] The aforementioned blind spot coordinates and defect characteristics are fed back to the dynamic evolution engine. The engine immediately calculates the scene entropy gradient field of the area, identifies the high information entropy characteristics of the heavy truck's merging behavior, and generates a device scheduling field force vector. Three nearby mobile data acquisition vehicles, driven by the field force, focus on the target merging area to enhance data acquisition, capturing in real time the steering wheel angle characteristics and throttle fluctuation patterns of truck drivers during emergency merging. Simultaneously, the virtual-real symbiotic scene generator receives defect characteristic parameters and generates a targeted adversarial scenario: setting up multiple trucks in the virtual merging area to continuously merge with a 32-35 degree merging angle and a speed difference of 18-20 km / h, and injecting parameters of the aggressive driving habits of drivers collected in reality.

[0138] The updated scenario library includes enhanced real-world data and generated adversarial scenarios, which were then tested on the test platform again. This time, the system revealed a new flaw: the path planning module frequently switched lane-keeping strategies in aggressive entry scenarios. This feature was fed back to the evolutionary engine, driving the generator to construct extreme adversarial scenarios involving truck convoys blocking lanes. After four rounds of closed-loop iteration, the system's false positive rate in merging zone testing decreased significantly.

[0139] Current technologies employ fixed scenario libraries for testing; for example, a mainstream solution uses 100 hours of pre-recorded highway data to generate a test set. When misjudging merging zones is discovered, similar scenarios can only be manually selected for retesting. Because the original data does not contain enough entry points of 30 degrees or more, effective adversarial scenarios cannot be generated. Engineers need to spend weeks re-collecting road data, and the supplementary scenario library is disconnected from the system testing process, making it difficult to continuously track the evolution of defects.

[0140] This embodiment employs a closed-loop feedback mechanism, enabling the defect features exposed during testing to directly drive the data acquisition equipment to migrate to high-entropy regions and generate adversarial scenarios targeting policy weaknesses in real time. This forms an automated iterative chain, where system test defects are instantly transformed into scenario library optimization instructions. The optimized scenario library then verifies the system improvement effect, overcoming the bottleneck of traditional open-loop testing where scenario updates lag behind system defects.

[0141] The number of devices and processing scale described herein are for the purpose of simplifying the description of the invention. Applications, modifications, and variations of the invention will be readily apparent to those skilled in the art.

[0142] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details.

Claims

1. A method for dynamic generation of an automatic driving scene library based on multi-modal data fusion, characterized in that, include: A distributed data acquisition layer is established, which accesses the data streams of the vehicle terminal autonomous driving system and V2X roadside equipment through a federated learning framework with time buffering, and uses a streaming processing pipeline to achieve time decoupling between data acquisition and model training. A multimodal semantic encoding layer is constructed, and multimodal data streams from vehicle terminals and V2X roadside devices are fused through a federated learning framework to extract cross-domain semantic features. Simultaneously, LiDAR point cloud data is converted into a spatiotemporal constraint matrix through a topology graph mapping engine. The semantic features and spatiotemporal constraint matrix are input into a driving strategy tree generation module. The driving strategy tree generation module constructs a hierarchical decision logic tree structure based on the semantic features and spatiotemporal constraint matrix. The root node corresponds to a high-level driving objective, branch nodes represent decision conditions, and leaf nodes correspond to execution actions. The spatiotemporal constraint matrix provides physical constraint information for the decision nodes. At the same time, the weights of the branch nodes can be dynamically adjusted with real-time data to reflect changes in environmental state. Finally, the entire decision path is encoded into a scene DNA encoding sequence. Deploy a dynamic evolution engine to map DNA coding sequences to traffic flow state parameters through traffic flow field theory, and calculate scene entropy gradient field based on Navier-Stokes equations. Based on the gradient field strength, the device scheduling field force vector is dynamically generated to drive the distributed data acquisition device to focus and migrate to the high information entropy region. The virtual-real symbiotic scene generator is activated, receiving the high-entropy region coordinates and physical interference intensity parameters output by the dynamic evolution engine. The physical interference intensity parameters include visibility value, precipitation intensity level and radar wave attenuation coefficient. A controlled physical interference test scene of the target area is generated using 3D scene rendering technology. At the same time, the uncertainty gradient of group behavior is calculated based on the dynamic evolution engine, and group interaction rule parameters that conform to the characteristics of traffic flow field are injected through a conditional adversarial generative network. The final output includes a dynamically updated scenario library and a risk assessment report.

2. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 1, characterized in that, The federated learning framework establishes a three-level privacy filtering mechanism to collaboratively process the data streams from the vehicle-mounted autonomous driving system and the V2X roadside equipment, specifically as follows: 1) Terminal-level noise injection stage: Locally at the vehicle terminal and V2X roadside equipment, adaptive Gaussian noise is injected into the original sensor data stream through the differential noise injection module to generate an encrypted feature stream; 2) Edge-level feature desensitization stage: The edge server receives encrypted feature streams from multiple terminals, uses knowledge distillation technology to remove device identity information and aggregate cross-domain features, and outputs the desensitized cross-domain scene feature stream to the multimodal semantic coding layer; 3) Cloud-level security update phase: The model parameters of each terminal are aggregated on the cloud platform to generate a global model. The global weight parameter stream is protected by homomorphic encryption algorithm and distributed to the vehicle terminal and V2X roadside equipment to dynamically optimize their data acquisition capabilities. The vehicle terminal includes LiDAR, camera and millimeter-wave radar; the V2X roadside equipment includes roadside unit, smart traffic light and weather station.

3. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 2, characterized in that, The driving strategy tree generation module includes a conflict resolution mechanism, specifically: When lidar point cloud data and smart traffic light phase data conflict in the same spatiotemporal unit, a confidence fusion algorithm based on Bayesian inference is activated to calculate the reliability weight of each data source in the spatiotemporal unit. The same spatiotemporal unit meets the following conditions: time tolerance ≤ 100ms and spatial tolerance ≤ 1m. Simultaneously, a dynamic topology correction engine is deployed. When temporary traffic control, road construction, or sudden factors at an accident scene are identified, the engine automatically reconstructs the road topology connection relationship and updates the branch weights of the driving strategy tree. Specifically, the dynamic topology correction engine reconstructs the road topology connection relationship by: a) loading a pre-set topology template based on the type of sudden factor; b) fusing real-time multimodal data to generate a semantic raster map; c) dynamically adjusting the topology connection matrix according to the semantic raster status; and d) updating the branch weight coefficients of the driving strategy tree according to the urgency of the event.

4. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 1, characterized in that, A 3D scene rendering technology and an environmental interference simulation model work together to generate a controlled physical interference scene that conforms to physical laws; the environmental interference simulation model specifically includes: Visual interference sub-model: Based on the physical properties of rain and fog particles, calculate the degree of light attenuation and color deviation caused by them to visible and near-infrared light; Radar jamming sub-model: Based on electromagnetic wave characteristics, predict the distortion and position offset of lidar point cloud data in rain and fog environments; Among them, the 3D scene rendering technology combines the pixel-level degradation parameters output by the visual interference sub-model with the point cloud distortion parameters output by the radar interference sub-model to generate a physically realistic controlled physical interference scene. Meanwhile, the conditional adversarial generative network introduces group behavior rules to adjust the interactive decision-making mechanism of traffic participants and generate group interaction rule parameters.

5. The method for dynamically generating an autonomous driving scenario library based on multimodal data fusion as described in claim 4, characterized in that, The group behavior rules generate group interaction rule parameters through the following steps: 1) Behavioral decision modeling: Extract the historical movement trajectories of traffic participants and analyze their speed changes, direction adjustments, and interactive avoidance habits; 2) Environmental Status Perception: Real-time monitoring of traffic light status, road congestion levels, and the location of surrounding road users; 3) Conflict resolution mechanism: When the movement paths of multiple traffic participants conflict, right-of-way is dynamically allocated according to the priority of traffic regulations; 4) Interaction parameter generation: Integrate behavioral decision-making habits with environmental conditions to output probability parameters of interaction actions such as acceleration, deceleration or turning.

6. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 1, characterized in that, The dynamic evolution engine includes a version compatibility verification layer to handle compatibility issues with non-traditional traffic participant data. When importing non-traditional traffic participant data, a scene migration verification test is first run in an isolated sandbox. By comparing the differences in the DNA encoding of the new and old versions of the scene, it is determined whether the topology mapping rules need to be reconstructed. If the difference exceeds a preset threshold, the incremental learning module is activated to fine-tune and optimize the encoder. At the same time, the parallel call interface of the historical version scene library is retained to ensure the continued compatibility of the old version scene library. The non-traditional traffic participants include smart road signs and robots.

7. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 1, characterized in that, The distributed data acquisition layer enables a data completion mechanism in areas with missing roadside data, specifically as follows: Based on the scene entropy distribution of the collected area, a spatiotemporal adversarial generative network is trained to synthesize simulated data that conforms to the traffic flow characteristics of the target area; Simultaneously, a data authenticity verification unit is set up. By comparing the distribution differences between generated data and real data at driving strategy decision nodes, the training parameters of the generated network are dynamically adjusted, thereby filling the data gaps in areas with weak infrastructure.

8. The method for dynamically generating an autonomous driving scenario library based on multimodal data fusion as described in claim 1, characterized in that, The specific steps involved in generating a risk assessment report include: 1) Perform multi-agent reinforcement learning risk assessment based on a dynamically updated scenario library, and load the driving strategy confusion index assessment model into the test sandbox; 2) Quantify the failure probability of autonomous driving systems under cognitive uncertainty; 3) Generate dynamically updated risk assessment reports based on failure probability.

9. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 8, characterized in that, The driving strategy confusion index assessment model includes a cognitive load monitoring module, which monitors three types of decision-making anomaly indicators in real time: the frequency of policy tree branch switching exceeds the safety threshold; the output of control commands fluctuates abnormally under the same input conditions; and the spatiotemporal constraint matrix satisfaction rate drops sharply. Specifically, when the cognitive load monitoring module detects any decision-making anomaly indicator that triggers an alarm, it automatically initiates a scenario backtracking mechanism to compare and analyze the current decision chain with the historical best decision strategy to locate the root cause of the decision anomaly.

10. The automatic driving scene library dynamic generation method based on multi-modal data fusion according to claim 1, characterized in that, The dynamically updated scenario library drives iterative optimization of the autonomous driving system through a closed-loop feedback mechanism of testing and verification, specifically including: The dynamically updated scenario library is input into the autonomous driving system test platform to perform stress tests and collect system response data; Based on system response data, coordinates of blind spots in scene coverage and characteristics of strategy defects are generated. The coordinates of the blind spots and the characteristics of strategy defects are fed back to the dynamic evolution engine; The dynamic evolution engine drives the distributed data acquisition device to enhance data acquisition in areas corresponding to the coordinates of the coverage blind spots and the characteristics of the strategy defects, and drives the virtual and real symbiotic scene generator to generate adversarial test scenarios targeting the characteristics of the strategy defects. Ultimately, this forms a closed-loop iterative chain encompassing scenario library generation, system testing, defect feedback, and dynamic optimization.

Citation Information

Patent Citations

  • Automatic driving scene library construction method and device

    CN117874454A

  • Intelligent automobile edge scene generation method based on hierarchical reinforcement architecture

    CN119167802A