Self-driving automobile design operation domain boundary delimiting method and device, terminal and medium
By combining the adaptive control variable method and active sampling strategy with the tensor coupled field model, the shortcomings of existing technologies in evaluating the performance boundary of autonomous driving systems under multi-element interactions are addressed. This enables comprehensive performance evaluation of complex driving scenarios and improves the safety and reliability of the system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to systematically assess the performance boundaries of autonomous driving systems under the interaction of various driving environment elements, leading to an increase in unknown and unsafe scenarios.
The boundary of a single element is determined by an adaptive control variable method and a threshold judgment condition. Combined with an active sampling strategy and a tensor coupled field model, the boundary of the design operation domain of multiple elements is delineated by a boundary symbol distance network model.
It enables comprehensive performance boundary assessment of complex real-world driving scenarios with multiple elements superimposed, improving the safety and reliability of autonomous driving systems in complex environments.
Smart Images

Figure CN121835008A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of automatic driving, in particular to a self-driving car design operation domain boundary demarcation method and device, terminal and medium. BACKGROUND
[0002] The design operation domain (Operational Design Domain, ODD) is like the "capability specification" or "safety boundary" of the automatic driving system, which contains the specific road types, weather conditions, traffic conditions and other environments in which the system can safely perform driving tasks. Clear ODD definition is to clarify these restrictions to prevent the system from "overcoming the odds" in scenarios beyond its processing capacity, thereby avoiding possible safety accidents.
[0003] Accurate definition of ODD is crucial for achieving the expected functional safety (SOTIF). The focus of SOTIF standards is not on traditional hardware failures, but on risk definition due to system performance limitations or human-machine interaction problems.
[0004] An important goal of SOTIF activities is to reduce "unknown unsafe scenarios". A clear definition of ODD can help the development team focus on analysis and systematically identify scenarios within the ODD that may become dangerous due to insufficient performance. This provides a clear target for subsequent testing and verification.
[0005] ODD is the basis for generating a test scenario library. In order to verify the SOTIF of the system, it is necessary to ensure that the test cases can fully cover various situations within the ODD, especially those "edge scenarios" or "critical scenarios" that are prone to challenge the performance boundaries of the system. For example, if the ODD contains "rain" conditions, the test must cover the system's performance under different rainfall intensities.
[0006] Many intelligent automobile manufacturers have defined ODD for their products, but there is a lack of systematic comprehensive evaluation. The existing ODD definition method is often insufficient for testing a single environmental factor, and there is a clear gap in the current industry in defining the interaction of multiple boundary elements. In actual driving environments, multiple ODD elements often appear at the same time (such as rainy + night + construction area), but existing technologies have not provided a systematic method to evaluate the comprehensive boundaries in these complex scenarios. SUMMARY
[0007] In view of the deficiencies in the prior art, the present application proposes a self-driving car design operation domain boundary demarcation method, device, terminal and medium, which aims to evaluate the comprehensive performance boundaries of complex real driving scenarios with multiple elements.
[0008] In a first aspect, the embodiments of the present application provide a self-driving car design operating domain boundary demarcation method, which comprises the following steps: obtaining operating domain element information, wherein the operating domain element information comprises a plurality of operating domain elements, a value range corresponding to each operating domain element, and a value that can be taken by each operating domain element; determining a single-element boundary threshold value of each operating domain element by using an adaptive control variable method and a threshold value determination condition, and obtaining a single-element boundary information set; obtaining a preliminary test step length according to the single-element boundary information set, and interacting with a controlled test field according to the preliminary test step length by using an active sampling strategy to obtain a test point data set; inputting the test point data set into a pre-trained tensor coupling field model as training data to train the tensor coupling field model and obtain a training data set, and inputting the training data set into a pre-trained boundary signed distance network model to train the boundary signed distance network model; demarcating a design operating domain boundary of a plurality of elements according to the test point data set and the boundary signed distance network model.
[0009] In a second aspect, the embodiments of the present application provide a self-driving car design operating domain boundary demarcation device, which comprises the following modules: An information obtaining module is configured to obtain operating domain element information, wherein the operating domain element information comprises a plurality of operating domain elements, a value range corresponding to each operating domain element, and a value that can be taken by each operating domain element; An element boundary determining module is configured to determine a single-element boundary threshold value of each operating domain element by using an adaptive control variable method and a threshold value determination condition, and obtain a single-element boundary information set; A test data determining module is configured to obtain a preliminary test step length according to the single-element boundary information set, and interact with a controlled test field according to the preliminary test step length by using an active sampling strategy to obtain a test point data set; A model determining module is configured to input the test point data set into a pre-trained tensor coupling field model as training data to train the tensor coupling field model and obtain a training data set, and input the training data set into a pre-trained boundary signed distance network model to train the boundary signed distance network model; An operating domain boundary determining module is configured to demarcate a design operating domain boundary of a plurality of elements according to the test point data set and the boundary signed distance network model.
[0010] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the self-driving car design operating domain boundary demarcation method according to any one of the first aspect.
[0011] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the self-driving car design operating domain boundary demarcation method according to any one of the first aspect.
[0012] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when executed on a terminal device, causes the terminal device to perform the self-driving car design operating domain boundary demarcation method according to any one of the first aspect.
[0013] In the embodiment of the present application, by obtaining the operating domain element information, the operating domain element information includes a plurality of operating domain elements, a value range corresponding to the operating domain elements, and a value that can be taken corresponding to the operating domain elements; the adaptive control variable method and the threshold determination condition are used to determine the single element boundary threshold value of the operating domain elements through the value range and the value that can be taken, to obtain a single element boundary information set; according to the single element boundary information set, a preliminary test step is obtained, and through an active sampling strategy, the preliminary test step is interacted with a controlled test field to obtain a test point data set; the test point data set is input into a pre-trained tensor coupling field model as training data for training, to obtain a tensor coupling field model and a training data set, and the training data set is input into a pre-trained boundary symbolic distance network model for training, to obtain a boundary symbolic distance network model; according to the test point data set and the boundary symbolic distance network model, a design operating domain boundary of a plurality of element superpositions is demarcated. The comprehensive performance boundary of the real driving scene of the complex superposition of a plurality of elements is evaluated. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed in the specific embodiments or prior art description will be briefly introduced below. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, each element or part is not necessarily drawn according to the actual scale.
[0015] Figure 1 is a flowchart of the first embodiment of the self-driving car design operating domain boundary demarcation method provided by an embodiment of the present application; Figure 2 is a structural schematic diagram of the self-driving car design operating domain boundary demarcation device provided by an embodiment of the present application; Figure 3 FIG. 1 is a structural schematic diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0016] Figure 1 FIG. 2 shows a flowchart of a first embodiment of a self-driving car design operation domain boundary demarcation method provided by an embodiment of the present application, which is shown by way of example but not limitation, and as shown in FIG. 2, the method can include: Figure 1 S10, obtaining operation domain element information, the operation domain element information including a plurality of operation domain elements, a value range corresponding to the operation domain elements, and a value that can be taken corresponding to the operation domain elements; Wherein, the value range and the value that can be taken include the measurement unit. If the operation domain element is rainfall intensity, the measurement unit is millimeter / hour (mm / h), and the value range is 5-150 mm / h; the value that can be taken can be 5, 10, 15, 20, …, 145, 150 mm / h.
[0017] Operation domain elements, ODD elements (operation domain elements) are divided into three categories, and the quantitative parameter of each element is defined: (1) landscape elements: quantified by high-precision maps and positioning systems, such as road curvature radius (≥100 m), slope (≤5%), lane width (≥2.8 m), etc.; (2) environmental conditions: quantified by vehicle-mounted sensors and meteorological data, such as rainfall intensity (0-150 mm / h), visibility (≥50 m), light intensity (0.1-100 k lux), etc.; (3) dynamic elements: quantified by perception systems, such as traffic density (0-1 vehicle / square meter), pedestrian density (0-0.1 person / square meter), etc.
[0018] Wherein, when defining the boundary of a single element, a step test method is used to determine the boundary threshold of a single ODD element (operation domain element). Taking rainfall boundary definition as an example: (1) test environment setting, using an artificial rainfall simulation system to generate different rainfall intensities from 5 mm / h to 150 mm / h in a closed test site; (2) performance index monitoring, the vehicle is tested and the key performance indicators are recorded under each rainfall level, including: ① perception system performance, target detection accuracy and maximum detection distance of camera, lidar, radar; ② positioning system accuracy: position error, heading error; ③ control system performance: lateral control error, longitudinal control stability; (2) threshold determination, when any key performance indicator decays more than 30% or the system issues a takeover request, it is determined that the boundary threshold of the element is reached.
[0019] A single ODD element boundary definition example is shown in Table 1.
[0020] Table 1 In defining the boundary of a single element, a meta-reinforcement learning-based active sampling strategy is adopted to replace the traditional exhaustive or orthogonal experiment method for the complex case of multiple ODD elements superimposed, thereby efficiently focusing on the test near the performance boundary.
[0021] S20, using the adaptive control variable method and threshold determination conditions, the boundary threshold of a single element is determined for the running domain element through the value range and the possible values, to obtain a single element boundary information set; As one implementation method, an adaptive control variable method and threshold determination criteria are used to determine the single-element boundary threshold of the running domain elements based on the value range and the possible values, resulting in a single-element boundary information set, including: A1, by setting the test environment for each running domain element according to the value range and the possible values, a test environment information set is obtained; As one implementation method, a test environment is set for each runtime element based on the value range and the possible values, resulting in a test environment information set, including: B1, set the test environment for the value range corresponding to a run domain element and the run domain element corresponding to the possible value pair, and obtain the test environment information; Taking the "rainfall intensity" test environment as an example: a closed test site is selected, equipped with an artificial rainfall simulation system capable of generating controllable rainfall of 5-150 mm / h. The test road should be a standard highway section (straight sections + curved sections) to avoid interference from other complex factors. The test vehicle is equipped with a complete autonomous driving system, including sensors such as cameras, LiDAR, and radar, as well as a complete computing and control unit. A benchmark measurement system (such as RTK-GPS, motion capture system, etc.) is also installed to obtain the vehicle's real-world status as a comparison benchmark.
[0022] B2 will combine the test environment information obtained by setting the test environment for all runtime domain elements to obtain a test environment information set.
[0023] A2, under control conditions, the test vehicle is operated to obtain benchmark information on performance indicators; Then, baseline testing was conducted: five tests were performed under rainless conditions, and the baseline values of various system performance indicators were recorded to obtain performance indicator baseline information. A3. Control the operation of the test vehicle according to each possible value corresponding to the first operating domain element in the test environment information set, and obtain the first key indicator information corresponding to the first operating domain element. That is, the test vehicle is controlled to run according to each possible value corresponding to the first running domain element in the test environment information set, and the first key indicator information corresponding to the first running domain element is obtained when the test vehicle is running. Taking the performance data collection of "rainfall intensity" as an example, the following data was collected to obtain the first key indicator information: At each rainfall level, the following tests were conducted and data recorded: perception performance test, standard target objects (vehicles, pedestrian dummies, etc.) were placed on the test road section, and the detection distance, recognition accuracy and false alarm rate of the system were measured; positioning performance test, the deviation between the vehicle position and the reference trajectory was measured; control performance test, lane keeping error, following distance error, etc. were measured. Repeated testing: Perform 3-5 repeated tests for each rainfall level to reduce the impact of random errors; Data recording: Record auxiliary parameters such as rainfall intensity, duration, ambient temperature, and wind speed for each test.
[0024] A4. Based on the threshold judgment conditions, the first key indicator information and the performance indicator benchmark information, the design operating domain boundary of the first operating domain element is determined, and the first boundary information corresponding to the first operating domain element is obtained. As one implementation method, the design operating domain boundary of the first operating domain element is determined based on threshold judgment conditions, first key indicator information, and performance indicator benchmark information, to obtain the first boundary information corresponding to the first operating domain element, which may include: C1. Based on the measured performance values in the key indicator information and the performance indicator benchmark information, obtain the first performance indicator decay value corresponding to the first running domain element. C2, the first system failure probability is calculated based on the first performance index decay value; C3, based on the threshold judgment conditions, the failure probability of the first system, and key indicator information, obtains the boundary information of the first element.
[0025] The detailed steps of steps C1 to C3 correspond to the detailed steps of steps D1 to D3.
[0026] A5. Control the operation of the test vehicle according to each possible value corresponding to the second operating domain element in the test environment information set, and obtain the second key indicator information corresponding to the second operating domain element. That is, the test vehicle is controlled to run according to each possible value corresponding to the second running domain element in the test environment information set, and the second key indicator information corresponding to the second running domain element is obtained when the test vehicle is running. A6. Based on the threshold judgment conditions, the second key indicator information, and the performance indicator benchmark information, the design operating domain boundary of the second operating domain element is determined, and the second boundary information corresponding to the second operating domain element is obtained. As one implementation method, the design operating domain boundary of the second operating domain element is determined based on the threshold determination condition, the second key indicator information, and the performance indicator benchmark information, to obtain the second boundary information corresponding to the second operating domain element, including: D1. Based on the measured performance values in the key indicator information and the performance indicator benchmark information, obtain the second performance indicator decay value corresponding to the second running domain element. The calculation steps for the performance metric degradation degree (second performance metric degradation value) are as follows: This parameter is used to quantify the percentage decrease in system performance relative to its optimal state. Its calculation is based on the test procedure established in the "Single Element Boundary Definition Implementation".
[0027] Determine performance metrics and baseline values: First, conduct baseline tests under undisturbed or optimal conditions (e.g., a daytime day without rain) to measure baseline values for key system performance metrics (such as sensing distance). Target recognition accuracy wait).
[0028] Measurements are taken under test conditions: Then, tests are conducted under specific ODD element conditions (such as specific rainfall intensity) to obtain the measured values of performance indicators under those conditions (e.g., , ).
[0029] Calculating the attenuation rate: The formula for calculating the degree of performance index attenuation (second performance index attenuation value) is as follows: For example, the degree of perceived distance attenuation. ; For metrics such as accuracy, the calculation formula can be used. ; D2, based on the decay value of the second performance index, calculates the failure probability of the second system; As one implementation method, calculating the second system failure probability based on the second performance index decay value may include: A multi-indicator comprehensive judgment combined with a risk probability model is adopted: in This represents the system failure probability (which can be either the second system failure probability or the first system failure probability). The degree of performance index decay (which can be the decay value of the second performance index or the decay value of the first performance index). As the baseline attenuation threshold, This is the sensitivity coefficient.
[0030] The determination of the baseline attenuation threshold (I0) follows these steps: This is a preset, maximum acceptable performance attenuation limit for the system. It is not calculated using a single formula, but rather is an engineering safety margin determined comprehensively based on system safety requirements, regulatory standards, historical test data, and expert review. For example, based on functional safety analysis (such as ISO 26262) and expected functional safety (SOTIF, ISO 21448) analysis, it can be determined that when the attenuation of a certain performance (such as forward target detection) exceeds 30%, the remaining risk of the system will exceed the acceptable range. In Table 1, "Example of Boundary Definition for a Single ODD Element," the "≤30% attenuation," "≥95%," etc., in the "Boundary Threshold" column represent preset I0 (or their equivalents) for different performance indicators. During boundary determination, the calculated real-time attenuation level I is compared with the preset I0.
[0031] D3. Based on the threshold judgment conditions, the failure probability of the second system, the safety score of the second system, and the key indicator information, the boundary information of the second element is obtained.
[0032] The boundary threshold is determined to be reached when any of the following conditions are met, that is, when the threshold determination condition is met: Key performance indicators degrade by more than 30%: for example, the sensing distance drops from a baseline of 100m to below 70m; System security score below threshold: According to the preset security evaluation model, the score is below the security limit; The system issues a takeover request: the autonomous driving system actively requests manual intervention; The probability of system failure exceeds a preset risk threshold (e.g., 0.1%).
[0033] The boundary threshold is determined as the lowest rainfall intensity at which performance is acceptable. For example, if the system performs well at 100 mm / h but fails at 150 mm / h, the boundary threshold is determined as a value between 100 and 150 mm / h, which can be determined through interpolation.
[0034] The system safety score S-Score is calculated as follows: based on a "layered fusion evaluation network." This network receives real-time raw data streams and intermediate state confidence scores from multiple modules, including perception, prediction, planning, control, and system monitoring. The input layer includes, but is not limited to: the detection confidence score and trajectory prediction uncertainty of the perception module for various target types (vehicles, pedestrians); the size of the covariance ellipse of the localization module; the trajectory cost function value and conflict detection results with traffic rules of the planning module; the tracking error of the control module; and the system's own CPU / memory load and sensor data validity flags. The processing layer maps each input indicator to a "sub-health" score between 0 and 1 using a non-linear normalization function, where 1 represents optimal. Crucially, this mapping function is not linear but designed based on the indicator's impact on safety; for example, a decrease in pedestrian detection confidence is penalized more severely. The fusion layer employs an attention-weighted fusion algorithm. The system dynamically adjusts the weights of each indicator according to the current scenario (e.g., urban roads, highways). For example, in urban road scenarios, the weights of "pedestrian detection health" and "traffic signal recognition health" are automatically increased. Finally, through weighted aggregation, a comprehensive S-Score between 0 and 100 is output. Output and Application: This real-time calculated S-Score is directly used for boundary determination. When the S-Score falls below a preset "safe operating threshold" (e.g., 80 points), the system determines that the current operating condition is approaching or exceeding its dynamic ODD boundary, triggering an early warning or degradation strategy. This enables the system to make a holistic and intelligent safety assessment of complex and coupled environmental impacts.
[0035] A7. After all the possible values in all the running domain elements have been tested by vehicle operation, the boundary information corresponding to all the running domain elements is combined to obtain a single element boundary information set.
[0036] S30, a preliminary test step size is obtained based on the single element boundary information set, and a test point dataset is obtained by interacting with the controlled test field based on the preliminary test step size through an active sampling strategy. The test step size includes a state space, an action space, and a reward function. The state space includes the currently explored parameter point, the system performance index corresponding to the currently explored parameter point, and an uncertainty estimate. The system performance index is predicted by the system performance model corresponding to the currently explored parameter point.
[0037] As one implementation method, a preliminary test step size is obtained based on the single element boundary information set, and a test point dataset is obtained by interacting with a controlled test field based on the preliminary test step size using an active sampling strategy. Specifically, this includes: E1, based on the set of boundary information of the individual elements, the preliminary state space and the preliminary action space are obtained; E2, based on the initial state space, uses the acquisition function to select the action with the largest acquisition function value in the initial action space as the initial optimal candidate action; E3: Run the initial optimal candidate action in the test environment and observe the initial system performance index; add the initial optimal candidate action and the initial system performance index as new data points to the currently explored parameter points to obtain the currently explored parameter points for the first iteration; update the state space and action space based on the explored parameter points of the first iteration and the boundary information set of individual elements to obtain the state space and action space of the first iteration; obtain the initial reward function based on the initial optimal candidate action, the initial state space, and the state space of the first iteration; and train the acquisition function using the initial reward function to obtain the acquisition function for the first iteration. When updating the state space, the action space A defines the combination of test parameters that the agent can choose (e.g., [rainfall intensity = 75 mm / h, light intensity = 20 lux]). This selection range cannot be infinite; it must be constrained by physical reality and test safety. The "single element boundary information set" provides a safe range of values or boundary thresholds for each ODD element (e.g., rainfall intensity, light intensity) determined through preliminary testing. When updating the action space, this information set is used to constrain the parameter values of candidate actions (at), ensuring that the agent explores only within meaningful and safe parameter ranges, avoiding the selection of parameter points that clearly exceed the system's capabilities and may lead to dangerous or ineffective testing.
[0038] When updating the state space, the state space... Includes "system performance indicators corresponding to each point" To evaluate a new test point performance Or its measured performance All of these require a benchmark value to determine whether their performance is "good" or "bad," "safe," or "close to the limit." This benchmark value (e.g., optimal performance under undisturbed conditions) is crucial. or safety threshold The information is obtained from the "single element boundary definition" stage and is summarized in the "single element boundary information set". Therefore, when updating the state space, calculating the degree of performance degradation, or determining whether it is close to the boundary, it is necessary to query the single element boundary information set.
[0039] E4, based on the state space of the first iteration, uses the acquisition function of the first iteration to select the action with the largest acquisition function value in the action space of the first iteration as the optimal candidate action for the first iteration; E5. Run the optimal candidate action of the first iteration in the test environment and observe the system performance index of the first iteration. Add the optimal candidate action and the system performance index of the first iteration as new data points to the currently explored parameter points to obtain the currently explored parameter points of the second iteration. Update the state space and action space of the first iteration based on the explored parameter points of the second iteration and the boundary information set of individual elements to obtain the state space and action space of the second iteration. Obtain the reward function of the first iteration based on the optimal candidate action, the state space of the first iteration, and the state space of the second iteration. Train the acquisition function of the first iteration using the reward function of the first iteration to obtain the acquisition function of the second iteration. E6, until the updated uncertainty estimate is less than the preset uncertainty threshold, all currently explored parameter points and the corresponding system performance indicators are used as test point datasets.
[0040] Among them, the uncertainty convergence condition of the model (such as Gaussian process) (the updated uncertainty estimate is less than the preset uncertainty threshold) specifically includes the following: (1) the average prediction variance in the entire action space. Or, within a specific sub-region of interest (such as near the boundary), the prediction variance of the model (such as a Gaussian process) for all candidate points. The average value is below a certain threshold (2) Maximum prediction variance: The largest prediction variance among all candidate points is below the threshold, which ensures that there are no cognitive blind spots; (3) Information gain saturation: In multiple consecutive iterations, the information gain brought by new sampling points (such as...) It becomes very small, below a certain threshold.
[0041] As one implementation method, until the updated uncertainty estimate is less than a preset uncertainty threshold, all currently explored parameter points and the corresponding system performance indicators are used as a test point dataset, which may include: Until the performance boundary clarity meets the preset accuracy, all currently explored parameter points and the corresponding system performance indicators are used as the test point dataset.
[0042] The condition for achieving satisfactory performance boundary sharpness (meeting the preset accuracy) is that the accuracy of boundary positioning meets the requirements. This can be achieved by checking the distance between the boundary symbols and the network. Judged by the accuracy of the prediction: (1) On a set of known boundary points of the validation set obtained by other methods (such as high-precision simulation or a small amount of dense sampling), Predicted signed distance Distance from reality The error (such as root mean square error RMSE) is below a predetermined threshold.
[0043] (2) Boundary surface The shape remains stable with minimal changes throughout multiple iterations.
[0044] As one implementation method, until the updated uncertainty estimate is less than a preset uncertainty threshold, all currently explored parameter points and the corresponding system performance indicators are used as a test point dataset, which may include: The training continues until all elements in the single-element boundary information set have completed multi-element training. All currently explored parameter points and the corresponding system performance indicators are then used as the test point dataset.
[0045] Real-world driving environments are complex systems resulting from the nonlinear superposition of multiple ODD elements (such as rainfall A, nighttime illumination B, and high traffic density C). Existing methods, such as static scene complexity assessment based on the Analytic Hierarchy Process (AHP), primarily focus on subjectively weighting and ranking pre-defined scene combinations. Essentially, this is a "post-event evaluation" of discrete scenes, rather than an "active definition and exploration" of the system performance "safety boundary" within a continuous parameter space. The core innovation of this application lies in abandoning this discrete and subjective evaluation approach and instead employing a systematic method based on high-dimensional performance mapping and implicit boundary surface extraction to directly define and generate the composite boundary of ODDs under multi-element superposition.
[0046] The core process of this method is divided into three levels: data-driven performance space sampling, construction and modeling of continuous performance fields, and implicit extraction and representation of safety boundary surfaces.
[0047] Among them, the data-driven performance space adaptive sampling strategy includes: To avoid exhaustive testing of the n-dimensional parameter space consisting of n ODD elements (the combinatorial explosion problem), this invention proposes a meta-reinforcement learning-guided active sampling strategy. This strategy models the testing process as a sequential decision problem, where the agent adaptively selects the next test point with the highest information gain by interacting with the environment (an autonomous driving system in a controlled testing environment). , , ,...) The state space S is: the currently explored parameter region, the system performance index P corresponding to each point, and its uncertainty estimate.
[0048] Action space A is: selecting the coordinates of the next test point in the n-dimensional ODD parameter space, which determines the next specific combination of ODD parameters for testing.
[0049] The reward function R is carefully designed to simultaneously encourage exploration into regions of high model prediction uncertainty and to utilize boundary regions where prediction performance approaches a critical threshold. in, It is the boundary proximity function, which has a significant impact on prediction performance. Approaching the safety threshold High rewards are given at times; It is the information gain function, used to calculate new states based on a Bayesian model (Gaussian process). Compared to existing datasets The cognitive uncertainty that can be reduced; and This is the balancing coefficient. By maximizing the cumulative reward, the agent can efficiently draw "contour lines" near the performance boundary, accurately capturing the boundary shape with far fewer tests than in orthogonal experiments.
[0050] In the active sampling strategy guided by meta-reinforcement learning, the state space S is a composite state that encapsulates the current "knowledge state" of the testing process, including: the explored parameter region, the system performance index P corresponding to each point, and the uncertainty estimate.
[0051] The explored parameter region refers to the set of all points (Ai, Bi, Ci, ...) that have been tested so far in a multidimensional parameter space composed of multiple ODD elements (such as rainfall A, illumination B, traffic density C). This defines the known "map" for intelligent vehicles.
[0052] Here, the system performance index P corresponding to each point is given. For each explored parameter point, there is a corresponding measurement result of one or more system performance indices (such as sensing distance, control error, etc., which can be summarized into the aforementioned "system safety score" or "failure probability"). This reflects the mapping relationship between parameters and performance.
[0053] Uncertainty estimation, a key aspect of active learning, quantifies the confidence level of an intelligent vehicle in predicting performance in unknown regions (untested combinations of parameters). Within the framework of a Gaussian process-based equal-probability model, this "uncertainty" can be represented as the variance of the predicted values or the width of the confidence interval. Regions with high uncertainty indicate that the model's performance predictions for those regions are inaccurate and require close attention.
[0054] Therefore, state At time t, the information consists of three parts: the set of tested points, the performance observations of these points, and the uncertainty distribution of the performance prediction over the entire parameter space.
[0055] Among them, the boundary proximity function Intended to predict performance Approaching the safety threshold At this time, the agent is given a higher reward to incentivize it to explore the vicinity of the boundary of the Operational Design Domain (ODD). The formula is: in, It is the system performance predicted at a specific parameter point by the current performance field model (such as a Gaussian process), which can be an indicator such as safety score, success rate or failure probability. It is a preset system performance security threshold, such as the minimum acceptable security score (e.g., 70 points).
[0056] This represents the absolute distance between the predicted performance and the safety threshold. The physical meaning of this value is "distance from the boundary"; the smaller the value, the closer the test point is to the system's performance boundary. It is an exponential decay function. Its function is to map the absolute distance mentioned above to a reward value between 0 and 1. When the distance is 0 (i.e., the predictive performance is exactly equal to the safety threshold), the function reaches its maximum value of 1, and the reward is the highest; as the distance increases, the reward value decays exponentially to near 0. It is a hyperparameter greater than 0, used as a scaling factor to control the rate at which rewards decay with distance. The larger the value, the more sensitive the reward is to approaching the boundary, and the faster it decays.
[0057] The steps for calculating the boundary proximity function include: (1) performance prediction, in time step Intelligent cars based on the current state (This includes the currently trained performance field model), for candidate actions The system is evaluated using the parameter coordinates of the next test point (i.e., the planned sampling point). The agent uses the model to predict the system's performance at the corresponding parameter point after executing the action, thus obtaining... (2) Calculate the boundary reward, and use the reward obtained in the previous step. With known Substitute into the formula to calculate the boundary proximity reward. (3) To synthesize the total reward, the calculated boundary proximity reward is weighted and summed with other reward signals during the exploration process (e.g., rewards based on information gain or model uncertainty) to form the total reward. This total reward is used to evaluate the quality of the action and guide the selection of the next optimal test point.
[0058] This design encodes the exploration goal of "approaching the performance boundary" directly into a quantifiable reward signal. Combined with rewards such as information gain, it can effectively guide the exploration of unknown areas and prioritize the exploration of critical areas that are crucial to system security.
[0059] In meta-reinforcement learning-guided active sampling strategies, uncertainty estimation is crucial for guiding agents to explore unknown regions. It is not a single, fixed formula, but a dynamic computational process based on a probabilistic model. It utilizes a model (such as Gaussian process regression) that outputs predicted values and their uncertainties to quantify the "degree of unknown" regarding the performance of untested parameter points. The computational steps for uncertainty estimation include: (1) Model selection and initialization: Gaussian process was selected as the surrogate model. During initialization, the model was trained based on a small number of initial test points and their corresponding measured performance values. These data served as the prior knowledge base of Gaussian process.
[0060] (2) Uncertainty calculation, for any untested parameter point (i.e., a vector consisting of multiple ODD element values), the Gaussian process will provide its performance prediction value. (Mean) and forecast uncertainty (Variance). This prediction variance. This is an uncertainty estimate, which directly reflects the degree of lack of understanding of the point by the model.
[0061] Its calculation formula originates from the derivation of the posterior distribution of Gaussian processes. Given a set of training data... ,in To observe noise, the Gaussian process hypothesis function It follows a prior distribution: in It is the mean function (set to zero). It is the covariance function (kernel function): here, It is the signal variance, which controls the amplitude of the function output; It is a length scale that controls the smoothness of the function.
[0062] For a new point Its posterior prediction distribution is also a Gaussian distribution, and its variance (i.e., uncertainty) has a closed-form solution as follows: The covariance of a new point itself is usually a constant. ).
[0063] New Point The covariance vector between all training points X, i.e. K: The covariance matrix between training points, whose elements .
[0064] The variance of the observed noise is learned from the data.
[0065] I: Identity matrix.
[0066] Physical meaning: Formula (1) clearly reveals the source of uncertainty. The first term... It is a priori uncertainty. The second term... Measured by the observation of training data And thus, reduced uncertainty. The larger the value, the more likely it is to be a new point. The less similar a point is to all tested points (i.e., the farther away it is or the more it is in a region that the model has not fully learned), the more uncertain the model's performance prediction for that point will be. These are precisely the areas that need to be prioritized for exploration to gather information and reduce cognitive blind spots.
[0067] (3) In the state space The embodiment and active sampling in meta-reinforcement learning In this context, "uncertainty estimation" is not a single numerical value, but rather a continuous predictive variance field or cognitive uncertainty map depicted by the current surrogate model (Gaussian process) over the entire high-dimensional parameter space.
[0068] At each decision step, the agent can query any candidate action (i.e., candidate test points). The corresponding forecast uncertainty In reward function design, this uncertainty is directly used as an important component of information gain or exploration reward. For example, a common reward term design is as follows: Or after normalization: in, It is a trade-off coefficient. It is a small constant used for numerical stability.
[0069] By maximizing the cumulative reward that includes such exploration rewards, agents are systematically guided to test the points where the current model is most uncertain, thereby efficiently reducing cognitive uncertainty about the entire performance field, especially the region near the performance boundary, with the fewest number of tests.
[0070] Among them, candidate actions are selected in the action space. It is a policy-based decision-making process generated by the agent through interaction with the environment (i.e., the performance of the autonomous driving system in a controlled test environment). It is not randomly selected, but generated through a specific optimization algorithm. This process is the core of the meta-reinforcement learning-guided active sampling strategy, aiming to efficiently explore the high-dimensional parameter space and approach the performance boundary with the fewest possible tests.
[0071] Candidate Action The essence is that it is made by Select a specific coordinate point in a continuous multidimensional parameter space composed of ODD elements (such as weather, illumination, road type, traffic density, etc.). For example, at a time step... A candidate action can be represented as: This vector defines the next specific combination of environmental parameters to be tested.
[0072] Candidate Action The selection method is based on the optimization of the acquisition function, which is a function that takes the current state as an input. (Including predictions and uncertainty estimates from the current performance field model) is mapped to each potential action. The "expected utility" is a function. The agent selects the action that maximizes the value of the acquisition function. ,Right now: in, It is the action space. This refers to the data acquisition function. A typical and widely used data acquisition function is the Upper Confidence Bound (UCB). Its formula is: : Predicted by the current performance field model in action The average performance at the corresponding parameter points (i.e., the "utilization" item).
[0073] In action The predicted standard deviation at the corresponding parameter points is the square root of the uncertainty (i.e., the "exploration" term).
[0074] : A hyperparameter greater than 0 used to balance "exploration" and "exploitation". A large value encourages exploration of regions with high uncertainty. big); If the value is small, then the use of regions with high predictive performance is encouraged. big).
[0075] Selection process: At each step, the agent searches for the optimal parameter within the feasible region of the parameter space using optimization algorithms (such as selecting the optimal parameter after random sampling or gradient-based optimization). The largest point This point is the selected candidate action. .
[0076] To determine a specific The following core parameters and inputs are required: (1) Current state This forms the basis for decision-making and includes: Explored point set: All tested parameter points .
[0077] Performance observations: Measured system performance values corresponding to the explored points. .
[0078] Performance prediction model: A model trained on current data (such as a Gaussian process) that can predict any candidate point. Provide predicted mean and standard deviation .
[0079] (2) Acquisition function and its hyperparameters: Define the exploration strategy. Besides UCB, there are expected improvement, probabilistic improvement, etc. The balance coefficient in the formula... It is a key hyperparameter that can be fixed or dynamically adjusted according to the training phase.
[0080] (3) Boundary constraints of the parameter space: the physical or logical test range of each ODD element. For example, rainfall intensity might be constrained to... Light intensity at Candidate actions, etc. Selection must be made within this constraint.
[0081] (4) Optimization algorithm: used in continuous, high-dimensional action space Internally, efficiently search for and maximize the acquisition function. The point. Common methods include: Random sampling and selection: Randomly sample a large number of points within the feasible region, calculate their sampling function values, and select the optimal one. This is simple but may not be efficient.
[0082] Gradient-based optimization: If the acquisition function is favorable to... Differentiability (e.g., when using a Gaussian process and the kernel function is differentiable) allows for more accurate optimization using gradient ascent.
[0083] Heuristic or evolutionary algorithms, such as CMA-ES, are suitable for non-convex or non-differentiable cases.
[0084] Candidate Action The selection process is not isolated, but rather embedded in a closed loop of "interaction-learning-re-decision": (1) Observation and modeling: The agent is based on the current state (Including historical test data and models), the optimal candidate action is calculated using a data acquisition function (such as UCB). .
[0085] (2) Execution and Interaction: Performing actions in the test environment. This means configuring the test scenario according to the parameter combination, running the autonomous driving system, and observing its performance indicators. .
[0086] (3) Update and learn: update new data points Add the dataset and retrain or update the performance prediction model (e.g., update the posterior distribution of the Gaussian process). This changes the model's perception of the space (i.e., updates the...). and ), thus entering the next state. .
[0087] (4) Loop: In a new state Next, repeat step 1 and select the next action. .
[0088] Among them, the total reward It is applied in the "policy update" step of the reinforcement learning training loop. Its core function is to evaluate the quality of the actions already performed and use this evaluation to update the agent's decision-making policy (i.e., action selection function), enabling it to make better decisions in the future. This process is the core mechanism by which reinforcement learning agents learn from the environment and optimize their behavioral policies.
[0089] (1) Total Reward Roles and Composition In the active sampling strategy, the total reward It is a carefully designed scalar signal used to quantize states. Next action The "good" or "bad" of a reward is determined by a weighted sum of multiple sub-rewards, guiding the agent to achieve multi-objective optimization. A typical configuration is as described previously: in: It uses boundary proximity rewards to encourage agents to explore prediction performance. Approaching the safety threshold The area.
[0090] It is an information gain reward that encourages agents to explore areas of high model cognitive uncertainty in order to quickly reduce the unknown.
[0091] and It is a hyperparameter that balances the weights of the two rewards.
[0092] This overall reward design transforms the high-level goal of "exploring performance boundaries efficiently and accurately" into specific signals that the agent can directly optimize at each time step.
[0093] Total Rewards The application is embedded in a standard reinforcement learning "interaction-learning" loop. We will explain the application process in detail using algorithms based on value functions or policy gradients as examples.
[0094] Step 1: Perform the action and observe the results. (At the time step...) The agent, based on its current policy (For example, an exploration strategy based on upper confidence bounds (UCB)) from the action space Select and execute an action In our scenario, this corresponds to, in the test environment, according to the action... The defined combination of ODD parameters (e.g., [rainfall intensity = 75 mm / h, light intensity = 20 lux]) runs an autonomous driving test. The environment (i.e., the autonomous driving system in a controlled test range) receives this action, executes the test, and returns two results: (1) Next state This includes updated information, such as new test points. and the measured system performance After adding the dataset, the new cognitive state is obtained by retraining or updating the performance field model.
[0095] (2) Instant rewards This immediate reward is the total reward calculated by the reward function we designed. It quantifies the value of this test—did it approach the boundary? Did it significantly reduce uncertainty? Step Two: Store the experience. Store the experience tuple from this interaction. The data is stored in a data structure called an ExperienceReplayBuffer. This buffer acts like a memory, storing the historical trajectory of the agent's interactions with the environment. Using experience replay can break the temporal correlation between data, improve learning efficiency, and reuse valuable interaction data.
[0096] Step 3: Policy Update. This is the core step where the total reward comes into play. The agent periodically (e.g., after collecting a certain amount of new experience) randomly samples a batch of experience data from the experience replay buffer. For each sampled experience, we use it to update a parameterized function called the Q-network. The network is used to estimate the state. Take action below The expected cumulative discount return that can be obtained.
[0097] (1) Calculate the target Q value: According to the Bellman optimality equation, the current state-action pair The "real" long-term value should equal the immediate reward plus the maximum future value of the next state. Therefore, we calculate the objective Q value. : in: It is the immediate reward obtained from sampling, i.e. .
[0098] It is a discount factor used to balance the importance of current rewards and future rewards.
[0099] It is an estimate of a target network, whose parameters Regularly from the main network It was copied and is used for stable training. Indicates the next state The maximum expected return that can be obtained from all possible actions.
[0100] (2) Calculate the loss function: The loss function measures the current prediction value of the Q network. With target value The difference between them is usually expressed as mean squared error (MSE): (3) Parameter update: Minimize this loss function using gradient descent (e.g., Adam optimizer). Calculate the loss function with respect to the Q-network parameters. gradient And update the parameters in the opposite direction of the gradient: in It's the learning rate. The total reward through this process... This information was backpropagated and encoded into the Q-network parameters. The network learned "what actions to take in what states will yield high rewards."
[0101] Step 4: Policy Improvement and Action Selection. As the Q-network continues to update, the agent's action selection strategy also improves. The most commonly used strategy is the ε-greedy strategy: (1) Based on probability Selection: Select the action that the Q-network currently considers optimal, i.e. (2) Based on probability Exploration: Randomly select an action to discover potentially better options. By continuously repeating the cycle of "interaction (performing actions) - storage (recording experience) - learning (updating the network)," the agent learns to prioritize actions that yield higher total rewards. The actions. In our scenario, this means that the agent will eventually tend to sample actions that are close to the performance boundary (high performance). (Rewards), and can also significantly reduce model uncertainty (high) Test points (rewards).
[0102] The state space (S), action space (A), and reward function (R) provide high-quality training data for constructing the tensor coupled field model. The tensor coupled field model is a data-driven model; its accuracy heavily depends on the quality and distribution of the training data. Traditional methods (such as grid sampling or random sampling) produce a "combinatorial explosion" in high-dimensional spaces, and most test points are located in "flat" regions where performance is good or completely ineffective, contributing little to the boundary definition and resulting in extremely low efficiency.
[0103] The role of reinforcement learning: The active sampling strategy, which consists of state space (S), action space (A), and reward function (R), aims to intelligently and adaptively select the test points that best help define the performance boundaries.
[0104] The state (S) records the explored areas and the uncertainties of the model, guiding the direction of the next exploration.
[0105] Action (A) determines the next specific combination of ODD parameters for testing.
[0106] award This also encourages exploration of regions with high uncertainty. (information gain) and the region near the performance boundary ( (Boundary proximity).
[0107] Data points collected through active sampling strategies are densely distributed near performance boundaries and in regions where model cognition is ambiguous. Using this data to train a tensor coupled field model enables the model to achieve extremely high fitting accuracy in boundary regions, thus more accurately characterizing boundary morphology. Therefore, reinforcement learning is a key tool for tensor coupled field models to obtain efficient and accurate training data.
[0108] S40, the test point dataset is used as training data to train the pre-trained tensor coupled field model, and the tensor coupled field model and training dataset are obtained. The training dataset is then used to train the pre-trained boundary symbol distance network model, and the boundary symbol distance network model is obtained. As one implementation method, the step of using the test point dataset as training data to train a pre-trained tensor coupled field model to obtain the tensor coupled field model and the training dataset includes: F1, based on the test point dataset, yields the basis field function. and higher-order tensor functions ; F2, based on the basis field function and the higher-order tensor function, yields the core coupling field. ; F3 inputs the core coupling field into the pre-trained tensor coupling field model to obtain coupling field loss data. ; ;in, For state space The predicted value; For state space The true value.
[0109] F4. The tensor coupled field model is trained based on the coupled field loss data to obtain the tensor coupled field model. F5, based on the test point dataset and the tensor coupled field model, obtain the training dataset. .
[0110] Based on the data points obtained from sampling We need to construct a continuous field model that can smoothly interpolate and predict system performance under arbitrary parameter combinations. This invention abandons simple linear or weighted summation models and proposes a tensor coupled field model.
[0111] This model will consider system performance. Represented as baseline performance In a coupled field composed of individual ODD elements Attenuation results under the action: in, This represents the element-wise performance degradation operation. It is a non-linear saturation function, ensuring that the decay is bounded. It is the core coupling field, the computation of which involves a learnable high-order tensor. and element feature mapping : in, A function that maps a single ODD element value x to a high-dimensional feature space to capture its nonlinear effects. It is an n-order tensor (where n is the number of ODD elements) that explicitly models the interaction coupling effects of any order between all elements. This indicates tensor contraction. It is a set of pre-defined basis functions (such as those based on radial basis functions or Fourier features) used to capture specific, known physical interaction patterns (such as the scattering effect of fog and light). These are the adaptive weights of the basis field function.
[0112] The model's parameters are learned by maximizing the likelihood probability on the test data. Its innovation lies in using tensors... and base field The combined effect of these technologies can both discover unknown complex interactions in a data-driven manner and embed domain knowledge, thereby achieving more reliable generalized predictions in areas with limited data.
[0113] Through the performance field model Solving the true signed distance field (SDF) using isosurfaces is a well-established approach in computer graphics, with underlying concepts (such as the Marching Cubes algorithm) already mature. However, applying it to solve the performance field of high-dimensional, non-uniform, and coupled autonomous driving ODDs, and serving the training of boundary networks, is an innovative application. This differs from traditional ODD boundary definition methods based on static thresholding or simple rule modeling.
[0114] High-dimensional sparse sampling and interpolation: The parameter space of ODD is typically 5-dimensional or even higher (weather, illumination, road type, traffic density, vehicle speed, etc.). Directly performing dense grid discretization in such a high-dimensional space is not feasible. We adopt an adaptive sparse grid sampling technique, which only performs dense sampling in regions where the gradient predicted by the performance field model P is large (near the boundary) and where uncertainty is high.
[0115] Isosurface Extraction Based on Neural Fields: We utilize a trained tensor-coupled neural field model P, which enables fast lookup and differentiation of arbitrary parameter points. A method combining ray tracing and Newton's iteration method is used to locate the boundary. Starting from a seed point in space, along the performance gradient It emits rays in a specific direction.
[0116] Utilizing neural fields The differentiability property allows for the rapid finding of the ray that satisfies the Newton-Raphson iteration method. of point.
[0117] By emitting rays in different directions multiple times, a large number of high-precision boundary points can be collected efficiently. }, without having to traverse the entire grid.
[0118] Signed distance calculation: for any sample point Find the nearest boundary point. Calculate the Euclidean distance. .
[0119] Symbols by and The relationship determines: if (Within the security domain), then Conversely, it is .
[0120] These The training boundary symbolic distance network is formed A high-quality dataset, also known as a training dataset.
[0121] As one implementation method, the training dataset is input into a pre-trained boundary symbol distance network model for training to obtain the boundary symbol distance network model, which may include: G1, based on the training dataset and the pre-trained boundary symbol distance network model, obtains boundary loss data; Obtaining the tensor-coupled field model Then, the boundary of the safe operating domain is satisfied. The set of points is an (n-1)-dimensional surface (or hypersurface) embedded in an n-dimensional parameter space. This invention employs implicit neural representations to efficiently define and query this complex boundary.
[0122] We train a lightweight boundary symbol distance network model. : The network learns to map any parameter point (A,B,C,...) to its signed distance d to the security boundary, where d>0 indicates that the point is inside the security region, d<0 indicates that the point is outside the security region, and points with d=0 are precisely located on the boundary.
[0123] Boundary Symbol Distance Network The training objective is: in, It can be derived from the performance field model Calculated (using numerical methods) (isosurfaces) It is a regularization term for the network gradient, used to ensure that the learned boundary surfaces are smooth.
[0124] G2 trains the boundary symbol distance network model based on the boundary loss data to obtain the boundary symbol distance network model.
[0125] Find neural network models by using training datasets. Parameters that best fit the training data .
[0126] Meaning: This represents a neural network model, where It is the set of all adjustable parameters of the network, such as weights and biases. (A,B,C,…) means that the input (A,B,C,…) is fed into the parameter. After the network is completed, the output value d is obtained.
[0127] The training process: (1) Prepare training data: Collect a set of training samples {( , )},in It is a point in the parameter space. It is the true signed distance from the point to the ODD boundary (calculated using the numerical method described in Question 5).
[0128] (2) Define the loss function The loss function measures the current parameters The predicted value of the lower network Compared with the true value The difference between them (mean squared error) is added to a regularization term to smooth the boundary surface.
[0129] (3) Optimization Solution: Use gradient descent to minimize the loss function. Calculate the gradient of the loss function with respect to the network parameters θ (backpropagation algorithm), and update the algorithm in the reverse direction of the gradient. After multiple iterations, the network prediction will get closer and closer to the true value.
[0130] (4) After training Training stops when the loss function decreases to a sufficiently small value or when the preset number of iterations is reached. The resulting parameter set is... That is to make the network It can accurately predict the final parameters of the signed distance. These are the "network parameters" that we need to store and use. ".
[0131] The system performance field model P is the implicit "data source" and "query tool" for extracting boundary surfaces.
[0132] Relationship: To extract the boundary surface, we first need to know "which points lie on the boundary". Mathematically, a boundary is defined as satisfying... The set of all points.
[0133] The specific application steps include: (1) Generate training data pairs. This is the most crucial step and requires the boundary symbol distance network to be generated. Prepare training data .in It is a point in the parameter space. It is the signed distance from the point to the true boundary.
[0134] (2) Utilization calculate For a sample point We utilize a pre-trained tensor-coupled field model Find the distance using numerical methods Recent, satisfied boundary points .
[0135] calculate arrive Euclidean distance .
[0136] according to and Relationship determination symbol: If (Within the security domain) Conversely .
[0137] (3) Training the boundary network Using the large amount generated in the previous step Data pairs for training networks This enables it to learn from parameters Direct mapping to signed distance .
[0138] The system performance field model P plays the role of implicit boundary definer and distance calculation. It is itself a continuous boundary function (isosurface), but we use it to generate data and train another, lighter, faster-querying, and more easily integrated and visualized network. To explicitly represent this boundary. It is a "teacher model" used to construct boundary representations, and This is the student model that is ultimately used for deployment and querying.
[0139] S50, Based on the test point dataset and the boundary symbol distance network model, delineate the boundary of the multi-element superposition design running domain.
[0140] As one implementation method, the step of defining the boundary of the multi-element superimposed design runtime domain based on the test point dataset and the boundary symbol distance network model includes: in, The design operation domain boundary for multiple elements superimposed; This is the test point dataset; This is a boundary symbol distance network model.
[0141] The final multidimensional ODD boundary, in a machine-readable and queryable form, is defined by this neural network. The zero isosurface: This representation method has significant advantages: Highly efficient storage: Only network parameters need to be stored. It is not a massive boundary point cloud.
[0142] Fast query: Determining whether any scene is within the ODD requires only one network forward computation. The symbol.
[0143] Easy to visualize: For any subspace consisting of two or three ODD elements, it can be evaluated on a grid. It also generates user-friendly 2D / 3D boundary maps by drawing contour lines.
[0144] Convenient dynamic updates: After an OTA upgrade, simply update the performance field model and boundary network with the new test data. This allows for the acquisition of new boundary surfaces, thus realizing the dynamic evolution of the ODD definition.
[0145] In summary, through the three steps of "active sampling - field modeling - implicit representation," the definition of multi-element ODD boundaries is transformed from a static problem based on discrete combinations and subjective evaluation into a dynamic computational problem based on continuous space exploration, data-driven modeling, and neural representation. It can not only accurately define the safety boundaries under arbitrarily complex combinations, but its output (boundary network) also... Furthermore, it can be directly integrated into the vehicle system as a core module for real-time ODD compliance checks, thus achieving a significant leap forward in both methodology and engineering practicality compared to existing technologies.
[0146] Once the initial ODD boundary definition is completed, it needs to be verified for authenticity and updated periodically.
[0147] (1) Authenticity verification: Select representative scenarios in real road environments for verification testing to confirm whether the defined boundaries are applicable in the actual environment. Collect at least 1000 hours of real road test data and analyze the system performance near the boundaries.
[0148] (2) OTA linkage update mechanism: When the autonomous driving system undergoes software OTA upgrades or sensor updates, the ODD boundary update process is initiated: Impact Analysis: Analyze the potential impact of the upgrade content on the ODD boundary; Targeted testing: Simplified testing for potentially changing boundary areas; Boundary Adjustment: Adjust the ODD boundary definition based on the test results; Documentation Update: Updated user manual and system configuration files.
[0149] Add real-time data-driven boundary prediction functionality: Establish a boundary adaptive model based on vehicle-mounted sensors and cloud data: in: For the baseline ODD boundary, To adjust the amount of data in real time based on onboard sensor data. This refers to the adjustment amount based on cloud-based meteorological and traffic data for forecasting.
[0150] The specific implementation includes: A. Sensor data fusion: Real-time monitoring of the performance indicators of cameras, radar, and lidar to calculate the actual perception capability in the current environment; B. Cloud data integration: Access to real-time data such as weather forecasts, traffic conditions, and road construction; C. Boundary prediction algorithm: Uses time series analysis to predict future changes in the ODD boundary.
[0151] in The trend of boundary changes is estimated using historical data.
[0152] Wherein, ODD defines the output. The final output includes two forms of ODD definitions: (1) Machine-readable format It uses XML or JSON format and contains the precise numerical range and combination rules for each ODD element, for use by autonomous driving systems for real-time monitoring.
[0153] (2) User-friendly format Using natural language and diagrams, the ODD boundaries are clearly displayed in the user manual and in-vehicle interface. For example, green / yellow / red areas are used to represent the system's capability range under different conditions, and typical scenario examples are provided.
[0154] Through the above embodiments, the present invention can provide accurate, complete and practical ODD boundary definitions for autonomous vehicles, significantly improving system safety and user experience.
[0155] This application transforms qualitative ODD elements into precise quantitative indicators. It defines measurable physical quantities and units for each type of element (such as landscape elements, environmental conditions, and dynamic elements). For example, "rainfall" is quantified as "rainfall intensity (0-150 mm / h)," and "road curvature" is quantified as "radius of curvature (≥100 m)." This quantification transforms the ODD boundary from a vague semantic description into a clear numerical range, laying the foundation for subsequent precise testing and verification.
[0156] This application enables the comprehensive performance boundary assessment of real-world driving scenarios involving complex superposition of multiple elements. Instead of exhaustive full-scenario testing (leading to combinatorial explosion), it scientifically selects key ODD elements and their levels to construct an efficient test matrix. Using test data, a mathematical model is established to describe the degree of system performance degradation when multiple elements are superimposed. The core value of this model lies in its ability to predict and define the system's safe operating boundary under arbitrary element combinations, i.e., generating a multi-dimensional ODD boundary surface, significantly improving the comprehensiveness and accuracy of safety assessments.
[0157] This application integrates reinforcement learning for adaptive testing, intelligently focusing on critical performance regions. Simultaneously, it introduces a risk-based probabilistic model as the boundary determination criterion, ensuring that the determination of the ODD boundary is based on a scientific probabilistic risk assessment, rather than a simple performance threshold judgment.
[0158] For various conditions defined within the ODD (especially "edge scenarios" near performance boundaries), targeted test cases can be generated. Through simulation (such as Software-in-the-Loop (SIL) and Hardware-in-the-Loop (HIL)) and real-vehicle testing, the performance of the system under critical conditions can be effectively verified, ensuring its safety redundancy.
[0159] While the vehicle is running, the system continuously monitors current environmental parameters (second parameter set) and compares them in real time with preset ODD boundaries (first parameter set). When the system detects that it is about to exceed or has already exceeded the ODD, it can trigger an early warning or activate a "minimum risk strategy" (such as gradual deceleration, pulling over, or requesting takeover) based on the SOTIF principle to prevent the system from suddenly failing at the capability boundary, thereby achieving safety degradation.
[0160] The ODD provides clear boundaries and priorities for testing. Testing teams can efficiently generate simulation and real-vehicle test cases based on this, focusing on covering key and boundary scenarios within the ODD and avoiding wasting resources on irrelevant scenarios. This scenario coverage matrix-based testing method can more fully expose system weaknesses at a controllable cost, improving the completeness and efficiency of verification.
[0161] In summary, this application obtains runtime element information, including several runtime elements, the value range corresponding to each runtime element, and the possible values corresponding to each runtime element. It then uses an adaptive control variable method and threshold judgment conditions to determine the single-element boundary threshold of each runtime element based on the value range and the possible values, obtaining a single-element boundary information set. Based on the single-element boundary information set, a preliminary test step size is obtained, and an active sampling strategy is used to interact with a controlled test field based on the preliminary test step size to obtain a test point dataset. This test point dataset is then used as training data to train a pre-trained tensor coupled field model, resulting in a tensor coupled field model and a training dataset. The training dataset is then input into a pre-trained boundary symbolic distance network model, resulting in a boundary symbolic distance network model. Finally, based on the test point dataset and the boundary symbolic distance network model, the design runtime boundary of a multi-element superposition is defined. This achieves the comprehensive performance boundary evaluation of real-world driving scenarios with complex superposition of multiple elements.
[0162] For those consistent with the above, please refer to Figure 2 , Figure 2 This application provides a schematic diagram of the structure of a device for delineating the boundary of the operating domain of a self-driving car, as illustrated in an embodiment. Figure 2 As shown, the device includes: The information acquisition module 201 is used to acquire runtime element information, which includes several runtime elements, the value range corresponding to the runtime elements, and the possible values corresponding to the runtime elements. The element boundary determination module 202 is used to determine the single element boundary threshold of the running domain element by using an adaptive control variable method and threshold determination conditions through the value range and the possible values, so as to obtain a single element boundary information set. The test data determination module 203 is used to obtain a preliminary test step size based on the single element boundary information set, and to interact with the controlled test field according to the preliminary test step size through an active sampling strategy to obtain a test point dataset; The model determination module 204 is used to input the test point dataset as training data into the pre-trained tensor coupled field model for training, to obtain the tensor coupled field model and the training dataset, and to input the training dataset into the pre-trained boundary symbolic distance network model for training, to obtain the boundary symbolic distance network model. The runtime boundary determination module 205 is used to delineate the design runtime boundary of multi-element superposition based on the test point dataset and the boundary symbol distance network model.
[0163] This application also provides a terminal device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor. When the processor executes the computer program, it implements the steps in the embodiment of the self-driving car design and operation domain boundary delineation method.
[0164] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the self-driving vehicle design operating domain boundary delineation methods described in the above method embodiments.
[0165] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program that causes a computer to perform some or all of the steps of any of the self-driving vehicle design operating domain boundary delineation methods described in the above method embodiments.
[0166] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium can include at least: any entity or device capable of carrying computer program code to a device / terminal equipment, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable storage media cannot be electrical carrier signals or telecommunication signals.
[0167] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0168] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0169] In the embodiments provided in this application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0170] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
Claims
1. A method for delineating the design and operation domain boundary of an autonomous vehicle, characterized in that, The method includes: Obtain runtime element information, which includes several runtime elements, the value range corresponding to the runtime elements, and the possible values corresponding to the runtime elements; The adaptive control variable method and threshold determination conditions are used to determine the single element boundary threshold of the running domain element by means of the value range and the possible values, so as to obtain the single element boundary information set; The initial test step size is obtained based on the boundary information set of the individual elements, and the test point dataset is obtained by interacting with the controlled test field according to the initial test step size through an active sampling strategy. The test point dataset is used as training data to train the pre-trained tensor coupled field model, resulting in the tensor coupled field model and the training dataset. The training dataset is then used as training data to train the pre-trained boundary symbol distance network model, resulting in the boundary symbol distance network model. Based on the test point dataset and the boundary symbol distance network model, the boundary of the multi-element superposition design running domain is defined.
2. The method for delineating the design and operation domain boundary of a self-driving car according to claim 1, characterized in that, The adaptive control variable method and threshold determination criteria are used to determine the single-element boundary threshold of the running domain elements based on the value range and the possible values, resulting in a single-element boundary information set, including: By setting the test environment for each runtime domain element according to the value range and the possible values, a test environment information set is obtained; The test vehicle was operated under control conditions to obtain benchmark information on performance indicators. Control the operation of the test vehicle according to each possible value corresponding to the first operating domain element in the test environment information set, and obtain the first key indicator information corresponding to the first operating domain element. Based on the threshold judgment conditions, the first key indicator information, and the performance indicator benchmark information, the design operating domain boundary of the first operating domain element is determined, and the first boundary information corresponding to the first operating domain element is obtained. Control the operation of the test vehicle according to each possible value corresponding to the second operating domain element in the test environment information set, and obtain the second key indicator information corresponding to the second operating domain element. The design operating domain boundary of the second operating domain element is determined based on the threshold judgment condition, the second key indicator information, and the performance indicator benchmark information, so as to obtain the second boundary information corresponding to the second operating domain element. After all possible values in all operating domain elements have been tested in vehicle operation, the boundary information corresponding to all operating domain elements is combined to obtain a single element boundary information set.
3. The method for delineating the design and operation domain boundary of a self-driving car according to claim 2, characterized in that, The step of determining the design operating domain boundary of the second operating domain element based on the threshold determination condition, the second key indicator information, and the performance indicator benchmark information, to obtain the second boundary information corresponding to the second operating domain element, includes: Based on the measured performance values in the key indicator information and the performance indicator benchmark information, the second performance indicator decay value corresponding to the second running domain element is obtained. The failure probability of the second system is calculated based on the decay value of the second performance index. Based on the threshold determination criteria, the failure probability of the second system, and key indicator information, the boundary information of the second element is obtained.
4. The method for delineating the design and operation domain boundary of an autonomous vehicle according to any one of claims 1 to 3, characterized in that, The test step size includes a state space, an action space, and a reward function. The state space includes the currently explored parameter point, the system performance index corresponding to the currently explored parameter point, and an uncertainty estimate. The system performance index is predicted by the system performance model corresponding to the currently explored parameter point. The preliminary test step size is obtained based on the individual element boundary information set, and a test point dataset is obtained by interacting with the controlled test field according to the preliminary test step size using an active sampling strategy. This dataset includes: The preliminary state space and preliminary action space are obtained based on the set of boundary information of the individual elements; Based on the initial state space, the action with the largest acquisition function value is selected as the initial optimal candidate action in the initial action space using the acquisition function. The initial optimal candidate action is run in the test environment, and the initial system performance index is observed. The initial optimal candidate action and the initial system performance index are added as new data points to the currently explored parameter points to obtain the currently explored parameter points for the first iteration. The state space and action space are updated according to the explored parameter points of the first iteration and the boundary information set of individual elements to obtain the state space and action space of the first iteration. The initial reward function is obtained according to the initial optimal candidate action, the initial state space, and the state space of the first iteration. The acquisition function is trained using the initial reward function to obtain the acquisition function of the first iteration. Based on the state space of the first iteration, the action with the largest acquisition function value in the action space of the first iteration is selected as the optimal candidate action for the first iteration using the acquisition function of the first iteration. The optimal candidate action of the first iteration is run in the test environment, and the system performance index of the first iteration is observed. The optimal candidate action and the system performance index of the first iteration are added as new data points to the currently explored parameter points to obtain the currently explored parameter points of the second iteration. The state space and action space of the first iteration are updated according to the explored parameter points of the second iteration and the boundary information set of individual elements to obtain the state space and action space of the second iteration. The reward function of the first iteration is obtained according to the optimal candidate action of the first iteration, the state space of the first iteration, and the state space of the second iteration. The acquisition function of the first iteration is trained using the reward function of the first iteration to obtain the acquisition function of the second iteration. Until the updated uncertainty estimate is less than the preset uncertainty threshold, all currently explored parameter points and the corresponding system performance indicators are used as the test point dataset.
5. The method for delineating the design and operation domain boundary of an autonomous vehicle according to any one of claims 1 to 3, characterized in that, The step of using the test point dataset as training data to train the pre-trained tensor coupled field model, resulting in the tensor coupled field model and the training dataset, includes: Based on the test point dataset, the basis field function and higher-order tensor function are obtained; The core coupling field is obtained based on the basis field function and the higher-order tensor function; The core coupling field is input into a pre-trained tensor coupling field model to obtain coupling field loss data; The tensor coupled field model is trained based on the coupled field loss data to obtain the tensor coupled field model; The training dataset is obtained based on the test point dataset and the tensor coupled field model.
6. The method for delineating the design and operation domain boundary of a self-driving vehicle according to claim 5, characterized in that, The step of inputting the training dataset into a pre-trained boundary symbol distance network model for training to obtain the boundary symbol distance network model includes: Based on the training dataset and the pre-trained boundary symbol distance network model, boundary loss data is obtained; The boundary symbol distance network model is trained based on the boundary loss data to obtain the boundary symbol distance network model.
7. The method for delineating the design and operation domain boundary of a self-driving vehicle according to claim 6, characterized in that, The step of defining the boundary of the multi-element superimposed design runtime domain based on the test point dataset and the boundary symbol distance network model includes: in, The design operation domain boundary for multiple elements superimposed; This is the test point dataset; This is a boundary symbol distance network model.
8. A device for delineating the boundary of the operating domain of a self-driving car, characterized in that, include: The information acquisition module is used to acquire runtime element information, which includes several runtime elements, the value range corresponding to the runtime elements, and the possible values corresponding to the runtime elements. The element boundary determination module is used to determine the single element boundary threshold of the running domain element by using an adaptive control variable method and threshold determination conditions through the value range and the possible values, so as to obtain a single element boundary information set. The test data determination module is used to obtain a preliminary test step size based on the single element boundary information set, and to interact with the controlled test field based on the preliminary test step size through an active sampling strategy to obtain a test point dataset; The model determination module is used to input the test point dataset as training data into the pre-trained tensor coupled field model for training, to obtain the tensor coupled field model and the training dataset, and to input the training dataset into the pre-trained boundary symbolic distance network model for training, to obtain the boundary symbolic distance network model. The runtime boundary determination module is used to delineate the design runtime boundary of multi-element superposition based on the test point dataset and the boundary symbol distance network model.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the self-driving vehicle design and operation domain boundary delineation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the self-driving vehicle design and operation domain boundary delineation method as described in any one of claims 1 to 7.