A method and system for autonomous flight decision-making and verification of unmanned intelligent agents

By optimizing the constraint contraction projection and completing the verification scenario driven by the coverage gap, combined with the consistency judgment of simulation and actual flight and the failure closed-loop update, the problem of insufficient safety constraints and verification in the autonomous flight of unmanned intelligent agents is solved, and safe and feasible control and continuous optimization are achieved.

CN122239495BActive Publication Date: 2026-07-31CHINA ORDNANCE SCI INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA ORDNANCE SCI INST
Filing Date
2026-05-21
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies struggle to reliably meet various safety constraints during autonomous flight of unmanned intelligent agents, have insufficient coverage of verification scenarios, lack consistency between simulation and actual flight, and fail results are difficult to reverse-drive decision-making and verification optimization.

Method used

The system employs constrained contraction projection optimization to generate safe and feasible controls, verifies scenario completion by filling in coverage gaps, and performs simulation-flight consistency determination and failure closed-loop updates to form a unified technology chain.

Benefits of technology

It improves the flight safety, verification adequacy, transfer reliability, and continuous optimization capabilities of unmanned intelligent agents, and enables safe and feasible flight and computational verification in complex and uncertain environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122239495B_ABST
    Figure CN122239495B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for autonomous flight decision-making and verification of unmanned intelligent agents. The method integrates mission objectives, flight states, environmental constraints, and model uncertainties to construct a decision state. A target-conditional-policy network outputs candidate controls, and feasible controls are obtained through constraint contraction projection optimization. Verification scenarios are completed based on coverage gaps and risk importance, and simulation-real-flight consistency judgment is performed on a unified scenario set. When inconsistencies or safety failures occur, failed segments are extracted to update policy parameters, constraint contraction parameters, and scenario sampling parameters. This scheme improves the safety, verification sufficiency, and closed-loop optimization capabilities of autonomous flight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous flight control and engineering verification technology for unmanned intelligent agents, and particularly to a method and system for autonomous flight decision-making and verification of unmanned intelligent agents. Specifically, it involves: generating candidate controls by combining environmental constraints and model uncertainties under the drive of mission objectives; obtaining safe and feasible controls through constraint contraction projection optimization; and combining verification scenario completion driven by coverage gaps, simulation-real flight consistency determination, and failure closed-loop update to achieve integrated collaboration of autonomous flight decision-making and verification of unmanned intelligent agents. Background Technology

[0002] With the increasing application of drones, unmanned aerial vehicles, and other unmanned intelligent agents in scenarios such as inspection, logistics, emergency search and rescue, border patrol, and environmental monitoring, goal-driven autonomous flight decision-making capabilities and engineering verifiability are gradually becoming key factors restricting practical implementation. Existing technologies often employ rule-based path planning methods, optimization-based trajectory generation methods, or control decision-making methods based on deep learning and reinforcement learning to generate flight control commands. However, in real-world flight scenarios, complex factors such as geofencing, no-fly zones, altitude restrictions, dynamic obstacles, communication degradation, wind field disturbances, energy constraints, and model mismatch often coexist. Existing technologies suffer from at least the following shortcomings: First, the control quantities directly output by the learning strategy are often difficult to stably satisfy multiple safety constraints. Simple pruning, saturation restrictions, or rule overlays fail to balance target advancement capability and constraint feasibility, resulting in insufficient real-time performance and safety in engineering. Second, most existing verification processes rely on experience-based scenario selection or manual scenario enumeration, lacking calculable coverage evaluation standards and automatic gap-filling mechanisms, making it difficult for verification results to fully reflect the real risks under complex task conditions. Third, simulation tests and actual flight tests typically only compare mean errors or empirical comparisons, lacking statistically significant consistency judgment mechanisms, making it difficult to support engineering acceptance and model migration reliability assessments. Fourth, when safety failures or inconsistencies between simulation and actual flight occur during verification, existing technologies often struggle to structurally reinject failed segments into the control decision-making, constraint handling, and scenario sampling stages, leading to a disconnect between "decision-verification-correction" and insufficient continuous system optimization capabilities. Therefore, there is an urgent need for an autonomous flight decision-making and verification method and system for unmanned intelligent agents that can organically couple safety and feasibility control, verification coverage completion, consistency judgment and failure closed-loop update, so as to improve flight safety, verification sufficiency, transfer reliability and continuous optimization capability. Summary of the Invention

[0003] The purpose of this invention is to provide an autonomous flight decision-making and verification method and system for unmanned intelligent agents, in order to solve the problems in the prior art, such as the difficulty in stably meeting safety constraints with learning control output, insufficient coverage of verification scenarios and lack of automatic completion mechanism, lack of statistically significant consistency judgment between simulation and actual flight, and difficulty in using failure results to drive subsequent decision-making and verification optimization. This enables unmanned intelligent agents to achieve safe and feasible flight, computable verification and closed-loop evolution in complex constraints and uncertain environments.

[0004] To achieve the above objectives, the present invention provides a method for autonomous flight decision-making and verification of an unmanned intelligent agent, comprising: at time... Receive task objectives issued by the task planning layer Synchronously acquire the unmanned intelligent agent at any time Flight status Environmental constraints and model uncertainty And merged into moments Decision-making status , decision state Input the target conditional policy network and obtain the time step. candidate control ; based on , , and Generate constraint matrix With constraint boundary vector And based on the control variables to be optimized ,time Deviation weighted matrix Constraint shrinkage coefficient Risk upper bound function Risk slack variables and risk penalty coefficient Construct a constrained contraction projection optimization problem and obtain the time step. Feasible control The constrained shrinking projection optimization problem satisfies: , ; Constructing a scenario to be verified Corresponding scene parameter vector and based on Get the scene Risk estimate Then based on the scenario Corresponding coverage gap , Positive smoothing term and risk emphasis coefficient Generate verification scenario sampling distribution The verification scenario set is completed based on the sampling distribution of the verification scenarios, and the sampling distribution of the verification scenarios satisfies: ; Obtain data based on the verification scenario set respectively. The simulation and flight performance indicators obtained are used to determine consistency according to the preset equivalent boundaries; When inconsistencies exist between simulation and actual flight, or when safety failures occur, the decision state, including at least the moment of failure, should be extracted. Candidate control Feasible control Minimum constraint margin and scenes Key Indicator Index Corresponding simulation-flight difference The failed segments are processed, and the parameters of the target conditional policy network are updated simultaneously. Constraint shrinkage parameter set and the set of sampling parameters for verification scenarios And update the target condition policy network parameters. Constraint shrinkage parameter set , Verification scenario sampling parameter set These factors respectively affect the generation of subsequent candidate controls and the constraint contraction coefficient. Generation and verification scenario sampling distribution The generation of .

[0005] To achieve the above objectives, the present invention also provides an autonomous flight decision-making and verification system for unmanned intelligent agents, comprising: The decision state construction module is used to receive the task objectives issued by the task planning layer, synchronously acquire the flight state, environmental constraints and model uncertainties of the unmanned intelligent agent, and integrate them into the decision state; The target condition policy module is used to input the decision state into the target condition policy network to obtain candidate controls and control confidence. The constraint contraction projection optimization module is used to generate constraint matrices and constraint boundary vectors based on the flight state, environmental constraints, model uncertainty and candidate controls, and generate deviation weighting matrices based on the control confidence. Then, based on the control variables to be optimized, the deviation weighting matrix, constraint contraction coefficients, risk upper bound function, risk relaxation variables and risk penalty coefficients, a constraint contraction projection optimization problem is constructed to obtain feasible control. The scene coverage completion module is used to construct the scene parameter vector corresponding to the scene to be verified, obtain the scene risk estimate based on the scene parameter vector, generate the verification scene sampling distribution based on the coverage gap corresponding to the scene, the scene risk estimate, the positive smoothing term and the risk emphasis coefficient, and complete the verification scene set based on the verification scene sampling distribution. The simulation-flight consistency determination module is used to obtain the simulation indicators and flight indicators obtained based on the feasible control execution on the verification scenario set, and perform consistency determination according to the preset equivalent boundary. The failure closed-loop update module is used to extract failure segments when there is a simulation-flight inconsistency or safety failure. These segments include at least the decision state at the time of failure, candidate control, feasible control, minimum constraint margin, and the simulation-flight difference corresponding to the key index of the scenario. Simultaneously, it updates the target condition policy network parameters, constraint contraction parameter set, and verification scenario sampling parameter set. The updated target condition policy network parameters, constraint contraction parameter set, and verification scenario sampling parameter set are then applied to the subsequent candidate control generation, constraint contraction coefficient generation, and verification scenario sampling distribution generation, respectively.

[0006] Compared with the prior art, the present invention has at least the following beneficial effects:

[0007] 1. Enhanced safety and controllability. This invention does not directly employ a learning strategy to output control quantities. Instead, it places candidate controls under a constraint contraction projection optimization framework for secondary correction. This ensures that fence constraints, height restrictions, collision interval constraints, energy budget constraints, and communication availability constraints are synergistically satisfied within a unified optimization model, thereby improving the safety, executability, and real-time performance of flight control.

[0008] 2. The verification coverage possesses computability and self-completion capabilities. This invention constructs a verification scenario sampling distribution by jointly considering coverage gaps and risk importance, enabling verification scenarios to no longer rely on manual experience-based enumeration but instead focus on areas with insufficient coverage. It automatically fills in high-risk areas, thereby improving the relevance, sufficiency and repeatability of the verification scenario construction.

[0009] 3. Simulation-flight consistency assessment is more reliable for engineering applications. This invention obtains simulation and flight performance indicators on the same set of verification scenarios and performs consistency assessments based on preset equivalence boundaries. This ensures that the relationship between simulation and flight is no longer limited to empirical comparisons but forms a statistically significant basis for engineering acceptance, thereby improving the reliability of model migration and system deployment.

[0010] 4. Establishing a closed-loop collaborative mechanism for decision-making, verification, and correction. Upon discovering inconsistencies or security failures, this invention not only extracts the failed segments but also simultaneously updates the policy network parameters, constraint contraction coefficients, and verification scenario sampling parameters. This allows the failure results to directly influence subsequent control generation, constraint processing, and verification sampling, thereby significantly improving the system's continuous learning capability, closed-loop optimization capability, and long-term operational stability.

[0011] 5. The technology chain is complete and the coupling relationship is clear. This invention constructs a unified technology chain of "target-driven decision-making - constraint contraction projection - scene coverage and completion - simulation consistency judgment - failure closed-loop update". There are clear data coupling relationships and parameter write-back relationships between the steps. This overcomes the problem of the separation between control, verification and correction in the prior art, and is more suitable for forming a feasible and engineered autonomous flight and verification system. Attached Figure Description

[0012] Figure 1 This is a flowchart of the autonomous flight decision-making and verification method for unmanned intelligent agents provided by the present invention.

[0013] Figure 2 This is a block diagram of the autonomous flight decision-making and verification system for unmanned intelligent agents provided by the present invention. Detailed Implementation

[0014] To enable those skilled in the art to more clearly understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be noted that in this specific embodiment, the same complete symbols represent the same technical meaning, and different complete symbols represent different technical meanings; wherein, complete symbols include not only basic letters, but also their superscripts, subscripts, parenthetical variables, caps, hyphens, transposes, and set symbols. The following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of protection of the present invention. All English abbreviations appearing for the first time in the text are given their Chinese meanings; for example, quadratic programming is denoted as QP (Quadratic Programming), and two-one-sided equivalence tests are denoted as TOST (Two-One-Sided Tests). The following description uses a multi-rotor UAV as a representative embodiment of an unmanned intelligent agent, but the present invention is not limited to multi-rotor UAVs, but is also applicable to fixed-wing UAVs, compound-wing UAVs, unmanned helicopters, and other unmanned intelligent agents with autonomous flight decision-making capabilities.

[0015] like Figure 1As shown, this invention revolves around the technical framework of "generating candidate controls driven by mission objectives—obtaining safe and feasible control through constraint contraction projection optimization—completing verification scenarios based on coverage gaps and risk importance—determining consistency between simulation and actual flight—closed-loop update in case of inconsistency or failure." The data output from the previous step serves as the input for the next step, and the updated parameters generated in the next step are then written back to the model parameters or adjustment parameters used in the preceding steps, thus forming a complete closed loop around the technical problem to be solved.

[0016] S1: Decision State Construction

[0017] In this step, the task objective issued by the task planning layer is first received at time t. Synchronously acquire the unmanned intelligent agent at any time Flight status Environmental constraints and model uncertainty The above information is then integrated into a unified decision state at time t. This is provided for subsequent network calls to target condition strategies.

[0018] In this embodiment, time Corresponding task target vector It can be represented as: , In the formula, Indicates the target location; Indicates the shape parameters of the target region; Indicates the target arrival tolerance; Indicates the time limit for achieving the target; Indicates the risk level of the task.

[0019] time Flight state vector It can be represented as: , In the formula, Indicates time Location; Indicates time The velocity vector; Indicates time The heading angle; Indicates time Height; Indicates time The remaining energy.

[0020] time Environmental constraints It can be represented as: , In the formula, Indicates time A collection of geographical fences or no-fly zones; Indicates time Height limit boundaries; Indicates time A dynamic set of obstacles; Indicates time The distribution of communication availability; Indicates time The wind field parameters.

[0021] In order to enable subsequent constraint contraction projection optimization to directly access the uncertainty information corresponding to each type of constraint, time step Model uncertainty vector Represented as: , In the formula, Indicates time The total uncertainty vector; Indicates time The relevant uncertainties regarding fence constraints; Indicates time The relevant uncertainties regarding height limit constraints; Indicates time The relevant uncertainties of the obstacle constraints; Indicates time The relevant uncertainties regarding energy constraints; Indicates time The relevant uncertainties of communication constraints.

[0022] After obtaining the above data, based on the time... Corresponding task target vector ,time Flight state vector ,time Environmental constraints and time uncertainty vector The decision state is fused according to the following formula to obtain the time step. Decision-making status : , In the formula, Represents the state construction function; This represents the encoding function for the task objective; Represents the flight state encoding function; Represents the environmental constraint encoding function; This indicates that vectors are concatenated column by column.

[0023] In this embodiment, , and All of these can be achieved using a normalization-based approach combined with multilayer perceptron mapping, allowing target information, flight state information, and environmental constraint information of different dimensions to be mapped to a unified feature space and then concatenated. Through this setup, S1 provides a unified state input for the subsequent policy network and, on the other hand, provides the basic data needed to generate the constraint matrix and constraint boundaries for subsequent constraint contraction projection optimization.

[0024] This step couples mission objectives, flight status, environmental constraints, and model uncertainties into the decision-making state. This avoids the information fragmentation problem caused by "using one set of states for decision-making and another set of states for constraint processing" in existing technologies.

[0025] S2: Generation of candidate controls based on target conditions and output of control confidence level

[0026] In this step, the time obtained in step S1 is used. Decision-making status Input the target conditional policy network, and output the time from the target conditional policy network. candidate control and control of credibility Among them, candidate control Used to characterize the intent control given by the policy network under the current task objectives and environmental conditions; control credibility. Used to characterize the candidate control The reliability of the current state is used to further adjust the deviation weighting matrix in subsequent constrained shrinkage projection optimization. .

[0027] In this embodiment, time Corresponding candidate control and control credibility satisfy: , In the formula, Represents the objective-conditional-policy network; The parameters represent the target conditional policy network; Indicates time Candidate control output by the objective conditional policy network; Indicates time Control credibility.

[0028] In this embodiment, time Corresponding candidate control It can be represented as: , In the formula, Indicates time Candidate speed commands; Indicates time The candidate heading angular velocity command; Indicates time Candidate rate of increase / decrease instructions.

[0029] To ensure the credibility of control Directly applied to subsequent constraint shrinkage projection optimization, in this embodiment, according to time... Control credibility Construction time Deviation weighted matrix : , , In the formula, This represents a diagonal matrix consisting of the elements within the brackets; Indicates time The velocity dimension deviates from the weight; Indicates time The deviation weight of the heading dimension; Indicates time The deviation weights of the ascending and descending dimensions; Indicates time The Deviation weights corresponding to each control dimension; Indicates the first The basic weights corresponding to each control dimension; Indicates the first The credibility adjustment coefficients corresponding to each control dimension; Indicates the control dimension Values ​​are taken from the velocity dimension, heading dimension, and elevation dimension. , .

[0030] From the above relationship, it can be seen that when At higher levels, the deviation from the weighted matrix Prefer to maintain candidate control ;when At lower levels, the deviation from the weighted matrix Improving subsequent constraint contraction projection optimization for candidate control The degree of freedom for modification is adjusted to prioritize the safety and feasibility of control output.

[0031] In this embodiment, the target conditional policy network adopts a training method of "offline pre-training + online closed-loop fine-tuning". During the offline pre-training phase, a pre-training dataset is constructed based on historical flight logs, high-fidelity simulation trajectories, and safety control sequences generated by the reference planner. , In the formula, This represents the pre-trained dataset; Indicates time The decision-making state; Indicates time Reference safety controls; Indicates time Control credibility labels.

[0032] In this embodiment, when the deviation between the reference safety control and the projected safety feasible control is less than a preset threshold, the control confidence label is activated. Set the value to 1; otherwise, control the credibility label. Set to 0. During training, the control output branch minimizes the mean square error between the candidate control and the reference safe control, while the confidence output branch minimizes the control confidence prediction error. In the online phase, the failure segment data extracted in step S6 is used to incrementally update the target conditional policy network to adapt to new failure modes.

[0033] Candidate control output in step S2 Controlling credibility and deviation weighted matrix It is directly used as the input to step S3, where the candidate control Used as a reference solution for constrained shrinkage projection optimization, controlling the confidence level. Used to adjust deviation from the weighted matrix The diagonal elements, deviating from the weighted matrix This is used to characterize the penalty intensity for the deviation of the candidate control in the subsequent optimization process.

[0034] This step encodes both the task advancement intent and the control credibility into the output of the target condition policy network, so that subsequent constraint contraction projection optimization is no longer a simple pruning of candidate controls, but an adaptive and safe correction of candidate controls under the guidance of control credibility, thereby improving the safety, executability and engineering applicability of the control output.

[0035] Furthermore, in this invention, control confidence is not an independently output auxiliary score, but rather a gating factor used in the construction of the deviation weighting matrix to determine the correction strength between candidate and feasible controls. When control confidence is high, constraint contraction projection optimization prioritizes preserving the task progression intent expressed by the candidate control; when control confidence is low, the deviation weighting matrix reduces the penalty for deviations in the candidate control, allowing subsequent optimization processes to safely correct the candidate control within a wider range of degrees of freedom. Through these settings, this invention does not perform fixed-amplitude pruning of the target condition policy network output, but rather adaptively allocates the trade-off between "intent preservation" and "safe correction" based on the reliability of the candidate control in the current state, thereby improving the safety and engineering applicability of the control output under complex and uncertain environments.

[0036] S3: Constrained contraction projection optimization, risk upper bound construction, and local reprogramming

[0037] In this step, based on the time of step S1 Flight status Environmental constraints and model uncertainty and the time of output in step S2 candidate control and deviation weighted matrix , building time The constrained shrinking projection optimization problem is obtained at time... Feasible control .

[0038] First, the fence constraint, height limit constraint, collision interval constraint, energy budget constraint, and communication availability constraint are respectively applied at time... By performing linearization or conservative approximation, a unified matrix constraint is formed: , In the formula, Indicates time The constraint matrix; Indicates time The control variables to be optimized; Indicates time The constraint boundary vector; Indicates time The constraint contraction coefficient; Indicates time The uncertainty vector; This represents the amount of shrinkage of the original constraint boundary.

[0039] To ensure that the risk penalty term and the matrix constraints both enter the same optimization problem, in this embodiment, the time step is... Lower control variables Corresponding upper bound function for risk Constructed into a quadratic upper bound form: , In the formula, Indicates time The risk quadratic term coefficient matrix; Indicates time The risk linear term coefficient vector; Indicates time The risk constant term.

[0040] Based on this, according to time candidate control Deviation weighted matrix constraint matrix Constraint boundary vectors Constraint shrinkage coefficient and uncertainty vector Solve the following constrained contraction projection optimization problem to obtain the time step. Feasible control : , , In the formula, Indicates time Feasible control obtained through constrained contraction projection optimization; Indicates time Risk slack variables; Indicates time Risk penalty coefficient.

[0041] To further illustrate the impact of uncertainty and the risk of loss of contact on the degree of constraint contraction, in this embodiment, time... Constraint contraction coefficient satisfy: , In the formula, Indicates the basal shrinkage; and These represent the weighting coefficients; Indicates time uncertainty vector The second norm; Indicates time The probability of communication failure.

[0042] At the moment of seeking Feasible control Then, the constraints at time t is further calculated. Constraint margin: , , In the formula, Indicates time The The constraint margin corresponding to each constraint; Indicates time constraint boundary vector The One component; Indicates time constraint matrix The OK; Indicates time The The contraction coefficient corresponding to each constraint; Indicates time The The uncertainty corresponding to each constraint; Indicates time The minimum constraint margin.

[0043] when Less than the preset threshold Time indicates the moment. Feasible control Although the constraints are met, the safety boundary is approaching. At this point, global replanning is not performed on the entire current trajectory, but only on the sub-segments of the current trajectory that intersect with the affected obstacle domain, fence boundary domain, or communication blind spot.

[0044] In this embodiment, first based on time Feasible control In the short time domain Internal prediction yields a set of candidate trajectory segments. And construct the affected sub-segment index set: , In the formula, Indicates the first One candidate trajectory segment; Indicates the length of the short-time prediction domain; Indicates the set of indexes for the affected subfields; This represents the affected area consisting of the obstacle domain, fence boundary domain, or communication blind zone corresponding to the minimum constraint margin; the symbol " "" indicates the intersection of sets; the symbol " " indicates the empty set.

[0045] Furthermore, in this invention, when the minimum constraint margin is lower than a preset threshold, global replanning is not performed uniformly on all trajectories. Instead, the affected region corresponding to the minimum constraint margin is used as an index basis, and local replanning is performed only on the trajectory sub-segments corresponding to the index set of affected sub-segments. Through the above settings, this invention further refines 'whether to replan' into 'which sub-segments to replan,' making the minimum constraint margin not only used for safety judgment but also for locating the local trajectory range that needs correction. This reduces the computational overhead and control jitter caused by overall replanning while maintaining the stability of unaffected trajectory sub-segments, and improves real-time performance and targeted local correction in dynamic scenarios.

[0046] Flight status output in step S1 Environmental constraints and model uncertainty Used to generate constraint matrix Constraint boundary vectors and constraint contraction terms; candidate control output from step S2 and deviation weighted matrix Used to determine the projection optimization objective; the feasible control output of this step Minimum constraint margin The results of local replanning will be used for subsequent flight execution, risk modeling, and failure segment construction, respectively.

[0047] This step optimizes the candidate control by performing projection with uncertainty shrinkage terms, so that the control output satisfies multiple safety constraints while maintaining the target's intended progress as much as possible. Furthermore, by performing local replanning only on the affected segments, the real-time computational burden in dynamic scenarios is reduced, thereby improving the system's security, executability, and engineering deployment capabilities.

[0048] S4: Training of the verification scenario completion and risk estimation model driven by the coverage gap

[0049] In this step, a validation scenario parameter vector is constructed based on the current policy parameters, constraint shrinkage parameters, and scenario distribution. Furthermore, based on one-dimensional quantile binning, two-dimensional joint quantile binning, coverage gaps, and risk importance, the verification scenario set is automatically completed, thereby improving the sufficiency of verification.

[0050] In this embodiment, for each scenario to be verified Construct scene parameter vector : , In the formula, Representing a scene The corresponding scene parameter vector; Representing a scene Obstacle density; Representing a scene The wind field intensity; Representing a scene The sensor noise level; Representing a scene The intensity of communication degradation; Representing a scene The target change range.

[0051] For scene parameter vector For each dimension of the parameter, construct a one-dimensional bin set according to the quantile. And statistically analyze the current set of verification scenarios. Sample count within each bin: , In the formula, Indicates the first Dimensional parameters in the first dimension The number of samples in each one-dimensional bin; Representing a scene In the Values ​​on the dimension; Indicates the first The corresponding dimension parameter is the first One-dimensional bin; Indicates the current set of verification scenarios; This represents an indicator function, which takes the value 1 when the condition within the parentheses is true, and 0 otherwise.

[0052] Furthermore, a two-dimensional joint quantile binning is established for the selected two-dimensional parameter combination, and the sample counts in the two-dimensional joint binning are counted: , In the formula, Indicates the first Dimensional parameters and the first Two-dimensional joint binning corresponding to the dimensional parameters The number of samples in; Representing a scene In the Values ​​on the dimension; Indicates the first The corresponding dimension parameter is the first One-dimensional bin.

[0053] In this embodiment, the minimum sample size for each one-dimensional binning and two-dimensional joint binning is set to a minimum. The corresponding coverage gaps satisfy the following conditions: , , In the formula, Indicates the first The corresponding dimension parameter is the first A gap in the coverage of a one-dimensional bin; Indicates the first Dimensional parameters and the first Two-dimensional joint binning corresponding to the dimensional parameters Coverage gaps; This indicates the lower limit of the minimum sample size.

[0054] In this embodiment, the minimum sample size limit is satisfied only if both the corresponding one-dimensional binning and the corresponding two-dimensional joint binning satisfy the minimum sample size limit. Only then is it determined that the corresponding part of the verification scenario coverage meets the standard.

[0055] To unify scenario coverage gaps and scenario risks under a single coverage rule, for each candidate scenario... Define its corresponding comprehensive coverage gap : , In the formula, Representing a scene The coverage gap of the integrated distribution box; Indicates the first The one-dimensional coverage gap weight coefficient corresponding to the dimensional parameter; Indicates the first Dimensional parameters and the first The two-dimensional joint coverage gap weight coefficients corresponding to the dimensional parameters; Representing a scene In the The one-dimensional binning index corresponding to the dimension parameter; Representing a scene In the Dimensional parameters and the first The two-dimensional joint binning index corresponding to the dimension parameter.

[0056] In this embodiment, the scenario risk estimate satisfy: , In the formula, Representing a scene Risk estimates; Representing a scene The average risk; Representing a scene The risk standard deviation; This represents the risk amplification factor.

[0057] Based on the comprehensive coverage gap and scenario risk estimates Construct the sampling distribution for the verification scenario: , In the formula, Representing a scene The sampling probability; symbol " "Indicates a proportional relationship; Represents a positive number smoothing term; Represents an exponential function; This indicates the risk emphasis coefficient.

[0058] Furthermore, in this invention, the sampling distribution of verification scenarios is not determined solely by historical sample frequency, a single risk score, or manual empirical rules, but rather by a combination of coverage gaps and scenario risk estimates. Specifically, one-dimensional quantile binning and two-dimensional joint quantile binning jointly characterize the sufficiency of the current verification scenario set's coverage in the parameter space, while the scenario risk estimate reflects both the scenario risk level and the degree of risk uncertainty. Through this setup, this invention enables the verification scenario completion process to simultaneously address both "areas that are not yet fully verified" and "areas more likely to expose failures," thereby avoiding the problems of merely supplementing coverage while ignoring high-risk scenarios, or solely pursuing high-risk scenarios and causing an imbalance in parameter space verification.

[0059] With the above settings, scenarios with larger coverage gaps and higher risk estimates will have a higher sampling probability, so that the completed set of verification scenarios can not only cover the areas that are currently under-verified, but also prioritize the exposure of high-risk conditions.

[0060] In this embodiment, the risk estimation model is trained using an ensemble regression model. First, a scenario risk training set is constructed using existing validation samples: , In the formula, This represents the training set for scenario risks; Representing a scene The scene parameter vector; Representing a scene Risk label.

[0061] In this embodiment, risk label It is generated jointly by the simulation execution results, the actual flight execution results, and the minimum constraint margin output in step S3. For example, the risk label This can be composed of a weighted combination of factors such as the degree of insufficient safety interval, the number of out-of-bounds occurrences, the duration of disconnection, and the degree of insufficient minimum constraint margin. Then, multiple basic regressors are constructed using a bootstrap sampling method, with each basic regressor minimizing the prediction error on its corresponding training subset. Finally, the mean of the outputs of the multiple basic regressors is used as the scene... Risk mean The standard deviations of multiple base regressors are used as the scenario. risk standard deviation .

[0062] Furthermore, in this invention, the scenario risk estimate is not represented by a single predicted value, but rather by incorporating both the risk mean and the risk standard deviation. That is, it considers not only the expected risk level of a scenario under the current model, but also the uncertainty of the model's risk assessment for that scenario. Through this setup, when the predicted risk mean for certain scenarios is not the highest but the prediction discrepancy is significant, its sampling probability can still be increased through the risk standard deviation amplification mechanism, thereby facilitating the earlier exposure of unknown failure modes and boundary conditions.

[0063] When step S5 or step S6 generates new scenario execution results, the corresponding new samples will be incorporated into the scenario risk training set. The risk estimation model is incrementally updated so that the sampling distribution of the verification scenario can continuously reflect the latest failure modes and risk distribution.

[0064] The minimum constraint margin output in step S3 can serve as an important component of the scenario risk label; the simulation-flight performance difference and failure flags generated in subsequent step S5 will continue to update the risk label; the verification scenario set output in this step... This serves as the unified scenario base for the subsequent step S5, which involves performing a simulation-flight consistency determination.

[0065] This step maps the two types of information, "insufficient coverage" and "high risk," into scenario sampling probabilities. This allows the verification scenario set to automatically fill in the test blind spots and prioritize the coverage of high-risk scenarios, thereby significantly improving the sufficiency, relevance, and ability to expose to extreme conditions.

[0066] S5: Simulation-Flight Consistency Determination Based on a Unified Verification Scenario Set

[0067] In this step, the unified verification scenario set obtained in step S4 is... The decision-making and control processes described in steps S1 to S3 are executed in both the simulation environment and the real flight environment to obtain the simulation indicators and the actual flight indicators under the same scenario, and to determine whether the two are consistent in a statistical sense.

[0068] For any scenario Obtain the scene respectively Corresponding simulation indicators and actual flight indicators Among them, the set of key indicators It can be represented as: , In the formula, Indicates arrival time; Indicates the path length; Indicates unit energy consumption; Indicates the minimum safe interval; Indicates the number of times the boundary was exceeded; Indicates the duration of the loss of contact.

[0069] For each key indicator Constructing a scene Simulation-to-real flight difference: , In the formula, Representing a scene Up indicators Simulation-to-real flight difference; Representing a scene Up indicators Simulation values; Representing a scene Up indicators The actual flight value.

[0070] Let the significance level be... The number of key indicators is The significance level after multiple comparison correction is... satisfy: , In the formula, This indicates the significance level after correction.

[0071] To ensure more stringent consistency requirements for high-risk tasks, the indicators... The equivalent boundary is determined according to the task risk level. The adaptive setting is: , In the formula, Indicators The equivalent boundary; Indicators The basic equivalent boundary; This represents the risk level adjustment coefficient; This indicates the task risk level. As can be seen from the above relationship, with the task risk level... The increase of the equivalent boundary The monotonically decreasing reduces the requirement for consistency between simulation and actual flight.

[0072] For each key indicator For the difference sample, perform a two-tailed one-sided equivalence test (TOST) to obtain the confidence interval of the mean difference. The criterion is determined only if the following formula is satisfied. They are statistically equivalent: , In the formula, Indicators The lower bound of the confidence interval for the mean difference; Indicators The upper bound of the confidence interval for the difference between the means.

[0073] In addition to statistical consistency criteria, hard safety thresholds are also set. For example, the minimum safety interval. The number of times the threshold is exceeded must not be lower than the preset threshold. The value must be zero; duration of being out of contact The limit must not be exceeded. If any security hard threshold is not met, the corresponding scenario verification is directly judged as failed; the scenario is judged to pass verification only when all security hard thresholds are met and all key indicators pass the TOST judgment.

[0074] The unified verification scenario set used in this step From step S4; the control strategy parameters, constraint contraction coefficients, and scenario sampling parameters used in this step are derived from the system state of the current round; the difference output in this step... The pass / fail flags and statistical results of the scenarios will serve as the direct basis for constructing the failed fragment dataset and performing closed-loop updates in the subsequent step S6.

[0075] This step employs a unified set of verification scenarios and statistically significant consistency testing rules, ensuring that the comparison between simulation and actual flight is no longer limited to the empirical level, but rather forms a reproducible, quantifiable, and acceptable engineering judgment criterion.

[0076] Furthermore, in this invention, the simulation-flight consistency determination is not merely a comparison of the mean error between simulation and flight values, nor is it based solely on empirical judgments of the closeness of single execution results. Instead, it performs a double-sided equivalence test on multiple key indicators across a unified verification scenario set, after multiple comparison corrections, and adaptively tightens the equivalence boundary in conjunction with the mission risk level. Through these settings, high-risk missions correspond to stricter consistency boundaries, while low-risk missions correspond to relatively lenient but still statistically significant judgment criteria. This makes the simulation-flight consistency determination both quantifiable and adaptable to mission risk.

[0077] S6: Failure fragment construction, priority update, and closed-loop acceptance determination

[0078] In this step, for the scenarios that are determined to be failed or inconsistent in step S5, the corresponding states, candidate controls, feasible controls, constraint margins and index differences are extracted to form a failure fragment dataset. The failure fragment dataset is then used to update the target condition policy network parameters, constraint contraction parameters and scenario sampling parameters in a linked manner.

[0079] For failure scenarios, construct a failure fragment dataset. : , In the formula, This represents a dataset of failed fragments; Indicates the moment of failure The decision-making state; Indicates the moment of failure Candidate control; Indicates the moment of failure Feasible control; Indicates the moment of failure The minimum constraint margin; Representing a scene Up indicators Simulation-to-real flight difference; This indicates a failure scenario.

[0080] To ensure that severely failed samples are prioritized during parameter updates, a failure index is defined in this embodiment. satisfy: , In the formula, Indicates the failure index; Indicates the number of key indicators; Indicators The corresponding weighting coefficients; Indicators The equivalent boundary; This represents the weighting coefficient corresponding to the constraint margin term; Indicates the constraint margin threshold; Representing a scene Up indicators Simulation-to-real flight difference; This represents the minimum constraint margin. With the above settings, when the simulation-to-real-flight difference of a certain index exceeds the corresponding equivalent boundary, or the minimum constraint margin is lower than the constraint margin threshold, the failure index... This will increase the priority of the corresponding failed samples in subsequent updates.

[0081] Furthermore, in this invention, the failure segment is not merely used as a retraining sample for the target conditional policy network, but rather as a unified data carrier for the coordinated updating of the target conditional policy network parameters, constraint contraction parameters, and scenario sampling parameters. That is, the same failure segment reflects decision-side mismatch through the deviation between candidate and feasible controls, reflects constraint handling-side risk exposure through minimum constraint margin, and reflects verification-side migration mismatch through the simulation-flight difference of key indicators. Through these settings, this invention simultaneously transforms a single failure event into a basis for correction across the three links of decision-making, constraint, and verification, thereby improving the targeting and consistency of closed-loop updates.

[0082] Based on the failure index After prioritizing and sampling the failed samples, a projected feasible control is employed. As a monitoring signal, update the target conditional policy network parameters. Its update objective satisfies: , In the formula, This indicates the updated objective-conditional-policy network in the decision state. Control of the output; This represents the failure penalty coefficient. By setting it as described above, the objective-conditional-policy network can be made to output outputs closer to feasible control during subsequent decision-making processes. Candidate controls are used to reduce the probability of failure recurring in similar scenarios.

[0083] In this embodiment, the target conditional policy network is updated using an incremental training method with priority replay, that is: from the failed segment dataset... China's failure index Small batches of samples are extracted, and a certain proportion of historical safe samples are mixed in to prevent the target conditional policy network from overfitting only to failure modes and impairing control performance in normal scenarios. After each update, the control confidence output is recalculated to ensure that the control confidence is consistent with the updated policy distribution.

[0084] In this embodiment, the main update of the parameter set is for the constraint shrinkage parameter. It can be represented as: , In the formula, Represents the set of constraint contraction parameters; Indicates the basal shrinkage; and These represent the weighting coefficients of the uncertainty term and the communication loss probability term in the constraint contraction coefficient, respectively. This is achieved by updating the parameter set. This allows the updated constraint contraction rules to remain more sensitive to uncertainty and communication loss risks in failure modes.

[0085] In this embodiment, the main update of the parameter set is for scene sampling parameters. It can be represented as: , In the formula, Represents the set of scene sampling parameters; Indicates the risk emphasis coefficient; This represents the risk amplification factor; Indicates the first The one-dimensional coverage gap weight coefficient corresponding to the dimensional parameter; Indicates the first Dimensional parameters and the first The two-dimensional joint coverage gap weight coefficients correspond to the dimensional parameters. Furthermore, the risk estimation model is retrained using newly added failed samples, thus focusing subsequent validation scenarios more on "high-risk + low-coverage" areas.

[0086] To avoid system performance degradation after updates, this embodiment performs a closed-loop acceptance judgment for each parameter update. Let the security satisfaction rates before and after the update be respectively... and The compliance rates for container coverage in the verification scenarios were as follows: and Then the parameter update will be accepted only if the following condition is met: , In the formula, This indicates the satisfaction rate of the hard safety constraints before the update; This indicates the updated safety hard constraint satisfaction rate; This indicates the bin coverage compliance rate of the verification scenario before the update; This indicates the updated bin coverage compliance rate for the verification scenario. If the updated hard constraint compliance rate or the verification scenario bin coverage compliance rate decreases, the parameter update will be rejected, and the system will revert to the previous parameter.

[0087] Furthermore, in this invention, the closed-loop acceptance decision does not only consider whether the loss function decreases after the update, or whether the control effect in a single scenario improves, but simultaneously considers the satisfaction rate of safety hard constraints and the compliance rate of validation scenario bin coverage before and after the update. The update result is accepted only if the satisfaction rate of safety hard constraints and the validation coverage level do not decrease after the update; otherwise, the update is rejected and the parameters revert to the pre-update state. Through the above settings, this invention avoids the situation where overfitting is performed only for a few failed samples, sacrificing overall safety or validation sufficiency, and gives the closed-loop update a constrained evolutionary characteristic of "correcting failures without compromising global safety and coverage."

[0088] Step S6 directly uses the failure flag and indicator difference output from step S5. Use the minimum constraint margin output in step S3 And use the candidate control output from step S2 Feasible control output from step S3 Construct the supervised update objective; step S6 updates the set of objective condition policy network parameters and constraint shrinkage parameters. and scene sampling parameter set Write back to steps S2, S3, and S4 to enter the next round of the "decision-verification-correction" closed loop.

[0089] This step does not stop at the evaluation level of "pass / fail" for the verification results, but structuring the failed segments into data objects that can be used for training and parameter tuning, and preventing update degradation through a closed-loop acceptance judgment mechanism, thereby achieving true closed-loop self-correction and improving the system's continuous learning ability, closed-loop optimization ability and long-term operational stability.

[0090] Corresponding to the above method embodiments, such as Figure 2 As shown, this invention also provides an autonomous flight decision-making and verification system for unmanned intelligent agents. The system can be implemented by a processor, memory, sensor interface, flight control interface, map interface, and communication interface, or by multiple functional software modules deployed on the same computing platform. The processor is used to call program instructions stored in the memory to execute decision state construction, target condition policy reasoning, constraint contraction projection optimization, verification scenario completion, simulation-real flight consistency determination, and failure closed-loop update.

[0091] The system comprises: a decision state construction module, a target condition strategy module, a constraint contraction projection optimization module, a scene coverage completion module, a simulation-real flight consistency determination module, and a failure closed-loop update module. Each module can be implemented using a combination of hardware circuits, firmware logic, and software programs, or it can be implemented using multiple functional units scheduled by the same processor. The modules are not isolated but rather form a cohesive data coupling relationship around a unified technology chain of "decision-projection-verification-update."

[0092] Among them, the decision state construction module is used at time 10:00. Receive task objectives issued by the task planning layer Synchronously acquire the unmanned intelligent agent at any time Flight status Environmental constraints and model uncertainty and the task objective Flight status Environmental constraints and model uncertainty Merging into Moments Decision-making status In one embodiment, the decision state construction module performs unified time alignment, normalized encoding, and vector concatenation on the data from the task planning, state estimation, environmental perception, and uncertainty assessment links to ensure that the subsequent target condition strategy module and constraint contraction projection optimization module call the data base at the same time.

[0093] The target conditional policy module is used to base decisions on the decision state. Output candidate control and control credibility Among them, the candidate control The confidence level of the control is used to characterize the intended control under the current mission objectives and environmental conditions. Used to characterize the candidate control The reliability under the current state. The target condition policy module does not directly output the final flight control command, but instead outputs candidate control commands. As a reference input for subsequent constrained shrinkage projection optimization, the confidence level will be controlled. This serves as the basis for adjusting the update of the deviation weighted matrix.

[0094] The constrained contraction projection optimization module is used to optimize the candidate control. Deviation weighted matrix constraint matrix Constraint boundary vectors and constraint shrinkage coefficient Solving for feasible control And when the minimum constraint margin is lower than a preset threshold, local replanning is performed. In one embodiment, the constraint contraction projection optimization module performs local replanning based on the flight state. Environmental constraints and model uncertainty Generate constraint matrix and constraint boundary vector Based on control credibility Adjusting the deviation weighting matrix Within a unified optimization framework, constraints such as fence requirements, height limits, collision intervals, energy budgets, and communication availability are comprehensively addressed to obtain feasible control that satisfies safety constraints. When the minimum constraint margin is below a threshold, the constraint contraction projection optimization module performs local replanning only on the affected trajectory segments to balance control safety and engineering real-time performance. The feasible control... On the one hand, the data is sent to the unmanned intelligent agent for execution via the flight control interface; on the other hand, the minimum constraint margin is input to subsequent modules.

[0095] The scene coverage and completion module is used to construct scene parameter vectors. Based on the scene parameter vector Get the scene Risk estimate And based on the coverage gap and scenario risk estimates Generate verification scenario sampling distribution The scenario coverage completion module performs one-dimensional quantile binning and two-dimensional joint quantile binning statistics on the scenario parameters to determine coverage gaps; simultaneously, it calls a risk estimation model to process the scenario parameter vector. Reasoning is performed to obtain the scenario risk estimate. Then, based on the coverage gaps and risk estimates, a sampling distribution is generated so that the completed verification scenario set can simultaneously cover the currently insufficient verification areas and high-risk areas.

[0096] The simulation-flight consistency determination module is used to acquire simulation indicators and flight indicators on the verification scenario set, and perform consistency determination according to a preset equivalence boundary. In one embodiment, the simulation-flight consistency determination module organizes simulation execution and flight execution for the same verification scenario, acquires the simulation values ​​and flight values ​​corresponding to key indicators, forms an indicator difference, and outputs the consistency result and pass / fail flag using a combination of double-sided equivalence verification and safety hard threshold determination. The pass / fail flag and indicator difference are then input to the failure closed-loop update module.

[0097] The failure closed-loop update module is used to update the parameters of the target condition strategy module and the constraint contraction coefficient based on the failure fragments. The parameters of the sampling distribution of the verification scenario are updated and written back to the target condition policy module, the constraint contraction projection optimization module, and the scenario coverage completion module, respectively. In one embodiment, the failure segment includes at least the decision state at the time of failure, candidate control, feasible control, minimum constraint margin, and simulation-flight difference in the scenario; the failure closed-loop update module constructs a failure index based on the failure segment and performs priority sorting, and performs linked updates to the policy network parameters in the target condition policy module, the constraint contraction parameters in the constraint contraction projection optimization module, and the scenario sampling parameters and risk estimation model in the scenario coverage completion module. To prevent system performance degradation after the update, the failure closed-loop update module also performs a closed-loop acceptance judgment, and only accepts the update result if the safety hard constraint satisfaction rate and the verification scenario bin coverage compliance rate do not decrease after the update.

[0098] Specifically, the data flow in the system is as follows: the decision state output by the decision state construction module. Input to the target condition policy module; candidate control output by the target condition policy module. and control credibility Input is given to the constrained contraction projection optimization module; the feasible control output of the constrained contraction projection optimization module is... On one hand, the data is sent to the flight control system for execution; on the other hand, it is input to the scenario coverage completion module and the failure closed-loop update module along with the minimum constraint margin. The verification scenario set output by the scenario coverage completion module is input to the simulation-flight consistency determination module. The index difference and pass / fail flag output by the simulation-flight consistency determination module are input to the failure closed-loop update module. The update parameters output by the failure closed-loop update module are then written back to the target condition strategy module, the constraint contraction projection optimization module, and the scenario coverage completion module, thus forming a system-level closed loop.

[0099] The technical solution of the present invention has been described in detail above with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that, without departing from the spirit and essence of the present invention, various equivalent substitutions, modifications, or combinations can be made to the specific implementation, module division, parameter settings, model structure, constraint construction method, scene parameter selection method, optimization solution method, statistical judgment method, and closed-loop update method of the present invention, but such equivalent substitutions, modifications, or combinations should all fall within the protection scope defined by the present invention.

Claims

1. A method for autonomous flight decision-making and verification of an unmanned intelligent agent, characterized in that, include: At any moment Receive task objectives issued by the task planning layer Synchronously acquire the unmanned intelligent agent at any time Flight status Environmental constraints and model uncertainty And merged into moments Decision-making status , decision state Input the target conditional policy network and obtain the time step. candidate control ; based on , , and Generate constraint matrix With constraint boundary vector And based on the control variables to be optimized ,time Deviation weighted matrix Constraint shrinkage coefficient Risk upper bound function Risk slack variables and risk penalty coefficient Construct a constrained contraction projection optimization problem and obtain the time step. Feasible control The constrained shrinking projection optimization problem satisfies: , Constructing a scenario to be verified Corresponding scene parameter vector and based on Get the scene Risk estimate Then based on the scenario Corresponding coverage gap , Positive smoothing term and risk emphasis coefficient Generate verification scenario sampling distribution The verification scenario set is completed based on the sampling distribution of the verification scenarios, and the sampling distribution of the verification scenarios satisfies: , Obtain data based on the verification scenario set respectively. The simulation and flight performance indicators obtained are used to determine consistency according to the preset equivalent boundaries; When inconsistencies exist between simulation and actual flight, or when safety failures occur, the decision state, including at least the moment of failure, should be extracted. Candidate control Feasible control Minimum constraint margin and scenes Key Indicator Index Corresponding simulation-flight difference The failed segments are processed, and the parameters of the target conditional policy network are updated simultaneously. Constraint shrinkage parameter set and the set of sampling parameters for verification scenarios And update the target condition policy network parameters. Constraint shrinkage parameter set , Verification scenario sampling parameter set These factors respectively affect the generation of subsequent candidate controls and the constraint contraction coefficient. Generation and verification scenario sampling distribution The generation of .

2. The autonomous flight decision-making and verification method for unmanned intelligent agents according to claim 1, characterized in that, The target condition policy network is based on time... Decision-making status As input, simultaneously output time. candidate control and control credibility ,satisfy: , in, The parameter is Target conditional policy network; The deviation weighting matrix For the deviation weights from the velocity dimension Heading deviation weight and deviation weights of rising and falling dimensions The diagonal matrix formed, the first Deviation weights corresponding to each control dimension Based on basic weights Credibility adjustment coefficient and control credibility Together, we determine that the following conditions must be met: , , Enhancing the credibility of control When the weight is reduced, the deviation weight decreases, thereby increasing the impact on the candidate control. The corrected degrees of freedom, , .

3. The autonomous flight decision-making and verification method for unmanned intelligent agents according to claim 1, characterized in that, time Constraint contraction coefficient From the basic shrinkage Weighting coefficients and ,time uncertainty vector The second norm and time Probability of communication loss Together, we determine that the following conditions must be met: 。 4. The method for autonomous flight decision-making and verification of unmanned intelligent agents according to claim 1, characterized in that, At the moment of seeking Feasible control Then, further calculations of time were performed. The constraint margins corresponding to each constraint and minimum constraint margin The constraint margin and minimum constraint margin satisfy: , , among which, time constraint boundary vector The Each component is ,time constraint matrix The Behavior ,time The The contraction coefficient corresponding to each constraint is: ,time The The uncertainty corresponding to each constraint is: ; When the minimum constraint margin Below the preset threshold At that time, according to the time Feasible control In the short time domain Internal prediction yields a set of candidate trajectory segments. And based on the affected area Construct the affected sub-segment index set The affected sub-segment index set satisfy: , Subsequently, only the affected sub-segment index set... The corresponding trajectory segments undergo local replanning, while the unaffected segments remain unchanged.

5. The method for autonomous flight decision-making and verification of unmanned intelligent agents according to claim 1, characterized in that, The coverage gap is also based on the scene parameter vector. One-dimensional quantile binning and two-dimensional joint quantile binning calculations were performed, with a minimum sample size limit set for each bin. Only when both the one-dimensional quantile binning and the corresponding two-dimensional joint quantile binning satisfy the minimum sample size lower limit. At that time, it is determined that the verification scenario coverage meets the standards.

6. The method for autonomous flight decision-making and verification of unmanned intelligent agents according to claim 1, characterized in that, The scenario risk estimate From the scene Risk mean Scene risk standard deviation and risk amplification factor Together, we determine that the following conditions must be met: 。 7. The autonomous flight decision-making and verification method for unmanned intelligent agents according to claim 1, characterized in that, Let the key indicator index be Key metrics include arrival time, path length, energy consumption per unit, minimum safety interval, number of boundary crossings, and duration of loss of contact. Get Scene Up indicators Simulation values And actual flight value and construct the scene Up indicators Simulation-Real Flight Difference Sample ,satisfy: , Let the significance level be... The number of key indicators is The significance level after multiple comparison correction is... satisfy: , index Equivalent boundary By indicators Basic Equivalent Boundary Risk level adjustment coefficient and task risk level Together, we determine that the following conditions must be met: , And for each key indicator Perform a two-tailed one-sided equivalence test on the difference samples to obtain the corresponding confidence intervals for the mean difference. Only when and At that time, determine the key indicators. Consistent, among which, Indicators The lower bound of the confidence interval for the mean difference. Indicators The upper bound of the confidence interval for the difference between the means.

8. The method for autonomous flight decision-making and verification of unmanned intelligent agents according to claim 1, characterized in that, The failed segments extracted in claim 1 are classified according to the failure index. Prioritize the failure index satisfy: , in, This serves as an index for key indicators. For key indicator quantity, As an indicator The corresponding weighting coefficients, For the scene Up indicators Simulation-real flight difference As an indicator The equivalent boundary, These are the weighting coefficients corresponding to the constraint margin term. To constrain the margin threshold, For a moment The minimum constraint margin.

9. The autonomous flight decision-making and verification method for unmanned intelligent agents according to claim 8, characterized in that, Feasible control after projection As a monitoring signal, and with parameters as The target conditional policy network in state Down-output control With the aforementioned feasible control The L2 norm squared error and the failure penalty coefficient Weighted failure index The parameters are constructed together. The update target satisfies: , And based on the failure index Synchronously adjust the set of constraint contraction parameters and the set of sampling parameters for the verification scenario Among them, the updated set of constraint shrinkage parameters Constraint contraction coefficients used to generate subsequent time steps Updated set of sampling parameters for verification scenarios Used to generate sampling distribution for subsequent verification scenarios ; Let the satisfaction rates of the hard safety constraints before and after the update be respectively. and The compliance rates for bin coverage in the verification scenarios before and after the update were respectively: and The update will be accepted only if the following condition is met: 。 10. An autonomous flight decision-making and verification system for unmanned intelligent agents, characterized in that, include: The decision state construction module is used to receive the task objectives issued by the task planning layer, synchronously acquire the flight state, environmental constraints and model uncertainties of the unmanned intelligent agent, and integrate them into the decision state; The target conditional policy module is used to input the decision state into the target conditional policy network to obtain candidate control and control confidence. The constrained contraction projection optimization module is used to generate a constraint matrix and constraint boundary vector based on the flight state, environmental constraints, model uncertainty and candidate control, and to generate a deviation weighting matrix based on the control confidence. Then, based on the control variable to be optimized, the deviation weighting matrix, the constraint contraction coefficient, the risk upper bound function, the risk relaxation variable and the risk penalty coefficient, a constrained contraction projection optimization problem is constructed to obtain feasible control. The scene coverage completion module is used to construct the scene parameter vector corresponding to the scene to be verified, obtain the scene risk estimate based on the scene parameter vector, generate the verification scene sampling distribution based on the coverage gap corresponding to the scene, the scene risk estimate, the positive smoothing term and the risk emphasis coefficient, and complete the verification scene set based on the verification scene sampling distribution. The simulation-flight consistency determination module is used to obtain the simulation indicators and flight indicators obtained based on the feasible control execution on the verification scenario set, and perform consistency determination according to the preset equivalent boundary. The failure closed-loop update module is used to extract failure segments when there is a simulation-flight inconsistency or safety failure. These segments include at least the decision state at the time of failure, candidate control, feasible control, minimum constraint margin, and the simulation-flight difference corresponding to the key index of the scenario. Simultaneously, it updates the target condition policy network parameters, constraint contraction parameter set, and verification scenario sampling parameter set. The updated target condition policy network parameters, constraint contraction parameter set, and verification scenario sampling parameter set are then applied to the subsequent candidate control generation, constraint contraction coefficient generation, and verification scenario sampling distribution generation, respectively.