An autonomous decision optimization method for an aerial robot

By calculating environmental impact parameters and constructing a hybrid control framework and hierarchical control architecture, the control weights and decision-making strategies of the aerial robot are dynamically adjusted, solving the environmental adaptability and safety issues in the decision optimization of the aerial robot and achieving more flexible and reliable flight control.

CN120871978BActive Publication Date: 2025-11-25ZHEJIANG FULIN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511373631.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2025-11-25
Estimated Expiration
2045-09-25

AI Technical Summary

Technical Problem

Existing decision optimization methods for aerial robots cannot adaptively adjust the granularity and focus of feature representations according to changes in the complexity of the aerial environment, resulting in the loss of key information or overload of redundant information, insufficient accuracy in path planning and inadequate safety assessment. Furthermore, fixed weight allocation strategies cannot dynamically adjust decision preferences, leading to inflexible and poorly adaptable decision results.

Method used

By acquiring flight environment data of aerial robots, obstacle density factors, weather impact factors, and altitude risk factors are calculated to generate environmental impact parameters. Parallel control branches in a hybrid control framework are constructed, control weights are dynamically adjusted, and a hierarchical control architecture and state-weight mapping controller are built to achieve autonomous decision optimization.

Benefits of technology

It improves the flight control flexibility and robustness of aerial robots, enhances the accuracy and safety of path planning, ensures the adaptability and reliability of decision-making, and improves the quality of flight mission completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871978B_ABST
    Figure CN120871978B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of robot navigation, in particular to an autonomous decision optimization method for an aerial robot. Flight environment data are acquired, an obstacle density factor, a weather influence factor and a height risk factor are calculated, and environment influence parameters are generated to be used for dynamically adjusting a control strategy; a hybrid control framework is constructed, containing two parallel branches of global path tracking control and local obstacle avoidance control, and the control weight is dynamically adjusted according to the environment influence parameters; a three-layer hierarchical control architecture is established, including a macro task constraint layer, a medium maneuver decision layer and a micro flight control layer, the upper layer constraints are converted into bottom layer control instructions through a control constraint propagation mechanism, and consistency and safety verification are carried out; a state-weight mapping controller is adopted, the multi-objective optimization weight coefficient is dynamically adjusted according to a real-time state vector, and optimal control instructions are generated. The application improves the autonomous decision-making ability and flight safety of the aerial robot in a dynamic environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot navigation, in particular to an autonomous decision optimization method for an aerial robot. BACKGROUND

[0002] The existing aerial robot decision optimization method mainly has the following technical defects:

[0003] The traditional deep neural network adopts a fixed feature extraction layer, which cannot adaptively adjust the granularity and focus of feature representation according to the complexity of the aerial environment. When facing different flight altitudes, weather conditions and obstacle densities, the fixed feature extraction method leads to loss of key information or overload of redundant information, affecting the decision accuracy.

[0004] The existing knowledge reasoning technology discretizes the continuous three-dimensional space into a fixed grid for symbolic representation. This coarse-grained discretization method is difficult to accurately express the continuity constraints of the aerial robot flight path and the real-time changes of dynamic obstacles, resulting in inaccurate path planning and insufficient safety evaluation.

[0005] The existing method adopts a preset fixed weight allocation strategy when dealing with multiple decision objectives such as safety, efficiency and energy consumption, which cannot dynamically adjust the decision preference according to the real-time flight state, task urgency and environmental risk, resulting in inflexible and poor adaptive decision results.

[0006] Therefore, an autonomous decision optimization method for an aerial robot is proposed. SUMMARY

[0007] The purpose of the present application is to provide an autonomous decision optimization method for an aerial robot.

[0008] To achieve the above purpose, the present application provides the following technical solutions:

[0009] An autonomous decision optimization method for an aerial robot, comprising:

[0010] Obtain flight environment data of the aerial robot, and generate environment impact parameters for dynamically adjusting control strategies by calculating obstacle density factor, weather influence factor and height risk factor;

[0011] According to the environment impact parameters, in a hybrid control framework containing global path tracking control and local obstacle avoidance control, the control weights of the two are dynamically adjusted through parallel control branches, the optimality of the global path is prioritized in low-impact environments, and the safety of local avoidance is prioritized in high-impact environments, realizing adaptive allocation of flight control focus;

[0012] In a hierarchical control architecture composed of a macro task constraint layer, a meso maneuver decision layer and a micro flight control layer, the flight constraints and allowed maneuver behaviors defined by the upper layer are converted into specific aircraft control instructions at the bottom layer through control constraint propagation, and the control instructions are verified for consistency and safety;

[0013] According to the real-time state vector of the aircraft, the weight coefficients of multiple optimization objectives in the flight control law are dynamically adjusted through a state-weight mapping controller, optimal control instructions are generated and applied to the aircraft.

[0014] Further, the flight environment data includes obstacle data, weather data and terrain airspace data.

[0015] Further, the process of generating environment impact parameters for dynamically adjusting control strategies by calculating obstacle density factor, weather influence factor and height risk factor is:

[0016] The obstacle density factor is calculated by evaluating the degree of interference of obstacles to path planning through a three-dimensional space occupancy grid;

[0017] The weather influence factor is calculated by evaluating the influence of wind speed, visibility and precipitation intensity on flight stability;

[0018] The height risk factor is calculated by combining airspace restrictions and terrain undulations to evaluate the risk of height selection;

[0019] Finally, these evaluation results are integrated into quantitative environmental impact parameters according to the control strategy of obstacle density priority, weather influence secondary and height risk supplementary.

[0020] Further, the process of realizing adaptive allocation of flight control focus through parallel control branches in the hybrid control framework is: through the small-range perception control branch, fine receptive field is used to extract control information for real-time obstacle avoidance, and through the large-range perception control branch, wide receptive field is used to extract navigation control information for global path tracking; The environmental impact parameters are used to adjust the contribution of each control branch information in the final control decision, when the environmental impact parameter is low, the large-range control branch weight is increased to focus on global navigation control, and when the environmental impact parameter is high, the small-range control branch weight is increased to focus on local obstacle avoidance control.

[0021] Further, the hierarchical control architecture composed of a macro task constraint layer, a meso maneuver decision layer and a micro flight control layer specifically includes:

[0022] a macro task constraint layer for defining the geometric boundaries of flight tasks, no-fly zones and airspace control rules, storing airspace control constraints in a geometric-semantic hybrid representation; a meso maneuver decision layer for providing a set of allowed flight maneuver control actions according to the macro control constraints, representing maneuver control decisions in fuzzy production rules; a micro flight control layer for accurately converting the selected maneuver control action into direct control parameters of the aircraft.

[0023] Further, the process of converting the control constraints of the macro task constraint layer into specific aircraft control instructions at the bottom layer by control constraint propagation, and verifying the consistency and safety of the control instructions, is as follows:

[0024] The control constraint conditions of the macro task constraint layer are used to filter the maneuver control action options in the meso maneuver decision layer, calculate the constraint violation degree of each control action and generate control action weights; the micro flight control layer maps the maneuver control action selected by the meso maneuver decision layer into a set of initial control parameters according to the current flight control state; before output to the actuator, the control parameters are checked for safety boundaries to ensure that they will not cause the aircraft to stall or exceed the attitude control limit, and safety control correction is performed.

[0025] Further, the process of dynamically adjusting the weight coefficients of multiple optimization objectives in the flight control law by the state-weight mapping controller according to the real-time state vector of the aircraft includes:

[0026] A real-time control state vector is constructed, which includes the energy state of the aircraft itself, the kinematic state, and the state related to the task and the environment; the energy state includes the battery power percentage, the kinematic state includes the flight height value and the current speed, and the task environment state includes the task urgency score, the weather risk level and the obstacle density value;

[0027] The state-weight mapping controller uses a control state encoder and a control weight generator to convert the control state vector into a four-dimensional control weight output of safety, efficiency, energy consumption and accuracy; the control state encoder converts the six-dimensional control state vector through the first control processing layer, the second control processing layer and the third control processing layer for dimension compression, and converts it into a control state feature representation; the control weight generator converts it into a four-dimensional control weight output through the fourth control processing layer and the output control processing layer.

[0028] Further, the dynamic adjustment mechanism of the state-weight mapping controller specifically includes:

[0029] The state-weight mapping controller takes the real-time control state vector of the aerial vehicle as input, and outputs a set of weight coefficients for balancing the four control objectives of safety, efficiency, energy consumption and accuracy; the controller has a built-in safety priority control mechanism that forces the weight of the safety control objective to be raised when high-risk control states such as low battery and far from the landing point and imminent collision are detected; at the same time, a control focus smooth switching mechanism is established, and the control weight is smoothed by a time sliding window.

[0030] Compared with the prior art, the beneficial effects of the present application are:

[0031] 1. By mixing the parallel control branches in the control framework, using small-range and large-range perception control branches to extract local obstacle avoidance and global navigation information respectively, and dynamically adjusting the weight contribution of each branch according to the environmental impact parameters, the adaptive allocation of flight control focus can be realized, so as to optimize global path tracking in low-risk environment and prioritize local obstacle avoidance safety in high-risk environment, thereby improving the flexibility and robustness of overall flight control of aerial robots.

[0032] 2. By constructing a hierarchical control architecture of macro task constraint layer, medium maneuver decision layer and micro flight control layer, task constraints, maneuver decisions and control parameter generation are processed respectively, which can decompose complex flight tasks into clear control levels, realize the orderly transmission from high-level constraints to low-level control, and improve the generation efficiency of control instructions and the structure and controllability of decisions; by converting control constraints into bottom-level control instructions step by step and verifying consistency and safety, high-level constraints can effectively constrain medium and low-level decisions, while safety boundary checking and correction can avoid aircraft stall or attitude loss of control, thereby improving the reliability of control instructions and the safety of flight process.

[0033] 3. By constructing a real-time state vector containing energy, kinematics and task environment state, and using a state-weight mapping controller to dynamically adjust the weight coefficients of safety, efficiency, energy consumption and accuracy, the control law can be optimized according to the real-time state of the aerial vehicle, so that the control instructions can better balance the multi-objective requirements, thereby improving the adaptability of decisions and the completion quality of flight tasks. BRIEF DESCRIPTION OF DRAWINGS

[0034] Fig. 1 The flowchart of the autonomous decision optimization method of an aerial robot of the present application;

[0035] Fig. 2 The flowchart of the obstacle density factor calculation process of the present application;

[0036] Fig. 3 The structure diagram of the state-weight mapping controller of the present application. DETAILED DESCRIPTION

[0037] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0038] Please refer to Figs. 1 to 3 The present application provides an autonomous decision optimization method for an aerial robot, and the technical solutions are as follows:

[0039] Embodiment one:

[0040] In order to improve the intelligence of the aerial robot, an enterprise uses an autonomous decision optimization method for an aerial robot proposed by the present application. The flow of the method can refer to Fig. 1 , and specifically includes:

[0041] Obtain flight environment data of the aerial robot, calculate obstacle density factor, weather influence factor and height risk factor, and generate environment influence parameters for dynamically adjusting control strategy;

[0042] Further, the flight environment data includes obstacle data, meteorological data and terrain airspace data.

[0043] Further, three-dimensional obstacle point cloud data, meteorological data and ground elevation data are obtained by the laser radar sensor, meteorological sensor and RTK-GPS / INS integrated navigation system carried by the aerial robot respectively. The meteorological data includes wind speed, visibility and precipitation intensity information. The terrain airspace data includes airspace classification rules, airspace control information and the like in addition to the ground elevation data.

[0044] By specifying that the flight environment data includes obstacle data, meteorological data and terrain airspace data, comprehensive environment information input is provided for the aerial robot, ensuring that the calculation of subsequent environment influence parameters can accurately reflect the complexity of the actual flight scene, thereby improving the reliability and adaptability of autonomous decision.

[0045] Further, the process of generating environment influence parameters for dynamically adjusting control strategy by calculating obstacle density factor, weather influence factor and height risk factor is as follows:

[0046] The interference degree of obstacles to path planning is evaluated by a three-dimensional space occupancy grid, and the obstacle density factor is calculated.

[0047] The influence of wind speed, visibility and precipitation intensity information on flight stability is evaluated, and the weather influence factor is calculated.

[0048] Assess the risks of altitude selection by combining airspace restrictions and terrain undulations, and calculate the altitude risk factor;

[0049] Finally, these assessment results were integrated into quantitative environmental impact parameters according to the control strategy of prioritizing obstacle density, secondary weather impact, and supplementing with high risk.

[0050] Furthermore, the flowchart for calculating the obstacle density factor is shown below. Fig. 2 As shown, specifically: a three-dimensional spatial occupancy grid map is established, dividing the flight space into cubic grid cells with a side length of 1m. Each grid cell stores the degree of interference of obstacles on trajectory planning. Point cloud data generated by lidar is mapped to the corresponding grid cells. For each grid cell, the degree of trajectory planning interference is represented as a weighted sum of an altitude factor, a path blocking factor, and a detour cost factor. The altitude factor uses the normalized value of the difference between the obstacle height and the flight altitude to assess whether the obstacle obstructs the trajectory within the flight altitude range. The path blocking factor is calculated... The obstacle's projection occupancy ratio in the grid is used to calculate the ratio of the offset distance required for detours using the detour cost factor. When the grid's interference level exceeds the threshold of 0.7, it is marked as a high-interference state. The basic occupancy rate of the obstacle density factor is calculated by dividing the number of high-interference grids by the total number of effective grids. The spatial distribution variance is analyzed by examining the standard deviation of the coordinate values ​​of high-interference grids on the x, y, and z axes to determine the uneven distribution of high-interference grids in the trajectory planning space. Finally, the obstacle density factor is calculated by multiplying the basic occupancy rate by 0.7 and adding the normalized spatial distribution variance multiplied by 0.3.

[0051] Furthermore, the weather impact factor is expressed as a weighted sum of the wind speed impact value, visibility impact value, and precipitation impact value, with the final result constrained to the range of 0 to 1 using a hyperbolic tangent function. Specifically, the wind speed impact value is obtained by mapping the current wind speed to the critical wind speed using a sigmoid function; the visibility impact value is the standard visibility value minus the current visibility value, divided by the standard visibility value; and the precipitation impact value is expressed as a weighted sum of the ratio of precipitation intensity to the ratio of snowfall intensity.

[0052] Furthermore, the calculation process for the altitude risk factor is as follows: Different altitude penalty coefficients are set according to airspace classification. For example, the coefficient is 0.8 for ultra-low altitude (0 to 50 meters), 0.3 for low altitude (50 to 150 meters), 0.1 for medium altitude (150 to 300 meters), and 0.5 for high altitude (above 300 meters). The terrain elevation is calculated by sampling terrain elevation points within a 5 km × 5 km radius around the aircraft, and the ratio of the elevation standard deviation to the average elevation is used to obtain the terrain undulation. When the terrain undulation is less than 0.1, it is considered flat terrain; when it is greater than 0.3, it is considered complex mountainous terrain. Therefore, the altitude risk factor can be expressed as the sum of the current altitude penalty coefficient multiplied by 1 and the terrain undulation.

[0053] Further, the calculation formula of the environmental impact parameter is the obstacle density factor multiplied by 0.4, plus the weather impact factor multiplied by 0.3, plus the height risk factor multiplied by 0.3; in order to ensure that the output is within the range of 0 to 1, the environmental impact parameter is normalized using the hyperbolic tangent function.

[0054] By calculating the obstacle density factor, the weather impact factor and the height risk factor, and integrating them into the quantitative environmental impact parameter according to the priority, the influence of different factors in the flight environment on the flight path planning and flight stability can be accurately evaluated, which provides a scientific basis for dynamically adjusting the control strategy, thereby improving the safety and decision efficiency of the aerial robot in complex environments.

[0055] According to the environmental impact parameter, in a hybrid control framework including global path tracking control and local obstacle avoidance control, the control weights of the two are dynamically adjusted through parallel control branches, the optimality of the global flight path is prioritized in low-impact environments, and the safety of local avoidance is prioritized in high-impact environments, realizing adaptive allocation of flight control focus;

[0056] Further, the process of realizing adaptive allocation of flight control focus through parallel control branches in the hybrid control framework is: the small-range perception control branch uses a fine receptive field to extract control information for real-time obstacle avoidance, and the large-range perception control branch uses a wide receptive field to extract navigation control information for global path tracking; the environmental impact parameter is used to adjust the contribution of information from each control branch in the final control decision, and when the environmental impact parameter is low, the weight of the large-range control branch is increased to focus on global navigation control, and when the environmental impact parameter is high, the weight of the small-range control branch is increased to focus on local obstacle avoidance control.

[0057] Further, the small-range perception control branch uses a fine receptive field to extract control information for real-time obstacle avoidance, uses a convolutional neural network with a 3x3 convolution kernel, has 256 input channels and 128 output channels, and the receptive field radius is set to 15m, with a processing frequency of 20Hz, which is specifically used to capture local obstacle features and accurate boundary information for local obstacle avoidance control; the large-range perception control branch uses a wide receptive field to extract navigation control information for global path tracking, uses a convolutional neural network with a 7x7 convolution kernel, has 256 input channels and 128 output channels, and the receptive field radius is set to 100m, with a processing frequency of 5Hz, which is used to capture global context information such as terrain trend and flight path planning for global path tracking control.

[0058] Further, the weight dynamic adjustment mechanism of the parallel control branch is as follows: the environmental influence parameter is used to adjust the contribution degree of the small-range control branch and the large-range control branch information in the final control decision; the small-range control branch weight calculation adopts an S-shaped function, and the difference between the environmental influence parameter and 0.5 is multiplied by 2 to obtain the result sent into the Sigmoid function for calculation; the large-range control branch weight is 1 minus the small-range control branch weight, ensuring that the sum of the two weights is always 1; when the environmental influence parameter is low and less than 0.3, it is a low-influence environment, and the large-range control branch weight is not less than 0.8, focusing on global navigation control; when the environmental influence parameter is high and greater than 0.7, it is a high-influence environment, and the small-range control branch weight is not less than 0.8, focusing on local obstacle avoidance control;

[0059] Further, the parallel control branch has a dynamic receptive field adaptive adjustment mechanism, specifically as follows: the receptive field radius of the small-range perception control branch and the large-range perception control branch is related to the basic radius, the speed adjustment coefficient, and the complexity adjustment coefficient; the speed adjustment coefficient is a combination of the ratio of the current speed to the maximum speed, a multiplicative coefficient, and an additive coefficient; the complexity adjustment coefficient is a combination of the environmental influence parameter, a multiplicative coefficient, and an additive coefficient; the coefficients in the speed adjustment coefficient and the complexity adjustment coefficient of the two branches are different, thereby realizing intelligent matching of the perception range and the flight state, and improving the safety and efficiency under different flight conditions.

[0060] Through the parallel control branch in the hybrid control framework, local obstacle avoidance and global navigation information are extracted by the small-range and large-range perception control branches respectively, and the weight contribution of each branch is dynamically adjusted according to the environmental influence parameter, which can realize adaptive allocation of flight control focus, thereby optimizing global path tracking in a low-risk environment and prioritizing local obstacle avoidance safety in a high-risk environment, improving the flexibility and robustness of overall flight control.

[0061] In a hierarchical control architecture composed of a macro task constraint layer, a meso maneuver decision layer, and a micro flight control layer, the flight constraints and allowed maneuver behaviors defined by the upper layer are converted into specific aircraft control instructions at the bottom layer through control constraint propagation, and consistency and safety verification are performed on the control instructions;

[0062] Further, the hierarchical control architecture composed of a macro task constraint layer, a meso maneuver decision layer, and a micro flight control layer specifically includes:

[0063] a macro task constraint layer for defining geometric boundaries of flight tasks, no-fly zones and airspace control rules, storing airspace control constraints in a geometric-semantic hybrid representation; a meso maneuver decision layer for providing a set of allowed flight maneuver control actions according to the macro control constraints, representing maneuver control decisions in fuzzy production rules; and a micro flight control layer for accurately converting the selected maneuver control action into direct control parameters of the aircraft;

[0064] Further, the airspace division precision is a 5m*5m grid, and the no-fly zone data is stored in an R-tree index structure; the airspace control constraints include a boundary point coordinate array, a height range, a constraint type and a constraint boundary value;

[0065] Further, the flight maneuver control actions provided by the meso maneuver decision layer include forward, backward, left turn, right turn, ascend, descend, hover, precise approach and rapid retreat, each of which is associated with 5 to 8 fuzzy rules;

[0066] Further, the control instruction update frequency of the micro flight control layer is 100Hz, and it supports control parameters such as pitch angle, roll angle, thrust and yaw angular velocity;

[0067] By constructing the hierarchical control architecture of the macro task constraint layer, the meso maneuver decision layer and the micro flight control layer, the task constraints, maneuver decisions and control parameter generation are processed respectively, which can decompose the complex flight task into clear control levels, realize the orderly transmission from high-level constraints to low-level control, and thus improve the generation efficiency of control instructions and the structuralization and controllability of decisions.

[0068] Further, the process of converting the control constraints into specific aircraft control instructions through control constraint propagation and verifying the consistency and safety of the control instructions is as follows:

[0069] The control constraint conditions of the macro task constraint layer are used to filter the maneuver control action options in the meso maneuver decision layer, the constraint violation degree of each control action is calculated and the control action weight is generated; the micro flight control layer maps the maneuver control action selected by the meso maneuver decision layer into a set of initial control parameters according to the current flight control state; before output to the actuator, the control parameters are checked for safety boundary to ensure that they will not cause the aircraft to stall or exceed the attitude control limit, and safety control correction is performed;

[0070] Further, the calculation formula of the constraint violation degree of each control behavior can represent the absolute value of the difference between the control parameter and the constraint boundary value divided by the constraint deviation tolerance; and the control behavior weight is calculated by using an exponential decay function, and the weight of a single control behavior can represent the calculation result of the exponential decay function of the constraint violation degree of the behavior divided by the summation result of the exponential decay functions of all behaviors, thereby being normalized; the sorting result of the control behavior weight is taken as the basis for the macro maneuver decision layer to select the maneuver control action;

[0071] Further, the micro flight control layer converts the maneuver control action selected by the macro maneuver decision layer into a set of initial control parameters through a parameter mapping table; taking "forward maneuver" as an example, the mapping parameters include setting the pitch angle to -10 degrees, setting the thrust coefficient to 0.5, keeping the roll angle to 0 degrees, and setting the yaw angular velocity to 0 degrees per second; the parameter mapping considers the current state correction, such as reducing the thrust coefficient to 0.3 when the current speed has reached 10 meters per second, and increasing the corresponding roll angle compensation when encountering crosswinds; before outputting to the actuator, three layers of safety boundary checks are performed: the first layer checks whether the attitude angle is within the safe range, the second layer checks whether the thrust exceeds the rated power, and the third layer checks whether the speed is within the safe range; when any parameter exceeds the safety boundary, the nearest projection method is used to adjust the parameter to within the safety boundary; the safety-corrected parameters are sent to the flight controller for execution.

[0072] By propagating control constraints to convert into bottom-level control instructions step by step and performing consistency and safety verification, it can ensure that high-level constraints effectively constrain middle and low-level decisions, while avoiding aircraft stall or attitude loss of control through safety boundary checking and correction, thereby improving the reliability of control instructions and the safety of flight processes.

[0073] According to the real-time state vector of the aircraft, the state-weight mapping controller dynamically adjusts the weight coefficients of multiple optimization objectives in the flight control law, generates optimal control instructions, and acts on the aircraft.

[0074] Further, the process of dynamically adjusting the weight coefficients of multiple optimization objectives in the flight control law according to the real-time state vector of the aircraft through the state-weight mapping controller includes:

[0075] A real-time control state vector is constructed, which includes the energy state, kinematic state, and task and environment related state of the aircraft itself; the energy state includes the battery power percentage, the kinematic state includes the flight altitude value and the current speed, and the task and environment state includes the task urgency score, weather risk level, and obstacle density value;

[0076] The structure of the state-weight mapping controller is as follows Fig. 3As shown, the control state vector is converted into four-dimensional control weight outputs of safety, efficiency, energy consumption and accuracy by using a control state encoder and a control weight generator; the control state encoder sequentially compresses the six-dimensional control state vector through a first control processing layer, a second control processing layer and a third control processing layer to convert it into a control state feature representation; the control weight generator converts it into a four-dimensional control weight output through a fourth control processing layer and an output control processing layer;

[0077] Further, the first control processing layer uses a fully connected layer and a ReLU activation function to expand the control state vector from 6 dimensions to 64 dimensions, the second control processing layer uses a fully connected layer and a ReLU activation function to compress from 64 dimensions to 32 dimensions, and the third control processing layer uses a fully connected layer and a ReLU activation function to compress from 32 dimensions to 16 dimensions; the fourth control processing layer uses a fully connected layer and a ReLU activation function to compress from 16 dimensions to 8 dimensions, and the output control processing layer uses a Softmax activation function to output from 8 dimensions to 4 dimensions, ensuring that the sum of the four control weights of safety, efficiency, energy consumption and accuracy is 1;

[0078] Further, a multi-head attention mechanism is additionally introduced before the first control processing layer of the control state encoder, and the process of identifying the importance correlation of different elements in the state vector is as follows: the original 6-dimensional state vector is reorganized into a sequence form and converted into a state feature matrix through linear transformation; the input state feature matrix is respectively passed through three different linear transformation layers to generate a query matrix, a key matrix and a value matrix; the correlation weight between state elements is calculated through a multi-head attention layer, and feature fusion is performed through a linear transformation layer to generate an enhanced state feature representation; in a multi-target conflict scenario, the weight distribution can be more accurately balanced to ensure the rationality of the weight coefficient adjustment of the optimization target.

[0079] By constructing a real-time state vector containing energy, kinematics and task environment state, and dynamically adjusting the weight coefficients of safety, efficiency, energy consumption and accuracy using the state-weight mapping controller, the control law can be optimized according to the real-time state of the aircraft, so that the control command can better balance the multi-objective demand, thereby improving the adaptability of the decision and the completion quality of the flight task.

[0080] Further, the dynamic adjustment mechanism of the state-weight mapping controller specifically includes:

[0081] The state-weight mapping controller takes the real-time control state vector of the aircraft as input and outputs a set of weight coefficients for balancing the four control objectives of safety, efficiency, energy consumption, and accuracy. The controller has a safety-first control mechanism that forces the weight of the safety control objective to increase when high-risk control states such as low battery level and far from the landing point are detected. At the same time, a control focus smooth switching mechanism is established to smooth the control weights through a time sliding window.

[0082] Further, the minimum value of each control target weight is 0.05 to ensure that all objectives are considered; the maximum value of a single control target weight is 0.70 to avoid excessive bias towards a single objective; the situational adaptability constraints include that the safety weight is not less than 0.50 when the risk level is greater than 0.8, the efficiency weight is not less than 0.40 when the urgency is greater than 0.9, and the energy consumption weight is not less than 0.45 when the battery level is less than 0.2; the weight values are not unique.

[0083] Further, the weight update strategy is as follows: use the Adam optimizer for online learning with a learning rate of 0.001 and maintain an experience buffer of the last 1000 decisions; use exponential moving average for weight smoothing, with the new weight being the current weight multiplied by 0.8 plus the predicted weight multiplied by 0.2, and the single weight change limit being 0.1 to avoid drastic changes;

[0084] Further, the model confidence is evaluated by the variance of multiple prediction results, the average accuracy of past predictions, and the coverage of the current state in the training data. When the confidence is less than 0.7, the state-weight mapping controller switches to a conservative mode, increasing the safety weight and reducing the exploration rate;

[0085] Further, the weight coefficients of the control objectives also have a situational awareness boundary adaptive adjustment mechanism, which is to dynamically adjust the weight boundary values according to the specific flight situation through the establishment of a situation-boundary mapping table, rather than using fixed weight coefficient ranges. The situation-boundary mapping table is as follows: the upper limit of the efficiency weight is increased to 0.8 and the lower limit of the energy consumption weight is decreased to 0.02 under emergency rescue tasks; the upper limit of the accuracy weight is increased to 0.8 and the lower limit of the efficiency weight is decreased to 0.03 under precision measurement tasks; the upper limit of the energy consumption weight is increased to 0.8 and the lower limit of the accuracy weight is decreased to 0.04 under long-distance cruising; the weights remain at the standard boundaries during training flights. This mechanism allows the weight distribution to be more consistent with the task characteristics and avoids unreasonable weight constraints.

[0086] Through the dynamic adjustment mechanism of the state-weight mapping controller, combined with the safety priority control and the control focus smooth switching mechanism, the safety weight can be forced to be improved in the high-risk state, and the smooth processing is introduced in the weight adjustment to avoid the control instability caused by the mutation, thereby further enhancing the safety and control smoothness of the aerial robot in the complex dynamic environment.

[0087] The embodiment proposes an autonomous decision optimization method for an aerial robot. Flight environment data is obtained, obstacle density factor, weather influence factor and height risk factor are calculated, and environment influence parameters are generated for dynamic adjustment of control strategy; a hybrid control framework is constructed, including two parallel branches of global path tracking control and local obstacle avoidance control, and the control weight is dynamically adjusted according to the environment influence parameters; a three-layer hierarchical control architecture is established, including a macro task constraint layer, a medium maneuver decision layer and a micro flight control layer, the upper layer constraints are converted into bottom layer control instructions through a control constraint propagation mechanism, and consistency and safety verification are performed; a state-weight mapping controller is adopted, and the multi-objective optimization weight coefficient is dynamically adjusted according to the real-time state vector to generate optimal control instructions. The autonomous decision-making ability and flight safety of the aerial robot in the dynamic environment are improved.

[0088] Embodiment two:

[0089] The embodiment takes a UAV inspection task in a certain urban area as an example, and details the implementation process of an autonomous decision optimization method for an aerial robot, and shows the application effect of the method in a dynamic environment.

[0090] Flight environment data of the aerial robot is obtained, and environment influence parameters for dynamic adjustment of control strategy are generated by calculating obstacle density factor, weather influence factor and height risk factor;

[0091] The spatial coordinates and geometric information of the obstacles such as buildings, high-voltage lines and trees in the urban area are collected by the airborne laser radar; the obstacle density factor is calculated by analyzing the distribution variance of the obstacles, and the obstacle density factor value is 0.65. The weather influence factor is calculated by the weighted sum of the wind speed influence value, the visibility influence value and the rainfall influence value, and the specific value is 0.45. The height risk factor is calculated by combining the current height penalty coefficient and the terrain undulation, and the value is 0.67.

[0092] The environment influence parameter value is 0.58, which is obtained by combining the obstacle density factor, the weather influence factor and the height risk factor, and is used for subsequent control weight adjustment.

[0093] According to the environmental impact parameters, in a hybrid control framework containing global path tracking control and local obstacle avoidance control, the control weights of the two are dynamically adjusted through parallel control branches, prioritizing the optimality of global path in low impact environment and the safety of local avoidance in high impact environment, to realize adaptive allocation of flight control focus.

[0094] Through parallel control branches, local obstacle information and global navigation information are extracted by small and large range control branches respectively, and finally the branch weights are dynamically adjusted according to the loop impact parameters to realize flexible flight control information allocation; among them, the global control branch weight and the local control branch weight are 0.62 and 0.85 respectively; in high-risk areas such as dense building area, the local avoidance weight is increased to 0.9 to ensure the safe detour of the UAV.

[0095] In a hierarchical control architecture composed of macro task constraint layer, meso maneuver decision layer and micro flight control layer, the flight constraints and allowed maneuver behaviors defined by the upper layer are converted into specific aircraft control instructions at the bottom layer through control constraint propagation, and the consistency and safety of the control instructions are verified;

[0096] The macro task constraint layer is used to define the geometric boundary of flight task, no-fly zone and airspace control rules, and the airspace control constraints are stored in a geometric-semantic hybrid representation; the meso maneuver decision layer is used to provide a set of allowed flight maneuver control actions according to the macro control constraints, and the maneuver control decision is represented by fuzzy production rules; the micro flight control layer is used to accurately convert the selected maneuver action into direct control parameters of the aircraft;

[0097] According to the real-time state vector of the aircraft, the weight coefficients of multiple optimization objectives in the flight control law are dynamically adjusted through the state-weight mapping controller to generate optimal control instructions and act on the aircraft.

[0098] The state vector is converted into safety, efficiency, energy consumption and accuracy control weight output through the state-weight mapping controller; the built-in safety priority mechanism forces the safety weight to increase to 0.65 when the power is less than 30% or the obstacle distance is less than 10m, and the other weights are correspondingly reduced, ensuring that the total weight is 1.

[0099] Time sliding window is used to smooth the weights to avoid control instability caused by sudden changes; finally, the optimization weight coefficients are generated, such as safety 0.42, efficiency 0.24, energy consumption 0.19 and accuracy 0.15.

[0100] While embodiments of the application have been shown and described, it is to be understood that the embodiments described are merely exemplary of the principles and application of the present application. Numerous modifications and adaptions can be effected without departing from the spirit and scope of the present application, which is not limited to the exact construction and arrangement described. It is intended, therefore, to cover all modifications and adaptions that fall within the scope of the claims and their equivalents.

Claims

1. An autonomous decision-making optimization method for an aerial robot, characterized in that, include: The system acquires flight environment data of aerial robots and generates environmental impact parameters for dynamically adjusting control strategies by calculating obstacle density factors, weather impact factors, and altitude risk factors. Based on the environmental impact parameters, in a hybrid control framework that includes global path tracking control and local obstacle avoidance control, the control weights of the two are dynamically adjusted through parallel control branches. In low-impact environments, priority is given to ensuring the optimality of the global trajectory, while in high-impact environments, priority is given to ensuring the safety of local avoidance, thereby achieving adaptive allocation of flight control focus. In a hierarchical control architecture consisting of a macro-level task constraint layer, a meso-level maneuver decision layer, and a micro-level flight control layer, the flight constraints and permitted maneuver behaviors defined at the upper level are transformed into specific aerial robot control commands at the lower level through control constraint propagation, and the consistency and safety of the control commands are verified. Based on the real-time state vector of the aerial robot, the weight coefficients of multiple optimization objectives in the flight control law are dynamically adjusted by the state-weight mapping controller. The process includes: constructing a real-time control state vector that includes the aerial robot's own energy state, kinematic state, and state related to the task and environment; the state-weight mapping controller uses a control state encoder and a control weight generator to convert the control state vector into a four-dimensional control weight output that considers safety, efficiency, energy consumption, and accuracy; and generating the optimal control command and applying it to the aerial robot.

2. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The flight environment data includes obstacle data, meteorological data, and terrain and airspace data.

3. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The process of generating environmental impact parameters for dynamically adjusting control strategies by calculating obstacle density factors, weather impact factors, and high risk factors is as follows: The degree of interference of obstacles on trajectory planning is assessed by using a three-dimensional spatial occupancy grid map, and the obstacle density factor is calculated. The impact of wind speed, visibility, and precipitation intensity information on flight stability is assessed, and weather impact factors are calculated. Assess the risks of altitude selection by combining airspace restrictions and terrain undulations, and calculate the altitude risk factor; Finally, these assessment results were integrated into quantitative environmental impact parameters according to the control strategy of prioritizing obstacle density, secondary weather impact, and supplementing with high risk.

4. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The process of adaptively allocating flight control focus through parallel control branches in a hybrid control framework is as follows: a small-range perception control branch uses a fine receptive field to extract control information for real-time obstacle avoidance, and a large-range perception control branch uses a wide-area receptive field to extract navigation control information for global path tracking. The environmental impact parameter is used to adjust the contribution of each control branch's information to the final control decision. When the environmental impact parameter is low, the weight of the large-range control branch is increased to focus on global navigation control, and when the environmental impact parameter is high, the weight of the small-range control branch is increased to focus on local obstacle avoidance control.

5. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The hierarchical control architecture, consisting of a macro-level mission constraint layer, a meso-level maneuver decision layer, and a micro-level flight control layer, specifically includes: The macro-level task constraint layer defines the geometric boundaries, no-fly zones, and airspace control rules for flight missions, and uses a geometric-semantic hybrid representation to store airspace control constraints. The meso-level maneuver decision layer provides a set of permissible flight maneuver control actions based on macro-level control constraints, and uses fuzzy production rules to represent maneuver control decisions. The micro-level flight control layer accurately converts selected maneuver actions into direct control parameters for the aerial robot.

6. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The process of propagating control constraints step by step into specific aerial robot control commands at the lower levels, and verifying the consistency and security of these control commands, is as follows: The control constraints of the macro-level task constraint layer are used to filter the maneuver control action options in the meso-level maneuver decision layer, calculate the degree of constraint violation for each control action, and generate the control action weight. The micro-level flight control layer maps the maneuver control actions selected by the meso-level maneuver decision layer to a set of initial control parameters based on the current flight control state. Before outputting to the actuator, the control parameters are checked for safety boundaries to ensure that they will not cause the aerial robot to stall or exceed the attitude control limits, and safety control corrections are made.

7. The autonomous decision-making optimization method for an aerial robot according to claim 1, characterized in that, The process of dynamically adjusting the weight coefficients of multiple optimization objectives in the flight control law based on the real-time state vector of the aerial robot, through a state-weight mapping controller, includes: Construct a real-time control state vector that includes the aerial robot's own energy state, kinematic state, and mission and environment-related states; where the energy state includes battery percentage, the kinematic state includes flight altitude and current speed, and the mission environment state includes mission urgency score, weather risk level, and obstacle density. The state-weight mapping controller employs a control state encoder and a control weight generator to convert the control state vector into a four-dimensional control weight output that considers safety, efficiency, energy consumption, and accuracy. Specifically, the control state encoder compresses the six-dimensional control state vector sequentially through the first, second, and third control processing layers to convert it into a control state feature representation. The control weight generator then converts it into a four-dimensional control weight output through the fourth control processing layer and the output control processing layer.

8. The autonomous decision-making optimization method for an aerial robot according to claim 7, characterized in that, The dynamic adjustment of the flight control law by the state-weight mapping controller specifically includes: The state-weight mapping controller takes the real-time control state vector of the aerial robot as input and outputs a set of weight coefficients to balance the four control objectives of safety, efficiency, energy consumption and accuracy. The controller has a built-in safety priority control mechanism, which forcibly increases the weight of the safety control objective when a high-risk control state is detected. At the same time, a smooth switching mechanism for control focus is established, which smooths the control weights through a time sliding window.

Citation Information

Patent Citations

  • Intelligent agent autonomous navigation method based on deep reinforcement learning

    CN112179367A

  • Large unmanned aerial vehicle intelligent flight control integrated system

    CN120406541A