Novel pulse reinforcement learning-driven unmanned system brain-like intelligent multivariate decision-making method

By building task-specific state space, action space and reward functions, designing pulse triggering mechanism and frequency model, combining time-sharing multiplexing structure and time slot borrowing strategy, the task adaptability problem of the drone cluster in complex environments is solved, the system's response ability and stability are improved, and the autonomous collaborative decision-making of multi-tasks is realized.

CN120491452APending Publication Date: 2025-08-15NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510577278.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing drone cluster technology is difficult to adapt to changing task requirements and complex environments, and lacks detailed descriptions of task structure and physical laws, resulting in lack of stability and interpretability in strategy output, making it difficult to ensure the continuous and effective operation of the system in complex environments.

Method used

A new type of unmanned system brain-like intelligent multivariate decision-making method is adopted to build task-specific state space, action space and reward functions, design pulse trigger mechanism and frequency model, combine time-sharing multiplexing structure and time slot borrowing strategy, integrate data-driven strategies and prior knowledge models, dynamically adjust control strategies and flight inertia parameters, quantify task priorities and generate unified control instructions.

Benefits of technology

It improves the response capability and stability of the drone cluster in a multi-task environment, ensures task timing independence and scheduling efficiency, enhances the controllability and stability of the system in complex environments, and realizes the comprehensive collaboration capability of multi-task autonomous control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120491452A_ABST
    Figure CN120491452A_ABST
Patent Text Reader

Abstract

The invention discloses a novel pulse reinforcement learning driven unmanned system brain-like intelligent multivariate decision-making method. The method comprises the following steps: respectively establishing a state space, an action space and a reward function of detection identification, material delivery and flight maneuvering tasks; designing a pulse event triggering mechanism according to task characteristics, establishing a pulse frequency adjustment model of various tasks, and constructing a scheduling strategy of time division multiplexing and elastic borrowing; establishing a detection identification model, a material delivery model and a flight maneuvering model; a multi-modal variance index is introduced to dynamically adjust a control signal source by fusing a data strategy and a decision-making mechanism of a prior model; a sliding window state matrix is constructed, task priorities are dynamically calculated, and a unified control instruction is generated in a weighted fusion mode. According to the method, the response capability of the system to the multi-source task is improved, the time sequence independence and the scheduling efficiency of the task are guaranteed, and the stability and the controllability of the unmanned system in a complex environment are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a novel pulse reinforcement learning-driven brain-like intelligent multi-decision-making method for unmanned systems, belonging to the technical field of multi-comprehensive decision-making for unmanned systems. Background Art

[0002] In recent years, drone technology has rapidly developed and is widely used in various scenarios. Drone swarms, owing to their high efficiency and flexibility in executing complex tasks, have become a hot topic in research and application. Drone swarms, through the coordinated operation of multiple drones, can excel in tasks such as rescue, reconnaissance, and monitoring. However, in practical applications, the dynamic collaborative optimization of drone swarms still faces numerous challenges in the face of changing mission requirements and complex environmental conditions. Currently, drone swarms must handle a wide variety of tasks, including target detection, material delivery, and regional patrol. These tasks differ significantly in target form, resource requirements, and execution methods, requiring the system to distinguish between task characteristics and implement differentiated strategy modeling. Furthermore, factors such as sensor degradation, wind field disturbances, and dynamic obstacle changes may arise during mission execution, making mission success and flight safety highly dependent on real-time perception and response to the environment. Furthermore, the drones in a swarm exhibit heterogeneity in terms of payload capacity, computing performance, and powertrain systems. Different tasks also have varying requirements for perception accuracy, control frequency, and response latency. Therefore, how to rationally allocate computing resources and control frequency under the conditions of multi-tasking concurrency is crucial to system efficiency.

[0003] Existing drone swarm technologies often rely on static strategies or centralized scheduling schemes, making them difficult to adapt to the real-time scheduling requirements of multi-task environments. Furthermore, task modeling lacks a detailed description of task structure and physical laws, and policy output lacks stability and interpretability, making it difficult to ensure the system's continued effective operation in complex environments. Summary of the Invention

[0004] Purpose of the invention: In order to overcome the deficiencies in the prior art, the present invention provides a novel brain-like intelligent multi-decision-making method for unmanned systems driven by pulse reinforcement learning. The present invention is oriented to multi-source task requirements in complex dynamic environments, and achieves fine modeling and differentiated strategy optimization of multiple types of tasks by introducing task-specific state space, action space, and reward function structure. In addition, by combining the pulse mechanism with the timing scheduling method, the control frequency and execution rhythm are uniformly managed to avoid control resource conflicts and response delays during task concurrency. Finally, by integrating data-driven policy output with prior knowledge model constraints, the stability and physical consistency of decisions are enhanced while ensuring policy flexibility.

[0005] Technical solution: To achieve the above purpose, the technical solution adopted by the present invention is:

[0006] A novel brain-inspired intelligent multi-decision-making method for unmanned systems driven by pulse reinforcement learning includes the following steps:

[0007] Step 1: Construct the corresponding state space, action space, and reward function based on the multiple tasks of detection and identification, material delivery, and flight maneuvering.

[0008] Step 2: Based on the state space, action space, and reward function corresponding to the multi-task, a task-specific pulse triggering mechanism and frequency model are constructed, and a time-sharing multiplexing structure and time slot borrowing strategy are designed to realize task timing and priority scheduling.

[0009] Step 3: Model the unmanned system multi-task prior knowledge mechanism model based on the state space, action space, and reward function corresponding to the multi-task, and establish physical models for detection and identification, material delivery, and flight maneuvering tasks respectively.

[0010] Step 4: Design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduce a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model.

[0011] Step 5: According to the physical model, a task decision matrix based on a sliding time window is established for detection and identification, material delivery, and flight maneuvering tasks respectively, which is used to store the dynamic state characteristics and strategy input history of different tasks.

[0012] Step 6: Construct a task association matrix based on the task decision matrix to quantify the mutual influence between multiple tasks, calculate the task priority based on the static weight and dynamic correction factor, and generate the final control instruction using a weighted fusion method.

[0013] Preferably, the method for designing a time-division multiplexing structure and a time slot borrowing strategy in step 2 to implement task timing and priority scheduling includes the following steps:

[0014] In step 221, the time is divided into periodic frames through global clock synchronization, and independent time slots for each task are allocated within each frame. The mathematical model of this design is:

[0015]

[0016] in, is the total length of the time frame, is the basic time slot width of each task, Represents the guard interval, is the number of multivariate tasks.

[0017] Step 222: A flexible time slot borrowing mechanism is introduced based on the fixed time-sharing structure:

[0018]

[0019] in, represents the remaining capacity of each task time slot, Representative tasks Maximum number of pulses in a single time slot. It is a pulse event indicator function, which is 1 if triggered and 0 otherwise.

[0020] Step 223: Dynamically assigning task priorities:

[0021]

[0022] in,, Indicates the priority dynamics of the task, Indicates a task exist Momentary status Next, execute the action The expected reward, Indicates a task exist Momentary status , Indicates a task exist Always perform actions ;and is the temperature coefficient.

[0023] In step 224, the time slot borrowing rule allocates resources according to the priority of the task, and the formula is:

[0024]

[0025] in, It's a task The final available time slot. Representative Moment Time slots belonging to other tasks, allowing tasks Borrow resources across time slots. A task can be borrowed only if its time slot has not used up the specified maximum number of pulses. Representative tasks has a higher priority than .

[0026] Preferably, the method for constructing a task-specific pulse trigger mechanism and frequency model in step 2 comprises the following steps:

[0027] Step 211: Detection and recognition tasks

[0028] For detection and recognition tasks, the pulse frequency needs to be adjusted according to the confidence level and heading deviation of the target. and heading angle deviation The dynamic design of pulse frequency is as follows:

[0029]

[0030] in, Indicates the pulse frequency, Indicates the pulse frequency of the system's regular detection tasks, is the confidence of target detection, is the confidence gain coefficient.

[0031] Step 212, material delivery task

[0032] In the material delivery task, in order to adjust the execution accuracy of the task according to the predicted landing point deviation, the pulse interval function is:

[0033]

[0034] in, represents the pulse interval function, is the reference pulse interval, and is the normalization factor. and Used to predict the lateral and longitudinal deviation between the landing point and the target. , represents the free fall time of the projectile, is the throwing height, is the acceleration due to gravity, Indicates the ambient wind speed.

[0035] Step 213, flight maneuver mission

[0036] The flight maneuvering task is to deal with obstacle avoidance requirements. The triggering of the pulse needs to be determined based on the distance and closing rate of the obstacle. By integrating the closing rate of the obstacle, the pulse triggering to avoid collision is as follows:

[0037]

[0038] in, Represents the obstacle avoidance pulse evaluation function, which avoids the pulse triggering of collision. Indicates the The distance of obstacles detected by a laser radar, is the safe approach rate threshold.

[0039] Preferably, step 3 includes the following steps:

[0040] Step 31: Detecting the prior knowledge mechanism model of the recognition task

[0041] To describe the performance degradation behavior of visual sensors under complex meteorological conditions, the following signal-to-noise ratio attenuation model is established:

[0042]

[0043] in, Indicates the signal-to-noise ratio attenuation factor of the camera. For visible light cameras, the change in SNR is related to the exposure time. The longer the exposure time, the more significant the signal attenuation. represents the nominal signal-to-noise ratio of the sensor under ideal conditions, is the quantum efficiency of light. represents the critical exposure time.

[0044] In addition, to enhance the impact of perception modeling on aircraft dynamic control, a confidence-based flight inertia dynamic adjustment model is designed:

[0045]

[0046] in, represents the inertia matrix, Indicates the mass of the aircraft, corrected mass item It means that as the confidence level decreases, the quality simulation of the aircraft increases. represents the confidence of target detection, Indicates the change in the mass of the aircraft, which is adjusted as the confidence level changes, simulating the dynamic behavior of the aircraft under different perception conditions. Represents the vehicle's moment of inertia about the vertical axis.

[0047] The performance indicator function is as follows:

[0048]

[0049] in, represents the performance indicator function, is the cost function of tracking accuracy, is the cost function for safe maneuvering. and For the start and end time.

[0050] Step 32: Delivering the rescue mission prior knowledge mechanism model

[0051] The prior knowledge mechanism model of the rescue delivery mission consists of the horizontal and vertical motion equations of the materials. The following throwing motion equation under air resistance control is established:

[0052]

[0053] in, For the quality of materials, is the drag coefficient, is the air density, is the windward area of the material, is the acceleration due to gravity.

[0054] In addition, in order to improve the accuracy, a closed-loop PID throwing controller based on the landing point error is designed:

[0055]

[0056] in, 、 、 are the proportional, integral, and derivative gains respectively. and is the correction term for the effect of wind speed on the throw. The controller calculates the target error to adjust the throwing speed.

[0057] Step 33: Flight maneuver mission prior knowledge mechanism model

[0058] In order to improve the active obstacle avoidance capability of UAVs in complex terrain, a dynamic obstacle modeling and safety margin calculation model based on LiDAR is constructed. Obstacle velocity field estimation:

[0059]

[0060] in, Represents the laser point cloud beam distribution function of the polar coordinate system, Represents the number of laser lines of the lidar, which affects the precision of the scan. is the scanning frequency. is the maximum detection distance. is the focus detection distance. is the range attenuation coefficient, which describes the attenuation of the laser signal as the distance increases.

[0061] By modeling the perceived safety margin, the speed of the target obstacle is estimated, providing a basis for the aircraft's obstacle avoidance decisions:

[0062]

[0063] in, represents the speed of the target obstacle, Indicates the current moment The coordinates of the points, are the coordinates of the nearest neighbor point between the previous moment and the current moment, is the time interval.

[0064] The covariance matrix is used to characterize the accuracy of velocity estimation and describe the motion characteristics of dynamic obstacles:

[0065]

[0066] in, represents the covariance matrix of velocity, is the noise intensity at the point cloud location. is the maximum speed of the obstacle. is the measuring range of the sensor.

[0067] Quantify the aircraft's reaction and braking capabilities after sensing an obstacle, and construct the following dynamic safety margin model for comprehensive perception effectiveness:

[0068]

[0069] in, The dynamic safety margin representing the comprehensive perceived effectiveness is The scanning frequency determines the number of times the lidar updates the environmental data per second. Indicates the flight speed of the aircraft. Indicates the focused detection distance of the laser radar, is the maximum lateral acceleration of the UAV.

[0070] To ensure that the aircraft can maneuver flexibly in space, a minimum turning radius constraint model closely related to lidar perception is designed:

[0071]

[0072] in, Indicates the minimum turning radius, Indicates the flight speed of the aircraft. For laser radar in focus detection distance The minimum obstacle size that can be detected is is the horizontal field of view of the laser radar, is the maximum deflection angle of the drone.

[0073] The maximum turning speed of a drone is closely related to its minimum turning radius. The following formula is used to calculate the maximum angular velocity of a drone when performing a sharp turn:

[0074]

[0075] in, Indicates the maximum angular velocity of the drone when performing a sharp turn. Indicates the aircraft's flight speed.

[0076] Preferably, step 4 includes the following steps:

[0077] Step 41: In the detection and recognition task, the target coordinate prediction is provided , while the actual observation value is , in order to evaluate the difference between the prediction and the actual, the variance is defined as the weighted covariance of the prediction residuals:

[0078]

[0079] Among them, the weight matrix ,in and Is the confidence The reciprocal of .

[0080] Step 42: In the material delivery task, the output provided includes the initial velocity of the throw and compensation angle , and the output provided by the prior knowledge mechanism model includes the safety parameters of the weighted trajectory and In order to measure the deviation between the two, the variance is defined as the adaptive fusion of multimodal prediction deviations, and the calculation formula is as follows:

[0081]

[0082] Step 43, in the flight maneuvering task, the acceleration instruction provided is ,in is the linear acceleration, is the angular acceleration. The safety envelope constraint provided by the prior knowledge mechanism model is The variance is defined as the contrast of the control spectrum energy, and the formula is:

[0083]

[0084] in, Represents the Fourier transform. Differences in high-frequency components usually indicate potential instability in the data method. , it indicates that the decision is unstable and the system needs to switch to the control of the prior knowledge mechanism model.

[0085] Preferably: the method of establishing a task decision matrix based on a sliding time window in step 5 includes the following steps:

[0086] Step 51, to describe the historical state evolution of the UAV during the target detection process, a sliding time window mechanism is used to record the state characteristics of consecutive moments to form a two-dimensional state matrix:

[0087]

[0088] The matrix contains the drone from arrive Coordinates, heading angles, and target detection confidence at different times.

[0089] Step 52: To accurately model the delivery control process, a state matrix containing the characteristics of the aircraft power and environmental disturbance is constructed:

[0090]

[0091] The matrix contains the drone from arrive Initial throwing velocity at different times, throwing angle compensation, landing point deviation estimation, ambient wind speed and mass of thrown objects.

[0092] Step 53: To capture the drone’s obstacle avoidance and trajectory adjustment capabilities in complex terrain, a matrix containing maneuvering control and perception states is constructed:

[0093]

[0094] in Represents the horizontal acceleration of the drone. is the yaw angular velocity, The safe obstacle avoidance distance is calculated based on real-time lidar ranging data: . Is a conflict status identifier, and the triggering conditions are as follows: Less than the safety threshold, Set to 1, otherwise 0.

[0095] Preferably, step 6 comprises the following steps:

[0096] Step 61: Based on the constructed decision matrix , , , build the task association matrix ,The task association matrix is given by the following formula:

[0097]

[0098] in, For the task The latest row vector of the decision matrix. Indicates the task The decision space projection of is a matrix The mean of the diagonal covariances, is the weight coefficient, Indicates a task Task The interference intensity,

[0099] In step 62, the calculation formula of the static basic weight of the task priority measures the stability and uncertainty of the task through the variance, specifically:

[0100]

[0101] in, represents the static priority weight, 、 represents the variance of task i and task j.

[0102] Step 63: Introduce dynamic environment correction factor , the calculation formula is as follows:

[0103]

[0104] in, Indicates a task In state Next select action of value, represents the pulse frequency weight, Indicates a task In a given state The set of actions selected for the next task.

[0105] Step 64, by setting the static priority weight and dynamic correction factors Combined, we get the final priority weight of the task :

[0106]

[0107] Comprehensive decision-making action Calculated by weighted average, the formula is as follows:

[0108]

[0109] in, represents a comprehensive decision action, Represents the action performed by task i.

[0110] Preferably, step 1 comprises the following steps:

[0111] Step 11: Detection and recognition tasks

[0112] The state space of the UAV for the detection and identification task is represented as:

[0113]

[0114] in, Represents the state space of the UAV for the detection and recognition task, Indicates the current position of the drone. is the heading angle of the UAV. are the coordinates of the target, is the confidence of target detection.

[0115] The action space of the drone for the detection and recognition task is represented as:

[0116]

[0117] in, The action space of the drone for detection and recognition tasks, The speed adjustment of the drone. is the heading correction amount. To detect the dynamic pulse frequency of the recognition task, is the number of training steps.

[0118] The reward function for the detection and recognition task is:

[0119]

[0120] in, For the detection and recognition task reward function, The cumulative time when no target is detected. is the heading angle deviation. The adjustment coefficient controls the intensity of various rewards and penalties.

[0121] Step 12, material delivery task

[0122] The state space of the drone for the material delivery mission is represented as:

[0123]

[0124] in, The coordinates of the delivery point. Indicates the ambient wind speed. The weight of the material.

[0125] The action space of the drone for the material delivery mission is expressed as:

[0126]

[0127] Among them, the initial velocity Control the speed of material delivery. Wind direction compensation angle The throwing angle is adjusted according to the real-time wind speed. It is the pulse interval function of the material delivery task.

[0128] The reward function for the material delivery task is:

[0129]

[0130] in, Indicates the horizontal and vertical deviation between the material landing point and the target point. Indicates the acceleration change of the delivery action. The weight coefficient is used to balance the effects of landing point error, wind compensation, and acceleration penalty.

[0131] Step 13, Flight Maneuvering Mission

[0132] The state space representation of a UAV flying a maneuvering mission is:

[0133]

[0134] in, is the horizontal speed of the drone The weight. Yaw rate indicates how fast the drone rotates around its vertical axis. Indicates the The distance to the obstacle detected by the LiDAR. Indicates the The azimuth angle of the laser beam.

[0135] The action space of a UAV flying a maneuvering mission is represented as:

[0136]

[0137] in, is the horizontal acceleration of the UAV. is the yaw acceleration of the UAV. It is the pulse trigger threshold for flight maneuvering tasks.

[0138] The flight maneuver task reward function is:

[0139]

[0140] in, Indicates the approach rate of an object. Indicates the change in yaw rate. The regulation factor ensures that the system can balance safety, speed and stability. It is an infinitesimal safety term to avoid the denominator being zero.

[0141] Another object of the present invention is to provide a novel impulse reinforcement learning-driven unmanned system brain-like intelligent multi-decision-making system, which is used to implement a novel impulse reinforcement learning-driven unmanned system brain-like intelligent multi-decision-making method, including an input unit, a state-action reward function construction unit, a task priority scheduling unit, a physical model unit, a dynamic deviation unit, a task decision unit, a task priority dynamic correction unit, and an output unit, wherein:

[0142] The input unit is used to input detection and identification mission information, material delivery mission information, and flight maneuver mission information.

[0143] The state-action reward function construction unit is used to construct corresponding state spaces, action spaces and reward functions according to the multi-tasks of detection and identification, material delivery and flight maneuvering.

[0144] The task priority scheduling unit is used to construct a task-specific pulse triggering mechanism and frequency model according to the state space, action space and reward function corresponding to the multi-task, and to design a time-sharing multiplexing structure and time slot borrowing strategy to realize task timing and priority scheduling.

[0145] The physical model unit is used to model the unmanned system multi-task prior knowledge mechanism model according to the state space, action space and reward function corresponding to the multi-task, and to establish physical models for detection and identification, material delivery and flight maneuvering tasks respectively.

[0146] The dynamic deviation unit is used to design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduces a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model.

[0147] The task decision unit is used to establish a task decision matrix based on a sliding time window for detection and identification, material delivery and flight maneuvering tasks according to the physical model, and is used to store the dynamic state characteristics and strategy input history of different tasks.

[0148] The task priority dynamic correction unit is used to construct a task association matrix based on the task decision matrix, quantify the mutual influence between multiple tasks, calculate the task priority based on static weights and dynamic correction factors, and generate the final control instructions using a weighted fusion method.

[0149] The output unit is used to output the final control instruction.

[0150] Compared with the prior art, the present invention has the following beneficial effects:

[0151] 1. This method primarily constructs state-space, action-space, and reward function models for multiple types of tasks (detection and identification, material delivery, and flight maneuvering). This allows for differentiated modeling and strategy optimization tailored to specific mission requirements, enhancing the system's responsiveness to multi-source tasks. It also introduces a pulse frequency-based scheduling and flexible time-slot borrowing mechanism to rationally coordinate control resource allocation across tasks during execution, ensuring task temporal independence and scheduling efficiency.

[0152] 2. This invention builds a mission-related prior knowledge mechanism model, taking into account dynamic environmental factors such as sensor degradation, wind field disturbances, and approaching obstacles. This modeled expression improves the physical consistency of mission execution. By combining sensory states with confidence indicators, it dynamically adjusts control strategies and flight inertia parameters, enhancing the stability and controllability of unmanned systems in complex environments.

[0153] 3. This invention integrates impulse reinforcement learning strategies with physical model outputs and introduces a multimodal variance indicator to dynamically adjust strategy weights. This allows for adaptive selection of optimal control paths in uncertain environments, ensuring strategy stability and safety. Furthermore, a task state matrix is constructed using a sliding time window, and a task priority fusion mechanism is designed based on task similarity and feedback information. This enables adaptive sequencing and unified action generation across multiple tasks, enhancing the comprehensive collaborative capabilities and decision-making effectiveness of unmanned systems under heterogeneous tasks, and providing theoretical and methodological support for the engineering application of multi-task autonomous control methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0154] Figure 1 Flowchart for time slot scheduling of pulse reinforcement learning tasks.

[0155] Figure 2 Schematic diagram of multi-dimensional integrated decision-making for data-knowledge hybrid unmanned systems.

[0156] Figure 3 This is a comparison chart of the Thunder curve for pulse reinforcement learning multi-task.

[0157] Figure 4 Scatter plot of safety threshold and collision number under different obstacle densities. DETAILED DESCRIPTION

[0158] The present invention is further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0159] This embodiment proposes a novel pulse reinforcement learning driven unmanned system brain-like intelligent multi-decision method, such as Figure 1 、 2As shown, the following steps are included:

[0160] Step 1: Construct the corresponding state space, action space, and reward function based on the multiple tasks of detection and identification, material delivery, and flight maneuvering.

[0161] The multi-tasks include but are not limited to: detection and identification tasks, material delivery tasks, and flight maneuvering tasks. For different task types, the corresponding state space, action space, and reward function structure are established respectively, as follows:

[0162] a. Detection and recognition tasks

[0163] The state space of the drone is represented as:

[0164]

[0165] in, Represents the two-dimensional spatial coordinates of the drone, indicating the current position of the drone. is the heading angle of the UAV. is the coordinate of the target, indicating the real-time position of the target in space. is the confidence of target detection.

[0166] The action space of the drone is represented as:

[0167]

[0168] in, The speed adjustment of the drone. is the heading correction amount. To detect the dynamic pulse frequency of the recognition task, is the number of training steps.

[0169] Design the detection and recognition task reward function, expressed as:

[0170]

[0171] in, The cumulative time when no target is detected. is the heading angle deviation. The adjustment coefficient controls the intensity of various rewards and penalties.

[0172] b. Material delivery mission

[0173] The state space of the drone is represented as:

[0174]

[0175] in, The coordinates of the delivery point. Indicates the ambient wind speed. The weight of the material.

[0176] The action space of the drone is represented as:

[0177]

[0178] Among them, the initial velocity Controls the speed of material delivery, affecting the distance and trajectory of material flight. Wind direction compensation angle The throwing angle is adjusted according to the real-time wind speed so that the materials can better offset the impact of the wind and be delivered accurately to the target location. It is the pulse interval function of the material delivery task.

[0179] Design the reward function for the material delivery task, expressed as:

[0180]

[0181] in, Indicates the horizontal and vertical deviation between the material landing point and the target point. Indicates the acceleration change of the delivery action, which is used to suppress excessive acceleration and avoid unnecessary motion jitter. The weight coefficient is used to balance the effects of landing point error, wind compensation, and acceleration penalty.

[0182] c. Flight maneuvering missions

[0183] The state space of the drone is represented as:

[0184]

[0185] in, is the horizontal speed of the drone The weight. Yaw angular velocity indicates the speed at which the drone rotates around the vertical axis. Using a multi-line lidar, we start from the front of the drone and circle clockwise. We take a lidar ranging sampling point every 9 degrees, so there are 40 sampling points in total. Indicates the The distance to the obstacle detected by the LiDAR. Indicates the The azimuth angle of a laser beam is used to indicate the emission angle of the laser beam and help obtain environmental information at different angles.

[0186] The action space of the drone is represented as:

[0187]

[0188] in, is the horizontal acceleration of the UAV. is the yaw acceleration of the UAV. It is the pulse trigger threshold for flight maneuvering tasks.

[0189] Design the flight maneuver task reward function, expressed as:

[0190]

[0191] in, Indicates the object's approach speed. The greater the approach speed, the greater the penalty, forcing the drone to adjust its flight path in time. Indicates the change in yaw rate. The regulation factor ensures that the system can balance safety, speed and stability. It is an infinitesimal safety term to avoid the denominator being zero.

[0192] Step 2: Based on the state space, action space, and reward function corresponding to the multi-task, construct a task-specific pulse trigger mechanism and frequency model, and design a time-sharing multiplexing structure and time slot borrowing strategy to achieve task timing and priority scheduling.

[0193] This embodiment designs a pulse event triggering mechanism based on task characteristics and constructs a pulse frequency dynamic adjustment model for multiple types of tasks. Three representative tasks are selected for pulse frequency modeling, as follows:

[0194] a. Detection and recognition tasks

[0195] For detection and recognition tasks, the pulse frequency needs to be adjusted according to the confidence level and heading deviation of the target. Here, the pulse frequency is related to the confidence level of target detection. and heading angle deviation The dynamic design of pulse frequency is as follows:

[0196]

[0197] in, It is the basic detection frequency, which indicates the pulse frequency of the system's routine detection tasks. is the confidence gain factor, which is used to increase the pulse frequency when the target confidence is high.

[0198] b. Material delivery mission

[0199] In the material delivery task, delivery accuracy is closely related to delivery deviation. In order to adjust the execution accuracy of the task according to the predicted landing point deviation, the following pulse interval function is designed:

[0200]

[0201] in, It is the reference pulse interval, used to set the pulse frequency of routine tasks; and is a normalization factor used to adjust the deviation correction in different directions. and Used to predict the lateral and longitudinal deviation between the landing point and the target. ., represents the free fall time of the projectile, is the throwing height, is the acceleration due to gravity.

[0202] c. Flight maneuvering mission

[0203] The flight maneuvering mission mainly deals with obstacle avoidance needs. The triggering of the pulse needs to be determined based on the distance and closing rate of the obstacle. By integrating the closing rate of the obstacle, the pulse triggering to avoid collision is as follows:

[0204]

[0205] in, is the safe approach rate threshold, which indicates the safe approach speed between the obstacle and the drone.

[0206] The multi-task time-sharing multiplexing frame structure design uses global clock synchronization to divide time into periodic frames, and allocates independent time slots to each task within each frame, thereby avoiding overlap of pulse events between tasks and ensuring efficient use of hardware resources. The mathematical model of this design is:

[0207]

[0208] in, is the total length of the time frame. Is the basic time slot width of each task. Guard interval Used to prevent crosstalk of pulse signals and ensure stable operation of the system. is the number of multivariate tasks.

[0209] In a fixed time-sharing multiplexing frame structure, although each task is allocated a fixed time slot resource, due to the dynamic fluctuation of task requirements, some time slot resources may be idle. To improve resource utilization, this invention introduces a flexible time slot borrowing mechanism based on the fixed time-sharing structure:

[0210]

[0211] in, Representative tasks Maximum number of pulses in a single time slot. It is a pulse event indicator function, which is 1 when triggered and 0 otherwise. The significance of calculating the remaining capacity is that it provides a basis for "borrowing" for high-priority tasks, ensuring that they can obtain additional time slices when resources are idle, thereby improving the real-time response of tasks.

[0212] For dynamic allocation of task priorities:

[0213]

[0214] in, Indicates a task exist Momentary status Next, execute the action The expected reward of is the temperature coefficient, which is used to adjust the distribution of priorities.

[0215] The time slot borrowing rule allocates resources according to the priority of the task, and the formula is:

[0216]

[0217] in, It's a task The final available time slot. Representative Moment Time slots belonging to other tasks, allowing tasks Borrow resources across time slots. A task can be borrowed only if its time slot has not used up the specified maximum number of pulses. Representative tasks has a higher priority than .

[0218] Step 3: Model the unmanned system multi-task prior knowledge mechanism model based on the state space, action space, and reward function corresponding to the multi-task, and establish physical models for detection and identification, material delivery, and flight maneuvering tasks respectively.

[0219] This embodiment constructs a priori knowledge mechanism model based on task characteristics. According to the core task type of the unmanned system, the following three types of physical models are constructed:

[0220] a. Prior knowledge mechanism model for detection and recognition tasks

[0221] To describe the performance degradation behavior of visual sensors under complex meteorological conditions, the following signal-to-noise ratio attenuation model is established:

[0222]

[0223] For visible light cameras, the change in SNR with exposure time The longer the exposure time, the more significant the signal attenuation. represents the nominal signal-to-noise ratio of the sensor under ideal conditions, is the quantum efficiency of light. represents the critical exposure time.

[0224] In addition, to enhance the impact of perception modeling on aircraft dynamic control, a confidence-based flight inertia dynamic adjustment model is designed:

[0225]

[0226] in, Indicates the mass of the aircraft, reflecting the inertial strength of the system's response to acceleration. It means that as the confidence level decreases, the mass simulation of the aircraft increases, that is, the aircraft shows greater inertial resistance and requires greater control input to maintain stability. In terms of rotational inertia adjustment, in This item ensures that the aircraft can make attitude adjustments more sensitively with high confidence.

[0227] In order to generate the optimal trajectory, the performance index of the aircraft combines confidence and speed. The performance index function is defined in the following form:

[0228]

[0229] in, is the cost function of tracking accuracy, is the cost function for safe maneuvering. and For the start and end time.

[0230] b. Prior knowledge mechanism model for rescue mission delivery

[0231] To predict the landing trajectory of aerial objects under wind disturbance, the model consists of the horizontal and vertical motion equations of the material. The following throwing motion equation under air resistance control is established:

[0232]

[0233] in, For the quality of materials, is the drag coefficient, is the air density, is the windward area of the material, is the acceleration due to gravity. These two equations can describe the dynamic trajectory of materials during flight under the combined effects of air resistance and gravity.

[0234] In addition, in order to improve the accuracy, a closed-loop PID throwing controller based on the landing point error is designed:

[0235]

[0236] in, 、 、 are the proportional, integral, and derivative gains respectively. and is the correction term for the effect of wind speed on the throw. The controller calculates the target error to adjust the throwing speed.

[0237] c. Flight maneuver mission prior knowledge mechanism model

[0238] In order to improve the active obstacle avoidance capability of UAVs in complex terrain, a dynamic obstacle modeling and safety margin calculation model based on LiDAR is constructed. Obstacle velocity field estimation:

[0239]

[0240] in, Represents the number of laser lines of the lidar, which affects the precision of the scan. is the scanning frequency. is the maximum detection distance. is the focus detection distance. is the range attenuation coefficient, which describes the attenuation of the laser signal as the distance increases.

[0241] By modeling the perceived safety margin, the speed of the target obstacle can be estimated, thus providing a basis for the aircraft's obstacle avoidance decision-making:

[0242]

[0243] in, Indicates the current moment The coordinates of the points, are the coordinates of the nearest neighbor point between the previous moment and the current moment, is the time interval.

[0244] The covariance matrix is used to characterize the accuracy of velocity estimation and describe the motion characteristics of dynamic obstacles:

[0245]

[0246] in, is the noise intensity at the point cloud location. is the maximum speed of the obstacle. is the measuring range of the sensor.

[0247] Quantify the aircraft's reaction and braking capabilities after sensing an obstacle, and construct the following dynamic safety margin model for comprehensive perception effectiveness:

[0248]

[0249] in, The scanning frequency determines the number of times the lidar updates the environmental data per second. The higher the frequency, the more timely the perception data is. The smaller it is, the higher the speed at which the aircraft can safely fly. is the maximum lateral acceleration of the UAV.

[0250] To ensure that the aircraft can maneuver flexibly in space, a minimum turning radius constraint model closely related to lidar perception is designed:

[0251]

[0252] in, For laser radar in focus detection distance The minimum obstacle size that can be detected is is the horizontal field of view of the laser radar, is the maximum deflection angle of the drone.

[0253] The maximum turning speed of a drone is closely related to its minimum turning radius. The following formula can be used to calculate the maximum angular velocity of a drone when performing a sharp turn:

[0254]

[0255] This ensures that the drone always maintains real-time awareness of obstacles ahead when performing maneuvers.

[0256] Step 4: Design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduce a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model.

[0257] This embodiment proposes a decision-making mechanism that integrates a data-driven approach with a priori knowledge mechanism model to dynamically adjust the contribution weight between the impulse reinforcement learning strategy and the physical model output.

[0258] To quantify the difference between the impulse reinforcement learning strategy and the physical model output, the variance indicator is introduced , the calculation method under different tasks is as follows:

[0259] In the detection and recognition task, the data method provides target coordinate prediction , while the actual observation value is , in order to evaluate the difference between the prediction and the actual, the variance is defined as the weighted covariance of the prediction residuals:

[0260]

[0261] Among them, the weight matrix ,in and Is the confidence The inverse of , used to strengthen the error penalty in the low confidence region.

[0262] In the material delivery mission, the output provided by the data method includes the initial velocity of the throw and compensation angle , and the output provided by the prior knowledge mechanism model includes the safety parameters of the weighted trajectory and In order to measure the deviation between the two, the variance is defined as the adaptive fusion of multimodal prediction deviations, and the calculation formula is as follows:

[0263]

[0264] In the flight maneuvering mission, the acceleration instruction provided by the data method is ,in is the linear acceleration, is the angular acceleration. The safety envelope constraint provided by the prior knowledge mechanism model is The variance is defined as the contrast of the control spectrum energy, and the formula is:

[0265]

[0266] in, Represents the Fourier transform. Differences in high-frequency components usually indicate potential instability in the data method. , it indicates that the decision is unstable and the system needs to switch to the control of the prior knowledge mechanism model.

[0267] Step 5: Based on the physical model, a task decision matrix based on a sliding time window is established for detection and identification, material delivery, and flight maneuvering tasks respectively, which is used to store the dynamic state characteristics and strategy input history of different tasks.

[0268] This example proposes a method for constructing a task decision matrix based on a sliding time window. This method constructs a structured state expression matrix for different tasks to carry the state history of data-driven strategy output and prior physical models, achieving a joint expression of temporal continuity and spatial information. Specifically, for three typical tasks: detection and identification, material delivery, and flight maneuvering, task-specific decision matrices are constructed:

[0269] In order to describe the historical state evolution of the UAV during the target detection process, a sliding time window mechanism is used to record the state characteristics of continuous moments to form a two-dimensional state matrix:

[0270]

[0271] The matrix contains the drone from arrive Coordinates, heading angles, and target detection confidence at different times.

[0272] To accurately model the launch control process, a state matrix containing the characteristics of the aircraft dynamics and environmental disturbances is constructed:

[0273]

[0274] The matrix contains the drone from arrive Initial throwing velocity at different times, throwing angle compensation, landing point deviation estimation, ambient wind speed and mass of thrown objects.

[0275] In order to capture the obstacle avoidance and trajectory adjustment capabilities of the UAV in complex terrain, a matrix containing maneuver control and perception states is constructed:

[0276] in Represents the horizontal acceleration of the drone. is the yaw angular velocity, The safe obstacle avoidance distance is calculated based on real-time lidar ranging data: . Is a conflict status identifier, and the triggering conditions are as follows: Less than the safety threshold, Set to 1, otherwise 0.

[0277] Step 6: Construct a task association matrix based on the task decision matrix to quantify the mutual influence between multiple tasks, calculate the task priority based on the static weight and dynamic correction factor, and use weighted fusion to generate the final control instructions.

[0278] This embodiment proposes a dynamic fusion strategy for comprehensive decision-making on multiple tasks. Based on the previously constructed task state matrix, it dynamically calculates the priority weights of each task, combining factors such as task similarity, execution stability, and strategy benefits. The final action instructions are generated based on a weighted fusion mechanism.

[0279] Based on the constructed decision matrix , , ,In order to quantify the relationship and influence between tasks, a task correlation matrix is constructed , used to clearly quantify the relationship between tasks. The matrix is given by the following formula:

[0280]

[0281] in, For the task The latest row vector of the decision matrix. Indicates the task The decision space projection of , namely the least squares mapping, is used to measure the degree of dependency between tasks. is a matrix The mean of the diagonal covariance, reflecting the task Its own stability and reliability. is the weight coefficient used to Regulation of self-stability. Indicates a task Task The interference intensity,

[0282] The calculation formula of the static basic weight of task priority measures the stability and uncertainty of the task through variance, specifically:

[0283]

[0284] in, It indicates the relative weight of the uncertainty of a task among all tasks. The larger the variance, the smaller the relative weight.

[0285] In multi-task decision making, static weights are calculated based on the stability and uncertainty of tasks, but the dynamic changes of the environment and tasks require real-time adjustment of the weights. To this end, a dynamic environment correction factor is introduced. , the calculation formula is as follows:

[0286]

[0287] in, :Task In state Next select action of value, The larger the value, the greater the contribution of the action to the successful execution of the task. : Pulse frequency weight, indicating the task Feedback strength in different states. The higher the pulse frequency, the stronger the feedback of the task and the greater the impact on the priority of the task. Indicates a task In a given state Next, the set of actions that a task can choose.

[0288] By adding static priority weights and dynamic correction factors Combined, we can get the final priority weight of the task :

[0289]

[0290] Comprehensive decision-making action Calculated by weighted average, the formula is as follows:

[0291]

[0292] This formula combines the priority weight and safety factor of the task to ensure that tasks with high priority and strong safety have a greater impact on the final decision.

[0293] In another embodiment of the present invention, a novel impulse reinforcement learning-driven unmanned system brain-inspired intelligent multi-decision-making system is provided, which is used to implement a novel impulse reinforcement learning-driven unmanned system brain-inspired intelligent multi-decision-making method, including an input unit, a state-action reward function construction unit, a task priority scheduling unit, a physical model unit, a dynamic deviation unit, a task decision unit, a task priority dynamic correction unit, and an output unit, wherein:

[0294] The input unit is used to input detection and identification mission information, material delivery mission information, and flight maneuver mission information.

[0295] The state-action reward function construction unit is used to construct corresponding state spaces, action spaces and reward functions according to the multi-tasks of detection and identification, material delivery and flight maneuvering.

[0296] The task priority scheduling unit is used to construct a task-specific pulse triggering mechanism and frequency model according to the state space, action space and reward function corresponding to the multi-task, and to design a time-sharing multiplexing structure and time slot borrowing strategy to realize task timing and priority scheduling.

[0297] The physical model unit is used to model the unmanned system multi-task prior knowledge mechanism model according to the state space, action space and reward function corresponding to the multi-task, and to establish physical models for detection and identification, material delivery and flight maneuvering tasks respectively.

[0298] The dynamic deviation unit is used to design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduces a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model.

[0299] The task decision unit is used to establish a task decision matrix based on a sliding time window for detection and identification, material delivery and flight maneuvering tasks according to the physical model, and is used to store the dynamic state characteristics and strategy input history of different tasks.

[0300] The task priority dynamic correction unit is used to construct a task association matrix based on the task decision matrix, quantify the mutual influence between multiple tasks, calculate the task priority based on static weights and dynamic correction factors, and generate the final control instructions using a weighted fusion method.

[0301] The output unit is used to output the final control instruction.

[0302] The comparison of multi-task Thunder curves of pulse reinforcement learning in this embodiment is as follows: Figure 3 As shown in the figure, the safety threshold and collision number scatter plots under different obstacle densities are as follows: Figure 4 As shown in the figure, the present invention can perform differentiated modeling and strategy optimization according to different mission requirements, and improve the system's responsiveness to multi-source tasks. The introduction of pulse frequency-based timing scheduling and flexible time slot borrowing mechanism can reasonably coordinate the allocation of control resources for each task during execution, and ensure the timing independence and scheduling efficiency of the tasks. Through the task-related prior knowledge mechanism model, dynamic environmental factors such as perception degradation, wind field disturbance, and approaching obstacles are taken into account, and the physical consistency of task execution is improved through modeling expression. Combining the perception state and confidence index, the control strategy and flight inertia parameters are dynamically adjusted to enhance the stability and controllability of the unmanned system in complex environments.

[0303] The present invention is suitable for autonomous decision-making scenarios of unmanned systems under multi-task concurrent conditions, and has good task adaptability and real-time scheduling capabilities.

[0304] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A novel brain-like intelligent multi-decision-making method for unmanned systems driven by pulse reinforcement learning, characterized by: The following steps are involved: Step 1: Construct the corresponding state space, action space, and reward function based on the multi-tasks of detection and identification, material delivery, and flight maneuvering; Step 2: Based on the state space, action space, and reward function corresponding to the multi-task, a task-specific pulse triggering mechanism and frequency model are constructed, and a time-sharing multiplexing structure and time slot borrowing strategy are designed to achieve task timing and priority scheduling; Step 3: Model the prior knowledge mechanism model of the multi-task of the unmanned system based on the state space, action space and reward function corresponding to the multi-task, and establish physical models for detection and identification, material delivery and flight maneuvering tasks respectively; Step 4: Design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduce a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model; Step 5: Based on the physical model, a task decision matrix based on a sliding time window is established for the detection and identification, material delivery, and flight maneuvering tasks respectively, which is used to store the dynamic state characteristics and strategy input history of different tasks; Step 6: Construct a task association matrix based on the task decision matrix to quantify the mutual influence between multiple tasks, calculate the task priority based on the static weight and dynamic correction factor, and generate the final control instruction using a weighted fusion method.

2. The novel pulse reinforcement learning-driven brain-inspired intelligent multi-level decision-making method for unmanned systems according to claim 1 is characterized by: The method for designing a time-division multiplexing structure and a time slot borrowing strategy in step 2 to implement task timing and priority scheduling includes the following steps: In step 221, the time is divided into periodic frames through global clock synchronization, and independent time slots for each task are allocated within each frame. The mathematical model of this design is: in, is the total length of the time frame, is the basic time slot width of each task, Represents the guard interval, is the number of multivariate tasks; Step 222: A flexible time slot borrowing mechanism is introduced based on the fixed time-sharing structure: in, represents the remaining capacity of each task time slot, Representative tasks Maximum number of pulses in a single time slot; It is a pulse event indicator function, which is 1 when triggered and 0 otherwise; Step 223: Dynamically assigning task priorities: in, Indicates the priority dynamics of the task, Indicates a task exist Momentary status Next, execute the action The expected reward, Indicates a task exist Momentary status , Indicates a task exist Always perform actions ;and is the temperature coefficient; In step 224, the time slot borrowing rule allocates resources according to the priority of the task, and the formula is: in, It's a task The final available time slot; Representative Moment Time slots belonging to other tasks, allowing tasks Borrowing resources across time slots; A time slot of a task can be borrowed only if the time slot of the task being borrowed has not used up the maximum number of pulses specified; Representative tasks has a higher priority than .

3. The novel brain-inspired intelligent multi-level decision-making method for unmanned systems driven by pulsed reinforcement learning according to claim 2 is characterized by: The method for constructing a task-specific pulse triggering mechanism and frequency model in step 2 includes the following steps: Step 211: Detection and recognition tasks For detection and recognition tasks, the pulse frequency needs to be adjusted according to the confidence level and heading deviation of the target; the pulse frequency is related to the confidence level of target detection. and heading angle deviation It forms a nonlinear relationship; the dynamic design of the pulse frequency is as follows: in, Indicates the pulse frequency, Indicates the pulse frequency of the system's routine detection tasks, is the confidence of target detection, is the confidence gain coefficient; Step 212, material delivery task In the material delivery task, in order to adjust the execution accuracy of the task according to the predicted landing point deviation, the pulse interval function is: in, represents the pulse interval function, is the reference pulse interval, and is the normalization factor; and Used to predict the lateral and longitudinal deviation between the landing point and the target; , represents the free fall time of the projectile, is the throwing height, is the acceleration due to gravity, Indicates the wind speed of the environment; Step 213, flight maneuver mission The flight maneuvering task is to deal with obstacle avoidance requirements. The triggering of the pulse needs to be determined based on the distance and closing rate of the obstacle. By integrating the closing rate of the obstacle, the pulse triggering to avoid collision is as follows: in, Represents the obstacle avoidance pulse evaluation function, which avoids the pulse triggering of collision. Indicates the The distance of obstacles detected by the laser radar, is the safe approach rate threshold.

4. The novel pulse reinforcement learning-driven brain-inspired intelligent multi-dimensional decision-making method for unmanned systems according to claim 3 is characterized by: Step 3 includes the following steps: Step 31: Detecting the prior knowledge mechanism model of the recognition task To describe the performance degradation behavior of visual sensors under complex meteorological conditions, the following signal-to-noise ratio attenuation model is established: in, Indicates the camera's signal-to-noise ratio attenuation factor. For visible light cameras, the change in SNR is related to the exposure time. Inversely proportional, the longer the exposure time, the more significant the signal attenuation; represents the nominal signal-to-noise ratio of the sensor under ideal conditions, is the quantum efficiency; represents the critical exposure time; In addition, to enhance the impact of perception modeling on aircraft dynamic control, a confidence-based flight inertia dynamic adjustment model is designed: in, represents the inertia matrix, Indicates the mass of the aircraft, corrected mass item It means that as the confidence level decreases, the quality simulation of the aircraft increases. represents the confidence of target detection, Indicates the change in the mass of the aircraft, which is adjusted as the confidence level changes, simulating the dynamic behavior of the aircraft under different perception conditions. Represents the moment of inertia of the aircraft about the vertical axis; The performance indicator function is as follows: in, represents the performance indicator function, is the cost function of tracking accuracy, is the cost function of safe maneuver; and for the start and end time; Step 32: Delivering the rescue mission prior knowledge mechanism model The prior knowledge mechanism model of the rescue delivery mission consists of the horizontal and vertical motion equations of the materials. The following throwing motion equation under air resistance control is established: in, For the quality of materials, is the drag coefficient, is the air density, is the windward area of the material, is the acceleration due to gravity; In addition, in order to improve the accuracy, a closed-loop PID throwing controller based on the landing point error is designed: in, 、 、 are proportional, integral, and derivative gains respectively; and is the correction term for the effect of wind speed on throwing; the controller calculates the target error To adjust the throwing speed; Step 33: Flight maneuver mission prior knowledge mechanism model To improve the active obstacle avoidance capability of UAVs in complex terrain, a dynamic obstacle modeling and safety margin calculation model based on LiDAR is constructed; obstacle velocity field estimation: in, Represents the laser point cloud beam distribution function of the polar coordinate system, Represents the number of laser lines of the lidar, which affects the precision of the scan; is the scanning frequency; is the maximum detection distance; is the focus detection distance; is the range attenuation coefficient, which describes the attenuation of the laser signal as the distance increases; By modeling the perceived safety margin, the speed of the target obstacle is estimated, providing a basis for the aircraft's obstacle avoidance decisions: in, represents the speed of the target obstacle, Indicates the current moment The coordinates of the points, are the coordinates of the nearest neighbor point between the previous moment and the current moment, is the time interval; The covariance matrix is used to characterize the accuracy of velocity estimation and describe the motion characteristics of dynamic obstacles: in, represents the covariance matrix of velocity, is the noise intensity at the point cloud location; is the maximum speed of the obstacle; is the range of the sensor; Quantify the aircraft's reaction and braking capabilities after sensing an obstacle, and construct the following dynamic safety margin model for comprehensive perception effectiveness: in, The dynamic safety margin representing the comprehensive perceived effectiveness is The scanning frequency determines the number of times the lidar updates the environmental data per second. Indicates the flight speed of the aircraft. Indicates the focused detection distance of the laser radar, is the maximum lateral acceleration of the UAV; To ensure that the aircraft can maneuver flexibly in space, a minimum turning radius constraint model closely related to lidar perception is designed: in, Indicates the minimum turning radius, Indicates the flight speed of the aircraft. For laser radar in focus detection distance The minimum obstacle size that can be detected is is the horizontal field of view of the laser radar, is the maximum deflection angle of the UAV; The upper limit of a drone's turning speed is closely related to its minimum turning radius. The following formula is used to calculate the maximum angular velocity of a drone when performing a sharp turn: in, Indicates the maximum angular velocity of the drone when performing a sharp turn. Indicates the aircraft's flight speed.

5. The novel pulse reinforcement learning-driven unmanned system brain-inspired intelligent multi-level decision-making method according to claim 4 is characterized by: Step 4 includes the following steps: Step 41: In the detection and recognition task, the target coordinates are predicted , while the actual observation value is , in order to evaluate the difference between the prediction and the actual, the variance is defined as the weighted covariance of the prediction residuals: Among them, the weight matrix ,in and Is the confidence The reciprocal of Step 42: In the material delivery task, the output provided includes the initial velocity of the throw and compensation angle , and the output provided by the prior knowledge mechanism model includes the safety parameters of the weighted trajectory and ; In order to measure the deviation between the two, the variance is defined as the adaptive fusion of multimodal prediction deviations, and the calculation formula is as follows: Step 43, in the flight maneuvering task, the acceleration instruction provided is ,in is the linear acceleration, is the angular acceleration; and the safety envelope constraint provided by the prior knowledge mechanism model is ; Define variance as the contrast of control spectrum energy, the formula is: in, Represents Fourier transform; high-frequency component differences usually indicate potential instability of the data method; if , it indicates that the decision is unstable and the system needs to switch to the control of the prior knowledge mechanism model.

6. The novel pulse reinforcement learning-driven brain-inspired intelligent multi-level decision-making method for unmanned systems according to claim 5 is characterized by: Step 5: The method for establishing a task decision matrix based on a sliding time window includes the following steps: Step 51, to describe the historical state evolution of the UAV during the target detection process, a sliding time window mechanism is used to record the state characteristics of consecutive moments to form a two-dimensional state matrix: The matrix contains the drone from arrive Coordinates, heading angles, and target detection confidence at different times; Step 52: To accurately model the delivery control process, a state matrix containing the characteristics of the aircraft power and environmental disturbance is constructed: The matrix contains the drone from arrive Initial throwing velocity at different times, throwing angle compensation, landing point deviation estimation, ambient wind speed and thrown object mass; Step 53: To capture the drone’s obstacle avoidance and trajectory adjustment capabilities in complex terrain, a matrix containing maneuvering control and perception states is constructed: in Represents the horizontal acceleration of the drone; is the yaw angular velocity, The safe obstacle avoidance distance is calculated based on real-time lidar ranging data: ; Is a conflict status identifier, and the triggering conditions are as follows: Less than the safety threshold, Set to 1, otherwise 0.

7. The novel pulse reinforcement learning-driven brain-inspired intelligent multi-level decision-making method for unmanned systems according to claim 6 is characterized by: Step 6 includes the following steps: Step 61: Based on the constructed decision matrix , , , build the task association matrix ,The task association matrix is given by the following formula: in, For the task The latest row vector of the decision matrix; Indicates the task The decision space projection of is a matrix The mean of the diagonal covariances, is the weight coefficient, Indicates a task Task The interference intensity, In step 62, the calculation formula of the static basic weight of the task priority measures the stability and uncertainty of the task through the variance, specifically: in, represents the static priority weight, 、 represents the variance of task i and task j; Step 63: Introduce dynamic environment correction factor , the calculation formula is as follows: in, Indicates a task In state Next select action of value, represents the pulse frequency weight, Indicates a task In a given state The set of actions for selecting the next task; Step 64, by setting the static priority weight and dynamic correction factors Combined, we get the final priority weight of the task : Comprehensive decision-making action Calculated by weighted average, the formula is as follows: in, represents a comprehensive decision action, Represents the action performed by task i.

8. The novel brain-inspired intelligent multi-level decision-making method for unmanned systems driven by pulse reinforcement learning according to claim 7 is characterized by: Step 1 includes the following steps: Step 11: Detection and recognition tasks The state space of the UAV for the detection and identification task is represented as: in, Represents the state space of the UAV for the detection and recognition task, Indicates the current position of the drone; is the heading angle of the UAV; are the coordinates of the target, is the confidence of target detection; The action space of the drone for the detection and recognition task is represented as: in, The action space of the drone for detection and recognition tasks, The speed adjustment of the drone; is the heading correction; To detect the dynamic pulse frequency of the recognition task, is the number of training steps; The reward function for the detection and recognition task is: in, For the detection and recognition task reward function, The cumulative time when no target is detected; is the heading angle deviation; The adjustment coefficient controls the intensity of various rewards and penalties; Step 12, material delivery task The state space of the drone for the material delivery mission is represented as: in, is the delivery point coordinate; Indicates the wind speed of the environment; is the weight of the materials; The action space of the drone for the material delivery mission is expressed as: Among them, the initial velocity Control the speed of material delivery; wind direction compensation angle The throwing angle is adjusted according to the real-time wind speed; It is the pulse interval function of the material delivery task; The reward function for the material delivery task is: in, Indicates the horizontal and vertical deviations between the landing point of the material and the target point; Indicates the acceleration change of the delivery action; The weight coefficient is used to balance the impact of landing point error, wind compensation and acceleration penalty; Step 13, Flight Maneuvering Mission The state space representation of a UAV flying a maneuvering mission is: in, is the horizontal speed of the drone The weight; The yaw rate indicates how fast the drone rotates around its vertical axis; Indicates the The distance of obstacles detected by a LiDAR; Indicates the The azimuth angle of the laser beam; The action space of a UAV flying a maneuvering mission is represented as: in, is the horizontal acceleration of the UAV; is the yaw acceleration of the UAV; It is the pulse trigger threshold of the flight maneuver mission; The flight maneuver task reward function is: in, Indicates the approach rate of an object; Indicates the change in yaw angular velocity; The regulation coefficient ensures that the system can balance safety, speed and stability; It is an infinitesimal safety term to avoid the denominator being zero.

9. A decision-making system for implementing the novel impulse reinforcement learning-driven unmanned system brain-inspired intelligent multi-decision-making method of claim 1, characterized by: It includes input unit, state-action reward function construction unit, task priority scheduling unit, physical model unit, dynamic deviation unit, task decision unit, task priority dynamic correction unit, and output unit, among which: The input unit is used to input detection and identification mission information, material delivery mission information, and flight maneuver mission information; The state-action-reward function construction unit is used to construct corresponding state spaces, action spaces and reward functions according to the multi-tasks of detection and identification, material delivery and flight maneuvering; The task priority scheduling unit is used to construct a task-specific pulse triggering mechanism and frequency model based on the state space, action space and reward function corresponding to the multi-task, and design a time-sharing multiplexing structure and time slot borrowing strategy to achieve task timing and priority scheduling; The physical model unit is used to model the unmanned system multi-task prior knowledge mechanism model according to the state space, action space and reward function corresponding to the multi-task, and establish physical models for detection and identification, material delivery and flight maneuvering tasks respectively; The dynamic deviation unit is used to design a decision-making mechanism that integrates data-driven strategies and prior knowledge models, and introduces a differential variance indicator based on task characteristics to dynamically measure the degree of deviation between the strategy output and the physical model; The task decision unit is used to establish a task decision matrix based on a sliding time window for detection and identification, material delivery, and flight maneuvering tasks according to the physical model, and is used to store the dynamic state characteristics and strategy input history of different tasks; The task priority dynamic correction unit is used to construct a task correlation matrix based on the task decision matrix, quantify the mutual influence between multiple tasks, calculate the task priority based on the static weight and the dynamic correction factor, and generate the final control instruction using a weighted fusion method; The output unit is used to output the final control instruction.

Citation Information

Cited By

  • Resistance servo adjusting method and system based on machine learning

    CN121300070A

  • Unmanned aerial vehicle group space structure regularity evaluation method based on brain-like calculation

    CN121765652A

  • Livestock and poultry behavior identification method and system based on brain-like network

    CN121881010A