Decision method of unmanned vehicle based on deep reinforcement learning

CN118468978BActive Publication Date: 2026-09-22BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410479204.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2026-09-22
Estimated Expiration
2044-04-22

AI Technical Summary

Technical Problem

[0004]而且,在不断演变的战场环境中,过去的经验可能不再适用,固定的模型也可能失效

Benefits of technology

[0130](1)本发明保证资源的高效率利用以及避免无人车频繁临机决策去对上层指挥系统进行频繁的任务建议使得上层任务执行系统受到干扰以及可能通信的拥堵情况,设置临机决策模块开启条件;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118468978B_ABST
    Figure CN118468978B_ABST
Patent Text Reader

Abstract

The application relates to a decision-making method of an unmanned vehicle based on deep reinforcement learning, in particular to an intelligent decision-making method of an unmanned vehicle based on deep reinforcement learning, and belongs to the technical field of intelligent decision-making of unmanned vehicles. The application provides a decision-making method of an unmanned vehicle based on deep reinforcement learning, establishes an unmanned vehicle carrying military equipment emergency decision-making opening strategy under various emergency situations, introduces various target fusion information such as target threat degree and target confidence degree, and appropriately sets an emergency decision-making opening condition; a multi-target grouping mechanism is established, a cluster envelope line is obtained by using a convex hull algorithm, and behavior decision-making information is adopted for a target cluster; and deep reinforcement learning technology is used for efficient and intelligent decision-making calling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology for autonomous vehicles, specifically to a decision-making method for autonomous vehicles based on deep reinforcement learning, and more particularly to an intelligent decision-making method for autonomous vehicles based on deep reinforcement learning. Background Technology

[0002] Unmanned systems play a crucial role in today's complex battlefield environments. The complexity of the battlefield situation presents significant challenges to these devices. Traditional group decision-making and combined decision-making methods often prove inadequate in such complex environments.

[0003] Autonomous vehicle systems need to make rapid decisions to respond to unexpected situations and changing battlefield scenarios. However, existing decision-making methods can lead to excessively long decision-making cycles. Current decision-making approaches often rely on experience or fixed formulas.

[0004] Moreover, in the ever-evolving battlefield environment, past experience may no longer be applicable, and fixed models may fail. Decisions made based on these methods may not achieve optimal operational effectiveness.

[0005] Unmanned systems face complex battlefield situations, and conventional group decision-making can lead to problems such as long decision-making cycles and the inability to achieve optimal combat effectiveness by relying on experience or fixed models. Summary of the Invention

[0006] In view of the above problems, this invention provides a decision-making method for unmanned vehicles based on deep reinforcement learning. It establishes a contingency decision-making activation strategy for unmanned vehicles carrying military equipment under various emergency situations, introduces multiple target fusion information such as target threat level and target confidence level, and appropriately sets contingency decision activation conditions; it establishes a multi-target clustering mechanism, and uses the convex hull algorithm to obtain the cluster envelope, and takes behavioral decision-making information for target clusters; it uses deep reinforcement learning technology to make efficient and intelligent decision-making calls.

[0007] This invention provides a decision-making method for autonomous vehicles based on deep reinforcement learning, specifically relating to an intelligent decision-making method for autonomous vehicles based on deep reinforcement learning, comprising:

[0008] S1. Obtain unmanned vehicle information and corresponding target information; obtain unmanned vehicle payload parameters based on unmanned vehicle information;

[0009] It is understood that the target is a target encountered by the unmanned vehicle when performing a certain task in a battlefield environment; the unmanned vehicle information includes the unmanned vehicle's position, speed, and angle information; the target information includes the target's three-dimensional position coordinates, target speed, and target value.

[0010] Preferably, the unmanned vehicle information expression is:

[0011] m=(x m ,y m ,z m ,v m ,θ m ,c m Z m D m );

[0012] Among them, (x m ,y m ,z m ) represents the three-dimensional spatial coordinates of the autonomous vehicle, v m θ represents the speed of the driverless car. m c represents the velocity-direction angle of the autonomous vehicle. m Z represents the cost of driverless cars. m This represents the sensing and detection payload of the autonomous vehicle; D m The impact damage payload representing the unmanned vehicle The i-th sensing and detection payload is damaged by an attack.

[0013] Preferably, the target information expression is:

[0014]

[0015] in, Represents the three-dimensional spatial coordinates of the j-th target. This represents the velocity of the j-th target. This represents the target value of the j-th objective; This represents the velocity direction angle of the j-th target. The confidence level for the j-th target reflects the system's assessment of the accuracy and reliability of target information, and the degree of certainty regarding the target's existence and classification. Let id represent the threat level of the j-th target, encompassing the target's inherent danger attributes, specifically considering a combination of factors such as its potential destructive capabilities, attack intent, likelihood of action, and the vulnerability of the defender. j This is the number of target j, used to identify different targets.

[0016] Furthermore, the expression for the damage to the i-th sensing and detection payload of the unmanned vehicle is:

[0017]

[0018] in, The damage coefficient represents the i-th sensing and detection payload of the autonomous vehicle, including the payload's type, mass, velocity, shape, material, and structure; The radius of the payload coverage area of ​​the i-th sensing and detection payload of the autonomous vehicle; This represents the cost of the i-th sensing and detection payload of the autonomous vehicle, which is a comprehensive parameter of the resource occupation and resource consumption of a single payload. The payload strike distance of the i-th sensing and detection payload of the unmanned vehicle refers to the maximum distance at which the unmanned vehicle system carrying a specific payload can effectively strike the target.

[0019] Furthermore, the expression for the perception and detection payload of the unmanned vehicle is:

[0020] in, This represents the i-th sensing and detection payload of the autonomous vehicle. r represents the line-of-sight range of the i-th sensing and detection payload of the autonomous vehicle. i m This represents the sensing angle of the i-th sensing and detection payload of the autonomous vehicle. For example, the sensing and detection angle of radar is a wide-angle of 90 degrees, while the sensing and detection angle of photoelectric sensors is around 40-60 degrees. The cost of the i-th sensing and detection payload of the autonomous vehicle is a comprehensive parameter representing the resource occupation and resource consumption of a single payload. The coefficient representing the sensing capability of the i-th sensing payload of the autonomous vehicle is denoted as .

[0021] Furthermore, the sensing and detection capability coefficient of the i-th sensing and detection payload of the unmanned vehicle Depending on the type of payload, the larger the sensing and detection capability coefficient, the higher the accuracy or resolution that the i-th sensing and detection payload can achieve when detecting a target. The relevant information obtained by calling the i-th sensing and detection payload once through intelligent decision-making is more detailed and accurate.

[0022] Furthermore, the threat level of the target includes: target capabilities, target intent, and distance;

[0023] The target confidence level includes the target capability and the source of target information;

[0024] Furthermore, the target capabilities include target mobility, target jamming capabilities, and target defense capabilities;

[0025] The objectives are intended to include counterattack, defense, and escape;

[0026] The information sources are categorized as reliable, generally reliable, and unreliable.

[0027] Furthermore, the threat level expression for the target is:

[0028]

[0029] Among them, wx t The representative factor t represents the degree of influence of the threat level, t = 1, 2, 3, 4, 5; μ is a quantitative indicator of target threat and distance; A1 is the target's maneuverability; A2 is the target's jamming capability; A3 is the target's defense capability; Y is the target's intention.

[0030] The confidence expression for the target is:

[0031]

[0032] Among them, zx a Let t represent the degree of influence of factor t on the confidence level, where t = 1, 2, 3, 4, and L represent the information sources of the factors influencing the assessment of the target confidence level.

[0033] S2. Based on the unmanned vehicle information and target information, determine whether the target is a single target or a cluster of multiple targets;

[0034] If it is a single target, the distance between the unmanned vehicle and the corresponding target and the line of sight of the corresponding target are obtained; the distance between the unmanned vehicle and the corresponding target and the line of sight of the corresponding target are discretized to obtain the discretized target information of the unmanned vehicle and the single target;

[0035] If it is a multi-target cluster, the distance between the unmanned vehicle and each target in the multi-target cluster, the line of sight angle between the unmanned vehicle and the target, and the target area of ​​the multi-target cluster are obtained; the distance between the unmanned vehicle and the corresponding target, the line of sight angle between the corresponding target, and the target area of ​​the multi-target cluster are discretized to obtain the discretized information of the unmanned vehicle and the target of the multi-target cluster.

[0036] It is understandable that, assuming an autonomous vehicle system is performing a task in a complex environment, it may encounter a situation where it faces multiple targets and needs to react quickly and intelligently. If the target is a single target, the autonomous vehicle does not need to make a choice, but can directly act on the single target and make a payload call decision after reaching a certain position.

[0037] When faced with a cluster of targets, i.e., multiple targets, the autonomous vehicle first performs a grouping operation. Grouping refers to the autonomous vehicle system intelligently classifying or grouping the multiple targets it perceives in order to allocate resources and formulate action strategies more effectively. Through grouping, the autonomous vehicle can more clearly identify the characteristics and priorities of different target groups, providing a basis for target allocation and selection decisions based on deep reinforcement learning.

[0038] Preferably, the specific steps for obtaining the corresponding target area include: identifying the corresponding target information to obtain an identified multi-target cluster; it is understood that the multi-target cluster includes multiple single targets;

[0039] The area of ​​the multi-target cluster is obtained. Further, the expression for the area of ​​the multi-target cluster is:

[0040]

[0041] Where S is the area of ​​the multi-object cluster, x k The x-coordinate and y-coordinate of the k-th target i The k-th target y Axis coordinates; k = 1, 2, 3…n; n is the number of targets.

[0042] In one embodiment of the present invention, the specific steps for obtaining the area of ​​the identified target region using the Graham scanning method include:

[0043] Step a: Establish a planar coordinate system with the location of the unmanned vehicle as the origin, the direction of the unmanned vehicle's velocity as the x-axis, and the direction of the unmanned vehicle's velocity rotated 90 degrees counterclockwise as the y-axis.

[0044] In a multi-target cluster, multiple coordinate points of multiple individual targets in a planar coordinate system are determined, and the leftmost coordinate point in the planar coordinate system is selected as the target pole. The multiple coordinate points of the multi-object cluster in the planar coordinate system form a convex hull;

[0045] Furthermore, the leftmost coordinate point among multiple coordinate points in the planar coordinate system is the target pole, which is a vertex on the convex hull;

[0046] Step b: Calculate the polar angle of the single target point remaining in the coordinate system after removing the target poles in the multi-target cluster, and obtain multiple polar angles; sort the multiple polar angles to obtain a point set sorted by polar angle;

[0047] Based on the point set sorted by polar angle, the remaining single target points in the multi-target cluster after removing the target poles are sorted to obtain the sorted target points.

[0048] Furthermore, the sorting rules are as follows: sort by size from smallest to largest; if the polar angles are the same, sort by distance from the target pole. Sort by proximity

[0049] It is understood that the polar angle is the angle between each individual target point and the x-axis;

[0050] Step c: Initialize the convex hull corresponding to the multi-target cluster in the planar coordinate system, and obtain the convex hull P. e ,:

[0051] Step d: Let e ​​= 1. When e = 1, it is the first target point in the sorted target points.

[0052] Step e: Target pole and the e-th target point among the sorted target points Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 ;

[0053] Step f: Compare the size of e and E, where E represents the total number of target points after sorting. If e < E, increment the value e by 1 and return to step d. If e = E, end the training and obtain the final convex hull, which is the envelope corresponding to the multi-target cluster.

[0054] Furthermore, step e describes targeting the pole. and the e-th target point among the sorted target points Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 The specific steps include:

[0055] Determine the e-th target point Is it in the convex hull P? e If it is inside, skip it; if it is outside, then the e-th target point... Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 ;

[0056] Furthermore, determine the e-th target point. Is it in the convex hull P? e The specific internal steps include: checking the convex hull P e The last two target points and the e-th target point Does the formed three points satisfy a clockwise or counterclockwise direction? If it is clockwise (or counterclockwise), then the e-th target point... In the convex hull P e Inside, otherwise in the convex hull P e external.

[0057] Methods for updating the convex hull include: from the convex hull P e The last target point Begin by sequentially searching forward for the target point within the planar coordinate system. In the convex hull P e The previous target point within The target point Target point and the e-th target point The three points formed are in a counter-clockwise (or clockwise) direction;

[0058] The preceding target point is located both in the planar coordinate system and within the convex hull P.e The target point within; the target point and target point There may or may not be multiple coordinate points within a planar coordinate system;

[0059] target point To the target point All target points between the convex hull P e Delete and remove the e-th target point Add to the convex hull and obtain the updated convex hull P. e+1 .

[0060] S3. Build a reinforcement learning architecture;

[0061] Design a reward function; establish a reinforcement learning model based on the reinforcement learning architecture and reward function;

[0062] The reinforcement learning model includes input and output;

[0063] Preferably, the input is the state vector s of the autonomous vehicle decision and the action vector a of the autonomous vehicle decision;

[0064] Preferably, the output is the state-action value Q(s,a);

[0065] The expression for the state action value is:

[0066] Q(s,a)←Q(s,a)+α[R+γmax a′ Q(s′,a′)-Q(s,a)]

[0067] Where Q(s,a) represents the current state action value; α is the learning rate; R is the update reward signal; and R+γmax is the maximum value of the reward. a′ Q(s′,a′) represents the state action value at the previous time step, and γ represents the discount factor; max a′ Q(s′,a′) represents the action vector a′ of the previous time step where the maximum state action value Q(s,a) is collected under the state vector s′ of the previous time step; s is the state vector of the current time step, and a is the action vector of the current time step; α(·) is the learning rate of the error between the state action value of the previous time step and the state action value of the current time step.

[0068] For example, the specific steps for designing the reward function include:

[0069] To obtain the task efficiency of autonomous vehicles in selecting targets;

[0070] Identify the factors influencing the target selection performance of the autonomous vehicle; obtain the target tracking probability based on the aforementioned factors;

[0071] An initial reward function is designed based on the target tracking probability, and the initial reward function is optimized using uncertainty to obtain the final reward function;

[0072] The reward function is globally optimized to obtain an updated reward function; decision variables are then obtained based on the updated reward function.

[0073] It is understood that the target selection efficiency of the unmanned vehicle is the degree to which the selected target can be eventually tracked and the next payload call can be executed;

[0074] Preferably, the factors affecting the target selection performance of the unmanned vehicle include: angle factors, distance factors, speed factors, and environmental factors;

[0075] Furthermore, the expression for the target tracking probability is:

[0076]

[0077] Where F is the target tracking probability, C1 is the weighting coefficient of the speed factor on the target tracking probability, C2 is the weighting coefficient of the distance factor on the target tracking probability, and C3 is the weighting coefficient of the angle factor on the target tracking probability. v For speed factors; For angle factors; F d For distance factors; F h Environmental factors.

[0078] Furthermore, the expression for the updated reward function is:

[0079] r = δr p +(1-α)r g ;

[0080] Where r is the updated reward signal, δ is the control factor balancing the weights of local and global rewards, δ∈[0,1]. If δ=1, the target selection strategy learned by the DQN network is equivalent to a greedy rule based on marginal reward. When δ=0, the target selection decision only focuses on the global task performance, and each decision randomly obtains the average reward signal; r p The reward is a local reward, which is the marginal return of the current single decision, r. g The global reward is the result of the autonomous vehicle randomly selecting multiple targets and making multiple parallel decisions under the current single decision condition.

[0081] Furthermore, the local reward expression is:

[0082]

[0083] Among them, Fj Let value be the target tracking probability of the j-th target. j Let E(·) represent the target value of the j-th objective; E(·) represents the task effectiveness. Represents the decision matrix;

[0084] Furthermore, the reward function includes reward function one and reward function two;

[0085] The reward function, given the uncertainty of the objective, has the following expression for its task performance:

[0086]

[0087] in, The reward function is a signal; β is the weighting factor, which takes a random value between [0,1]. For reward function two, E ρ (X) is the initial reward function 2;

[0088] The performance of the second reward function under the uncertainty of the vehicle's survival environment is expressed as follows:

[0089]

[0090] Where fa represents the uncertainty factor of the autonomous vehicle based on its living environment, and c m This represents the cost of driverless cars.

[0091] Furthermore, the uncertainties include: uncertainties based on the vehicle's living environment and uncertainties based on the target;

[0092] Furthermore, the initial reward function includes initial reward function one and initial reward function two; the expression for initial reward function one is:

[0093]

[0094] Among them, E σ (X) represents the reward signal F considering the uncertainty of the environment in real-world conditions. j Let value be the target tracking probability of the j-th target. j Let be the target value of the j-th objective.

[0095] The second initial reward function is expressed as follows:

[0096]

[0097] Among them, E ρ (X) represents the reward signal considering target uncertainty, ρ is the target uncertainty factor, and takes a random value between [0,1]; fb jLet be the uncertainty parameter of the autonomous vehicle's target-based protective measures for the j-th target.

[0098] Preferably, the reinforcement learning architecture further includes a state space and an action space;

[0099] Furthermore, based on the state vector s, the state space is designed to obtain the state space, which is expressed as:

[0100]

[0101] in, The action in the previous state is represented by dis, the distance between the discretized autonomous vehicle and the corresponding target is represented by va, the line of sight between the discretized autonomous vehicle and the corresponding target is represented by area, and the area of ​​the discretized autonomous vehicle and the corresponding target is represented by area.

[0102] The action space includes the action space of the sensing load and the action space of the equipment load;

[0103] The expression for the action space a1 of the sensing load is:

[0104] a1 = (g1, g2, g3, g4)

[0105] Among them, g1 represents the call to the vehicle camera, g2 represents the call to the vehicle radar, g3 represents the call to the vehicle infrared, and g4 represents the call to the vehicle depth camera.

[0106] The expression for the motion space a2 of the equipment load is:

[0107] a2 = (u1, u2, u3)

[0108] Among them, u1 represents the use of vehicle-mounted water bombs, u2 represents the use of vehicle-mounted rubber bombs, and u3 represents the use of vehicle-mounted laser irradiation.

[0109] S4. Obtain the step rewards for the perception load and equipment load decision based on the unmanned vehicle load parameters and target information, respectively.

[0110] Preferably, step S4, which involves obtaining the perception load and equipment load decision-making steps based on the unmanned vehicle's load parameters and target-related information, specifically includes:

[0111] Based on the information of the unmanned vehicle's payload parameters and the target, the sensor payload matching degree and equipment payload matching degree between the unmanned vehicle and the corresponding target are obtained; based on the sensor payload matching degree and equipment payload matching degree, the step reward for the sensor payload and equipment payload decision is obtained.

[0112] Preferably, the sensing load matching degree includes: distance matching degree, angle matching degree, and sensing load performance matching degree;

[0113] The equipment load matching degree includes: distance matching degree, angle matching degree, and equipment load performance matching degree;

[0114] It is understood that the distance matching degree refers to the matching degree between the distance between the autonomous vehicle and the target and the line-of-sight distance of the sensing payload;

[0115] The angle matching degree between the autonomous vehicle and the target is the degree of matching between the angle of the autonomous vehicle and the target and the detection angle of the sensing payload;

[0116] Furthermore, the step reward for the perceived load is expressed as:

[0117] R1=p1·P l +p2·P r +p3·P c

[0118] Where R1 is the step reward of the sensing payload, p1 is the weight coefficient of the distance matching degree, p2 is the weight coefficient of the angle matching degree between the autonomous vehicle and the target, p3 is the weight coefficient of the performance matching degree of the sensing payload, p1+p2+p3=1, and 0≤p i ≤1. P l For distance matching degree, P r P represents the angle matching degree between the autonomous vehicle and the target. c To ensure the matching degree of the sensing load performance;

[0119] The step reward expression for the equipment load is:

[0120] R2=q1·P dis +q2·P ran +q3·P dp

[0121] Where R2 is the step reward of the equipment load, q1 is the weight coefficient of the distance matching degree of the equipment load, q2 is the weight coefficient of the range matching degree of the equipment load and the target, q3 is the weight coefficient of the performance matching degree of the equipment load, q1+q2+q3=1, and 0≤q i ≤1, P dis For the distance matching degree of equipment load, P ran To ensure the range matching between equipment payload and target, P dp For equipment load performance matching;

[0122] S5. Input the discrete unmanned vehicle information and target information obtained in step S2 into the reinforcement learning model, update the state action value based on the step reward, and obtain the optimal state action value; obtain the decision result based on the optimal state action value, and apply the payload to the target in the environment according to the decision result until the termination condition is met, and the decision stops.

[0123] Preferably, step S5, which involves invoking payloads to targets in the environment based on the decision result until the termination condition is met, includes the following specific steps:

[0124] The information of the discrete unmanned vehicle and the target described in step S2 is input into the reinforcement learning model, and the state action value is updated using the update reward function and step reward to obtain the optimal state action value.

[0125] Based on the optimal state action value, a decision result is obtained. The corresponding load is invoked according to the decision result. The decision stops when the termination condition is met.

[0126] Preferably, the termination condition is that the target threat level disappears or the target threat is confirmed.

[0127] This invention considers the volatile and unpredictable battlefield environment and establishes a contingency decision-making strategy for unmanned vehicles carrying military equipment under various unforeseen circumstances. It introduces multiple target fusion information, such as target threat level and target confidence level, and appropriately sets contingency decision-making activation conditions through comprehensive judgment of the target fusion information. This invention also establishes a multi-target clustering mechanism. For multiple targets in a battlefield environment, it classifies targets into groups using multiple fusion information and uses a convex hull algorithm to obtain the cluster envelope, taking action decisions based on the target clusters. Furthermore, this invention utilizes deep reinforcement learning technology to efficiently and intelligently make decisions based on the different information attributes of multi-target groups, significantly improving the autonomous decision-making and action capabilities of unmanned vehicles in complex environments.

[0128] This invention addresses the goal of intelligent unmanned vehicle operations in combat scenarios by employing reinforcement learning algorithms to solve the problem of real-time intelligent and rapid decision-making on the battlefield. Intelligent decision-making in battlefield combat requires unmanned vehicles to determine whether to initiate ad hoc decision-making based on the situation in the battlefield environment and to immediately mobilize payloads or actions to respond, thereby meeting the requirements of battlefield immediacy and decision-making accuracy.

[0129] Compared with the prior art, the present invention has at least the following beneficial effects:

[0130] (1) This invention ensures efficient use of resources and avoids frequent ad-hoc decision-making by unmanned vehicles to make frequent task suggestions to the upper command system, which may cause interference to the upper task execution system and communication congestion. It sets the conditions for opening the ad-hoc decision-making module.

[0131] (2) This invention designs a multi-target cluster reinforcement learning model to realize the rapid target selection of unmanned vehicles in the mission combat allocation of unmanned vehicles, establishes a multi-target grouping mechanism, and uses the convex hull algorithm to obtain the cluster envelope and take action decision information for the target cluster.

[0132] (2) This invention fully considers the characteristics of multiple task links, multiple load scheduling and multiple execution actions in the context of ad hoc decision-making scenarios. The rapid decision-making technology realizes the ad hoc deployment in this solution, thereby meeting the requirements of battlefield immediacy and decision correctness. Attached Figure Description

[0133] The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of the invention.

[0134] Figure 1 This is a schematic diagram of the behavior decision-making process of an autonomous vehicle based on reinforcement learning in an embodiment of the present invention;

[0135] Figure 2 This is a schematic diagram illustrating the target identification by an unmanned vehicle in an embodiment of the present invention;

[0136] Figure 3 This is a schematic diagram of the reinforcement learning architecture in an embodiment of the present invention;

[0137] Figure 4 This is a schematic diagram of the payload invocation decision-making process based on a reinforcement learning model in an embodiment of the present invention. Detailed Implementation

[0138] To better understand the above-described objectives, features, and advantages of the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments of the present invention and the features thereof can be combined with each other. Furthermore, the present invention can be implemented in other ways different from those described herein; therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0139] A specific embodiment of the present invention, such as Figure 1-4 A decision-making method for autonomous vehicles based on deep reinforcement learning is disclosed.

[0140] To illustrate the effectiveness of the method proposed in this invention, the following detailed description of the above technical solution is provided through a specific embodiment. The specific implementation steps are as follows:

[0141] S1. Obtain unmanned vehicle information and corresponding target information; obtain unmanned vehicle payload parameters based on unmanned vehicle information;

[0142] It is understood that the target is a target encountered by the unmanned vehicle when performing a certain task in a battlefield environment; the unmanned vehicle information includes the unmanned vehicle's position, speed, and angle information; the target information includes the target's three-dimensional position coordinates, target speed, and target value.

[0143] Preferably, the unmanned vehicle information expression is:

[0144] m=(x m ,y m ,z m ,v m ,θ m ,c m Z m D m );

[0145] Among them, (x m ,y m ,z m ) represents the three-dimensional spatial coordinates of the autonomous vehicle, vm represents the speed of the autonomous vehicle, and θ represents the speed of the autonomous vehicle. m c represents the velocity-direction angle of the autonomous vehicle. m Z represents the cost of driverless cars. m This represents the sensing and detection payload of the autonomous vehicle; D m The impact damage payload representing the unmanned vehicle The i-th sensing and detection payload is damaged by an attack.

[0146] Preferably, the target information expression is:

[0147]

[0148] in, Represents the three-dimensional spatial coordinates of the j-th target. This represents the velocity of the j-th target. This represents the target value of the j-th objective; This represents the velocity direction angle of the j-th target. The confidence level for the j-th target reflects the system's assessment of the accuracy and reliability of target information, and the degree of certainty regarding the target's existence and classification. Let id represent the threat level of the j-th target, encompassing the target's inherent danger attributes, specifically considering a combination of factors such as its potential destructive capabilities, attack intent, likelihood of action, and the vulnerability of the defender. j This is the number of target j, used to identify different targets.

[0149] Furthermore, the expression for the damage to the i-th sensing and detection payload of the unmanned vehicle is:

[0150]

[0151] in, The damage coefficient represents the i-th sensing and detection payload of the autonomous vehicle, including the payload's type, mass, velocity, shape, material, and structure; The radius of the payload coverage area of ​​the i-th sensing and detection payload of the autonomous vehicle; This represents the cost of the i-th sensing and detection payload of the autonomous vehicle, which is a comprehensive parameter of the resource occupation and resource consumption of a single payload. The payload strike distance of the i-th sensing and detection payload of the unmanned vehicle refers to the maximum distance at which the unmanned vehicle system carrying a specific payload can effectively strike the target.

[0152] Furthermore, the expression for the perception and detection payload of the unmanned vehicle is:

[0153] in, This represents the i-th sensing and detection payload of the autonomous vehicle. r represents the line-of-sight range of the i-th sensing and detection payload of the autonomous vehicle. i m This represents the sensing angle of the i-th sensing and detection payload of the autonomous vehicle. For example, the sensing and detection angle of radar is a wide-angle of 90 degrees, while the sensing and detection angle of photoelectric sensors is around 40-60 degrees. The cost of the i-th sensing and detection payload of the autonomous vehicle is a comprehensive parameter representing the resource occupation and resource consumption of a single payload. The coefficient representing the sensing capability of the i-th sensing payload of the autonomous vehicle is denoted as .

[0154] Furthermore, the sensing and detection capability coefficient of the i-th sensing and detection payload of the unmanned vehicle Depending on the type of payload, the larger the sensing and detection capability coefficient, the higher the accuracy or resolution that the i-th sensing and detection payload can achieve when detecting a target. The relevant information obtained by calling the i-th sensing and detection payload once through intelligent decision-making is more detailed and accurate.

[0155] Furthermore, the threat level of the target includes: target capabilities, target intent, and distance;

[0156] The target confidence level includes the target capability and the source of target information;

[0157] Furthermore, the target capabilities include target mobility, target jamming capabilities, and target defense capabilities;

[0158] The objectives are intended to include counterattack, defense, and escape;

[0159] The information sources are categorized as reliable, generally reliable, and unreliable.

[0160] Furthermore, the threat level expression for the target is:

[0161]

[0162] Among them, wx tThe representative factor t represents the degree of influence of the threat level, t = 1, 2, 3, 4, 5; μ is a quantitative indicator of target threat and distance; A1 is the target's maneuverability; A2 is the target's jamming capability; A3 is the target's defense capability; Y is the target's intention.

[0163] The confidence expression for the target is:

[0164]

[0165] Among them, zx a Let t represent the degree of influence of factor t on the confidence level, where t = 1, 2, 3, 4, and L represent the information sources of the factors influencing the assessment of the target confidence level.

[0166] S2. Based on the unmanned vehicle information and target information, determine whether the target is a single target or a cluster of multiple targets;

[0167] If it is a single target, the distance between the unmanned vehicle and the corresponding target and the line of sight of the corresponding target are obtained; the distance between the unmanned vehicle and the corresponding target and the line of sight of the corresponding target are discretized to obtain the discretized target information of the unmanned vehicle and the single target;

[0168] If it is a multi-target cluster, the distance between the unmanned vehicle and each target in the multi-target cluster, the line of sight angle between the unmanned vehicle and the target, and the target area of ​​the multi-target cluster are obtained; the distance between the unmanned vehicle and the corresponding target, the line of sight angle between the corresponding target, and the target area of ​​the multi-target cluster are discretized to obtain the discretized information of the unmanned vehicle and the target of the multi-target cluster.

[0169] It is understandable that, assuming an autonomous vehicle system is performing a task in a complex environment, it may encounter a situation where it faces multiple targets and needs to react quickly and intelligently. If the target is a single target, the autonomous vehicle does not need to make a choice, but can directly act on the single target and make a payload call decision after reaching a certain position.

[0170] When faced with a cluster of targets, i.e., multiple targets, the autonomous vehicle first performs a grouping operation. Grouping refers to the autonomous vehicle system intelligently classifying or grouping the multiple targets it perceives in order to allocate resources and formulate action strategies more effectively. Through grouping, the autonomous vehicle can more clearly identify the characteristics and priorities of different target groups, providing a basis for target allocation and selection decisions based on deep reinforcement learning.

[0171] Preferably, the specific steps for obtaining the corresponding target area include: identifying the corresponding target information to obtain an identified multi-target cluster; it is understood that the multi-target cluster includes multiple single targets;

[0172] The area of ​​the multi-target cluster is obtained. Further, the expression for the area of ​​the multi-target cluster is:

[0173]

[0174] Where S is the area of ​​the multi-object cluster, x k The x-coordinate and y-coordinate of the k-th target i The k-th target y Axis coordinates; k = 1, 2, 3…n; n is the number of targets.

[0175] In one embodiment of the present invention, the specific steps for obtaining the area of ​​the identified target region using the Graham scanning method include:

[0176] Step a: Establish a planar coordinate system with the location of the unmanned vehicle as the origin, the direction of the unmanned vehicle's velocity as the x-axis, and the direction of the unmanned vehicle's velocity rotated 90 degrees counterclockwise as the y-axis.

[0177] In a multi-target cluster, multiple coordinate points of multiple individual targets in a planar coordinate system are determined, and the leftmost coordinate point in the planar coordinate system is selected as the target pole. The multiple coordinate points of the multi-object cluster in the planar coordinate system form a convex hull;

[0178] Furthermore, the leftmost coordinate point among multiple coordinate points in the planar coordinate system is the target pole, which is a vertex on the convex hull;

[0179] Step b: Calculate the polar angle of the single target point remaining in the coordinate system after removing the target poles in the multi-target cluster, and obtain multiple polar angles; sort the multiple polar angles to obtain a point set sorted by polar angle;

[0180] Based on the point set sorted by polar angle, the remaining single target points in the multi-target cluster after removing the target poles are sorted to obtain the sorted target points.

[0181] Furthermore, the sorting rules are as follows: sort by size from smallest to largest; if the polar angles are the same, sort by distance from the target pole. Sort by proximity

[0182] It is understood that the polar angle is the angle between each individual target point and the x-axis;

[0183] Step c: Initialize the convex hull corresponding to the multi-target cluster in the planar coordinate system, and obtain the convex hull P. e ,:

[0184] Step d: Let e ​​= 1. When e = 1, it is the first target point in the sorted target points.

[0185] Step e: Target pole and the e-th target point among the sorted target points Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 ;

[0186] Step f: Compare the size of e and E, where E represents the total number of target points after sorting. If e < E, increment the value e by 1 and return to step d. If e = E, end the training and obtain the final convex hull, which is the envelope corresponding to the multi-target cluster.

[0187] Furthermore, step e describes targeting the pole. and the e-th target point among the sorted target points Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 The specific steps include:

[0188] Determine the e-th target point Is it in the convex hull P? e If it is inside, skip it; if it is outside, then the e-th target point... Place the convex hull P e In the process, we obtain the convex hull P after the e-th update. e+1 ;

[0189] Furthermore, determine the e-th target point. Is it in the convex hull P? e The specific internal steps include: checking the convex hull P e The last two target points and the e-th target point Does the formed three points satisfy a clockwise or counterclockwise direction? If it is clockwise (or counterclockwise), then the e-th target point... In the convex hull P e Inside, otherwise in the convex hull P e external.

[0190] Methods for updating the convex hull include: from the convex hull P e The last target point Begin by sequentially searching forward for the target point within the planar coordinate system. In the convex hull P e The previous target point within The target point Target point and the e-th target point The three points formed are in a counter-clockwise (or clockwise) direction;

[0191] The preceding target point is located both in the planar coordinate system and within the convex hull P.e The target point within; the target point and target point There may or may not be multiple coordinate points within a planar coordinate system;

[0192] target point To the target point All target points between the convex hull P e Delete and remove the e-th target point Add to the convex hull and obtain the updated convex hull P. e+1 .

[0193] S3. Build a reinforcement learning architecture;

[0194] Design a reward function; establish a reinforcement learning model based on the reinforcement learning architecture and reward function;

[0195] The reinforcement learning model includes input and output;

[0196] Preferably, the input is the state vector s of the autonomous vehicle decision and the action vector a of the autonomous vehicle decision;

[0197] Preferably, the output is the state-action value Q(s,a);

[0198] The expression for the state action value is:

[0199] Q(s,a)←Q(s,a)+α[R+γmax a′ Q(s′,a′)-Q(s,a)]

[0200] Where Q(s,a) represents the current state action value; α is the learning rate; R is the update reward signal; and R+γmax is the maximum value of the reward. a′ Q(s′,a′) represents the state action value at the previous time step, and γ represents the discount factor; max a′ Q(s′,a′) represents the action vector a′ of the previous time step where the maximum state action value Q(s,a) is collected under the state vector s′ of the previous time step; s is the state vector of the current time step, and a is the action vector of the current time step; α(·) is the learning rate of the error between the state action value of the previous time step and the state action value of the current time step.

[0201] For example, the specific steps for designing the reward function include:

[0202] To obtain the task efficiency of autonomous vehicles in selecting targets;

[0203] Identify the factors influencing the target selection performance of the autonomous vehicle; obtain the target tracking probability based on the aforementioned factors;

[0204] An initial reward function is designed based on the target tracking probability, and the initial reward function is optimized using uncertainty to obtain the final reward function;

[0205] The reward function is globally optimized to obtain an updated reward function; decision variables are then obtained based on the updated reward function.

[0206] It is understood that the target selection efficiency of the unmanned vehicle is the degree to which the selected target can be eventually tracked and the next payload call can be executed;

[0207] Preferably, the factors affecting the target selection performance of the unmanned vehicle include: angle factors, distance factors, speed factors, and environmental factors;

[0208] Furthermore, the expression for the target tracking probability is:

[0209]

[0210] Where F is the target tracking probability, C1 is the weighting coefficient of the speed factor on the target tracking probability, C2 is the weighting coefficient of the distance factor on the target tracking probability, and C3 is the weighting coefficient of the angle factor on the target tracking probability. v For speed factors; For angle factors; F d For distance factors; F h Environmental factors.

[0211] Furthermore, the expression for the updated reward function is:

[0212] r = δr p +(1-α)r g ;

[0213] Where r is the updated reward signal, δ is the control factor balancing the weights of local and global rewards, δ∈[0,1]. If δ=1, the target selection strategy learned by the DQN network is equivalent to a greedy rule based on marginal reward. When δ=0, the target selection decision only focuses on the global task performance, and each decision randomly obtains the average reward signal; r p The reward is a local reward, which is the marginal return of the current single decision, r. g The global reward is the result of the autonomous vehicle randomly selecting multiple targets and making multiple parallel decisions under the current single decision condition.

[0214] Furthermore, the local reward expression is:

[0215]

[0216] Among them, Fj Let value be the target tracking probability of the j-th target. j Let E(·) represent the target value of the j-th objective; E(·) represents the task effectiveness. Represents the decision matrix;

[0217] Furthermore, the reward function includes reward function one and reward function two;

[0218] The reward function, given the uncertainty of the objective, has the following expression for its task performance:

[0219]

[0220] in, The reward function is a signal; β is the weighting factor, which takes a random value between [0,1]. For reward function two, E ρ (X) is the initial reward function 2;

[0221] The performance of the second reward function under the uncertainty of the vehicle's survival environment is expressed as follows:

[0222]

[0223] Where fa represents the uncertainty factor of the autonomous vehicle based on its living environment, and c m This represents the cost of driverless cars.

[0224] Furthermore, the uncertainties include: uncertainties based on the vehicle's living environment and uncertainties based on the target;

[0225] Furthermore, the initial reward function includes initial reward function one and initial reward function two; the expression for initial reward function one is:

[0226]

[0227] Among them, E σ (X) represents the reward signal F considering the uncertainty of the environment in real-world conditions. j Let value be the target tracking probability of the j-th target. j Let be the target value of the j-th objective.

[0228] The second initial reward function is expressed as follows:

[0229]

[0230] Among them, E ρ (X) represents the reward signal considering target uncertainty, ρ is the target uncertainty factor, and takes a random value between [0,1]; fb jLet be the uncertainty parameter of the autonomous vehicle's target-based protective measures for the j-th target.

[0231] Preferably, the reinforcement learning architecture further includes a state space and an action space;

[0232] Furthermore, based on the state vector s, the state space is designed to obtain the state space, which is expressed as:

[0233]

[0234] in, The action in the previous state is represented by dis, the distance between the discretized autonomous vehicle and the corresponding target is represented by va, the line of sight between the discretized autonomous vehicle and the corresponding target is represented by area, and the area of ​​the discretized autonomous vehicle and the corresponding target is represented by area.

[0235] The action space includes the action space of the sensing load and the action space of the equipment load;

[0236] The expression for the action space a1 of the sensing load is:

[0237] a1 = (g1, g2, g3, g4)

[0238] Among them, g1 represents the call to the vehicle camera, g2 represents the call to the vehicle radar, g3 represents the call to the vehicle infrared, and g4 represents the call to the vehicle depth camera.

[0239] The expression for the motion space a2 of the equipment load is:

[0240] a2 = (u1, u2, u3)

[0241] Among them, u1 represents the use of vehicle-mounted water bombs, u2 represents the use of vehicle-mounted rubber bombs, and u3 represents the use of vehicle-mounted laser irradiation.

[0242] S4. Obtain the step rewards for the perception load and equipment load decision based on the unmanned vehicle load parameters and target information, respectively.

[0243] Preferably, step S4, which involves obtaining the perception load and equipment load decision-making steps based on the unmanned vehicle's load parameters and target-related information, specifically includes:

[0244] Based on the information of the unmanned vehicle's payload parameters and the target, the sensor payload matching degree and equipment payload matching degree between the unmanned vehicle and the corresponding target are obtained; based on the sensor payload matching degree and equipment payload matching degree, the step reward for the sensor payload and equipment payload decision is obtained.

[0245] Preferably, the sensing load matching degree includes: distance matching degree, angle matching degree, and sensing load performance matching degree;

[0246] The equipment load matching degree includes: distance matching degree, angle matching degree, and equipment load performance matching degree;

[0247] It is understood that the distance matching degree refers to the matching degree between the distance between the autonomous vehicle and the target and the line-of-sight distance of the sensing payload;

[0248] The angle matching degree between the autonomous vehicle and the target is the degree of matching between the angle of the autonomous vehicle and the target and the detection angle of the sensing payload;

[0249] Furthermore, the step reward for the perceived load is expressed as:

[0250] R1=p1·P l +p2·P r +p3·P c

[0251] Where R1 is the step reward of the sensing payload, p1 is the weight coefficient of the distance matching degree, p2 is the weight coefficient of the angle matching degree between the autonomous vehicle and the target, p3 is the weight coefficient of the performance matching degree of the sensing payload, p1+p2+p3=1, and 0≤p i ≤1. P l For distance matching degree, P r P represents the angle matching degree between the autonomous vehicle and the target. c To ensure the matching degree of the sensing load performance;

[0252] The step reward expression for the equipment load is:

[0253] R2=q1·P dis +q2·P ran +q3·R dp

[0254] Where R2 is the step reward of the equipment load, q1 is the weight coefficient of the distance matching degree of the equipment load, q2 is the weight coefficient of the range matching degree of the equipment load and the target, q3 is the weight coefficient of the performance matching degree of the equipment load, q1+q2+q3=1, and 0≤q i ≤1, P dis For the distance matching degree of equipment load, P ran To ensure the range matching between equipment payload and target, P dp For equipment load performance matching;

[0255] S5. Input the discrete unmanned vehicle information and target information obtained in step S2 into the reinforcement learning model, update the state action value based on the step reward, and obtain the optimal state action value; obtain the decision result based on the optimal state action value, and apply the payload to the target in the environment according to the decision result until the termination condition is met, and the decision stops.

[0256] Preferably, step S5, which involves invoking payloads to targets in the environment based on the decision result until the termination condition is met, includes the following specific steps:

[0257] The information of the discrete unmanned vehicle and the target described in step S2 is input into the reinforcement learning model, and the state action value is updated using the update reward function and step reward to obtain the optimal state action value.

[0258] Based on the optimal state action value, a decision result is obtained. The corresponding load is invoked according to the decision result. The decision stops when the termination condition is met.

[0259] Preferably, the termination condition is that the target threat level disappears or the target threat is confirmed.

[0260] This invention considers the volatile and unpredictable battlefield environment and establishes a contingency decision-making strategy for unmanned vehicles carrying military equipment under various unforeseen circumstances. It introduces multiple target fusion information, such as target threat level and target confidence level, and appropriately sets contingency decision-making activation conditions through comprehensive judgment of the target fusion information. This invention also establishes a multi-target clustering mechanism. For multiple targets in a battlefield environment, it classifies targets into groups using multiple fusion information and uses a convex hull algorithm to obtain the cluster envelope, taking action decisions based on the target clusters. Furthermore, this invention utilizes deep reinforcement learning technology to efficiently and intelligently make decisions based on the different information attributes of multi-target groups, significantly improving the autonomous decision-making and action capabilities of unmanned vehicles in complex environments.

[0261] This invention addresses the goal of intelligent unmanned vehicle operations in combat scenarios by employing reinforcement learning algorithms to solve the problem of real-time intelligent and rapid decision-making on the battlefield. Intelligent decision-making in battlefield combat requires unmanned vehicles to determine whether to initiate ad hoc decision-making based on the situation in the battlefield environment and to immediately mobilize payloads or actions to respond, thereby meeting the requirements of battlefield immediacy and decision-making accuracy.

[0262] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A decision-making method for autonomous vehicles based on deep reinforcement learning, characterized in that, include: S1. Obtain information about the unmanned vehicle and the corresponding target information; Obtain the payload parameters of the unmanned vehicle based on the unmanned vehicle information; S2. Based on the unmanned vehicle information and target information, discretize the information to obtain the discretized unmanned vehicle information and target information. Specific steps include: Based on the unmanned vehicle information and target information, it is determined whether the target is a single target or a cluster of multiple targets; if it is a single target, the distance between the unmanned vehicle and the corresponding target and the line of sight angle of the corresponding target are obtained; the distance between the unmanned vehicle and the corresponding target and the line of sight angle of the corresponding target are discretized to obtain the discretized target information of the unmanned vehicle and the single target. If it is a multi-target cluster, the distance between the unmanned vehicle and each target in the multi-target cluster, the line of sight angle between the unmanned vehicle and the target, and the target area of ​​the multi-target cluster are obtained; the distance between the unmanned vehicle and the corresponding target, the line of sight angle between the corresponding target, and the target area of ​​the multi-target cluster are discretized to obtain the discretized information of the unmanned vehicle and the target of the multi-target cluster. S3. Construct a reinforcement learning architecture; design a reward function; establish a reinforcement learning model based on the reinforcement learning architecture and reward function; the output of the reinforcement learning model is a state-action value; S4. Obtain the step rewards for the perception load and equipment load decision based on the unmanned vehicle load parameters and target information, respectively. S5. Input the discrete unmanned vehicle information and target information obtained in step S2 into the reinforcement learning model, update the state action value based on the step reward, and obtain the optimal state action value; obtain the decision result based on the optimal state action value, and apply the payload to the target in the environment according to the decision result until the termination condition is met, and the decision stops.

2. The decision-making method according to claim 1, characterized in that, The target is the target encountered by the unmanned vehicle when performing a mission in a battlefield environment; the unmanned vehicle information includes the unmanned vehicle's position, speed, and angle information; the target information includes the target's three-dimensional position coordinates, target speed, and target value.

3. The decision-making method according to any one of claims 1-2, characterized in that, The expression for the area of ​​the multi-objective cluster is: , in, S The area of ​​the multi-object cluster. No. k One goal x Axis coordinates No. k The y-axis coordinates of the target; k =1,2,3… n ; n The number of targets.

4. The decision-making method according to claim 3, characterized in that, The input to the reinforcement learning model is the state vector of the autonomous vehicle's decision-making process. Action vectors for autonomous vehicle decision-making .

5. The decision-making method according to claim 4, characterized in that, The expression for the state action value is: in, This indicates the current state / action value. To update the reward signal, This represents the state / action value from the previous moment. Indicates the discount factor; Represents the state vector at the previous time step. Collect the maximum state action value The action vector of the previous moment ; s Let this be the current state vector. a This is the action vector at the current moment; (•) represents the learning rate of the error between the state / action value at the previous time step and the state / action value at the current time step.

6. The decision-making method according to claim 5, characterized in that, The specific steps for designing the reward function include: The task performance of the autonomous vehicle in selecting targets is obtained; the task performance of the autonomous vehicle in selecting targets is the degree to which the selected target can be eventually tracked and the next payload call can be executed. Identify the factors influencing the target selection performance of the autonomous vehicle; obtain the target tracking probability based on the aforementioned factors; An initial reward function is designed based on the target tracking probability, and the initial reward function is optimized using uncertainty to obtain the final reward function; The reward function is globally optimized to obtain an updated reward function; decision variables are then obtained based on the updated reward function. The expression for the updated reward function is: ; in, r To update the reward signal, To balance the weights of local and global rewards, ,like The target selection strategy learned by the DQN network is equivalent to a greedy rule based on marginal reward, and when In this case, the goal selection decision only focuses on the overall task effectiveness, and each decision randomly obtains the average reward signal; This is a local reward, which is the marginal reward of the current individual decision. The global reward is the result of the autonomous vehicle randomly selecting multiple targets and making multiple parallel decisions under the current single decision condition. The local reward expression is: in, For the first j The target tracking probability of each target For the first j The target value of each objective; To indicate task effectiveness; Represents the decision matrix; The reward function includes reward function one and reward function two; The reward function, given the uncertainty of the objective, has the following expression for its task performance: in, The reward function is a signal; The weighting factor is a random value. between; For reward function two, The second initial reward function; The performance of the second reward function under the uncertainty of the vehicle's survival environment is expressed as follows: in, For autonomous vehicles, based on the uncertainties of their operating environment, This represents the cost of driverless cars; The uncertainties include: uncertainties based on the vehicle's operating environment and uncertainties based on the target; The initial reward function includes initial reward function one and initial reward function two; the expression for initial reward function one is: in, To consider reward signals under real-world environmental uncertainties For the first j The target tracking probability of each target For the first j The target value of each objective; The second initial reward function is expressed as follows: in, To account for reward signals under conditions of target uncertainty, As the uncertainty factor of the objective, a random value is taken in... between; For driverless cars for the first j Uncertainty parameters of each target based on target protective measures; The reinforcement learning architecture also includes a state space and an action space.

7. The decision-making method according to claim 6, characterized in that, The expression for the target tracking probability is: in, F For target tracking probability, The weighting coefficient for the impact of speed on target tracking probability. This represents the weighting coefficient for the impact of distance on the target tracking probability. These are the weighting coefficients for the angle factor and its impact on the target tracking probability. For speed factors; Angle factor; Distance factor; Environmental factors.

8. The decision-making method according to claim 7, characterized in that, The expression for the step reward of the sensing payload is: in, Rewards for the steps of sensing the load. This is the weighting coefficient for distance matching. This represents the weighting coefficient for the angle matching degree between the autonomous vehicle and the target. The weighting coefficients for the perceived load performance matching degree, , For distance matching degree, To improve the angle matching between the autonomous vehicle and the target, To sense the load performance matching degree.

9. The decision-making method according to claim 1, characterized in that, The termination condition in step S5 is that the target threat level disappears or the target threat is confirmed.

Citation Information

Patent Citations

  • Multi-unmanned aerial vehicle task decision-making method based on MADDPG

    CN111880563A

  • Unmanned chariot team firepower distribution method based on deep reinforcement learning

    CN112364972A