Genetic optimization intelligent decision method and device based on hierarchical fuzzy reasoning

By adopting a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning, the problem of intelligent decision-making in the field of cyber warfare has been solved, enabling covert, precise, and efficient strikes against the enemy, and improving the efficiency and scope of mission completion.

CN115423100BActive Publication Date: 2026-03-27INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

At present, research on intelligent technology in the field of cyber warfare is still in its infancy, lacking effective intelligent decision-making methods, and thus unable to achieve covert, precise, and efficient strikes against the enemy.

Method used

A genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is adopted. By fuzzifying each state variable and decision variable, a tactical rule database is formed. The global adversarial reward function is used to perform tactical optimization of the genetic optimization algorithm, so as to realize action decision-making in the network electrical domain and physical domain.

Benefits of technology

It effectively reduces the number of tactical rules and the amount of reasoning computation, enabling the initial exploratory application of intelligent decision-making in the field of cyber warfare, and improving the efficiency and scope of mission completion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115423100B_ABST
    Figure CN115423100B_ABST
Patent Text Reader

Abstract

The application provides a genetic optimization intelligent decision-making method and device based on hierarchical fuzzy reasoning, which comprises the following steps: fuzzifying each state variable affecting the decision output in network electricity confrontation, each output variable of network electricity layer decision and each output variable of physical layer decision in their respective domains to obtain multiple fuzzy sets, performing digital symbol coding on the fuzzy sets to obtain the antecedent and the consequent of fuzzy reasoning, and forming a tactical rule database; in each battle round, the following steps are repeatedly executed until the termination condition is met: determining the input variables and their dynamic weights of the current decision period, reasoning the action decision of the network electricity domain and the action decision of the physical domain according to the tactical rule database, and updating the input variables of the next decision period; calculating the global confrontation return value based on the global confrontation return function, performing tactical optimization of the genetic optimization algorithm, and updating the tactical rule database. The application can realize the preliminary exploratory application of intelligent decision-making in the field of network electricity confrontation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of collaborative technology for unmanned swarm systems, and in particular to a genetic optimization intelligent decision-making method and device based on hierarchical fuzzy reasoning. Background Technology

[0002] Cyber-electronic warfare is an emerging field of combat, characterized by its ability to engage in combat across land, sea, air, and space. It refers to the empowerment of physical domain warfare through various means, including wireless and wired reconnaissance, attack, and defense, thereby gaining valuable time windows for physical domain energy strikes and enhancing their effectiveness. Compared to electronic warfare, cyber-electronic warfare offers greater precision and more efficient energy utilization; however, research into intelligent technologies in this field is currently lacking.

[0003] In military fields such as joint cross-domain collaborative operations, large-area battlefield reconnaissance, electronic warfare, and multi-domain collaborative information sharing, cyber-electronic warfare has extremely important application prospects. Unmanned swarms are increasingly replacing humans in infiltrating enemy territory through self-organized swarm collaboration, conducting covert reconnaissance and surprise attacks. This can solve problems that a single unmanned system cannot accomplish or solve, greatly improving the efficiency and scope of mission completion, significantly empowering various scenarios, and changing the traditional working methods of single individuals. Drawing on this form of warfare, exploring its application in the field of cyber-electronic warfare is of great significance for achieving covert, precise, and efficient strikes against the enemy. Summary of the Invention

[0004] This invention provides a genetic optimization intelligent decision-making method and device based on hierarchical fuzzy reasoning, which addresses the current lack of research on intelligent technology in the field of cyber warfare, and enables the preliminary exploratory application of intelligent decision-making in the field of cyber warfare.

[0005] This invention provides a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning, comprising:

[0006] Each state variable affecting decision output in cyber-electronic warfare, each output variable of cyber-electronic layer decision, and each output variable of physical layer decision are fuzzified within their respective universes of discourse to obtain multiple fuzzy sets. These multiple fuzzy sets are then digitally symbolically encoded to obtain the antecedent and consequent of fuzzy inference, forming a tactical rule database.

[0007] In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables of the current decision cycle and the dynamic weight of each input variable are determined; according to the tactical rule database, the action decisions of the network and electrical domains and the action decisions of the physical domains are fuzzily inferred hierarchically according to the network and electrical domains and the physical domains, and the input variables of the next decision cycle are updated.

[0008] The global adversarial reward value corresponding to the current battle situation is calculated based on the global adversarial reward function, and the tactical optimization of the genetic optimization algorithm is performed based on the global adversarial reward value to update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0009] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided. In the cyber-electronic warfare, the state variables affecting the decision output include: the relative distance between the agent and obstacles and targets, the line-of-sight angle between the agent and targets, and the state of target components; the output variables of the cyber-electronic layer decision in the cyber-electronic warfare include: radar power variables and radar frequency variables; the output variables of the physical layer decision in the cyber-electronic warfare include: the agent's desired position variables and the agent's yaw angle variables.

[0010] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein determining the input variables for the current decision cycle includes:

[0011] Acquire the state information of each agent, each target, and each obstacle in the group; wherein, each agent in the group includes at least one first agent and at least one second agent, and each target includes at least one first target and at least one second target;

[0012] The distance between the first intelligent agent and each obstacle is calculated based on the state information of the first intelligent agent and the state information of each obstacle.

[0013] If the distance between the first agent and at least one of the obstacles is less than a safe distance, calculate the gravitational force between the first agent and the first target and the repulsive force between the first agent and the obstacle;

[0014] The first intelligent agent is controlled to avoid obstacles based on the gravity and the repulsion, and the desired position of the first intelligent agent after obstacle avoidance is determined.

[0015] The signal-to-interference ratio (SIR) of the second target is calculated, and the SIR of the second target is used to determine whether the communication link between the first agent and the second target is interfered with.

[0016] The state information of the first agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target, and the expected position of the first agent after obstacle avoidance are determined as the input variables for the current decision cycle.

[0017] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein the action decision in the network electrical domain and the action decision in the physical domain are derived hierarchically according to the tactical rule database and updated to the input variables for the next decision cycle, including:

[0018] Based on the input variables of the current decision cycle and the dynamic weight of each input variable, the action decision of the network domain is fuzzily inferred according to the tactical rule database.

[0019] Control the first intelligent agent to execute action decisions in the network domain, and evaluate the execution effect;

[0020] Based on the input variables of the current decision cycle and the dynamic weight of each input variable, the action decision of the physical domain is fuzzily inferred according to the execution effect and the tactical rule database.

[0021] The action decision of the physical domain and the expected position after the first agent avoids obstacles are fused to obtain the expected position of the next decision cycle.

[0022] The input variables for the next decision cycle are updated based on the expected position for the next decision cycle.

[0023] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein calculating the information-to-interference ratio of the second objective includes:

[0024] The power loss of the second target is calculated according to the following expression:

[0025] proloss j =32.44+20*ω f1 *log 10 d j +20*ω f2 *log 10 f j ;

[0026] The power loss of the first target is calculated according to the following expression:

[0027] proloss i =32.44+20*ω f1 *log 10 d i +20*ω f2*log 10 f i ;

[0028] The signal-to-interference ratio of the second target is calculated according to the following expression:

[0029] itf = proloss j -proloss i

[0030] Where, ω f1 and ω f2 All represent parameters related to the effectiveness of wireless attacks, d j d represents the straight-line distance between the first agent and the second target. i f represents the straight-line distance between the first agent and the first target. j f represents the operating frequency of the second target. i The proloss represents the operating frequency of the first target. j Proloss represents the power loss of the second target. i denoted by , itf represents the power loss of the first target, and itf represents the signal-to-interference ratio of the second target.

[0031] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein obtaining the state information of each agent, each target, and each obstacle in the population includes:

[0032] Obtain state information of each agent in the swarm, including position, velocity, and heading;

[0033] Acquire status information for each target, including position, velocity, and heading;

[0034] Obtain status information, including the position, of each obstacle.

[0035] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein the global adversarial reward value corresponding to the current battle situation is calculated based on the global adversarial reward function, and tactical optimization of the genetic optimization algorithm is performed based on the global adversarial reward value to update the tactical rule database, including:

[0036] A hierarchical evaluation approach is adopted, incorporating adaptive dynamic parameters and a decay factor to construct a global adversarial reward function;

[0037] The global adversarial reward value corresponding to the battle situation in the current battle round is calculated based on the global adversarial reward function.

[0038] Repeat the following steps until the size of the next generation population reaches the size of the parent population: copy individuals from the parent population whose scores are higher than a preset score to the next generation population; select the remaining individuals from the next generation population through a binary tournament; randomly select two individuals from the parent population; and retain the individual with the higher global adversarial reward value from the two individuals to the next generation population. The individuals are each set of tactical rules in the tactical rule database.

[0039] The next generation population is updated based on a crossover function with adaptive crossover probability;

[0040] The updated next-generation population is subjected to individual mutations based on simulated annealing to increase individual climbing ability;

[0041] The tactical rules database is updated based on the next generation population after individual mutations.

[0042] According to the present invention, a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning is provided, wherein the crossover function based on adaptive crossover probability updates the next generation population, comprising:

[0043] The crossover function with adaptive crossover probability is constructed using the following expression:

[0044]

[0045] Where pc1 = 0.9, pc2 = 0.6, pc3 = 0.3, f avg f represents the average fitness value of the next generation population. max f represents the maximum fitness value of the next generation population. avg f' represents the minimum fitness value of the next generation population, f′ represents the fitness value of the individual with the larger fitness value among the two individuals to be crossovered, and pc represents the crossover function of the adaptive crossover probability;

[0046] The next generation population is updated based on the crossover function with the adaptive crossover probability.

[0047] The present invention also provides a genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning, comprising:

[0048] The forming module is used to fuzzify each state variable that affects the decision output in the network and electronic warfare, each output variable of the network and electronic layer decision and each output variable of the physical layer decision in their respective universes of discourse to obtain multiple fuzzy sets, and to encode the multiple fuzzy sets with digital symbols to obtain the antecedent and consequent of fuzzy inference, thus forming a tactical rule database.

[0049] The reasoning module is used to repeatedly execute the following steps in each battle round until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, determine the input variables of the current decision cycle and the dynamic weight of each input variable; according to the tactical rule database, perform hierarchical fuzzy reasoning of the action decision in the network and electrical domains and the action decision in the physical domains, and update the input variables for the next decision cycle.

[0050] The update module is used to calculate the global adversarial reward value corresponding to the current battle situation based on the global adversarial reward function, and to perform tactical optimization using a genetic optimization algorithm based on the global adversarial reward value, and update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0051] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning as described above.

[0052] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning as described above.

[0053] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning as described above.

[0054] The present invention provides a genetic optimization intelligent decision-making method and device based on hierarchical fuzzy reasoning. First, a tactical rule database is constructed by fuzzifying the state variables affecting decision output in cyber-electronic warfare, the output variables of cyber-electronic layer decisions, and the output variables of physical layer decisions within their respective domains, resulting in multiple fuzzy sets. These fuzzy sets are then digitally coded to obtain the antecedents and consequents of fuzzy reasoning, forming the tactical rule database. Next, a battle is conducted between the two sides. In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables and dynamic weights of each input variable for the current decision cycle are determined. Based on the tactical rule database, according to… The system uses hierarchical fuzzy reasoning in the cyber-electronic and physical domains to derive action decisions for both domains, and updates the input variables for the next decision cycle. By introducing dynamic weights for each input variable, the system decouples the input variables for the current decision cycle, effectively reducing the number of tactical rules and the computational burden of reasoning. Finally, based on the global adversarial reward function, the system calculates the global adversarial reward value corresponding to the current battle situation. Tactical optimization using a genetic optimization algorithm is then performed based on this global adversarial reward value, updating the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations according to a hierarchical evaluation method. The genetic optimization algorithm can be used for tactical optimization, enabling a preliminary exploratory application of intelligent decision-making in the cyber-electronic adversarial field. Attached Figure Description

[0055] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0056] Figure 1 This is a flowchart illustrating the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning provided by the present invention.

[0057] Figure 2 This is a schematic diagram of network countermeasures provided by the present invention;

[0058] Figure 3 This is a schematic diagram of the structure of the genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning provided by the present invention;

[0059] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0061] The following is combined Figures 1-2 This invention describes a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning.

[0062] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning provided by the present invention. Figure 1 As shown, the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning provided by this invention may include the following steps:

[0063] Step 101: Fuzzify each state variable that affects the decision output in the cyber-electronic confrontation, each output variable of the cyber-electronic layer decision, and each output variable of the physical layer decision within their respective universes of discourse to obtain multiple fuzzy sets. Then, encode the multiple fuzzy sets with digital symbols to obtain the antecedent and consequent of fuzzy inference, forming a tactical rule database.

[0064] Step 102: In each battle round, determine whether the conditions for one side to win are met or the maximum step length of a battle round is reached; if yes, proceed to step 104; if no, proceed to step 103.

[0065] Step 103: For each decision cycle, determine the input variables and dynamic weights of each input variable for the current decision cycle. Based on the tactical rules database, perform hierarchical fuzzy reasoning in the network and physical domains to determine the action decisions for the network and physical domains, and update the input variables for the next decision cycle. Return to step 102.

[0066] Step 104: Calculate the global adversarial reward value corresponding to the current battle situation based on the global adversarial reward function, and perform tactical optimization of the genetic optimization algorithm based on the global adversarial reward value to update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0067] In step 101, optionally, the state variables affecting the decision output in cyber-electronic warfare include: the relative distance between the agent and obstacles and targets, the line-of-sight angle between the agent and targets, and the state of target components; the output variables of cyber-electronic layer decision in cyber-electronic warfare include: radar power variables and radar frequency variables; the output variables of physical layer decision in cyber-electronic warfare include: the agent's desired position variables and the agent's yaw angle variables.

[0068] Specifically, the intelligent agent can be a drone, an unmanned vessel, etc., the obstacle can be an island or reef, and the target can be the location that the intelligent agent ultimately wants to reach, such as a designated airport or port. Among them, drones and unmanned vessels are both regarded as ideal point masses.

[0069] The state variable space of the unmanned vessel is:

[0070] S UUV =[x,y,θ,V] (1)

[0071] Where x and y represent the position of the unmanned surface vessel (USV), V represents the nose-direction velocity of the USV, and θ represents the angle between the nose-direction velocity of the USV and due north.

[0072] The kinematic equations of the unmanned vessel are:

[0073]

[0074] Among them, V x V y Let x and y represent the velocities of the unmanned surface vessel in the x and y directions, respectively, and let x0 and y0 represent the initial positions of the unmanned surface vessel. Δt represents the x-axis velocity and y-axis velocity of the unmanned vessel, respectively, and Δt represents the step size.

[0075] The state variable space of the drone is:

[0076]

[0077] Where x and y represent the position of the UAV, V represents the nose-direction velocity of the UAV, and γ represents the flight path tilt angle of the UAV. This indicates the azimuth angle of the drone's flight path.

[0078] The kinematic equations of the unmanned aerial vehicle are:

[0079]

[0080] Where x and y represent the positions of the drone. This represents the drone's velocity along the x-axis and y-axis.

[0081] The drone tracker model is as follows:

[0082]

[0083] Among them, S i S represents the actual state variable. c Represents the desired state variable, u i Let e ​​represent the controlled variable. i k represents the error. i Δt represents the tracking coefficient, and Δt represents the tracking step size.

[0084] Specifically, step 101 may include the following sub-steps:

[0085] Step 201: Fuzzify the state variables ΔI = {D, LOS, S} that affect the decision output in the electronic warfare scenario within their universe of discourse, obtaining D = [N, Z, P], LOS = [N, Z, P], and S = [N, P]. Here, D represents the relative distance between the agent and obstacles and the target, LOS represents the line-of-sight angle between the agent and the target, and S represents the state of the target component.

[0086] Step 202: Calculate the output variables of the network layer decision-making process in network-to-network confrontation: Δu w The expression {Δp, Δf} is fuzzified within its universe of discourse to obtain Δp = [NB, NS, Z, PS, PB] and Δf = [NB, NS, Z, PS, PB]. Here, Δp represents the radar power variable, and Δf represents the radar frequency variable.

[0087] Step 203: Calculate the output variables of the physical layer decision in cyber warfare: Δu d The expression {Δx, Δy, Δθ} is fuzzified within its universe of discourse to obtain Δx = [N, Z, P], Δy = [N, Z, P], and Δθ = [N, Z, P]. Here, Δx and Δy represent the agent's desired position variables, and Δθ represents the agent's yaw angle variable.

[0088] Where NB represents negative large, NS represents negative small, ZO represents zero, PS represents positive small, and PB represents positive large; N represents negative, Z represents zero, and P represents positive.

[0089] Step 204: Encode the multiple fuzzy sets obtained in steps 1011-1013 using digital symbols according to the following rules: NB corresponds to -2, NS corresponds to -1, ZO corresponds to 0, PS corresponds to 1, PB corresponds to 2, N corresponds to -1, Z corresponds to 0, and P corresponds to 1. Generate a total of n tactical sets, forming a unique mapping from the fuzzified state of the input variables to the output fuzzy variables, thus creating a rule base for fuzzy inference, i.e., a tactical rule database.

[0090] Taking a single variable as an example, the tactical rule database can be described as follows:

[0091]

[0092] Each set of tactical rules forms an individual in the genetic optimization algorithm.

[0093] In step 102, a combat round refers to a round in which the two intelligent agents attack each other in the cyber and physical domains. A winning condition for one side is when one intelligent agent successfully interferes with the other's target.

[0094] In each battle round, step 103 is repeatedly executed, and the status of the target radar and electromagnetic components in each decision cycle is recorded until the conditions for one side to win are met or the maximum step length of a battle round is reached.

[0095] In step 103, the decision cycle refers to the cycle of completing an action decision in the electrical domain and an action decision in the physical domain.

[0096] The network's action decision includes the power and frequency at which the first agent interferes with the second target. In this embodiment, the first agent is the red team's agent, the second agent is the blue team's agent, the first target is the red team's target, i.e., the location the first agent ultimately wants to reach, and the second target is the blue team's target, i.e., the blue team's threat platform unit.

[0097] Action decisions in the physical domain include the desired position and yaw angle of the first agent.

[0098] The input variables for the current decision cycle include: the state information of the first agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target, and the expected position of the first agent after obstacle avoidance.

[0099] Consequence variable Δu i The result of reasoning It can be obtained through any defuzzification method. The i-th input variable x i The consequent variable Δu of a single input rule group i The result of reasoning for:

[0100]

[0101] in, In the j-th fuzzy rule, x represents... i membership function, Denotes the j-th fuzzy rule △u i membership function, m i This represents a fuzzy inference engine.

[0102] In practical reasoning systems, different input variables have varying effects on the same output variable; some input variables have a greater impact on the output, while others have a smaller impact. Therefore, to differentiate between input variables, we need to define a method for handling each input variable x. i Design a dynamic weight for each of (1,2,...,n).

[0103]

[0104] in, B represents the base value. i This represents the dynamic amplitude, which can be set based on experience. Represents a dynamic variable.

[0105] In this step, for each decision cycle, based on the input variables of the current decision cycle and the dynamic weights of each input variable, and according to the tactical rules database, the action decision in the network domain is first fuzzily inferred in two levels: the network domain and the physical domain. Then, the action decision in the physical domain is fuzzily inferred, and the input variables for the next decision cycle are updated.

[0106] Optionally, the input variables for the current decision-making cycle are determined, including the following sub-steps:

[0107] Step 301: Obtain the state information of each agent, each target, and each obstacle in the group; wherein, each agent in the group includes: at least one first agent and at least one second agent, and each target includes: at least one first target and at least one second target;

[0108] Step 302: Calculate the distance between the first intelligent agent and each obstacle based on the state information of the first intelligent agent and the state information of each obstacle;

[0109] Step 303: If the distance between the first agent and at least one obstacle is less than the safe distance, calculate the gravitational force between the first agent and the first target and the repulsive force between the first agent and the obstacle;

[0110] Step 304: Control the first intelligent agent to avoid obstacles based on gravity and repulsion, and determine the desired position of the first intelligent agent after obstacle avoidance;

[0111] Step 305: Calculate the signal-to-interference ratio (SIR) of the second target. The SIR of the second target is used to determine whether the communication link between the first agent and the second target is interfered with.

[0112] Step 306: Determine the state information of the first intelligent agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target, and the expected position of the first intelligent agent after obstacle avoidance as the input variables for the current decision cycle.

[0113] In step 301, optionally, state information including position, velocity and heading of each agent in the group is obtained, state information including position, velocity and heading of each target is obtained, and state information including position of each obstacle is obtained, thereby completing the acquisition of state information of each agent, each target and each obstacle in the group.

[0114] In step 302, the relative distance between the first agent and the obstacle is calculated based on the position of the first agent and the position of each obstacle.

[0115] In steps 303 and 304, if there is only one obstacle such that the distance between the first agent and the obstacle is less than the safe distance, the attraction between the first agent and the first target and the repulsion between the first agent and the obstacle are calculated based on the artificial potential field method. The first agent is controlled to avoid the obstacle based on the attraction and repulsion, and the desired position of the first agent after avoiding the obstacle is determined.

[0116] If there are multiple obstacles such that the distance between the first agent and each of the obstacles is less than the safe distance, the attraction between the first agent and the first target and the repulsion between the first agent and each obstacle are calculated based on the artificial potential field method. The first agent is controlled to avoid obstacles based on the attraction and repulsion, and the desired position of the first agent after obstacle avoidance is determined.

[0117] In step 305, the signal-to-interference ratio (SIR) of the second target is used to determine whether the communication link between the first agent and the second target is interfered with. If the SIR of the second target is less than or equal to 10, then the communication link between the first agent and the second target is interfered with.

[0118] Optionally, step 305 includes the following sub-steps:

[0119] Calculate the power loss of the second objective (i.e., objective j):

[0120] proloss j =32.44+20*ω f1 *log 10 d j +20*ω f2 *log 10 f j (9)

[0121] Calculate the power loss of the first target (i.e., target i):

[0122] proloss i =32.44+20*ω f1 *log 10 d i +20*ωf2 *log 10 f i (10)

[0123] Calculate the signal-to-interference ratio (SIR) of the second target:

[0124] itf = proloss j -proloss i (11)

[0125] Where, ω f1 and ω f2 All represent parameters related to the effectiveness of wireless attacks, d j d represents the straight-line distance between the first agent and the second target. i f represents the straight-line distance between the first agent and the first target. j f represents the operating frequency of the second target. i The proloss represents the operating frequency of the first objective. j Proloss represents the power loss of the second objective. i The first target's power loss is represented by , and the second target's signal-to-interference ratio is represented by .

[0126] In step 306, the obtained state information of the first agent, the state information of the first target, the state information of the second target, the calculated signal-to-interference ratio of the second target, and the expected position of the first agent after obstacle avoidance are determined as the input variables of the current decision cycle.

[0127] In this embodiment, the state information of each agent, target, and obstacle in the group is obtained. When the distance between the first agent and the obstacle is less than the safe distance, the first agent is controlled to avoid the obstacle. The signal-to-interference ratio of the second target and the expected position of the first agent after avoiding the obstacle are calculated to obtain the input variables of the current decision cycle.

[0128] In step 103, optionally, based on the tactical rules database, hierarchical fuzzy reasoning is performed to derive action decisions in the network and physical domains, and the input variables for the next decision cycle are updated, including the following sub-steps:

[0129] Step 307: Based on the input variables of the current decision cycle and the dynamic weight of each input variable, fuzzy inference is made in the network domain action decision according to the tactical rule database;

[0130] Step 308: Control the first intelligent agent to execute action decisions in the network domain and evaluate the execution effect;

[0131] Step 309: Based on the input variables of the current decision cycle and the dynamic weight of each input variable, fuzzy inference is made about the action decision of the physical domain according to the execution effect and the tactical rule database.

[0132] Step 310: Fuse the action decision in the physical domain with the expected position after the first agent avoids obstacles to obtain the expected position for the next decision cycle;

[0133] Step 311: Update the input variables for the next decision cycle based on the expected position for the next decision cycle.

[0134] In step 307, the tactical rules database can be described as follows:

[0135]

[0136] Where SIRMi represents the i-th fuzzy inference system, x i This represents the i-th input to the fuzzy inference system. This represents the j-th tactical rule in the i-th fuzzy inference system. Indicates input x i The obfuscated encoding, i.e., the antecedent of the tactical rule, Indicates the output variable △u i The obfuscated encoding is the consequent of the tactical rule.

[0137] Action decisions in the network domain are derived from fuzzy reasoning in the network domain:

[0138]

[0139]

[0140]

[0141] in, express Dynamic weights, express Dynamic weights, express Dynamic weights, This represents the change in the relative distance between the agent and obstacles and target i. Δp represents the change in the line-of-sight angle between the agent and target i, Δf represents the change in radar power variable, and Δf represents the change in radar frequency variable.

[0142] In step 308, the first intelligent agent is controlled to interfere with the second target using the power and frequency in the action decision of the network electrical domain, and the success of the interference is evaluated.

[0143] In step 309, the action decision of the physical domain is derived according to the fuzzy reasoning of the electrical domain:

[0144]

[0145]

[0146]

[0147] Where Δx and Δy represent the changes in the desired position of the first agent, and Δθ represents the changes in the yaw angle of the first agent. m+n-1 represents m+n-1 sets of fuzzy inference systems.

[0148] The line-of-sight angle between the i-th agent of the red team and the target i of the red team is calculated using the following expression:

[0149]

[0150]

[0151]

[0152]

[0153] Where x and y are the OX of the drone. g OY g Position in the axial direction, V, Let represent the flight speed and azimuth angle of the UAV, respectively; D represents the relative distance between the i-th agent of the red team and the target of the blue team; r represents the i-th agent of the red team; and b represents the j-th agent of the blue team.

[0154] In step 310, Δx and Δy are both fused with the desired position of the first agent after obstacle avoidance to obtain the final ΔX and ΔY:

[0155]

[0156]

[0157] Where Δxp represents the change in the expected position of the first agent in the x-direction after obstacle avoidance, Δx represents the change in the expected position of the first agent in the x-direction during the current decision period, and ΔX represents the change in the expected position of the first agent in the x-direction during the next decision period. Δyp represents the change in the expected position of the first agent in the y-direction after obstacle avoidance, Δy represents the change in the expected position of the first agent in the y-direction during the current decision period, and ΔY represents the change in the expected position of the first agent in the y-direction during the next decision period.

[0158] Optionally, step 104 includes the following sub-steps:

[0159] Step 401: Using a hierarchical evaluation approach, an adaptive dynamic parameter and a decay factor are introduced to construct a global adversarial reward function;

[0160] Step 402: Calculate the global adversarial reward value corresponding to the battle situation in the current battle round based on the global adversarial reward function;

[0161] Step 403: Repeat the following steps until the size of the next generation population reaches the size of the parent population. The number of repetitions is equal to the size of the parent population: Copy individuals from the parent population whose scores are higher than the preset score to the next generation population. Select the remaining individuals from the next generation population through a binary tournament. Randomly select two individuals from the parent population and keep the individual with the higher global adversarial reward value from the two individuals to the next generation population. The individuals are each set of tactical rules in the tactical rule database.

[0162] Step 404: Update the next generation population based on the crossover function with adaptive crossover probability;

[0163] Step 405: Perform individual mutations on the updated next generation population based on simulated annealing to increase individual climbing ability;

[0164] Step 406: Update the tactical rules database based on the next generation population after individual mutations.

[0165] In step 401, a global adversarial reward function is constructed:

[0166]

[0167] in, V(k) represents the state transition probability, V(k) represents the heuristic function, s_num represents the total action reward value, n represents the action step size, gen represents the population size, ada_w1 and ada_w2 represent the weight coefficients, total_step represents the step size limit for one combat round, and real_step represents the actual step size.

[0168] In step 402, it can be seen from the global adversarial reward function that different calculation expressions are used to calculate the global adversarial reward value for different battle situations, such as success, failure, and draw.

[0169] Optionally, step 404 includes the following sub-steps:

[0170] Construct a crossover function with adaptive crossover probabilities:

[0171]

[0172] Where pc1 = 0.9, pc2 = 0.6, pc3 = 0.3, f avg f represents the average fitness value of the next generation population. max f represents the maximum fitness value of the next generation population.avg f' represents the minimum fitness value of the next generation population, f′ represents the fitness value of the individual with the larger fitness value among the two individuals to be crossovered, and pc represents the crossover function with adaptive crossover probability;

[0173] The next generation of the population is updated using a crossover function based on adaptive crossover probability.

[0174] This embodiment uses an adaptive crossover probability method to update the next generation population, which can improve the "premature maturity" phenomenon of the next generation population and improve the global optimization performance of GA.

[0175] In step 405, based on the idea of ​​simulated annealing, individual mutation is performed to increase the individual's hill-climbing ability. This step calculates the difference between the new fitness after mutation and the fitness before mutation. If the fitness after mutation is greater, the mutation is accepted; if the fitness after mutation is smaller, a certain annealing probability exp(δ / T) is used to determine whether mutation should occur. Here, δ is the difference in fitness before and after mutation, and T is the annealing temperature, which decreases linearly with the number of generations. This operation improves the efficiency of the genetic algorithm and can also avoid premature convergence to some extent.

[0176] In step 406, each individual in the next generation of the mutated population is updated with each set of tactical rules in the tactical rule database.

[0177] In summary, the genetic optimization intelligent decision-making method and apparatus based on hierarchical fuzzy reasoning provided by this invention first constructs a tactical rule database. This involves fuzzifying each state variable affecting decision output in cyber-electronic warfare, each output variable of cyber-electronic layer decision, and each output variable of physical layer decision within their respective domains to obtain multiple fuzzy sets. These fuzzy sets are then digitally coded to obtain the antecedents and consequents of fuzzy reasoning, forming the tactical rule database. Next, a battle is conducted between the two sides. In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables of the current decision cycle and the dynamic weights of each input variable are determined. Based on the tactical rule database, ... The system employs hierarchical fuzzy reasoning to derive action decisions in the network and physical domains, updating the input variables for the next decision cycle. By introducing dynamic weights for each input variable, the system decouples the input variables of the current decision cycle, effectively reducing the number of tactical rules and the computational burden. Finally, based on the global adversarial reward function, the system calculates the global adversarial reward value corresponding to the current battle situation. Tactical optimization using a genetic optimization algorithm is then performed based on this global adversarial reward value, updating the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations using a hierarchical evaluation method. The genetic optimization algorithm can be used for tactical optimization, enabling a preliminary exploratory application of intelligent decision-making in the field of network and electronic warfare.

[0178] The following describes the genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning provided by the present invention. The genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning described below can be referred to in correspondence with the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning described above.

[0179] Please refer to Figure 3 , Figure 3 This is a schematic diagram of the structure of the genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning provided by the present invention. Figure 3 As shown, the genetic optimization intelligent decision-making device based on hierarchical fuzzy reasoning provided by the present invention includes:

[0180] The forming module 10 is used to fuzzify each state variable that affects the decision output in the network and electronic warfare, each output variable of the network and electronic layer decision and each output variable of the physical layer decision in their respective universes of discourse to obtain multiple fuzzy sets, and to encode the multiple fuzzy sets with digital symbols to obtain the antecedent and consequent of fuzzy inference, thus forming a tactical rule database.

[0181] The reasoning module 20 is used to repeatedly execute the following steps in each battle round until the conditions for one side to win are met or the maximum step length of a battle round is reached: for each decision cycle, determine the input variables of the current decision cycle and the dynamic weight of each input variable; according to the tactical rule database, perform hierarchical fuzzy reasoning of the action decision in the network and electrical domains and the action decision in the physical domains, and update the input variables of the next decision cycle.

[0182] The update module 30 is used to calculate the global confrontation reward value corresponding to the current battle situation based on the global confrontation reward function, and to perform tactical optimization of the genetic optimization algorithm based on the global confrontation reward value, and update the tactical rule database. The global confrontation reward function is used to calculate the global confrontation reward value of different battle situations in a hierarchical evaluation manner.

[0183] Optionally, the state variables affecting the decision output in the cyber-electronic warfare include: the relative distance between the agent and obstacles and targets, the line-of-sight angle between the agent and targets, and the state of target components; the output variables of the cyber-electronic layer decision in the cyber-electronic warfare include: radar power variables and radar frequency variables; the output variables of the physical layer decision in the cyber-electronic warfare include: the agent's desired position variables and the agent's yaw angle variables.

[0184] Optionally, the inference module 20 includes:

[0185] An acquisition unit is used to acquire the state information of each agent, each target, and each obstacle in the group; wherein, each agent in the group includes at least one first agent and at least one second agent, and each target includes at least one first target and at least one second target;

[0186] The first calculation unit is used to calculate the distance between the first intelligent agent and each obstacle based on the state information of the first intelligent agent and the state information of each obstacle;

[0187] The second computing unit is used to calculate the gravitational force between the first intelligent agent and the first target and the repulsive force between the first intelligent agent and the obstacle when the distance between the first intelligent agent and at least one of the obstacles is less than a safe distance.

[0188] An obstacle avoidance unit is used to control the first intelligent agent to avoid obstacles based on the attraction and the repulsion, and to determine the desired position of the first intelligent agent after obstacle avoidance.

[0189] The third calculation unit is used to calculate the signal-to-interference ratio (SIR) of the second target, which is used to determine whether the communication link between the first agent and the second target is interfered with.

[0190] The determining unit is used to determine the state information of the first intelligent agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target, and the expected position of the first intelligent agent after obstacle avoidance as input variables for the current decision cycle.

[0191] Optionally, the inference module 20 further includes:

[0192] The first reasoning unit is used to fuzzily infer the action decision of the network domain based on the input variables of the current decision cycle and the dynamic weight of each input variable according to the tactical rule database.

[0193] An execution unit is used to control the first intelligent agent to execute action decisions in the network domain and to evaluate the execution effect;

[0194] The second reasoning module is used to fuzzily infer the action decision of the physical domain based on the input variables of the current decision cycle and the dynamic weight of each input variable, according to the execution effect and the tactical rule database.

[0195] The fusion unit is used to fuse the action decision of the physical domain with the expected position of the first agent after obstacle avoidance to obtain the expected position of the next decision cycle.

[0196] The update unit is used to update the input variables of the next decision cycle based on the expected position of the next decision cycle.

[0197] Optionally, the third computing unit is specifically used for:

[0198] The power loss of the second target is calculated according to the following expression:

[0199] proloss j =32.44+20*ω f1 *log 10 d j +20*ω f2 *log 10 f j ;

[0200] The power loss of the first target is calculated according to the following expression:

[0201] proloss i =32.44+20*ω f1 *log 10 d i +20*ω f2 *log 10 f i ;

[0202] The signal-to-interference ratio of the second target is calculated according to the following expression:

[0203] itf = proloss j -proloss i

[0204] Where, ω f1 and ω f2 All represent parameters related to the effectiveness of wireless attacks, d j d represents the straight-line distance between the first agent and the second target. i f represents the straight-line distance between the first agent and the first target. j f represents the operating frequency of the second target. i The proloss represents the operating frequency of the first target. j Proloss represents the power loss of the second target. i The first target's power loss is represented by , and ift represents the signal-to-interference ratio of the second target.

[0205] Optionally, the acquisition unit is specifically used for:

[0206] Obtain state information of each agent in the swarm, including position, velocity, and heading;

[0207] Acquire status information for each target, including position, velocity, and heading;

[0208] Obtain status information, including the position, of each obstacle.

[0209] Optionally, the update module 30 includes:

[0210] The building unit is used to construct a global adversarial reward function by introducing adaptive dynamic parameters and a decay factor in a hierarchical evaluation approach.

[0211] The fourth calculation unit is used to calculate the global confrontation reward value corresponding to the battle situation in the current battle round based on the global confrontation reward function;

[0212] The elite selection unit is used to repeatedly perform the following steps until the size of the next generation population reaches the size of the parent population: copy individuals with scores higher than a preset score from the parent population to the next generation population; select the remaining individuals from the next generation population through a binary tournament; randomly select two individuals from the parent population; and retain the individual with the higher global adversarial reward value from the two individuals to the next generation population. The individuals are each set of tactical rules in the tactical rule database.

[0213] A population update unit is used to update the next generation population based on a crossover function with adaptive crossover probability.

[0214] Individual mutation unit, used to perform individual mutation on the updated next generation population based on simulated annealing method to increase individual climbing ability;

[0215] The rule base update unit is used to update the tactical rule database based on the next generation population after individual mutations.

[0216] Optionally, the population update unit is specifically used for:

[0217] The crossover function with adaptive crossover probability is constructed using the following expression:

[0218]

[0219] Where pc1 = 0.9, pc2 = 0.6, pc3 = 0.3, f avg f represents the average fitness value of the next generation population. max f represents the maximum fitness value of the next generation population. avg f' represents the minimum fitness value of the next generation population, f′ represents the fitness value of the individual with the larger fitness value among the two individuals to be crossovered, and pc represents the crossover function of the adaptive crossover probability;

[0220] The next generation population is updated based on the crossover function with the adaptive crossover probability.

[0221] Figure 4 An example is a schematic diagram of the structure of an electronic device, such as... Figure 4 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communication interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute a genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning, which includes:

[0222] Each state variable affecting decision output in cyber-electronic warfare, each output variable of cyber-electronic layer decision, and each output variable of physical layer decision are fuzzified within their respective universes of discourse to obtain multiple fuzzy sets. These multiple fuzzy sets are then digitally symbolically encoded to obtain the antecedent and consequent of fuzzy inference, forming a tactical rule database.

[0223] In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables of the current decision cycle and the dynamic weight of each input variable are determined; according to the tactical rule database, the action decisions of the network and electrical domains and the action decisions of the physical domains are fuzzily inferred hierarchically according to the network and electrical domains and the physical domains, and the input variables of the next decision cycle are updated.

[0224] The global adversarial reward value corresponding to the current battle situation is calculated based on the global adversarial reward function, and the tactical optimization of the genetic optimization algorithm is performed based on the global adversarial reward value to update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0225] Furthermore, the logical instructions in the aforementioned memory 830 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0226] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to execute the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning provided by the above methods, the method comprising:

[0227] Each state variable affecting decision output in cyber-electronic warfare, each output variable of cyber-electronic layer decision, and each output variable of physical layer decision are fuzzified within their respective universes of discourse to obtain multiple fuzzy sets. These multiple fuzzy sets are then digitally symbolically encoded to obtain the antecedent and consequent of fuzzy inference, forming a tactical rule database.

[0228] In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables of the current decision cycle and the dynamic weight of each input variable are determined; according to the tactical rule database, the action decisions of the network and electrical domains and the action decisions of the physical domains are fuzzily inferred hierarchically according to the network and electrical domains and the physical domains, and the input variables of the next decision cycle are updated.

[0229] The global adversarial reward value corresponding to the current battle situation is calculated based on the global adversarial reward function, and the tactical optimization of the genetic optimization algorithm is performed based on the global adversarial reward value to update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0230] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning provided by the methods described above, the method comprising:

[0231] Each state variable affecting decision output in cyber-electronic warfare, each output variable of cyber-electronic layer decision, and each output variable of physical layer decision are fuzzified within their respective universes of discourse to obtain multiple fuzzy sets. These multiple fuzzy sets are then digitally symbolically encoded to obtain the antecedent and consequent of fuzzy inference, forming a tactical rule database.

[0232] In each battle round, the following steps are repeated until the conditions for one side to win are met or the maximum step length of a battle round is reached: For each decision cycle, the input variables of the current decision cycle and the dynamic weight of each input variable are determined; according to the tactical rule database, the action decisions of the network and electrical domains and the action decisions of the physical domains are fuzzily inferred hierarchically according to the network and electrical domains and the physical domains, and the input variables of the next decision cycle are updated.

[0233] The global adversarial reward value corresponding to the current battle situation is calculated based on the global adversarial reward function, and the tactical optimization of the genetic optimization algorithm is performed based on the global adversarial reward value to update the tactical rule database. The global adversarial reward function is used to calculate the global adversarial reward value for different battle situations in a hierarchical evaluation manner.

[0234] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0235] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A genetic optimization intelligent decision making method based on hierarchical fuzzy reasoning, characterized in that, The method comprises the following steps: In each battle round, the following steps are repeatedly executed until the winning condition of one party is met or the maximum step of a battle round is reached: for each decision period, determining the input variables of the current decision period and the dynamic weight of each input variable, and according to the tactical rule database, hierarchically fuzzy reasoning the action decision of the cyber domain and the action decision of the physical domain, and updating the input variables of the next decision period; The method comprises the following steps: Obtaining state information of each agent, each target and each obstacle in the group; wherein the agents in the group comprise at least one first agent and at least one second agent, and the targets comprise at least one first target and at least one second target; Calculating the distance between the first agent and each obstacle based on the state information of the first agent and the state information of each obstacle, respectively; In the case that the distance between the first agent and at least one obstacle is less than a safe distance, calculating the attractive force between the first agent and the first target and the repulsive force between the first agent and the obstacle; Controlling the first agent to avoid obstacles based on the attractive force and the repulsive force, and determining the expected position of the first agent after avoiding obstacles; Calculating the signal-to-interference ratio of the second target, which is used to determine whether the communication link between the first agent and the second target is interfered; Determining the state information of the first agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target and the expected position of the first agent after avoiding obstacles as the input variables of the current decision period; Based on the global counter-reward function, the global counter-reward value corresponding to the battle situation of the current battle round is calculated, and the tactical optimization of the genetic optimization algorithm is performed based on the global counter-reward value to update the tactical rule database. The global counter-reward function is used to calculate the global counter-reward value of different battle situations in a hierarchical evaluation manner. The state variables affecting the decision output in the cyber confrontation include the relative distance between the agent and the obstacle and the target, the line-of-sight angle between the agent and the target, and the state of the target component; the output variables of the cyber layer decision in the cyber confrontation include the radar power variable and the radar frequency variable; and the output variables of the physical layer decision in the cyber confrontation include the expected position variable of the agent and the yaw angle variable of the agent.

2. The method of claim 1, wherein, The method comprises the following steps:

3. The method of claim 1, wherein the hierarchical fuzzy inference-based genetic optimization intelligent decision method is characterized by, ​ fuzzy reasoning of the action decision of the cyberspace domain according to the tactical rule database based on the input variables of the current decision cycle and the dynamic weight of each input variable; controlling the first agent to execute the action decision of the cyberspace domain and evaluating the execution effect; fuzzy reasoning of the action decision of the physical domain according to the execution effect and the tactical rule database based on the input variables of the current decision cycle and the dynamic weight of each input variable; fusing the action decision of the physical domain and the expected position of the first agent after obstacle avoidance to obtain the expected position of the next decision cycle; updating the input variables of the next decision cycle based on the expected position of the next decision cycle.

4. The method of claim 1, wherein the hierarchical fuzzy inference-based genetic optimization intelligent decision method is characterized by, the calculation of the signal-to-interference ratio of the second target comprises: the power loss of the second target is calculated according to the following expression: ; the power loss of the first target is calculated according to the following expression: ; the signal-to-interference ratio of the second target is calculated according to the following expression: ; wherein, and each represent a wireless attack effect parameter, represents a straight-line distance of the first agent from the second target, represents a straight-line distance of the first agent from the first target, represents an operating frequency of the second target, represents an operating frequency of the first target, represents a power loss of the second target, represents a power loss of the first target, represents a signal-to-interference ratio of the second target.

5. The method of claim 1, wherein, the acquisition of the state information of each agent, each target and each obstacle in the group comprises: the state information of each agent in the group including position, speed and heading is acquired; the state information of each target including position, speed and heading is acquired; the state information of each obstacle including position is acquired.

6. The method of claim 1, wherein, the calculation of the global confrontation return value corresponding to the situation of the current battle round based on the global confrontation return function, and the tactical optimization of the genetic optimization algorithm based on the global confrontation return value, and the update of the tactical rule database, comprise: a hierarchical evaluation method is adopted to introduce adaptive dynamic parameters and decay factors to construct a global confrontation return function; the global confrontation return value corresponding to the situation of the current battle round is calculated based on the global confrontation return function; the following steps are repeatedly executed until the size of the next generation population reaches the size of the parent population: copying individuals with a score higher than a preset score in the parent population to the next generation population, selecting the remaining individuals in the next generation population through a binary tournament selection, randomly selecting two individuals in the parent population, and retaining the individual with a higher global confrontation return value in the two individuals to the next generation population, the individual being each set of tactical rules in the tactical rule database; the next generation population is updated based on an adaptive crossover probability crossover function; individual mutation is performed on the updated next generation population based on a simulated annealing method to increase the mountain climbing ability of the individual; the tactical rule database is updated based on the next generation population after individual mutation.

7. The method of claim 6, wherein the hierarchical fuzzy inference-based genetic optimization intelligent decision method is characterized by, the update of the next generation population based on the adaptive crossover probability crossover function comprises: an adaptive crossover probability crossover function is constructed through the following expression: ; wherein, , , , denotes the average of the fitness values of the next generation population, denotes the maximum of the fitness values of the next generation population, denotes the minimum of the fitness values of the next generation population, denotes the fitness value of the individual with the greater fitness value of the two individuals to be crossed, denotes a crossing function of the self-adaptive crossing probability; the next generation population is updated based on the adaptive crossover probability crossover function.

8. A genetic optimization intelligent decision device based on hierarchical fuzzy reasoning, characterized by, comprise: a forming module is configured to fuzz each state variable affecting the decision output in cyberspace confrontation, each output variable of cyberspace layer decision and each output variable of physical layer decision in their respective domains to obtain a plurality of fuzzy sets, and to perform digital symbol coding on the plurality of fuzzy sets to obtain the antecedent and the consequent of fuzzy reasoning, thereby forming a tactical rule database; The reasoning module is configured to repeatedly perform the following steps in each battle round until a winning condition is met or a maximum step of the battle round is reached: determining input variables of a current decision period and dynamic weights of each of the input variables, performing fuzzy reasoning in a cyber domain and a physical domain according to the tactical rule database to obtain action decisions in the cyber domain and action decisions in the physical domain, and updating input variables of a next decision period for each decision period. The determination of the input variables of the current decision period comprises: obtaining state information of each agent, each target and each obstacle in the group; wherein the agents in the group comprise at least one first agent and at least one second agent, and the targets comprise at least one first target and at least one second target; calculating distances between the first agent and each of the obstacles based on state information of the first agent and state information of each of the obstacles; in the case that the distance between the first agent and at least one of the obstacles is less than a safety distance, calculating an attractive force between the first agent and the first target and a repulsive force between the first agent and the obstacle; controlling the first agent to avoid obstacles based on the attractive force and the repulsive force, and determining a desired position of the first agent after obstacle avoidance; calculating a signal-to-interference ratio of the second target, which is used to determine whether a communication link between the first agent and the second target is interfered; and determining the state information of the first agent, the state information of the first target, the state information of the second target, the signal-to-interference ratio of the second target and the desired position of the first agent after obstacle avoidance as the input variables of the current decision period; The updating module is configured to calculate a global confrontation return value corresponding to a battle situation of the current battle round based on a global confrontation return function, and perform tactical optimization of a genetic optimization algorithm based on the global confrontation return value to update the tactical rule database, wherein the global confrontation return function is configured to calculate global confrontation return values of different battle situations in a hierarchical evaluation manner.

9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning according to any one of claims 1 to 7 when executing the program. 10.A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program implements the genetic optimization intelligent decision-making method based on hierarchical fuzzy reasoning according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Patent Citations

  • Fuzzy inference tree lifting method and device based on genetic algorithm

    CN110689129A

  • Layered multi-agent reinforcement learning method for multi-element joint command and control

    CN114330651A

  • Air combat maneuver decision design method based on fuzzy reasoning

    CN114492805A