Battle network disintegration method and device based on deep reinforcement learning
By constructing combat networks and weapon information based on deep reinforcement learning methods and using deep reinforcement learning networks to solve the objective function, the problem of difficult balance between efficiency and effectiveness in the collapse of combat networks is solved, and a more efficient combat system destruction effect is achieved.
Patent Information
- Application Number
- CN202410331883.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-23
AI Technical Summary
Existing combat network disintegration methods cannot effectively balance efficiency and effectiveness under limited weapon resources, and it is difficult to maximize the decomposition effect of the combat system.
A method based on deep reinforcement learning is used to construct a combat network and obtain weapon information. The objective function is solved through the deep reinforcement learning network to determine the optimal node removal strategy. The encoder and decoder in the deep reinforcement learning network are combined with the attention mechanism to select the attack target and generate a node output sequence that minimizes the connectivity of the combat network.
The speed and quality of solving combat network disintegration have been improved, which can more effectively destroy the enemy's combat system under limited weapon resources.
Smart Images

Figure CN120688757A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of combat network technology, and specifically to a combat network disruption method and device based on deep reinforcement learning. Background Art
[0002] With the rapid development of advanced information technologies such as data links and satellite communications, the connections and interactions between systems have become increasingly complex and diverse. A combat system is a group of operational entities working together to achieve common missions and objectives. More and more combat entities are transcending the limitations of the operational domain, forming vast operational networks through complex interactions. This has shifted the combat model from confrontations based on high-tech platforms to confrontations based on combat systems. In traditional high-tech platform warfare, all combat operations are organized around these platforms. However, in combat system warfare, destroying the combat system is an effective way to win the war.
[0003] The core of combat system destruction is attacking key entities within it, aiming to disrupt the structure and achieve victory. In recent years, many researchers have used complex network theory to study combat systems, abstracting them into combat networks with distinct nodes and edges. Therefore, some research refers to the process of combat system destruction as combat network disintegration. Regarding network disintegration, the key strategy lies in removing a set of the most critical nodes in the network to achieve maximum disintegration effects, such as preventing virus transmission, disrupting terrorism, and disrupting rumor spread. Within combat systems, some work has also analyzed network structure to identify key nodes or links, such as homogeneous single-layer networks, homogeneous multi-layer networks, and heterogeneous networks. However, little work has considered how to maximize the effectiveness of combat system disintegration with limited weapon resources from an offensive perspective. The attacker's immediate goal is to utilize combat resources (such as tanks, artillery, and other weapons) to destroy the opponent's combat system. Therefore, incorporating weapons into the combat network disintegration problem in real battlefield scenarios is of practical significance.
[0004] The combat system destruction problem is regarded as a network disintegration problem, and discrete weapon allocation is modeled using nonlinear optimization programming. By observing the characteristics of limited weapon resources and the requirements of different node removal capabilities, it is found that existing solutions cannot achieve a good balance between effectiveness and efficiency. Summary of the Invention
[0005] In view of the above problems, an embodiment of the present invention provides a method and device for combat network disintegration based on deep reinforcement learning, which overcomes the above problems or at least partially solves the above problems.
[0006] According to one aspect of an embodiment of the present invention, a method for disrupting a combat network based on deep reinforcement learning is provided, the method comprising: constructing a combat network, obtaining weapon information of at least one weapon, and combining any node in the combat network with any weapon to generate an object; constructing an objective function for characterizing the degree of disruption of the combat network based on all the objects in the combination of the combat network and each of the weapons, the degree of disruption being the number of closed loops with different path lengths among all nodes in the combat network after disruption; inputting all the objects into a preset deep reinforcement learning network, and solving, through the deep reinforcement learning network, a node output sequence after the combat network is disrupted by each of the weapons that minimizes the objective function.
[0007] In an optional manner, the construction of a combat network and obtaining weapon information of at least one weapon include: defining the combat network as G = (V, E), where V represents a set of nodes of combat entities, and E is a set of edges between nodes, representing the interaction between combat entities; and setting any node v j Described as {x j ,y j ,z j ,k j}, where (x j ,y j ,z j ) represents node v j The position coordinates, k j Represents node v j The degree of; any weapon w i With tuple w i :={x i ,y i ,z i ,c i ,r i} description, where (x i ,y i ,z i ) indicates weapon w i The position coordinates, c i Indicates weapon w i Attack ability, r i Indicates weapon w i attack range.
[0008] In an optional manner, constructing an objective function for characterizing the degree of collapse of the combat network based on all the objects in the combat network and the weapon combinations includes: calculating the Euclidean distance between the node in any of the objects and the weapon; determining an indicator function value of whether the weapon can be used to attack the node based on the Euclidean distance, and obtaining a first binary variable indicating whether the weapon attacks the node; determining multiple constraints based on the first binary variable and the indicator function value; calculating a total damage value of each weapon to any node in the combat network based on the weapon information and the indicator function value; determining a removal threshold for any node in the combat network, and determining a second binary variable of whether the node is removed based on the first binary variable and the removal threshold; determining a removal strategy based on the second binary variable of each node in the combat network, and calculating the number of closed loops with different path lengths of all nodes in the collapsed combat network corresponding to the removal strategy, to obtain an objective function for characterizing the degree of collapse of the combat network.
[0009] In an optional manner, the multiple constraints are determined based on the first binary variable and the indicator function value, including: based on the fact that any weapon can only attack once and can only attack at most one node, determining that the sum of the first binary variables of all the weapons for any of the nodes is less than or equal to 1; based on the fact that any weapon can only be used to attack nodes within the attack range, determining that the first binary variable is less than or equal to the indicator function value.
[0010] Optionally, before solving the node output sequence of the combat network after being disrupted by each weapon so as to minimize the objective function through the deep reinforcement learning network, the method includes: applying an actor-critic algorithm to train the deep reinforcement learning network.
[0011] Optionally, the node output sequence of the combat network after being disrupted by each weapon that minimizes the objective function is solved by the deep reinforcement learning network, including: applying the encoder in the deep reinforcement learning network to extract the feature vector of each object to form a vector matrix; applying the decoder in the deep reinforcement learning network in combination with the attention mechanism to calculate the attention value of unselected objects, the agent greedily selects the object with the maximum attention value, repeats the iteration until the maximum constraint or the maximum number of steps of the weapon attack capability is reached, and obtains the selected object group; mapping the objects in the selected object group to the weapon's attack on the node, deriving a group of removed nodes, and calculating the value of the objective function.
[0012] Optionally, the decoder in the deep reinforcement learning network is combined with the attention mechanism to calculate the attention value of the unselected objects, and the agent selects the object with the maximum attention value in a greedy manner, including: taking the feature vector of each object as input, decoding the current state into a high-dimensional hidden state in turn, applying the decoding network to decode, and obtaining the hidden layer state; feeding the weighted sum of the hidden layer state and the feature vector into the tanh activation function, and using a binary mask to determine whether the object is valid in the current time period; based on the premise of knowing the current state and the selected object, selecting the conditional probability distribution of the next object, and selecting the object with the highest probability through a greedy method.
[0013] Based on the same inventive concept, a combat network disruption device based on deep reinforcement learning is provided, comprising: a network construction unit, for constructing a combat network, obtaining weapon information of at least one weapon, and combining any node in the combat network with any weapon to generate an object; an objective function construction unit, for constructing an objective function for characterizing the degree of disruption of the combat network based on all the objects combined with the combat network and each of the weapons, wherein the degree of disruption is the number of closed loops with different path lengths among all nodes in the combat network after disruption; a network disruption unit, for inputting all the objects into a preset deep reinforcement learning network, and solving, through the deep reinforcement learning network, a node output sequence after the combat network is disrupted by each of the weapons that minimizes the objective function.
[0014] Based on the same inventive concept, an embodiment of the present invention further proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned method when executing the program.
[0015] Based on the same inventive concept, an embodiment of the present invention further proposes a computer storage medium, in which at least one executable instruction is stored, and the executable instruction enables a processor to execute the aforementioned method.
[0016] An embodiment of the present invention constructs a combat network and obtains weapon information of at least one weapon, combines any node in the combat network with any weapon to generate an object; based on all the objects in the combination of the combat network and each weapon, constructs an objective function for characterizing the degree of collapse of the combat network, where the degree of collapse is the number of closed loops with different path lengths among all nodes in the combat network after collapse; inputs all the objects into a preset deep reinforcement learning network, and solves the node output sequence of the combat network after collapse by each weapon that minimizes the objective function through the deep reinforcement learning network, thereby improving the speed and quality of solving the combat network collapse.
[0017] The above description is only an overview of the technical solutions of the embodiments of the present invention. In order to more clearly understand the technical means of the embodiments of the present invention, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:
[0019] Figure 1 A schematic diagram of a process for a combat network disruption method based on deep reinforcement learning provided by an embodiment of the present invention is shown;
[0020] Figure 2 A schematic diagram showing an object of a node and weapon combination according to an embodiment of the present invention;
[0021] Figure 3 An example diagram showing an illustrative scenario of weapon allocation in a combat system destruction according to an embodiment of the present invention;
[0022] Figure 4 An example diagram showing the effects of different destruction strategies on the combat network provided by the embodiment of the present invention is shown;
[0023] Figure 5 An example diagram showing a method for disrupting a combat network provided by an embodiment of the present invention is shown;
[0024] Figure 6 shows an architecture diagram of deep reinforcement learning provided by an embodiment of the present invention;
[0025] Figure 7 A schematic diagram showing the disintegration effects of different algorithms provided by an embodiment of the present invention on two different networks;
[0026] Figure 8 A schematic diagram showing the solution time of different algorithms provided by an embodiment of the present invention for collapsing an ER network of the same scale is shown;
[0027] Figure 9 A schematic diagram showing the solution time of different algorithms provided by an embodiment of the present invention for collapsing an SF network of the same scale;
[0028] Figure 10 A schematic diagram showing the correlation between the solution speed and network scale of different algorithms provided by an embodiment of the present invention;
[0029] Figure 11 A schematic diagram showing the generalization capabilities of different algorithms provided by embodiments of the present invention is shown;
[0030] Figure 12 A schematic diagram of the structure of a combat network disruption device based on deep reinforcement learning provided by an embodiment of the present invention is shown;
[0031] Figure 13 A schematic diagram of an electronic device in an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0032] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0033] Figure 1 The flowchart of the method for disrupting a combat network based on deep reinforcement learning provided by an embodiment of the present invention is shown. Figure 1 As shown, the combat network disruption method based on deep reinforcement learning includes:
[0034] Step S11: construct a combat network, obtain weapon information of at least one weapon, and combine any node in the combat network with any of the weapons to generate an object.
[0035] In the embodiment of the present invention, the combat network is defined as G = (V, E), where V represents the set of nodes of the combat entity, V = {v1, v2, ..., v N}, E is the set of edges between nodes, representing the interaction between combat entities, E = {e1, e2, ..., e T}. Set any node v j Described as {x j ,y j ,z j ,k j}, where (x j ,y j ,z j ) represents node v j The position coordinates, k j Represents node v j degree, k j The value of node v j The adjacency matrix A(G) of the combat network G is (a ij ) N×N The definition is as follows: If the node v iand node v j is connected, then a ij =1; otherwise, a ij =0.
[0036] Assume that the attacker has M weapon resources, and can use the set W = {w1, w2, ..., w M} to indicate that any weapon w i With tuple w i :={x i ,y i ,z i ,c i ,r i} description, where (x i ,y i ,z i ) indicates weapon w i The position coordinates, c i Indicates weapon w i Attack ability, r i Indicates weapon w i attack range.
[0037] In the embodiment of the present invention, any node in the combat network is combined with any weapon to generate an object. M×N objects are generated by Cartesian product. Figure 2 As shown, each object contains attributes from weapons and network nodes. Assume k th The object is i th Weapons and j th The combination of nodes, then the object o k It can be expressed as o k := <w i ,v j >, which means weapon w i Used to attack node v j Each input vector o i is encoded as an embedding vector e i , forming a (M*N)×d h The encoding vector of dimension E={e1,e2,…,e MN}, where d h is the dimension of the target vector. For example, in the embodiment of the present invention, the dimension of the target vector is d h For weapons w i The 5-dimensional and node v j The 4-dimensional combination, d h =9. Of course, in other embodiments of the present invention, weapons w can also be set separately as needed. i and node v j The dimension of the target vector is then determined, without limiting its specific value.
[0038] Step S12: Based on all the objects of the combat network and each weapon combination, construct an objective function for characterizing the degree of collapse of the combat network, where the degree of collapse is the number of closed loops with different path lengths among all nodes in the combat network after collapse.
[0039] In the embodiment of the present invention, optionally, the Euclidean distance between the node in any of the objects and the weapon is calculated. The weapon can be used to attack the node only when the node is within the attack range of the weapon. ij To define weapon w i and node v j Distance between:
[0040]
[0041] Then, the indicator function value of whether the weapon can be used to attack the node is determined based on the Euclidean distance, and a first binary variable of whether the weapon attacks the node is obtained. ij ) to indicate weapon w i Can it be used to attack node v j :
[0042]
[0043] A plurality of constraints are further determined based on the first binary variable and the indicator function value. One constraint is that, based on the fact that any weapon can only attack once and can only attack one node at most, the sum of the first binary variables of all the weapons on any node is determined to be less than or equal to 1. ii To determine the weapon w i Is it possible to attack node v i If the weapon w i Attacking node v i , then x ij =1; otherwise x ij = 0. With constraints:
[0044]
[0045] That is, for any node in the combat network, all weapons have the first binary variable x of that node ij The sum is less than or equal to 1.
[0046] According to the fact that any weapon can only be used to attack nodes within its attack range, the first binary variable is determined to be less than or equal to the indicator function value. In order to ensure that the weapon can only be used to attack nodes within its attack range, another constraint is that
[0047]
[0048] That is, for any object composed of any node and any weapon, the first binary variable is less than or equal to the indicator function value.
[0049] The embodiment of the present invention further calculates the total damage value of each weapon to any node in the combat network based on the weapon information and the indicator function value. A node can only be removed if the damage effect of a weapon exceeds its removal threshold. In the embodiment of the present invention, it is assumed that the damage effect of multiple weapons on the same weapon is the linear sum of their respective damage values. Let u i Represents all weapons against node v j The sum of the damage values,
[0050]
[0051] In this embodiment of the present invention, the destructive effect of a weapon on a node depends on two aspects: the weapon's attack capability and the node's centrality. The higher the weapon's attack capability, the greater the damage to the node. In addition, the difficulty of removing a node increases with the increase of the node's centrality. Here, the node v j The centrality can be measured using its degree k j To measure. Assume that the cost of removing a node is proportional to the number of nodes. The embodiment of the present invention determines the removal threshold of any node in the combat network, and determines the second binary variable of whether the node is removed based on the first binary variable and the removal threshold. Set the node removal threshold q i Defined as removing node v i The required weapon attack ability value, assuming q i is node v i degree k i The function of
[0052] q j =αk j ,j=1,…,N,α>0 (6)
[0053] q j The higher the value of , the greater the weapon attack capability required to remove the node. α is a random perturbation value, which represents the impact of external factors, such as electromagnetic interference, on the node removal threshold. α can be randomly set between 0 and 1.
[0054] Using the second binary variable p j To represent the node v j Can it be removed? If the node v j can be removed, then p j is 1; otherwise p jis 0, which can be expressed as:
[0055]
[0056] The embodiment of the present invention also determines a removal strategy based on the second binary variable of each node in the combat network, calculates the number of closed loops with different path lengths of all nodes in the combat network after the collapse corresponding to the removal strategy, and obtains an objective function for characterizing the degree of collapse of the combat network.
[0057] In an embodiment of the present invention, weapons are distributed in an illustrative scenario of a combat system destruction, such as Figure 3 As shown, the node v6 has the highest degree, which usually means that it is a key node. The combat system is the combat network in the embodiment of the present invention. The attacker has deployed four weapons {w1, w2, w3, w4}, each of which has its own attack range, represented by the large gray circle. It is easy to find that the attacker cannot destroy the node v6 with the highest degree. It can be seen that the destruction of the combat system must not only consider the network structure itself, but also the attack effect of the weapon. In fact, each node in the combat system can have a certain defense capability, and each weapon can cause limited damage. As shown in the figure, the attacker has deployed four weapons {w1, w2, w3, w4}, each of which has its own attack range, represented by the large gray circle. It is easy to find that the attacker cannot destroy the node v6 with the highest degree. Figure 4 As shown in Figure 1, there is a combat network consisting of 13 nodes. The attacker has deployed four weapons {w1, w2, w3, w4} around the combat network. Each weapon has an attack range, represented by a light gray circle. Node v1 can be attacked by weapon w1, node v3 can be attacked by weapon w3, and node v4 can be attacked by both weapons w3 and w4. Each weapon has different damage-dealing capabilities, and each node in the combat network has a different removal threshold. Without loss of generality, the attack capabilities of each weapon {w1, w2, w3, w4} are {5, 8, 5, 2}, respectively. The removal thresholds of each node {v1, v2, v3, v4} are {5, 9, 5, 6}, respectively. Figure 4 There are two different weapon allocation strategies for disrupting the combat network. In strategy 1, weapon w1 attacks node v1 (5 ≥ 5), weapon w3 attacks node v3 (5 ≥ 5), and weapon w4 attacks node v4 (2 ≤ 6). In this way, nodes {v1, v3} can be successfully destroyed and removed from the combat network. In strategy 2, weapon w1 attacks node v1 (5 ≥ 5), and weapons {w3, w4} attack node v4 (5 + 2 ≥ 6). In this way, nodes {v1, v4} can be successfully removed from the combat network.
[0058] The effectiveness of network disintegration can be measured using natural connectivity, which describes the number of closed loops with different path lengths between all nodes in the network. Let G denote the network topology and A denote the adjacency matrix of G. The natural connectivity of G, or Γ(G), can be calculated as follows:
[0059]
[0060] Among them, λ i is the i-th largest eigenvalue of the adjacency matrix.
[0061] According to the above formula, the network connectivity of the combat system can be easily calculated after executing strategies 1 and 2. Figure 4 ,The network connectivity in strategy 1 is 2.12, while the network connectivity in strategy 2 is 1.77.
[0062] In the embodiment of the present invention, the combat network collapse problem can be defined as: given a combat network, that is, G = (V, E), and multiple weapons W = {w1, w2, ..., w M Each node in V has location information and a removal threshold. Each weapon in W has location information, attack range, and attack value. The combat network disruption problem is then to distribute weapons to appropriate nodes to minimize the network connectivity of the combat network G.
[0063] In the embodiment of the present invention, Indicates the removed node, and the network after removing the node is Then the removal strategy can be Y=[y1,…,y N ] means that if y j =1 then As described in Section 1.1, the effect of the removal strategy can be expressed as The goal of combat system destruction is to find a set of removal strategies Y * , in order to minimize the network connectivity of the combat system. MinΦ(Y=[y1,…,y N ]). Thus, the optimization problem of the embodiment of the present invention can be expressed as follows:
[0064] MinΦ(Y=[y1,…,y N ])#(9a)
[0065] Satisfies (2) to (8) # (9b)
[0066] It can be seen that once a node is removed, the natural connectivity will be strictly monotonically decreasing. Therefore, a lower value of Φ means a more destructive collapse strategy.
[0067] Step S13: All the objects are input into a preset deep reinforcement learning network, and the deep reinforcement learning network is used to solve the node output sequence after the combat network is disintegrated by each weapon so as to minimize the objective function.
[0068] The combat network disruption method of the present invention is generated through deep reinforcement learning: through interaction with the environment, the agent acquires a strategy that can output a solution with the maximum reward. In terms of combat network disruption, the focus is on how to design the state space, action space, and reward function. The state S is defined as the set of selected objects. At time t, if the agent has selected t-1 objects, then the state information at time t can be represented as s t ={o1,o2,...,o t-1}, where s t ∈S. The action space is a set of all possible actions that an agent can perform in a given state. This embodiment of the present invention considers the selection of an object as an action, so the dimension of the action space is M × N. The reward definition for the solution is the same as the objective function, which is the network performance of the combat system after the node is removed.
[0069] The embodiment of the present invention adopts a deep reinforcement learning (DRL) network. DRL is an intelligent agent modeling method that combines the feature extraction capability of deep learning with the sequential decision-making capability of reinforcement learning. For complex network disintegration problems, they can be regarded as Markov decision processes and DRL can be applied to solve them. The processing process of the deep reinforcement learning network is divided into three stages: combination, selection and mapping. Figure 5 As shown, in the combination phase, weapons and nodes are combined to form new objects through a Cartesian product. The selection phase includes an encoding process and multiple decoding processes, and the decoding process outputs the probability of each object being selected at each step based on the state of the environment. The mapping phase maps the selected object to the weapon attack node.
[0070] In an embodiment of the present invention, optionally, the encoder in the deep reinforcement learning network is first applied to extract the feature vectors of each object to form a vector matrix. Figure 6This is an architectural diagram of a deep reinforcement learning network according to an embodiment of the present invention, including encoding, decoding and attention modules. The encoder extracts the features of the input object through the embedding layer. The decoder is used to store the decoded sequence information. The attention module uses the attention mechanism to output the probability distribution of subsequent inputs based on the embedded information and the hidden layer state of the decoding network. The encoder is designed to map the state information of the input sequence so that the agent can understand the representation of each object. Regarding the network collapse problem, the order of the input sequence does not contain any useful information, so there is no need to consider the encoding order of the input vectors. Regarding the network decomposition problem, the order of the input sequence does not contain any useful information, so there is no need to consider the encoding order of the input vectors. A one-dimensional convolutional neural network (CNN) can be selected as the encoding network to reduce the complexity of the model. Use the encoder to extract the features of all objects and output their embeddings (high-dimensional vectors), which are then passed to the decoding neural network for decoding. Each input vector o i is encoded as an embedding vector e i , forming a (M*N)×d h The encoding vector of dimension E={e1,e2,…,e MN}, where d h is the dimension of the target vector.
[0071] Then, the decoder in the deep reinforcement learning network is applied in combination with the attention mechanism to calculate the attention value of the unselected objects. The agent selects the object with the maximum attention value in a greedy manner, and iterates repeatedly until the maximum constraint or the maximum number of steps of the weapon attack capability is reached to obtain the selected object group. Optionally, the feature vector of each object is used as input, and the current state is decoded into a high-dimensional hidden state in sequence, and the decoding network is applied for decoding to obtain the hidden layer state; the weighted sum of the hidden layer state and the feature vector is fed into the tanh activation function, and a binary mask is used to determine whether the object is valid in the current time period; based on the known current state and the selected object, the conditional probability distribution of the next object is selected, and the object with the highest probability is selected by a greedy method. That is, in an embodiment of the present invention, the decoder uses the feature vector generated by the encoder as input to decode the current state into a high-dimensional hidden state in sequence. Using a recursive neural network (RNN) with a memory storage function as a decoding network, the hidden layer state d t , where d t The query contains the attention layer output by the decoder before step t. The agent greedily selects the object with the largest attention value. Afterwards, the decoding process described above is repeated to build a complete solution. The iterative process stops until the maximum constraint on the weapon's attack capability or the maximum number of steps is reached.
[0072] In an embodiment of the present invention, the output dimension of the solution is dynamically determined based on the dimension of the input information. When the number of weapons and network nodes changes, the output dimension will also change accordingly. The output dimension of the traditional Seq2Seq model is fixed and cannot solve the problem of dynamic output dimension. To solve this problem, Vinyals introduced the attention mechanism into the Seq2Seq model and achieved good results. In this way, the attention mechanism can be introduced into the neural network architecture to deal with this problem. As shown in the following equation, at step t, the decoded hidden layer state d is calculated t and the embedding vector e j The weighted sum of is then fed into the tanh activation function.
[0073]
[0074] Among them, v and W a and W b are trainable parameters.
[0075] A masking mechanism is used at the output of the neural network to reduce the policy space. Specifically, a binary mask is used to determine the object o j Is it still valid at time t as follows:
[0076]
[0077] Given the current state and the selected action, the conditional probability distribution of choosing the next action can be expressed as follows:
[0078] p(y t |y0,y1,…,y t-1 ,s t )=softmax(u t +mask t )#(12)
[0079] In the application phase, the action with the highest probability is selected through a greedy method. In the training phase, the next action can be selected through importance sampling, so that actions with very low probability can also be selected.
[0080] After obtaining the selected object group, the objects in the selected object group are mapped to the attacks of the weapons on the nodes, a set of removed nodes is derived, and the value of the objective function is calculated. The deep reinforcement learning network is then used to solve the sequence of node outputs of the combat network after being disrupted by each weapon that minimizes the objective function.
[0081] In an embodiment of the present invention, before step S13, the actor-critic algorithm is applied to train the deep reinforcement learning network. That is, the actor-critic algorithm is applied to train the parameters of the deep reinforcement learning network of the embodiment of the present invention. The actor-critic algorithm consists of two networks: the actor network consists of an encoder and a decoder, which is used to generate the probability distribution of selecting an action in the current state; the critic network estimates the state value of a given problem instance, and its network structure is similar to the encoder of the actor network. During the training phase, the reward value of the solution is calculated according to the designed reward function, and then used for backpropagation and adjustment of network parameters. When the loss value of the network parameters is stable and the reward value meets the expectations, a well-trained network model is obtained. During the testing phase, the trained network model can be used to quickly find a high-quality disruption method based on the input of weapons and enemy combat network information.
[0082] The following is an experimental comparison of the combat network disintegration method based on deep reinforcement learning of an embodiment of the present invention. Assume that the battlefield is a square area, and the coordinates of the network nodes and weapons are randomly generated. The attack range of the randomly generated weapons is uniformly distributed in [0,2], and the attack capability value of the weapons is uniformly distributed in the range of [0,10]. Random perturbations are added when setting the α value, which is used to indicate the influence of other factors on the weapon capability required to remove the node. Two classic synthetic network structures are considered as the architecture of the combat network, including scale-free (SF) network and (ER) Random Networks.
[0083] Two DRL models, DRL-25 and DRL-40, were trained. The DRL-25 model was trained with 10 weapons and 15 network nodes, while the DRL-40 model was trained with 15 weapons and 25 network nodes. Each model instance was trained using 1 million data points.
[0084] For the DRL model, the encoder of the actor network embeds the object information into a 128-dimensional vector through a 1D convolutional network with one layer, while the decoder is a GRU recurrent neural network with 128 hidden units, with a dropout of 0.1. The critic network consists of multiple 1D convolutional networks, where the output of the last layer is set to 1. The model uses the Adam optimizer for gradient optimization. The batch size is 128 and the learning rate is 10 -4 .
[0085] For the baseline algorithm used for comparison, the combat network disruption problem can be viewed as a combinatorial optimization problem with two subproblems: selecting several nodes from a candidate set of network nodes and assigning appropriate weapons to each selected node to attack it. The classic node centrality-based disruption strategy consists of only one subproblem: selecting several nodes from a candidate set of network nodes, but without specifying which weapons to use to attack these nodes. Therefore, only heuristic algorithms are considered as baseline algorithms. Heuristic algorithms are a flexible and effective approach for solving combinatorial optimization problems. For example, genetic algorithms (GAs) and differential evolution (DE) algorithms seek optimal solutions by simulating natural selection, inheritance, and evolutionary processes, demonstrating simple, robust, and powerful global search capabilities. By simulating genetic processes such as selection, crossover, and mutation, GAs gradually evolve solutions that better suit the given problem. Differential evolution algorithms are intelligent optimization search algorithms that emerge through cooperation and competition among individuals within a population.
[0086] The number of weapons was fixed at 40, and experiments were conducted on network nodes of different sizes ranging from 40 to 150. The solution quality of the DRL model instance was compared with the baseline algorithm. The embodiment of the present invention randomly generated 10 problem instances for the two networks; then, the average disintegration effect of each type of network at a specific scale was calculated. The disintegration effect of different algorithms on two different networks is shown in Figure 2. Figure 7 As shown in Figure 2. The smaller the target value Φ is, the better the solution quality of the algorithm is. In the ER network ( Figure 7 a) and SF network ( Figure 7 In b), the quality of solutions generated by the DRL algorithm is superior to that of the baseline methods, with a particularly significant advantage in SF networks. Specifically, for ER networks, when the problem size is 40, the average decomposition performance of the DRL-40 algorithm is slightly better than that of the other algorithms, while the DRL-25 algorithm achieves the best decomposition results when the problem size is 80, 120, and 160. For SF networks, the decomposition performance of the DRL-40 algorithm is significantly better than that of the GA and DE when the problem size is 40 and 160. Furthermore, the DRL-25 algorithm also outperforms the GA and DE for problem sizes of 80 and 120.
[0087] Solving the battle network disassembly problem in a timely manner is crucial. Rapidly generating disassembly methods based on weapon and operational SoS data is crucial for seizing the initiative in warfare. Because the DRL used in this embodiment of the present invention is an end-to-end model, compared to baseline algorithms, only the testing time of the application phase is considered, not the training time. Figure 8 Schematic diagram of the solution time of different algorithms for collapsing ER networks of the same scale. Figure 9 Schematic diagram of the solution time for different algorithms to collapse SF networks of the same scale, where: Figure 8 a- Figure 8 The scales in d are 40, 80, 120 and 160 respectively. Figure 9 a- Figure 9 The scales in d are also 40, 80, 120 and 160 respectively. Figure 8 and Figure 9 As shown in Figure 2, a box plot of the solution time of different algorithms is drawn when facing the same size synthetic network decomposition problem. It can be seen that regardless of the size of the network nodes, the solution time of the DRL algorithm is much shorter than that of other baseline algorithms. In order to further analyze the correlation between the algorithm and the scale of the network decomposition problem, a heat map of the average solution time is drawn. Figure 10 As shown, for the ER network ( Figure 10 a) and SF network ( Figure 10 b),As the problem size increases, it can be seen that the solution time of the GA and DE algorithms increases significantly.,In contrast, the color associated with the grid of the DRL algorithm remains relatively constant,,indicating that its solution time is not affected by changes in problem size.
[0088] DRL's generalization capability refers to the model's performance in unseen situations. Strong generalization enables the model to adapt to new environments, exceeding its performance on the training data alone. During the training phase, the DRL model is trained using data with fixed weapon capabilities and attack ranges. However, in actual combat scenarios, changes in the external environment, such as enemy electromagnetic interference and terrain changes, can affect the weapon's attack capability and attack range. In terms of both solution speed and solution quality, DRL demonstrates good generalization capabilities for problems of varying sizes. Next, we will evaluate DRL's performance under varying weapon capability and attack range conditions.
[0089] We rescale the weapon's capabilities and attack ranges, and then generate 10 problem instances of different sizes to evaluate the performance of DRL when the weapon's capabilities and attack ranges are not fixed. Figure 11 It can be seen that on both synthetic networks, DRL outperforms the baseline algorithm in both solution quality and speed.
[0090] In summary, the combat network disruption method based on deep reinforcement learning in an embodiment of the present invention constructs a combat network and obtains weapon information of at least one weapon, and combines any node in the combat network with any weapon to generate an object; based on all the objects combined with the combat network and each weapon, constructs an objective function for characterizing the degree of disruption of the combat network, where the degree of disruption is the number of closed loops with different path lengths among all nodes in the combat network after disruption; all the objects are input into a preset deep reinforcement learning network, and the node output sequence after the combat network is disrupted by each weapon that minimizes the objective function is solved by the deep reinforcement learning network, thereby improving the solution speed and quality of combat network disruption.
[0091] The foregoing description is of specific embodiments of the present invention. In some cases, the actions or steps described in the embodiments of the present invention may be performed in an order different from that shown in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0092] Based on the same concept, the embodiment of the present invention also provides a combat network disruption device based on deep reinforcement learning. Applied to the server. Figure 12 As shown in FIG, the combat network disintegration based on deep reinforcement learning includes: a network construction unit, an objective function construction unit, and a network disintegration unit.
[0093] a network construction unit, configured to construct a combat network, obtain weapon information of at least one weapon, and combine any node in the combat network with any of the weapons to generate an object;
[0094] an objective function construction unit, configured to construct an objective function for characterizing a degree of disruption of the combat network based on all the objects in the combat network and each of the weapon combinations, wherein the degree of disruption is the number of closed loops with different path lengths among all nodes in the combat network after the disruption;
[0095] The network collapse unit is used to input all the objects into a preset deep reinforcement learning network, and solve the node output sequence of the combat network after being collapsed by each weapon so as to minimize the objective function through the deep reinforcement learning network.
[0096] For the convenience of description, the above devices are described as being divided into various modules according to their functions. Of course, when implementing the embodiments of the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0097] The apparatus of the above embodiment is applied to the corresponding method of the above embodiment and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0098] Based on the same inventive concept, an embodiment of the present invention further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the method described in any one of the above embodiments is implemented.
[0099] An embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores at least one executable instruction, and the computer executable instruction can execute the method described in any one of the above embodiments.
[0100] Figure 13 13 shows a more specific hardware structure diagram of an electronic device provided in this embodiment. The device may include: a processor 1301, a memory 1302, an input / output interface 1303, a communication interface 1304, and a bus 1305. The processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 are communicatively connected to each other within the device via the bus 1305.
[0101] The processor 1301 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the method embodiments of the present invention.
[0102] The memory 1302 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1302 can store an operating system and other application programs. When the technical solutions provided by the method embodiments of the present invention are implemented through software or firmware, the relevant program codes are stored in the memory 1302 and called and executed by the processor 1301.
[0103] The input / output interface 1303 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0104] The communication interface 1304 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WIFI, Bluetooth, etc.).
[0105] The bus 1305 comprises a pathway for transmitting information between the various components of the device (eg, the processor 1301 , the memory 1302 , the input / output interface 1303 , and the communication interface 1304 ).
[0106] It should be noted that although the above device only shows the processor 1301, the memory 1302, the input / output interface 1303, the communication interface 1304, and the bus 1305, in a specific implementation, the device may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the above device may only include the components necessary to implement the embodiments of the present invention, and does not necessarily include all the components shown in the figure.
[0107] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure is limited to these examples. Based on the concept of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.
[0108] This application is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of all embodiments. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of this disclosure.
Claims
1. A combat network disruption method based on deep reinforcement learning, characterized in that: The method comprises: Building a combat network, obtaining weapon information of at least one weapon, and combining any node in the combat network with any of the weapons to generate an object; Based on all the objects of the combat network and each weapon combination, constructing an objective function for characterizing the degree of collapse of the combat network, wherein the degree of collapse is the number of closed loops with different path lengths among all nodes in the combat network after collapse; All of the objects are input into a preset deep reinforcement learning network, and the deep reinforcement learning network is used to solve the node output sequence after the combat network is disintegrated by each weapon so as to minimize the objective function.
2. The method according to claim 1, characterized in that The step of constructing a combat network and obtaining weapon information of at least one weapon includes: The combat network is defined as G = (V, E), where V represents the set of nodes of combat entities and E is the set of edges between nodes, representing the interactions between combat entities. Any node v j Described as {x j ,y j ,z j ,k j }, where (x j ,y j ,z j ) represents node v j The position coordinates, k j Represents node v j degree; Use any weapon w i With tuple w i :={x i ,y i ,z i ,c i ,r i } description, where (x i ,y i ,z i ) indicates weapon w i The position coordinates, c i Indicates weapon w i Attack ability, r i Indicates weapon w i attack range.
3. The method according to claim 2, characterized in that The objective function for characterizing the degree of collapse of the combat network based on all the objects in combination with the combat network and each weapon is constructed, including: calculating the Euclidean distance between the node in any one of the objects and the weapon; determining an indicator function value of whether the weapon can be used to attack the node according to the Euclidean distance, and obtaining a first binary variable of whether the weapon attacks the node; determining a plurality of constraints based on the first binary variable and the indicator function value; Calculate the total damage value of each weapon to any node in the combat network according to the weapon information and the indicator function value; determining a removal threshold for any of the nodes in the combat network, and determining a second binary variable indicating whether the node is removed based on the first binary variable and the removal threshold; A removal strategy is determined based on the second binary variable of each node in the combat network, and the number of closed loops with different path lengths of all nodes in the combat network after collapse corresponding to the removal strategy is calculated to obtain an objective function for characterizing the degree of collapse of the combat network.
4. The method according to claim 3, characterized in that The determining of a plurality of constraint conditions according to the first binary variable and the indicator function value comprises: According to the principle that any weapon can only attack once and can only attack one node at most, determining that the sum of the first binary variables of all the weapons on any of the nodes is less than or equal to 1; According to the fact that any weapon can only be used to attack nodes within an attack range, it is determined that the first binary variable is less than or equal to the indicator function value.
5. The method according to claim 1, wherein Before solving the objective function by the deep reinforcement learning network to minimize the node output sequence of the combat network after being disintegrated by each weapon, the method includes: The actor-critic algorithm is applied to train the deep reinforcement learning network.
6. The method according to claim 5, characterized in that The node output sequence after the combat network is disintegrated by each weapon, which is solved by the deep reinforcement learning network to minimize the objective function, includes: Applying the encoder in the deep reinforcement learning network to extract the feature vectors of each object to form a vector matrix; Applying the decoder in the deep reinforcement learning network in combination with the attention mechanism to calculate the attention values of unselected objects, the agent greedily selects the object with the largest attention value, and iterates repeatedly until the maximum constraint of the weapon attack capability or the maximum number of steps is reached, thereby obtaining a selected group of objects; The objects in the selected object group are mapped to weapon attacks on nodes, a set of removed nodes is derived, and a value of the objective function is calculated.
7. The method according to claim 1, characterized in that The decoder in the deep reinforcement learning network is applied in combination with an attention mechanism to calculate the attention value of unselected objects, and the agent selects the object with the maximum attention value in a greedy manner, including: Taking the feature vector of each object as input, decoding the current state into a high-dimensional hidden state in sequence, applying a decoding network to decode, and obtaining a hidden layer state; The weighted sum of the hidden layer state and feature vector is fed into the tanh activation function, and a binary mask is used to determine whether the object is valid in the current time period; Based on the known current state and the selected object, the conditional probability distribution of the next object is selected, and the object with the highest probability is selected through a greedy method.
8. A combat network disruption device based on deep reinforcement learning, characterized by: The device comprises: a network construction unit, configured to construct a combat network, obtain weapon information of at least one weapon, and combine any node in the combat network with any of the weapons to generate an object; an objective function construction unit, configured to construct an objective function for characterizing a degree of disruption of the combat network based on all the objects in the combat network and each of the weapon combinations, wherein the degree of disruption is the number of closed loops with different path lengths among all nodes in the combat network after the disruption; The network collapse unit is used to input all the objects into a preset deep reinforcement learning network, and solve the node output sequence of the combat network after being collapsed by each weapon so as to minimize the objective function through the deep reinforcement learning network.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A computer storage medium, characterized in that: The storage medium stores at least one executable instruction, and the executable instruction enables the processor to execute the method according to any one of claims 1 to 7.