Network recovery method and device, storage medium and program product

By constructing an intelligent agent and training a deep model of target classification distribution, the flexibility problem of traditional network recovery technology under complex network attacks is solved, and efficient recovery is achieved in different network environments.

CN121125149APending Publication Date: 2025-12-12CHINA MOBILE GROUP JILIN BRANCH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510652704.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Traditional network backup and recovery technologies struggle to adapt to changing network environments when faced with complex and ever-changing network attacks, resulting in low recovery flexibility and an inability to quickly and effectively restore network systems.

Method used

An intelligent agent is constructed to perform network recovery actions under different network states and obtain feedback results. The initial classification distribution depth model is trained to obtain the target classification distribution depth model, so as to determine and execute the target network recovery action to restore the network system.

Benefits of technology

It improves the flexibility of network recovery, enabling network recovery actions to adapt to the needs of different network environments and achieve effective recovery in the event of network anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125149A_ABST
    Figure CN121125149A_ABST
Patent Text Reader

Abstract

The invention discloses a network recovery method and device, a storage medium and a program product, and relates to the technical field of network security, and the method comprises the steps: obtaining an intelligent agent constructed based on a network system, and obtaining the network environment interaction data of the intelligent agent, the network environment interaction data comprises network recovery actions respectively executed by the intelligent agent in different network states of the network system and an execution feedback result after each network recovery action is executed; and training a preset initial classification distribution depth model through the network environment interaction data to obtain a trained target classification distribution depth model, so as to determine a target network recovery action of the network system when the network is abnormal through the target classification distribution depth model, and controlling the intelligent agent to execute the target network recovery action so as to recover the network system. The technical problem of low flexibility of network recovery is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network security technology, and in particular to network recovery methods, devices, storage media, and program products. Background Technology

[0002] Traditional network backup and recovery technologies rely on periodically backing up network data and configuration information to recover after a failure or attack. However, with increasingly complex and varied network attacks, attacks can significantly alter the network environment, making backup data difficult to adapt to the new environment and leading to network recovery failure. Therefore, current technologies suffer from low flexibility in network recovery.

[0003] The above content is only used to help understand the technical solutions of the embodiments of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a network recovery method, device, storage medium, and program product, aiming to solve the technical problem of low flexibility in network recovery.

[0005] To achieve the above objectives, embodiments of this application provide a network recovery method, the method comprising:

[0006] A smart agent constructed based on a network system is obtained, and network environment interaction data of the smart agent is obtained. The network environment interaction data includes network recovery actions performed by the smart agent under different network states of the network system, and execution feedback results after each network recovery action is performed.

[0007] By using the network environment interaction data, a preset initial classification distribution depth model is trained to obtain a trained target classification distribution depth model. The target classification distribution depth model is then used to determine the target network recovery action of the network system when the network is abnormal, and the agent is controlled to execute the target network recovery action to restore the network system.

[0008] In one embodiment, the agent includes a state space and an action space; the step of obtaining the agent constructed based on the network system includes:

[0009] The system network state of the network system is obtained, wherein the system network state includes the resources to be allocated in the network system at a preset time, the node resource usage state, effective state, and node threat event type of each network node in the network system at the preset time; the resource allocation vector of the resources to be allocated, the node resource vector of each network node at the preset time, the effective vector of the effective state, and the threat event vector of the node threat event type are determined; a state space is constructed based on the resource allocation vector, the node resource vector of each network node at the preset time, the effective vector, and the threat event vector; multiple network recovery actions supported by the agent are constructed to obtain an action space constructed by each network recovery action.

[0010] In one embodiment, the resources to be allocated include multiple resource types corresponding to sub-resources to be allocated, the network recovery action includes at least one node recovery action of a network node, and the node recovery action includes at least one sub-recovery action; the step of constructing all network recovery actions supported by the agent includes: obtaining the total node degree of all network nodes in the network system;

[0011] For any target network node among the network nodes, for each target network node, the target node degree of the target network node is obtained, and the target effective number of valid network nodes and the target invalid number of invalid network nodes are determined in the target network node and the network nodes connected to the target network node; based on the target node degree, the total node degree, and the preset degree correlation coefficient, the target degree correlation allocation coefficient is determined, and the ratio of the target effective number to the target invalid number is taken as the target effective ratio; for each target sub-resource to be allocated, the product of the target sub-resource to be allocated, the target degree correlation allocation coefficient, the target effective ratio, and the preset allocation coefficient is taken as the sub-recovery action to allocate the target sub-resource to the target network node.

[0012] In one embodiment, the initial classification distribution depth model includes an initial current neural network and an initial target neural network, the target classification distribution depth model includes a current neural network and a target neural network, and the execution feedback result includes the training network recovery state of the network system and the training action reward for the agent performing the training network recovery action; the step of training the preset initial classification distribution depth model through the network environment interaction data to obtain the trained target classification distribution depth model includes: determining the training network recovery action, training network recovery state, and training action reward for any training network state from the network environment interaction data; inputting the training network state and the training network recovery action into the initial current neural network; outputting the training action reward rate distribution after performing the training network recovery action in the training network state through the initial current neural network; inputting the training action reward and the training network recovery state into the initial target neural network; determining the target action reward rate distribution that the initial current neural network needs to achieve through the initial target neural network; calculating the training loss between the training action reward rate distribution and the target action reward rate distribution through a preset loss function; and obtaining the trained current neural network and target neural network when the training loss is less than or equal to the preset loss.

[0013] In one embodiment, the method further includes: determining a reward coefficient for the agent to perform the training network recovery action based on the training network recovery state and the training network state; for each network node in the network system, obtaining the node degree of the network node and the amount of recovery resources required for the network node to recover, and using the ratio of the node degree to the amount of recovery resources as the resource recovery ratio; for each network node, performing a power operation with the resource recovery ratio as the base and a preset reward priority parameter as the exponent to obtain the node reward priority coefficient of the network node; obtaining the abnormal start time when the network system enters the training network state, the original network performance of the network system before the abnormal training network state, and the recovery time when the network system enters the training network recovery state; using the difference between the recovery time and the abnormal start time as the recovery duration, and determining the network elasticity of the network system based on the original network performance and the network performance at each time between the abnormal start time and the recovery time; and determining the training action reward based on the reward coefficient, the recovery duration, the network elasticity, and the node reward priority coefficient corresponding to each network node.

[0014] In one embodiment, the method further includes:

[0015] If an anomaly is detected in the current network state of the network system, the current network state is input into the target classification distribution depth model. The target classification distribution depth model determines the action reward rate distribution corresponding to each network recovery action performed by the agent under the current network state. Based on a preset greedy random probability, a greedy network recovery action is determined among the network recovery actions. Based on the action reward rate distribution, the initial network recovery action with the highest expected action reward rate distribution is determined among the network recovery actions, and the difference between the preset normalized probability and the preset greedy random probability is used as the execution probability of the initial network recovery action.

[0016] The agent is invoked to select the action with the highest probability from the greedy network recovery action and the initial network recovery action, based on the preset greedy random probability and the execution probability. The agent is then invoked to determine the target node recovery action corresponding to each network node in the network system from the target network recovery action, and the network system is restored according to each target node recovery action to obtain the network recovery status of the network system.

[0017] In one embodiment, after the step of restoring the network system according to the recovery actions of each target node to obtain the network recovery state of the network system, the method further includes: calculating the action reward of the agent after performing the target network recovery action, and feeding back the action reward and the network recovery state to the agent; storing the current network state, the target network recovery action, the network recovery state, and the action reward as training samples in a preset experience replay pool through the agent; and optimizing the target classification distribution depth model through the training samples in the preset experience replay pool.

[0018] Furthermore, to achieve the above objectives, embodiments of this application provide a network recovery device, the device comprising:

[0019] The acquisition module is used to acquire the intelligent agent constructed based on the network system and to acquire the network environment interaction data of the intelligent agent. The network environment interaction data includes the network recovery actions performed by the intelligent agent under different network states of the network system, and the execution feedback results after each network recovery action is performed.

[0020] The training module is used to train a preset initial classification distribution depth model through the network environment interaction data to obtain a trained target classification distribution depth model. The target classification distribution depth model is used to determine the target network recovery action of the network system when the network is abnormal, and to control the agent to execute the target network recovery action to restore the network system.

[0021] In addition, to achieve the above objectives, this application embodiment also provides a network recovery device, which includes: a memory, a processor, and a program of the network recovery method stored in the memory and executable on the processor. When the program of the network recovery method is executed by the processor, it can implement the steps of the network recovery method as described above.

[0022] In addition, to achieve the above objectives, embodiments of this application also provide a computer-readable storage medium storing a program implementing a network recovery method, wherein when the program is executed by a processor, it implements the steps of the network recovery method as described above.

[0023] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the network recovery method described above.

[0024] One or more technical solutions proposed in this application have at least the following technical effects: This application can obtain an agent built based on a network system, thereby obtaining network environment interaction data of the agent. The network environment interaction data includes network recovery actions performed by the agent under different network states of the network system and the execution feedback results after each network recovery action is performed. Thus, the initial classification distribution depth model can be trained using each network recovery action and each execution feedback result to obtain a trained target classification distribution depth model. The target network recovery action of the network system when the network is abnormal can be determined through the target classification distribution depth model, and the agent can be controlled to perform the target network recovery action to restore the network system.

[0025] Since the target classification distribution depth model is trained on the initial classification distribution depth model by the network recovery actions executed by the agent under different network states and the execution feedback results after each network recovery action, it is convenient to continuously optimize the network recovery actions corresponding to different network states by combining the execution feedback results during the training process. This allows the network recovery actions to adapt to the corresponding network states and effectively restore the network system. In turn, the trained target classification distribution depth model can determine the target network recovery actions to be used when the network is abnormal, so that the output target network recovery actions can adapt to the network environment when the network system is abnormal. This allows the agent to be controlled to execute the target network recovery actions to restore the network system. Therefore, through the target classification distribution depth model in this application, the target network recovery actions corresponding to the network system under different network environments can be output, which facilitates the adaptation to the network recovery needs of different network environments and improves the flexibility of network recovery. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with those described herein and, together with the specification, serve to explain the principles of those embodiments.

[0027] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 This is a flowchart illustrating one embodiment of the network recovery method of this application;

[0029] Figure 2 This is a flowchart illustrating an example of a network recovery method in an embodiment of this application.

[0030] Figure 3 This is a schematic diagram of the modules of a network recovery system, which is another example of the network recovery method in the embodiments of this application.

[0031] Figure 4 This is a schematic diagram of the module structure of the network recovery device according to an embodiment of this application;

[0032] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the network recovery method in this application embodiment.

[0033] The objectives, features, and advantages of the embodiments described in this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0034] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of the embodiments of this application and are not intended to limit the embodiments of this application.

[0035] To better understand the technical solutions of the embodiments of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0036] Traditional network backup and recovery technologies rely on regularly backing up network data and configuration information, using this backup data to restore the network after an attack or failure. However, with the accelerating pace of digital transformation and the rapid development of new network technologies such as the Internet of Things (IoT), network attacks are becoming increasingly intelligent and automated. Attackers have more freedom in their attack times and methods, making network attacks more destructive and difficult to defend against. As network attacks become more complex, relying solely on backup data may not be sufficient to restore critical network functions and services in a timely and effective manner, because attacks can cause significant changes in the network environment, and backup data may not be fully adapted to the new situation.

[0037] Rule-based intrusion detection and prevention (IRP) technologies identify and block cyberattacks using pre-defined rules, thus protecting network operations to some extent. However, as cyberattack techniques evolve, new attack methods may become undetectable by existing rules, leading to defense failures. This rule-based approach lacks adaptability to changes in the network environment and struggles to cope with increasingly complex cybersecurity threats. For example, traditional backup and recovery technologies rely on fixed backup data, and rule-based detection technologies, constrained by pre-defined rules, lack flexibility in choosing recovery actions, making it difficult to make quick and effective decisions based on real-time changes in the network environment, and thus unable to adapt to the dynamic and complex changes in the modern network environment.

[0038] Therefore, this application provides a network recovery method. By constructing an agent for the network system and training an initial classification distribution depth model using network recovery actions executed by the agent under different network states and the execution feedback results after each network recovery action, a target classification distribution depth model is obtained. This allows the trained target classification distribution depth model to determine the target network recovery actions to be taken when the network is abnormal, and the output target network recovery actions can adapt to the network environment when the network system is abnormal. This enables the agent to execute the target network recovery actions to restore the network system. Furthermore, through the target classification distribution depth model in this application, the target network recovery actions corresponding to the network system under different network environments can be output, thereby facilitating the adaptation to the network recovery needs of different network environments and improving the flexibility of network recovery.

[0039] Based on this, the embodiments of this application provide a network recovery method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the network recovery method according to this application. The network recovery method includes steps S10 to S20:

[0040] Step S10: Obtain the intelligent agent constructed based on the network system, and obtain the network environment interaction data of the intelligent agent. The network environment interaction data includes the network recovery actions performed by the intelligent agent under different network states of the network system, and the execution feedback results after each network recovery action is performed.

[0041] It should be noted that a network system is a collection of multiple interconnected network nodes and communication links (such as network cables, wireless signals, and other transmission media). These components work together to achieve information transmission, sharing, and processing. A network system includes multiple network nodes, which can be devices such as computers, servers, and switches; this embodiment does not impose specific limitations on this.

[0042] An intelligent agent reflects the system framework of a network system, and includes each network node within the system. By constructing an intelligent agent based on the network system, the network system can be abstracted into a form that deep learning models can process. This provides a clear and understandable framework for deep learning models, enabling them to determine network recovery actions based on the network system reflected by the intelligent agent. For example, an intelligent agent for the network system can be constructed based on the system's network state, which reflects the node states of each network node. Network environment interaction data can include network recovery actions generated by the intelligent agent under different network states, as well as the execution feedback results after each network recovery action is performed.

[0043] Network recovery actions are defined as actions taken to restore the network system from an abnormal state to a normal state. An abnormal state refers to a network system that has been attacked or malfunctioning. When a network system is in an abnormal state, there are failed network nodes, and cascading failures may also occur. A failed network node is a node that is faulty and unable to perform its corresponding network function. A normal network state indicates that all network nodes in the system are operating normally. A cascading failure refers to a chain reaction triggered by the failure of a single network node; for example, it may cause other nodes associated with the failed node to also fail.

[0044] The feedback results include the network recovery status of the network system after performing a network recovery action, and the action reward for the agent after performing the action. Each network recovery action will have its own corresponding network recovery status and action reward. The network recovery status reflects the recovery status of the network system and can include the node status of each network node after recovery. The action reward represents the feedback to the agent after performing the network recovery action. The action reward can include positive rewards and negative rewards. A positive reward indicates that the network system's status has returned to normal or is close to normal after performing the network recovery action, while a negative reward indicates that the network performance of the network system has decreased or the network environment has not improved after performing the network recovery action.

[0045] Network environment interaction data can be obtained from a preset replay experience pool. Each time the agent performs a network recovery action, it stores the network recovery action, the corresponding action reward, the network recovery state, and the network state before the action as a training sample in the preset replay experience pool. The preset replay experience pool is used to store training samples. The network environment interaction data includes multiple sets of training samples. Each set of training samples includes the network recovery action performed by the agent in the network state, and the execution feedback result after the network recovery action. The network states corresponding to different training samples can be the same or different; this embodiment does not specifically limit this. The network environment interaction data includes training samples of different network states. When the network environment interaction data includes multiple sets of training samples, it represents the network recovery actions performed by the agent in different network states of the network system, and the execution feedback result after each network recovery action.

[0046] For example, the system network state of the network system is obtained, and an agent is constructed based on the system network state. The network recovery actions performed by the agent under different network states of the network system are obtained from a preset replay experience pool, along with the execution feedback results after each network recovery action. Generally, network recovery actions are only needed to restore the network system when the network state of the network system is abnormal.

[0047] In a feasible embodiment, the intelligent agent includes a state space and an action space; step S10 further includes steps S11 to S14:

[0048] Step S11: Obtain the system network status of the network system, wherein the system network status includes the resources to be allocated in the network system at a preset time, the node resource usage status, effective status and node threat event types of each network node in the network system at the preset time.

[0049] It should be noted that the state space can include resources to be allocated in the network system, as well as the node states corresponding to each network node. In the network system, the network state at different times is not necessarily the same; the system network state characterizes the network state at a preset time. The preset time is determined based on the actual situation; for example, the preset time can be the current time or any historical time, etc., and this embodiment does not specifically limit it. The node state of each network node includes the node resource usage status, effective status, and node threat event type.

[0050] The node resource usage status is characterized by the resource usage of a network node. The node resource usage status may include the use of multiple resource types of node sub-resources by the network node, such as network resources, computing resources, storage resources, and backup data recovery, etc. This embodiment does not make specific limitations on this.

[0051] The valid state characterizes the number of valid network nodes connected to a network node, which is the proportion of the total number of other network nodes connected to the network node.

[0052] Threat event types can reflect the types of cyberattacks suffered by network nodes. These types of cyberattacks can be analyzed through network system alarms, system logs, and network traffic logs.

[0053] The resources to be allocated represent the resources that can be allocated in the network nodes. The resources to be allocated include sub-resources to be allocated corresponding to multiple resource types. These multiple resource types can be network resources, computing resources, storage resources, network element load adjustment, authentication enhancement, backup data recovery, network parameter reconfiguration, node isolation, etc. This embodiment does not make specific limitations on these.

[0054] Step S12: Determine the resource allocation vector of the resources to be allocated, the node resource vector of each network node at a preset time, the effective vector of the effective state, and the threat event vector of the node threat event type.

[0055] Step S13: Construct the state space based on the resource allocation vector, the node resource vector of each network node at a preset time, the effective vector, and the threat event vector;

[0056] It should be noted that the resource allocation vector is a vector of resources to be allocated, and the node state of each network node can also be represented by a vector. For example, the node resource vector is a vector of the node's resource usage state, the effective vector is a vector of effective states, and the threat event vector is a vector of the node's threat event types. Constructing these vectors yields the state space. The node resource vector can include node sub-vectors representing node sub-resources corresponding to multiple resource types.

[0057] For example, the node resource vectors corresponding to each network node can be concatenated to obtain a node resource array, the effective vectors corresponding to each network node can be concatenated to obtain an effective array, and the threat event vectors can be concatenated to obtain a threat event array. Then, the node resource array, effective array, threat event array, and resource allocation vector can be concatenated to obtain the state space of the agent.

[0058] For example, the state space at a preset time t can be represented as:

[0059] S = [L(t), U(t), E(t), R(t)]

[0060] Where S represents the state space of the agent, which can represent the system network state of the network system; where L(t) is the node resource array at time t, which includes the node resource vectors corresponding to multiple network nodes; U(t) is the effective array at time t, which includes the effective vectors corresponding to multiple network nodes; E(t) is the threat event array at time t, which includes the threat event vectors corresponding to multiple network nodes; R(t) is the resource allocation vector at time t, where t is a preset time.

[0061] For example, a node resource array can be represented as:

[0062]

[0063] Where l1(t) is the node resource vector of the first network node at time t, and l1(t) to l n (t) represents the node resource vector from the first network node to the nth network node at time t, and T represents the transpose of the matrix. arrive This represents the node sub-vector of the first resource type in the node resource vector of the first network node to the node sub-vector of the i-th resource type of the first network node. Let be the node sub-vector of the i-th resource type for the n-th network node. n and i are positive integers, where i is the total number of resource types and n is the total number of network nodes. t is a preset time point. i and n can be determined based on actual conditions; this embodiment does not impose specific limitations on them. T represents the transpose of the matrix.

[0064] For example, a valid array can be represented as:

[0065]

[0066] in, k is the total number of valid network nodes connected to network node l. lThis represents the total number of all other network nodes connected to the network node. For network node l, the effective vector is... arrive The effective vector from the first network node to the nth network node is denoted as l, where n and l are positive integers and can be determined based on the actual situation. This embodiment does not impose any specific limitations on this.

[0067] For example, an array of threat events can be represented as:

[0068] E(t) = [e1(t), e2(t), ..., e n (t)] T

[0069] Where, e1(t) to e n (t) represents the threat event vector from the first network node to the nth network node at time t, and T is the transpose of the matrix.

[0070] For example, a resource allocation vector can be represented as:

[0071] R(t) = [R1(t), R2(t), ..., R i (t)]

[0072] Where R1(t) to R n (t) represents the vector of the unallocated sub-resources of the first type of resource at time t to the vector of the unallocated sub-resources of the i-th type of resource.

[0073] Step S14: Construct multiple network recovery actions that the agent can execute, and obtain the action space constructed by each network recovery action.

[0074] It should be noted that all network recovery actions supported by the agent refer to all network recovery actions that the agent can take. For example, the agent's action space can store all possible network recovery actions. Network recovery actions can be used to restore the network system.

[0075] Each network recovery action includes multiple executable node recovery actions corresponding to different network nodes, and an agent can have multiple network recovery actions. A node recovery action is characterized by allocating resources to be allocated in the network system to network nodes.

[0076] For example, all network recovery actions that the agent can execute can be constructed, resulting in an action space built from these network recovery actions. The network recovery actions in the action space can be represented as:

[0077] A = [a1(t), a2(t), ..., a n (t)]

[0078] Where A represents the network recovery action, and a1(t) to a n (t) represents the node recovery action from the 1st network node to the nth network node.

[0079] This embodiment constructs an agent by defining the state space and action space, thereby enabling the agent to more clearly reflect the state of the network system, which in turn facilitates the subsequent target classification distribution deep model to quickly understand the state of the network system.

[0080] In a feasible embodiment, the resources to be allocated include multiple resource types corresponding to sub-resources to be allocated, the network recovery action includes at least one node recovery action of a network node, and the node recovery action includes at least one sub-recovery action; step S14 further includes steps S141 to S144:

[0081] Step S141: Obtain the total node degree of all network nodes in the network system;

[0082] Step S142: For any target network node among all network nodes, obtain the target node degree of the target network node, and determine the target valid number of valid network nodes and the target invalid number of invalid network nodes among the target network node and each network node connected to the target network node.

[0083] It should be noted that the resources to be allocated include multiple resource types, each corresponding to a sub-resource to be allocated. Similarly, network recovery actions can include multiple or all node recovery actions corresponding to each network node. When constructing the agent's action space, at least one node recovery action for a network node can be constructed. When constructing a node recovery action, at least one sub-recovery action can also be constructed. However, it is best to construct node recovery actions corresponding to all network nodes, and simultaneously construct all sub-recovery actions corresponding to each network node. This allows the agent to employ various recovery actions, facilitating adaptation to different network environments and improving the efficiency and flexibility of network recovery. The resource types of the sub-recovery actions that can be allocated to each sub-recovery action of the same network node are different, and each sub-recovery action of the same network node corresponds one-to-one with its corresponding sub-resource to be allocated.

[0084] Each sub-recovery action can be used to allocate the corresponding unallocated sub-resource to a network node. For example, sub-recovery action B1 for network node A can allocate unallocated sub-resources of network resource type to network node A, and sub-recovery action B2 for network node A can allocate unallocated sub-resources of storage resource type to network node A, etc. This embodiment does not specifically limit this. Node recovery actions can be used to adjust the network resources of network nodes, adjust network element load, strengthen authentication, initiate backup and recovery, reconfigure network parameters, and isolate attacked nodes, etc. This embodiment does not specifically limit this.

[0085] The total node degree is the sum of the node degrees of all network nodes in the network system. The node degree is the number of other network nodes connecting to the target network node. The target valid number represents the total number of valid network nodes among the target network node and all network nodes connecting to the target network node. The target invalid number represents the total number of invalid network nodes among the target network node and all network nodes connecting to the target network node.

[0086] Step S143: Based on the target node degree, total node degree, and preset degree correlation coefficient, determine the target degree correlation allocation coefficient, and use the ratio of the effective number of targets to the ineffective number of targets as the target effective ratio;

[0087] Step S144: For each target sub-resource to be allocated, the product of the target sub-resource to be allocated, the target degree-related allocation coefficient, the target effective ratio, and the preset allocation coefficient is used as a sub-recovery action to allocate the target sub-resource to the target network node.

[0088] It should be noted that when constructing the action space of the intelligent agent, the action space can include the node recovery action corresponding to each network node, and the node recovery action can include the sub-recovery action corresponding to each sub-resource to be allocated. The preset degree correlation coefficient can be determined based on the actual situation, and this embodiment does not impose a specific limitation on it. The preset degree correlation coefficient is adjustable and can be selected from 0 to 1. For example, when the preset degree correlation coefficient is 0, it means that the resources to be allocated are evenly distributed among the network nodes. When the preset degree correlation coefficient is 1, it means that the allocation of the resources to be allocated is proportional to the node degree of the network node, that is, the higher the node degree of the network node, the more resources are allocated. The target sub-resource to be allocated is the sub-resource to be allocated among the resources to be allocated.

[0089] The first degree allocation coefficient is obtained by exponentiating the target node degree with the preset degree correlation coefficient. The second degree allocation coefficient is obtained by exponentiating the total node degree with the preset degree correlation coefficient. The ratio of the first degree allocation coefficient to the second degree allocation coefficient is used as the target degree correlation allocation coefficient. The preset allocation coefficient can be predetermined. For the same network node, the preset allocation coefficient may be different or the same for different types of target sub-resources to be allocated. The preset allocation coefficients for different network nodes and different resource types may be the same or different. This embodiment does not impose specific limitations on this.

[0090] For example, the sub-recovery action of the target sub-resource to be allocated can be represented as:

[0091]

[0092] Among them, a mc (t) represents the sub-recovery action that allocates resources to the m-th network node for the c-th time at time t. The m-th network node can be taken as the target network node in this embodiment. Then k m Let m be the target node degree of the m-th network node. For the first degree of allocation coefficient, For the second degree of allocation, The total node degree of the network system. Let the target degree correlation assignment coefficient be the m-th network node. For the target effective quantity, R represents the target invalid quantity. c Z(t) represents the target sub-resource to be allocated in the c-th allocation at time t. c (t) represents the allocation coefficients for allocating various types of sub-resources to the m-th network node. It is understandable that in Z... c Z(t) includes the allocation coefficients for each type of sub-resource to be allocated corresponding to the m-th network node, but when allocating sub-resources to the target network node for the c-th time, Z c (t) can take the allocation coefficient of the sub-resource to be allocated corresponding to the c-th time, that is, Z c (t) can take the preset allocation coefficient.

[0093] This embodiment considers both failed and valid network nodes when constructing sub-recovery actions. Therefore, in the event of cascading failure, it can not only recover the failed network nodes, but also protect the valid network nodes, thus realizing a prevention mechanism and protecting the valid network nodes.

[0094] Step S20: Train the preset initial classification distribution depth model through network environment interaction data to obtain the trained target classification distribution depth model. Use the target classification distribution depth model to determine the target network recovery action when the network system is abnormal, and control the agent to execute the target network recovery action to restore the network system.

[0095] It should be noted that the initial classification distribution deep model is a deep learning model, which can be a Categorical-DQN model (Categorical Distributional Deep Q-Network), also known as C51. The target classification distribution deep model is the trained Categorical-DQN model. The target classification distribution deep model can be used to determine the target network recovery actions when the network system experiences network anomalies. These target network recovery actions are the actions identified by the target classification distribution deep model for restoring the network system. Network anomalies indicate that the network system is abnormal, such as potential attacks, faults, or abnormal performance degradation. This embodiment does not specifically limit these actions. In this embodiment, the target classification distribution deep model can determine the target network recovery actions corresponding to different network anomaly states, thereby facilitating network system recovery under different network conditions. Network anomaly states represent network malfunctions in the network system.

[0096] The network system can be restored by the intelligent agent performing target network recovery actions, thus bringing the network system back to normal from an abnormal state. For example, based on the network environment recovery data, an initial classification distribution depth model can be trained to obtain a trained target classification distribution depth model.

[0097] Since the target classification distribution depth model is trained on the initial classification distribution depth model by combining the network recovery actions executed by the agent under different network states and the execution feedback results after each network recovery action, it is convenient to continuously optimize the network recovery actions corresponding to different network states during the training process by combining the execution feedback results. This allows the network recovery actions to adapt to the corresponding network states and effectively restore the network system. In turn, the trained target classification distribution depth model can determine the target network recovery actions to be used when the network is abnormal, so that the output target network recovery actions can adapt to the network environment when the network system is abnormal. This allows the agent to be controlled to execute the target network recovery actions to restore the network system. Thus, through the target classification distribution depth model in this embodiment, the target network recovery actions corresponding to the network system under different network environments can be output, thereby facilitating the adaptation to the network recovery needs of different network environments and improving the flexibility of network recovery.

[0098] In a feasible embodiment, the initial classification distribution deep model includes an initial current neural network and an initial target neural network, the target classification distribution deep model includes a current neural network and a target neural network, and the execution feedback result includes the training network recovery state of the network system and the training action reward for the agent performing the training network recovery action; step S20 further includes: steps S21 to S23:

[0099] Step S21: Determine the training network recovery action, training network recovery state, and training action reward for any training network state from the network environment interaction data. Input the training network state and training network recovery action into the initial current neural network. Output the training action reward rate distribution after executing the training network recovery action in the training network state through the initial current neural network. Each training network state can have its own corresponding execution feedback result, and the execution feedback results corresponding to different training network states are not necessarily the same.

[0100] It should be noted that the training network state can be any network state in the network environment interaction data. The training network recovery action is the network recovery action corresponding to the training network state. The training network recovery state is the state of the network system after the training network recovery action is used to restore the network system. The training action reward is the action reward determined after the agent executes the training network recovery action. The training action reward can be a positive reward or a negative reward, which can be determined based on the training network recovery state. For example, when the training network recovery state represents that the network system has returned to normal or near normal, the training action reward can be a positive reward. When the performance of the training network recovery state deteriorates relative to the training network state, then the training action reward is a negative reward. The training action reward can also include no reward, which means that the network state of the network system has not improved after the training network recovery action. For example, the performance of the training network recovery state is still consistent with the training network state. The training network recovery action can also include recovery actions for multiple training nodes.

[0101] The initial classification distribution deep model can be a Categorical-DQN model. Therefore, the initial classification distribution deep model can include an initial current neural network and an initial target neural network. The initial current neural network outputs the training action reward distribution for the training network's recovery actions, and the initial target neural network outputs the target action reward distribution. The training action reward distribution refers to the Q-value (action value) corresponding to the recovery actions of each training node in the training network after the recovery actions are executed. In other words, the training action reward is a Q-value distribution, which is the maximum Q-value distribution that the initial classification distribution deep model can output. The Q-value can be used to measure the long-term cumulative reward expectation of taking recovery actions of training nodes in the training network state. This facilitates the subsequent selection of the network recovery action with the largest Q-value distribution as the target network recovery action, thereby improving the recovery efficiency and performance of the network system. The training action reward distribution can also be represented as a Z(S, a) distribution. It can be understood that the Z(S, a) distribution is a Z distribution, where S represents the training network state and a can represent the recovery actions of each training node.

[0102] Step S22: Input the training action reward and the training network recovery state into the initial target neural network, and determine the target action reward rate distribution that the initial current neural network needs to achieve through the initial target neural network;

[0103] Step S23: Calculate the training loss between the training action reward rate distribution and the target action reward rate distribution using a preset loss function. When the training loss is less than or equal to the preset loss, obtain the current neural network and the target neural network that have been trained.

[0104] It should be noted that the target action reward distribution represents the target action reward output of the initial current neural network in its training state. The initial target neural network can be used to stabilize the training process of the initial classification distribution deep model. Since the training network recovery state reflects the actual state of the network system after the recovery action is executed, and the training action reward reflects the actual effect after the recovery action is executed, inputting the training network recovery action and training action reward into the initial target neural network allows the determination of the target action reward distribution that the initial current neural network needs to achieve. The output target action reward distribution is the true action reward distribution. This allows for more effective and stable training of the initial classification distribution deep model.

[0105] The preset loss function can be pre-set, such as the pre-set cross-entropy loss function. This embodiment does not impose specific limitations on this. The training loss represents the difference between the training action reward rate and the target action reward rate distribution. The larger the difference, the larger the training loss; the smaller the difference, the smaller the loss. The preset loss can be set based on actual conditions, and this embodiment does not impose specific limitations on this. For example, the preset loss can be a value such as 0.

[0106] For example, the training network recovery action, training network recovery state, and training action reward for any training network state are determined from the network environment interaction data. The training network state and training network recovery action are input into the initial current neural network, and the initial current neural network outputs the training action reward distribution after performing the training network recovery action in the training network state. The training action reward and training network recovery state are input into the initial target neural network, and the target action reward distribution that the initial current neural network needs to achieve is determined by the target neural network. The training loss between the training action reward distribution and the target action reward distribution is calculated using a preset loss function. If the training loss is greater than the preset loss, a new training network state, the corresponding training network recovery state, and the training action reward are obtained, and the initial classification distribution depth model is retrained. At the same time, the initial current neural network can also be updated using the training loss. If the training loss is less than or equal to the preset loss, it means that the initial classification distribution depth model has been trained successfully, and the target classification distribution depth model can be obtained. The target classification distribution depth model includes the trained current neural network and the target neural network.

[0107] During the initial training of the deep learning model for the classification distribution, the parameters of the initial current neural network can be periodically copied to the initial target neural network to maintain the stability of the initial target neural network. Additionally, it should be noted that since the Q-value distribution is discrete, it is necessary to project the target action reward distribution to match the discrete points in the training action reward distribution.

[0108] This embodiment can ensure the stability and accuracy of training during the training process, and the target classification distribution depth model obtained through training facilitates the subsequent determination of target network recovery actions, thereby improving the flexibility and efficiency of network recovery.

[0109] In a feasible embodiment, the network recovery method further includes steps X10 to X70, where steps X10 to X60 can be used to determine the reward for the training action:

[0110] Step X10: Based on the training network recovery state and the training network state, determine the reward coefficient for the agent to perform the training network recovery action;

[0111] It's important to note that the training network recovery state reflects the network system's state after executing the training network recovery action; the training network recovery state is the state of the network system after recovery. The reward coefficient measures whether the training action reward is positive, negative, or non-rewarding. The reward coefficient can be determined by comparing the training network recovery state and the training network state. A positive reward is given when the agent successfully executes the training network recovery action, restoring the network system from a faulty state to normal or near-normal, in which case the reward coefficient is greater than 0. If the training network recovery action leads to a decrease in network performance, a negative reward is given, in which case the reward coefficient is less than 0. If the training network recovery action brings no improvement and does not lead to a decrease in network performance, there is no reward, in which case the reward coefficient is equal to 0. When network performance deteriorates, the negative value of the reward coefficient can be determined based on the degree of performance degradation: the greater the performance degradation, the smaller the reward coefficient; the smaller the performance degradation, the larger the reward coefficient, but it cannot exceed 0.

[0112] When an agent successfully performs a training network recovery action, enabling the network system to recover from a fault to normal or near normal, the reward coefficient can be determined based on the degree of recovery. The higher the degree of recovery, the larger the reward coefficient, and the lower the degree of recovery, the smaller the reward coefficient, but the reward coefficient is still greater than 0.

[0113] For example, the recovery state of the training network can be compared with the initial network performance. The reward coefficient can be determined based on the difference between the recovered and initial network performance. For instance, if the initial network performance is less than the recovered network performance, the reward coefficient is positive; if the initial network performance is greater than the recovered network performance, the reward coefficient is negative. Alternatively, the degree of recovery from the training network state to the recovery state can be compared. The degree of recovery can be determined by the number of failed network nodes recovered in the training network state. For instance, if all failed network nodes are recovered, the reward coefficient is positive, and it can be a set maximum value. If only some failed network nodes are recovered, the reward coefficient can also be positive. The value of the reward coefficient can be specifically determined based on the degree of recovery; this embodiment does not impose specific limitations on this.

[0114] Step X20: For each network node in the network system, obtain the node degree of the network node and the amount of recovery resources required for the network node to recover, and use the ratio of node degree to recovery resources as the resource recovery ratio.

[0115] Step X30: For each network node, with the resource recovery ratio as the base and the preset reward priority parameter as the exponent, perform an exponentiation operation to obtain the node reward priority coefficient of the network node.

[0116] It should be noted that node degree refers to the number of other network nodes directly connected to a network node, the amount of resources restored is the amount of resources required for a network node to return to normal, and the resource recovery ratio is the ratio of node degree to the amount of resources restored. The preset reward priority parameter can be set based on the actual situation. This embodiment does not make a specific setting for it. For example, the preset reward priority parameter can be determined according to the importance of the network node and the amount of resources required. The preset reward priority parameter can be a value of 0 to 1. For example, when the preset reward priority parameter is 0, it means that the reward priority of each network node is the same. When the preset reward priority parameter is 1, it means that the reward priority is directly proportional to the node degree of the network node and inversely proportional to the amount of resources restored. That is, the higher the node degree of the network node and the smaller the amount of resources restored, the greater the reward.

[0117] Each network node has its own corresponding node reward priority coefficient, which can be used to measure the size of the reward for the network node.

[0118] For example, for each network node in a network system, the node degree and the amount of recovery resources required for the node to recover can be obtained. The ratio of node degree to recovery resources is used as the resource recovery ratio. Using the resource recovery ratio as the base and a preset reward priority parameter as the exponent, a power operation is performed to obtain the node reward priority coefficient of the network node. Alternatively, the node reward priority coefficients corresponding to each network node can be summed to obtain the total node priority coefficient.

[0119] This embodiment sets a preset reward priority parameter, which takes into account situations where multiple network nodes with lower resource requirements are more important than a single network node with higher resource requirements. This allows for the allocation of more rewards to more important network nodes with lower resource requirements, facilitating priority resource allocation to these nodes during subsequent network recovery operations. Traditional backup and recovery technologies typically use average or fixed allocation, failing to fully consider network resource utilization efficiency. This can lead to unnecessary resource waste during recovery, increasing network operating costs and management complexity. This embodiment introduces a node reward priority coefficient to optimize resource allocation strategies, differentiating allocation based on the importance and resource requirements of network nodes. This reduces resource waste during recovery and improves network resource utilization efficiency.

[0120] Step X40: Obtain the abnormal start time when the network system enters the training network state, the original network performance of the network system before the abnormal training network state, and the recovery time when the network system enters the training network recovery state.

[0121] Step X50: The difference between the recovery time and the anomaly start time is taken as the recovery time. Based on the original network performance and the network performance at each time between the anomaly start time and the recovery time, the network resilience of the network system is determined.

[0122] Step X60: Determine the training action reward based on the reward coefficient, recovery time, network elasticity, and the node reward priority coefficient corresponding to each network node.

[0123] It should be noted that the training network state can be any network state when the network system is abnormal. The abnormal start time can be the moment when the network system enters the training network state, and the recovery time refers to the moment when the network system enters the training network recovery state. The original network performance is the performance of the training network state, and the recovery time represents the time from the training network state to the training network recovery state. Network elasticity can be used to reflect the change of network performance over time from the abnormal start time to the recovery time. The time-based network performance is the network performance between the abnormal start time and the recovery time. The time-based network performance may be different or the same for different times, depending on the actual situation. This embodiment does not make specific limitations on this. For example, the sum of the reward priority coefficients of each node can be calculated to obtain the total node priority coefficient. The product of the total node priority coefficient and the reward coefficient can be calculated to obtain the initial reward. The initial reward can be used as the numerator, and the product of the network elasticity recovery time can be used as the denominator to obtain the training action reward.

[0124] For example, the formula for calculating the reward of a training action can be:

[0125]

[0126] Where R is the training action reward, r is the reward coefficient, and r can also be... e is the base of the natural logarithm, and x can be the reward factor in the reward coefficient. When the reward factor is greater than 0, the reward coefficient is positive; when the reward factor is less than 0, the reward coefficient is negative; and when the reward factor is equal to 0, the reward coefficient is 0. P is the total node priority coefficient. m P represents the node reward priority coefficient for the m-th network node, where m is a positive integer ranging from 1 to n, and n is the total number of network nodes. m Specifically, it can also be for k m Let r be the degree of the m-th network node. m To restore resource levels. To preset reward priority parameters, For network resilience, t1 is the time of the anomaly's onset, and t2 is the time of recovery. N(t) s ) for t s The network performance at any given time, t s The range is from t1 to t4, N0 is the original network performance, (t4-t1) is the recovery time. The original network performance is a constant and can be directly obtained from the network system. The network performance at each time point can also be obtained from the network system.

[0127] In this embodiment, when determining the reward for training actions, the recovery time is considered, as well as the degree of network performance recovery and network resilience. This allows training to be conducted with the goal of minimizing the recovery time and maximizing the degree of network performance recovery, thereby improving the performance of network recovery.

[0128] In a feasible embodiment, the network recovery method further includes steps A10 to A50:

[0129] Step A10: If an anomaly is detected in the current network state of the network system, the current network state is input into the target classification distribution depth model. The target classification distribution depth model is used to determine the action reward rate distribution corresponding to each network recovery action performed by the agent in the current network state.

[0130] It should be noted that the network state of the network system can be detected in real time. When an anomaly is detected in the network state, the target classification distribution depth model can be used for prediction. When an anomaly is detected in the current network state, the current network state can be converted into data that the target classification distribution depth model can recognize and input into the target classification distribution depth model. For example, the agent corresponding to the current network state can be determined. For example, the state space of the agent can be updated by the current network state so that the state space can represent the current network state. Then, the agent corresponding to the current network state can be input into the target classification distribution depth model. The target classification distribution depth model can output the action reward rate distribution corresponding to each network recovery action that the agent can perform in the current network state. Thus, the recovery effect of the network recovery action on the current network state can be determined by the action reward rate distribution, which facilitates the subsequent determination of the target network recovery action.

[0131] Step A20: Based on the preset greedy random probability, determine the greedy network recovery action among the various network recovery actions;

[0132] It should be noted that the preset greedy random probability is an ε-Greedy strategy, with a preset greedy random probability of ε. The greedy network recovery action is selected from among various network recovery actions based on the preset greedy random probability. The greedy network recovery action is selected without considering past experience.

[0133] Step A30: Based on the distribution of reward rates for each action, determine the initial network recovery action with the largest distribution of reward rates among all network recovery actions, and use the difference between the preset normalized probability and the preset greedy random probability as the execution probability of the initial network recovery action.

[0134] Step A40: Invoke the agent to select the action with the highest probability from the greedy network recovery action and the initial network recovery action by using preset greedy random probability and execution probability.

[0135] It should be noted that the preset normalization probability is 1, and the sum of the execution probability and the preset greedy random probability is the preset normalization probability. The initial network recovery action is the one with the largest action reward distribution, that is, the network recovery action with the best recovery effect. In order to avoid local optima, it is necessary to determine the target network recovery action among the greedy network recovery action and the initial network recovery action, so that the target network recovery action can better restore the network system.

[0136] For example, if the preset greedy random probability is greater than the execution probability, the greedy network recovery action is used as the target network recovery action. If the preset greedy random probability is less than the execution probability, the initial network recovery action is used as the target network recovery action. If the preset greedy random probability is equal to the execution probability, either the initial network recovery action or the greedy network recovery action is determined as the target network recovery action. The specific value of the preset greedy random probability can also be obtained after training the initial classification distribution depth model. This is because training the initial classification distribution depth model is a holistic process. For example, it requires the training network recovery actions executed by the agent, the corresponding training network state, the training network recovery state, and the training action reward. The training network recovery action that the agent needs to execute in the training network state is determined by the network recovery action with the highest reward rate in the training network state from the initial current neural network output, and the training greedy network recovery action selected based on the training greedy random probability determined during training. Therefore, each time the parameters of the initial current neural network are updated during training, the updated parameters include the training greedy random probability, thus obtaining the preset greedy random probability after training.

[0137] Step A50: Invoke the intelligent agent to determine the target node recovery action corresponding to each network node in the network system from the target network recovery action, and restore the network system according to the target node recovery action to obtain the network recovery status of the network system.

[0138] It should be noted that the target network recovery action includes target node recovery actions for each network node, and recovery is performed on the corresponding network nodes based on these target node recovery actions. The network recovery status is the network state after the target network recovery action is executed in the network system.

[0139] For example, an intelligent agent can be invoked to determine the target node recovery action corresponding to each network node in the network system from the target network recovery action. For each network node, the network node is recovered in the network system according to the target node recovery action of that network node, so as to complete the recovery of the entire network system.

[0140] This embodiment determines the initial network recovery action with the highest action reward rate through the target classification distribution depth model, thereby determining the network recovery action with the best recovery effect. In order to avoid local optima, a greedy network recovery action is also selected with a preset greedy random probability. In this way, the target network recovery action can be determined from the initial network recovery action and the greedy network recovery action, so as to truly improve the recovery efficiency of the network system and improve the flexibility of network system recovery.

[0141] Additionally, it should be noted that the target network recovery action includes target node recovery actions for each network node. The number of sub-recovery actions included in the target node recovery action may differ across different network nodes. Different network nodes are affected differently by attacks, and the types of resources that need to be recovered may also differ. Therefore, the required sub-recovery actions and their numbers may vary for different network nodes. For different sub-recovery actions within the same network node, each sub-recovery action is used to allocate different unallocated sub-resources to the network node.

[0142] In another feasible embodiment, steps A51 to A53 are included after step A50:

[0143] Step A51: Calculate the action reward of the agent after it has performed the target network recovery action, and feed back the action reward and network recovery status to the agent;

[0144] Step A52: The agent stores the current network state, the target network recovery action, the network recovery state, and the action reward as training samples into a preset experience replay pool.

[0145] Step A53: Optimize the target classification distribution depth model using training samples from the preset experience replay pool.

[0146] It should be noted that the target classification distribution depth model can be continuously optimized, thereby improving its accuracy. A pre-set experience replay pool can be used to store training samples. After each target network recovery action is performed by the agent, the target network recovery action, the corresponding network recovery state, the action reward, and the current network state before performing the action can be used as training samples. Furthermore, during optimization, multiple training samples can be randomly selected from the pre-set experience replay pool for optimization, breaking the temporal correlation between samples and improving training stability. During training, training samples can also be randomly selected from the pre-set experience replay pool for training; the network environment interaction data is extracted from this pool. The steps for calculating the action reward after the agent performs the target network recovery action are the same as those for calculating the training action reward, and will not be repeated in this embodiment.

[0147] For example, after executing a target network recovery action, the action reward can be determined, and the action reward and network recovery status can be fed back to the agent. The calculation of the action reward is not performed by the agent; it can be calculated within the network environment where the network system resides. Therefore, the action reward needs to be fed back to the agent. However, the target network recovery action operates within the network environment, so the network recovery status of the network system after the target network recovery action is executed needs to be fed back to the agent. This embodiment continuously optimizes the target classification distribution depth model, thereby facilitating more accurate output of the target network recovery action, and thus rapidly improving network recovery efficiency and flexibility.

[0148] To better understand this embodiment, please refer to Figure 2The process of restoring the network system in this embodiment is briefly described as follows: C51 is a target classification distribution depth model, Agent is an intelligent agent. Under network state S, a greedy network restoration action Ax1 is selected with a preset greedy random probability ε. Ax1 includes multiple greedy node restoration actions a1 to an, where n is the total number of network nodes. The current neural network determines the target network restoration action Ax2 with the largest output action reward distribution under network state S with an execution probability of 1-ε. argmaxZ(s,a) represents the largest action reward distribution. The target network restoration action Af can be determined from Ax1 and Ax2. Af has multiple target node restoration actions. The network environment in the network system is restored through the target network restoration action, thereby determining the network restoration state Sˋ after the network environment of the network system is restored, and determining the action reward R after the target network restoration action is executed. The network restoration state Sˋ and action reward R can be fed back to the intelligent agent. The intelligent agent can store the network state S, the target network restoration action Af, the action reward R, and the network restoration state Sˋ as training samples in a preset experience playback unit. The network recovery state Sˋ and action reward R can be output to the target neural network, which outputs the target action reward distribution y. The network state S and the target network recovery action Af can be input to the current neural network, which outputs the action reward distribution of the target network recovery action under network state S. The loss between this action reward distribution and the target action reward distribution y can then be calculated, and the learning parameters θ of the current neural network can be optimized based on this loss. Simultaneously, the learning parameters of the current neural network can be periodically copied to the target neural network, facilitating periodic synchronization between the target and current neural networks to improve the accuracy of the target classification distribution depth model. Furthermore, the network state S can be updated based on the network recovery state Sˋ, allowing the current neural network to determine the target network recovery action using the new network state.

[0149] In addition, in this embodiment, the network recovery method is applied to a network recovery system. The network recovery system may include a network modeling representation module, a model building module, an experience playback training module, a policy execution module, and a feedback module. The network modeling representation module includes a state space representation unit and an action space definition unit. The state space representation unit is used to determine the state space corresponding to the system network state of the network system, and the action space definition unit is used to determine all network recovery actions that the agent can execute. Thus, the agent can be constructed through the state space and the action space.

[0150] The model building module includes a value distribution modeling unit, a model structure unit, and a loss function determination unit. The value distribution modeling unit determines that the Categorical-DQN outputs a Q-value distribution, not a single Q-value. That is, the target classification distribution deep model outputs the Q-values ​​of each node's recovery action in the network recovery action, thus obtaining the Q-value distribution of the network recovery actions. Furthermore, the target classification distribution deep model can output the Q-value distributions corresponding to multiple network recovery actions. The model structure unit determines that the target classification distribution deep model includes the current neural network and the target neural network. The loss function determination unit determines the calculation of the training loss using a preset cross-entropy loss function.

[0151] The experience replay training module includes an experience replay storage unit and a reward determination unit. The experience replay storage unit stores training samples and is also used to randomly select training samples to train the initial classification distribution depth model and / or optimize the target classification distribution depth model to break the temporal correlation between samples and improve stability. The reward determination unit is used to calculate action rewards and can also define the training objective of the initial classification distribution depth model in this embodiment as aiming for the shortest network recovery time and the highest recovery degree.

[0152] The strategy execution module includes a strategy execution unit and a dynamic recovery unit. The strategy execution unit is used to determine the target network recovery action based on the initial network recovery action corresponding to the maximum action reward distribution of the current neural network output and the greedy network recovery action, and then execute the target network recovery action. The dynamic recovery unit is used to dynamically adjust the recovery action according to the real-time changes of the network system to improve network performance. For example, it can determine the target network recovery action corresponding to the network state of the network system in real time so that the recovery action can be adjusted in real time.

[0153] The feedback module includes a performance evaluation unit and an optimization unit. The performance evaluation unit evaluates the effectiveness of the target network recovery actions, such as the recovery degree, recovery time, and post-recovery network performance. The optimization unit optimizes the target classification distribution depth model, for example, by using training samples from a pre-defined experience replay pool to optimize the target classification distribution depth model. This enables the target classification distribution depth model to determine recovery actions for the target network under more different network states.

[0154] This application also provides a network recovery device; please refer to... Figure 5 The device includes:

[0155] The acquisition module 10 is used to acquire the intelligent agent constructed based on the network system and to acquire the network environment interaction data of the intelligent agent. The network environment interaction data includes the network recovery actions performed by the intelligent agent under different network states of the network system, and the execution feedback results after each network recovery action is performed.

[0156] The training module 20 is used to train a preset initial classification distribution depth model through network environment interaction data to obtain a trained target classification distribution depth model. The target classification distribution depth model is used to determine the target network recovery action of the network system when the network is abnormal, and to control the agent to execute the target network recovery action to restore the network system.

[0157] The network recovery apparatus provided in this application adopts the network recovery method described in the above embodiments, aiming to solve the technical problem of low flexibility in network recovery. Compared with the prior art, the beneficial effects of the network recovery method provided in this application are the same as those of the network recovery method provided in the above embodiments, and other technical features in this network recovery apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here. This application provides a network recovery device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the network recovery method in the first embodiment described above.

[0158] The following is for reference. Figure 5 The diagram illustrates a structural schematic suitable for implementing the network recovery device in the embodiments of this application. The network recovery device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The network recovery device may also be a server or similar device. Figure 5 The network recovery device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0159] like Figure 5As shown, the network recovery device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in the read-only memory 1002 or a program loaded from the storage device 1003 into the random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the network recovery device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the network recovery device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows network recovery devices with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented alternatively.

[0160] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0161] The network recovery device provided in this application, employing the network recovery method described in the above embodiments, can solve the technical problem of low flexibility in network recovery. Compared with the prior art, the beneficial effects of the network recovery device provided in this application are the same as those of the network recovery method provided in the above embodiments, and other technical features in this network recovery device are the same as those disclosed in the previous embodiment method, and will not be repeated here. It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or combinations thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0162] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0163] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to execute the network recovery method in Embodiment 1 above.

[0164] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable EPROM (Electrical Programmable Read Only Memory) or flash memory, optical fiber, portable compact disk CD-ROM (compact discread-only memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution device, apparatus, or apparatus. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof. The aforementioned computer-readable storage medium may be included in a network recovery device; or it may exist independently and not assembled into a network recovery device.

[0165] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by the network recovery device, the network recovery device: acquires an intelligent agent constructed based on the network system, and acquires network environment interaction data of the intelligent agent, wherein the network environment interaction data includes network recovery actions performed by the intelligent agent under different network states of the network system, and the execution feedback results after each network recovery action is executed; trains a preset initial classification distribution depth model through the network environment interaction data to obtain a trained target classification distribution depth model, so as to determine the target network recovery action of the network system when the network is abnormal through the target classification distribution depth model, and controls the intelligent agent to execute the target network recovery action to restore the network system.

[0166] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a LAN (local area network) or WAN (wide area network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0167] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based device that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0168] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the specific unit. The computer-readable storage medium provided in this application stores computer-readable program instructions for executing the above-described network recovery method, aiming to solve the technical problem of low flexibility in network recovery. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the network recovery method provided in the above embodiments, and will not be repeated here.

[0169] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the network recovery method described above. The computer program product provided in this application aims to solve the technical problem of low flexibility in network recovery. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the network recovery method provided in the above embodiments, and will not be repeated here.

[0170] The above are merely preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structural or procedural transformations made using the description and drawings of the present application, or direct or indirect applications in other related technical fields, are similarly included within the patent processing scope of the present application.

Claims

1. A network recovery method, characterized in that, The method includes: A smart agent constructed based on a network system is obtained, and network environment interaction data of the smart agent is obtained. The network environment interaction data includes network recovery actions performed by the smart agent under different network states of the network system, and execution feedback results after each network recovery action is performed. By using the network environment interaction data, a preset initial classification distribution depth model is trained to obtain a trained target classification distribution depth model. The target classification distribution depth model is then used to determine the target network recovery action of the network system when the network is abnormal, and the agent is controlled to execute the target network recovery action to restore the network system.

2. The network recovery method as described in claim 1, characterized in that, The intelligent agent includes a state space and an action space; the steps for obtaining the intelligent agent constructed based on the network system include: The system network status of the network system is obtained, wherein the system network status includes the resources to be allocated in the network system at a preset time, the node resource usage status, validity status and node threat event type of each network node in the network system at the preset time; Determine the resource allocation vector of the resource to be allocated, the node resource vector of each network node at the preset time of the node resource usage status, the effective vector of the effective status, and the threat event vector of the node threat event type; The state space is constructed based on the resource allocation vector, the node resource vector of each network node at the preset time, the effective vector, and the threat event vector. Multiple network recovery actions that the agent can execute are constructed to obtain the action space constructed by each of the network recovery actions.

3. The network recovery method as described in claim 2, characterized in that, The resources to be allocated include multiple resource types corresponding to sub-resources to be allocated; the network recovery action includes at least one node recovery action for a network node; and the node recovery action includes at least one sub-recovery action. The steps for constructing the agent to support all network recovery actions include: Obtain the total node degree of all network nodes in the network system; For any target network node among the network nodes, obtain the target node degree of the target network node, and determine the target valid number of valid network nodes and the target invalid number of invalid network nodes among the target network node and each network node connected to the target network node; Based on the target node degree, the total node degree, and the preset degree correlation coefficient, the target degree correlation allocation coefficient is determined, and the ratio of the target effective quantity to the target ineffective quantity is taken as the target effective ratio; For each target sub-resource to be allocated, the product of the target sub-resource to be allocated, the target degree-related allocation coefficient, the target effective ratio, and the preset allocation coefficient is used as a sub-recovery action to allocate the target sub-resource to the target network node.

4. The network recovery method as described in claim 1, characterized in that, The initial classification distribution depth model includes an initial current neural network and an initial target neural network, the target classification distribution depth model includes a current neural network and a target neural network, and the execution feedback result includes the training network recovery state of the network system and the training action reward for the agent to perform the training network recovery action; The step of training a preset initial classification distribution depth model using data exchanged through the network environment to obtain a trained target classification distribution depth model includes: The training network recovery action, training network recovery state, and training action reward for any training network state are determined from the network environment interaction data. The training network state and the training network recovery action are input into the initial current neural network. The initial current neural network outputs the training action reward rate distribution after the training network recovery action is performed in the training network state. The training action reward and the training network recovery state are input into the initial target neural network, and the target action reward rate distribution that the initial current neural network needs to achieve is determined through the initial target neural network. The training loss between the training action reward rate distribution and the target action reward rate distribution is calculated using a preset loss function. When the training loss is less than or equal to the preset loss, the current neural network and the target neural network that have been trained are obtained.

5. The network recovery method as described in claim 4, characterized in that, The method further includes: Based on the training network recovery state and the training network state, determine the reward coefficient for the agent to perform the training network recovery action; For each network node in the network system, obtain the node degree of the network node and the amount of recovery resources required for the network node to recover, and use the ratio of the node degree to the amount of recovery resources as the resource recovery ratio; For each network node, the node reward priority coefficient is obtained by exponentiation with the resource recovery ratio as the base and the preset reward priority parameter as the exponent. The abnormal start time when the network system enters the training network state, the original network performance of the network system before the abnormal training network state, and the recovery time when the network system enters the training network recovery state are obtained. The difference between the recovery time and the anomaly start time is used as the recovery duration, and the network resilience of the network system is determined based on the original network performance and the network performance at each time between the anomaly start time and the recovery time. The training action reward is determined based on the reward coefficient, the recovery time, the network elasticity, and the node reward priority coefficient corresponding to each of the network nodes.

6. The network recovery method as described in claim 1, characterized in that, The method further includes: If an anomaly is detected in the current network state of the network system, the current network state is input into the target classification distribution depth model. The target classification distribution depth model determines the action reward rate distribution corresponding to each network recovery action performed by the agent in the current network state. Based on a preset greedy random probability, a greedy network recovery action is determined among the various network recovery actions; Based on the action reward rate distributions, the initial network recovery action with the highest expected action reward rate distribution is determined among the network recovery actions, and the difference between the preset normalized probability and the preset greedy random probability is used as the execution probability of the initial network recovery action. The agent is invoked to select the action with the highest probability from the greedy network recovery action and the initial network recovery action, based on the preset greedy random probability and the execution probability; The intelligent agent is invoked to determine the target node recovery actions corresponding to each network node in the network system from the target network recovery actions, and the network system is restored according to each target node recovery action to obtain the network recovery status of the network system.

7. The network recovery method as described in claim 6, characterized in that, After the step of restoring the network system according to the recovery actions of each target node to obtain the network recovery status of the network system, the method further includes: Calculate the action reward of the agent after it has performed the target network recovery action, and feed back the action reward and the network recovery status to the agent. The agent stores the current network state, the target network recovery action, the network recovery state, and the action reward as training samples in a preset experience replay pool. The target classification distribution depth model is optimized using training samples from the preset experience replay pool.

8. A network recovery device, characterized in that, The network recovery device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of the network recovery method according to any one of claims 1 to 7.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and the computer-readable storage medium stores a program that implements the network recovery method, the program that implements the network recovery method being executed by a processor to implement the steps of the network recovery method as described in any one of claims 1 to 7.

10. A program product, characterized in that, The program product is a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the network recovery method as described in any one of claims 1 to 7.