Distributed large-scale anti-traceability elastic network intelligent routing method and system

Through the distributed proxy architecture and the QMIX multi-agent collaborative learning framework, combined with the lightweight graph attention mechanism, scalability bottlenecks and anonymity defects in large-scale networks are solved, efficient and anonymous routing strategies are realized, and transmission efficiency and resource allocation balance are improved.

CN120528852AInactive Publication Date: 2025-08-22BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510748137.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing technology faces the scalability bottlenecks under dynamic topology and hyper-large-scale networks, the anonymity defects caused by centralized dependence and cross-session associations, and the lack of global optimization of the distributed inter-agent collaboration mechanism in large-scale networks, resulting in low transmission efficiency, insufficient anonymity and imbalance in resource allocation.

Method used

Adopting a distributed large-scale anti-traceability elastic network intelligent routing method, by building a multi-agent distributed proxy architecture, combining a lightweight graph attention mechanism and a QMIX multi-agent collaborative learning framework, the nonlinear fusion of local Q values ​​and global graph states is achieved, dynamic topology perception and anti-traceability mechanism are supported, and an action entropy-driven adaptive exploration mechanism is designed, and rapid model fine-tuning when topology changes dramatically.

Benefits of technology

It realizes efficient routing strategies in a dynamic topological network environment, blocks path-associated attacks, ensures anonymity and resource allocation balance, and improves transmission efficiency and network elastic adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120528852A_ABST
    Figure CN120528852A_ABST
Patent Text Reader

Abstract

The invention provides a distributed large-scale anti-traceability elastic network intelligent routing method and system, belongs to the field of network communication and network security, and is suitable for intelligent routing decision optimization in a dynamic network environment. The method is based on a multi-agent reinforcement learning framework, network nodes are mapped into independent agents, and path traceability risks are blocked through local information constraints; dynamic feature aggregation of a neighbor link state is realized by adopting a lightweight graph attention network, and the local sensing efficiency of a large-scale network is improved; a QMIX algorithm is introduced, and network parameters are optimized by nonlinear fusion of a local Q value and global graph state representation through a hybrid network; and in combination with a self-adaptive exploration mechanism driven by action entropy, the sudden change scene strategy response capability is enhanced. According to the system, in military anonymous communication, dark network data transmission and cross-border sensitive services, the anti-traceability and transmission efficiency balance can be remarkably improved, the characteristics of high concealment, high elasticity and low resource consumption are achieved, and a systematic routing solution is provided for a dynamic network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of network communication technology. Specifically, the present invention relates to a distributed large-scale reverse tracing elastic network intelligent routing method and system. Background Art

[0002] With the rapid development of the internet, network scale continues to expand, the number of nodes increases dramatically, and network topologies become increasingly complex. Traditional centralized routing mechanisms, when used in large-scale networks, suffer from problems such as high routing latency, low bandwidth utilization, and a high risk of single points of failure. Furthermore, while existing anonymous communication networks utilize dynamic circuit mechanisms, paths remain fixed within a single session, and entry node selection strategies suffer from long-term immutability. This leads to the accumulation of traffic fingerprints and the risk of cross-circuit timing correlation attacks, making true dynamic reverse tracing difficult.

[0003] Existing distributed routing algorithms mitigate the scalability issues of centralized control through local decision-making. However, in dynamic topologies and large-scale node scenarios, they still face challenges such as frequent routing oscillations and a lack of global coordination in path optimization. Distributed Q-learning routing methods rely on nodes to independently learn forwarding strategies, leading to the accumulation of local optimal solutions and difficulty balancing network load and end-to-end latency. In the field of anonymous communication, the mainstream Tor network uses multi-layer relay encryption to hide communication endpoints. However, its fixed entry nodes and periodic rerouting mechanism still reveal user identities due to traffic timing characteristics. Attackers can use long-term traffic fingerprints to correlate the entry nodes of different sessions and conduct cross-circuit tracing attacks. Although existing dynamic routing anonymity technologies use randomized path selection, path generation relies on a centralized coordinator, which poses a single point of failure risk and does not address the dynamic coordinated optimization of multi-hop paths in large-scale networks. Furthermore, routing methods based on graph neural networks (GNNs) improve routing decision efficiency by modeling network topology relationships. However, traditional GNNs require message passing across the entire graph topology, resulting in computational complexity that increases exponentially with the number of nodes, making them unsuitable for ultra-large-scale networks.

[0004] Currently, distributed, large-scale, anti-traceable, elastic network routing technology faces the following challenges: First, scalability bottlenecks in dynamic topologies and ultra-large-scale networks. Traditional centralized / semi-centralized architectures cannot efficiently support dynamic node scaling. Furthermore, GNN routing methods based on global graph computations face exponential computational complexity, making them difficult to adapt to the real-time changes in ultra-large-scale network topologies. Second, anonymity flaws stemming from centralized reliance and cross-session associations. Existing anonymous routing mechanisms, due to fixed entry node selection, centralized path coordination, and exposed traffic timing characteristics, face the risk of cross-circuit traceability attacks and single points of failure, making it difficult to achieve dynamic anti-traceability goals. Third, the lack of global optimization of the coordination mechanism between distributed agents makes it difficult to coordinate local routing strategies with global network goals, leading to unbalanced resource allocation and suboptimal path optimization. Therefore, existing technologies cannot effectively support the combined requirements of efficient routing, elastic adaptation, and anti-traceability in dynamic, anonymous networks. An innovative distributed routing approach is urgently needed. Therefore, a multi-agent reinforcement learning framework combining QMIX and the Neighborhood Sampling Graph Attention Network (NSGAT) was introduced into distributed routing research. Local subgraph sampling was used to reduce the complexity of large-scale topological calculations and adapt to dynamic scaling scenarios. A distributed agent architecture and local information decision-making mechanism were designed to block path associations and achieve lightweight anti-traceability. The local Q value and the global graph state were dynamically integrated through the hybrid network to collaboratively optimize end-to-end latency and throughput, breaking through the global collaboration bottleneck of distributed agents. Summary of the Invention

[0005] In view of this, the present invention proposes a distributed large-scale anti-traceability elastic network intelligent routing method and system, aiming to solve the problems of low transmission efficiency, insufficient anonymity and unbalanced resource allocation caused by node scale expansion, path correlation attack threats and suboptimal local decision-making in a dynamic topology network environment.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A distributed large-scale anti-traceability elastic network intelligent routing method and system, through the design of distributed elastic architecture, maps network nodes into independent agents, makes decisions based on local state, and blocks path-related traceability attacks; combines a lightweight graph attention mechanism to dynamically aggregate neighbor link states to achieve local topology awareness; adopts the QMIX multi-agent collaborative learning framework, dynamically generates weight matrices through hybrid networks, and nonlinearly integrates local Q values ​​into global optimization goals to ensure that routing strategies take into account low latency, high throughput, and load balancing; introduces an action entropy-driven adaptive exploration mechanism to support rapid model fine-tuning when topology changes drastically, ensuring service continuity. The routing method includes the following steps:

[0008] Step 1: Build a multi-agent distributed anti-traceability elastic network model. Set up an agent for each network node. Agents communicate with each other through directly adjacent agents and exchange network status information. Distributed agents only know the previous and next hop information of the data and have anti-traceability capabilities.

[0009] Step 2: Message communication between adjacent agents is implemented using a graph attention network. Through neighbor sampling, full-graph sampling is optimized to small-batch neighbor sampling centered on some nodes. Using the attention mechanism to distinguish the importance of neighbors, GAT is adaptable to large-scale network scenarios.

[0010] Step 3: Introduce the QMIX algorithm to implement collaborative learning of agent intelligent routing. QMIX improves on the Deep Q-Network (DQN) and processes the joint action value function in a multi-agent environment through the value function decomposition method. Each network node acts as an independent agent and calculates a local Q value based on its own state and neighbor messages obtained by GAT through a local Q network to represent its local utility for routing decisions. The QMIX hybrid network dynamically integrates the local Q values ​​of all nodes through a supernetwork. Combined with the global graph state representation constructed by GAT, it nonlinearly generates a global Q value to optimize network objectives such as end-to-end delay and throughput.

[0011] Step 4: Training is divided into three steps: distributed experience collection, centralized training, and distributed execution. Each node selects an action based on an improved entropy-driven adaptive exploration mechanism, executes routing decisions, and observes the utilization of local reward links. Through periodic exchange of messages through GAT, an approximation of the local state s and the global state S is constructed. The hybrid network calculates the global Q value based on the global state, and optimizes the parameters of the local Q network and the hybrid network through TD error, with the goal of minimizing the prediction error of the global Q value. After training is completed, each node selects an action based solely on the local Q value and neighbor messages obtained by GAT, without the participation of the hybrid network. The distributed reverse-source elastic network intelligent routing gives the best next hop for the network node, and after combining them, the routing sequence of the service flow is obtained.

[0012] Furthermore, the distributed elastic network model constructed in step 1 is specifically:

[0013] A one-to-one mapping tightly coupled agent architecture is proposed to model the distributed large-scale reverse traceability elastic network environment as a graph. Represents the Anti-Traceability device set and represents the terminal link set, Represents the agent set, and the device set One-to-one mapping, the agent can only obtain local status through the bound device, avoiding global topology exposure and blocking tracing attacks based on path association. and Represent the global state and action space respectively. Assume that there are N agents in the system and the local observation state of each agent v is represented by s v , then the global state space can be defined as the Cartesian product of these local states. Specifically, if the local observation state s of each agent v is v From their respective local state space S v , define the global state space is the Cartesian product of the local state space of each agent The local observation state s v Contains device node d v queue depth, link bandwidth utilization, and global action space is the local action a of each agent v ∈A v Cartesian product of The local action a v Select the next routing path;

[0014] Furthermore, the message communication between adjacent agents in step 2 realizes cross-agent collaborative decision-making through the GAT functional component, and each agent The internally deployed GAT module includes the following core functional units:

[0015] Dynamic neighbor sampling and feature extraction: Each agent extracts local information based on its observations. Each agent then shares these features with neighboring agents in the graph and combines the received neighbor features with its own observations to enhance decision-making capabilities. A dynamic neighbor random sampling strategy is proposed to avoid the traditional GAT requiring all nodes in the graph to participate in the calculation. Combining the link delay changes, bandwidth utilization, and load balancing of the network topology, only some high-value neighbors are aggregated. Each agent only extracts some high-value neighbor nodes from the distributed large-scale reverse traceability elastic network instead of processing the entire graph data, reducing the computational load of distributed nodes and adapting to the dynamic nature of large-scale networks. Through the above-mentioned lightweight feature extraction unit, the agent encodes the queue load, adjacency table, bandwidth, and delay into feature vectors as the input basis of the GAT.

[0016] By combining the attention mechanism with a distributed, large-scale, and resilient back-tracing network through attention weight aggregation, the agent distinguishes the importance of different neighbors. GAT assigns higher attention weights to neighboring nodes with high link quality and low load. This process does not require preset rules, but is dynamically adjusted through learnable parameters. The agent aggregates the feature vectors of neighboring nodes according to the attention weights. The aggregated information contains the context required for global collaboration.

[0017] The above design enables the GAT component to extract network topology features while reducing communication overhead. The attention mechanism enables the system to flexibly respond to the dynamic changes of distributed large-scale back-traceable elastic networks, while the modular functional units facilitate distributed deployment and expansion, ultimately supporting an efficient, hidden, and highly available elastic network.

[0018] Furthermore, the collaborative learning mechanism of distributed elastic routing in step 3 uses the QMIX algorithm to implement multi-agent value function decomposition and global strategy optimization, which specifically includes the following technical features:

[0019] Distributed local Q network architecture, for each agent Deploy a local Q network v (s v ,a v θ v ), whose input is the hidden state h generated by the GAT module containing aggregated neighbor messages v (K) and local state representation of link indicators s v With optional action a v ∈A v , the output is the local Q value of the corresponding action, which represents the expected future reward of forwarding to the next hop in the current state of the local network;

[0020] Design the QMIX network to give different contribution ratios to different agents, integrate the hybrid network with each local agent network, and the hybrid network Q tot (s,a;ψ) generates a dynamic weight matrix W(s;ψ w ) and the bias vector b(s;ψ b ), the local Q value Q of each agent v The nonlinear combination is the global Q value, and each agent’s local Q value Q v Satisfy the monotonicity constraint The specific form is Q tot =W(s)·σ([Q1,Q2,...,Q |V| ])+b(s), where σ is the Exponential Linear Unit (ELU) activation function, W(s)∈R 1×|V| and b(s)∈R is generated by the global graph state representation s constructed based on GAT, where s is the node-level aggregate feature;

[0021] Distributed training mechanism, using the Centralized Training and Decentralized Execution (CTDE) framework, during the training phase, each agent v regularly uploads local experience data (s v ,a v ,r v,s v ′) to the regional coordinator, which is based on the global reward r tot =α·Δt+β·Θ-γ·L, calculate the temporal difference (TD) error δ=r tot +γ·(max a 'Q tot (s′,a′;ψ target )-Q tot (s, a; ψ)), where Δt represents the end-to-end delay, Θ represents the throughput, and L represents the packet loss rate. Finally, the local Q network parameters θ are synchronously updated through backpropagation. v and the mixed network parameters ψ, while freezing the target network parameters ψ target Training with stability;

[0022] Furthermore, the distributed experience collection described in step 4 includes the following steps:

[0023] Design an action entropy driven adaptive annealing mechanism to calculate the information entropy of the historical action distribution of each agent v Each agent Based on the exploration rate ε v Dynamic annealing adjustment based on node historical action entropy If the action exploration of agent v is sufficient and the entropy is high, ε is quickly reduced to give priority to the high reward path; if the action exploration is insufficient and the entropy is low, the decay rate of ε is slowed down to encourage diverse exploration;

[0024] Agent v performs the selected action a v After that, the state is transferred to s v ′, by collecting local immediate rewards r v =w1·(1-U)+w2·(1-Q)-w3·Δt, where U represents link utilization, Q represents queue occupancy, and Δt represents link delay. Using the GAT message passing mechanism, with T as the period, the hidden state differential is exchanged with the neighboring agent. Message, build local state Each agent v calculates the global graph feature vector based on the aggregated neighbor messages Traditional empirical data cannot explicitly reflect the real-time state changes of neighbor nodes such as bandwidth drops and proxy nodes joining / exiting. Therefore, it is considered to add the hidden state difference Δh of neighbor nodes in GAT multi-layer propagation to the empirical tuple. b(k) To perceive the sudden change in bandwidth utilization and delay jitter caused by link quality fluctuations, dynamic changes in topology such as neighbor nodes going offline or new proxy nodes being added, and changes in neighbor queue load trends caused by traffic pattern changes, the regional coordinator adopts an asynchronous parallel gradient update mechanism to receive the experience tuples uploaded by each proxy (s v ,a v ,r v ,S′,Δh b (k),b∈B(v)), build a global experience pool D=(S,a v ,R tot ,S′), where the global reward r v represents the local immediate reward of agent v, Θ global represents the global throughput, represents the variance of end-to-end delay;

[0025] Furthermore, the centralized training described in step 4 includes the following steps:

[0026] During the training and learning process, the QMIX hybrid network nonlinearly merges the local Q functions of the single agents. The hybrid weights are dynamically generated by the super network and assisted by the global state information. The hybrid network is conditioned on the global state S and generates a dynamic weight matrix W(S)∈R through the super network. d×|V| With the bias vector b(S)∈R d , the QMIX algorithm converts the local Q value Q v (s v ,a v )The nonlinear combination is the global Q value Q tot =W(S)·tanh([Q1,Q2,...,Q |V| ])+b(S), and calculate the time series difference target y=R tot +γ·max a 'Q tot (S′,a′;ψ target ), where ψ target is the target hybrid network parameter;

[0027] Minimize the loss function through a dual-delayed deep deterministic policy gradient framework Where β is the regularization coefficient, which suppresses the local Q value deviation. At the same time, the parameter delay synchronization strategy is adopted. θ Step synchronization of the local Q network parameters θ v To all agents;

[0028] Furthermore, the distributed execution described in step 4 includes the following steps:

[0029] Each agent v deploys a local Q network and GAT module with frozen parameters, and only relies on local observations s during execution v Neighbor messages aggregated with GAT, through greedy strategy Generate routing decisions. The routing sequence of the service flow is generated by an iterative path construction algorithm. The source node agent selects the optimal next hop node based on the Q value. The relay node agent recursively executes the same decision until it reaches the destination node, forming a complete routing path P = v0 → v1 → ... → v d , traditional methods require global retraining when the performance degrades due to topology changes. This method uses a path feedback mechanism to backpropagate routing delay and packet loss rate indicators to path nodes, triggering incremental fine-tuning of the local Q network, updating the model without interrupting service, and supporting online adaptive updates. When it is detected that the number of added / removed nodes exceeds ΔN, resulting in a drastic change in the topology structure, the federated learning protocol is started, and each agent calculates the gradient based on local experience Update the global model through multi-party aggregation;

[0030] Furthermore, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the distributed large-scale reverse tracing elastic network intelligent routing method and system as described in any one of the above.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] First, this paper designs a dynamic topology perception and anti-tracing mechanism based on a lightweight graph attention network. It uses a dynamic neighbor sampling strategy to select highly stable and anonymous links, combined with multi-head attention feature aggregation, to perceive network state changes in real time and block traceability attacks based on path association.

[0033] Secondly, the present invention proposes a multi-agent collaborative decision-making architecture based on QMIX, which dynamically fuses local Q values ​​through a hybrid network, constructs a global state, and models routing decisions as a multi-agent collaborative optimization problem, while ensuring end-to-end delay and high throughput, and achieving optimization of load balancing and link congestion rate.

[0034] Furthermore, the present invention proposes a dynamic empirical collaboration mechanism based on hidden state differential perception. By recording the hidden state differences of neighboring nodes in multi-layer GAT propagation, it constructs a multidimensional feature encoding that encompasses link quality fluctuations, topology changes, and traffic pattern changes. Combined with a dual-delay deep deterministic policy gradient framework, it minimizes the temporal difference loss function, suppresses local Q-value deviations, and ensures consistency between distributed decision-making and global policy.

[0035] Finally, based on the practical needs of large-scale dynamic networks, this paper proposes an online adaptive update and elastic collaboration mechanism. By deeply integrating local agents with global federated learning, it supports dynamic optimization of routing decisions and real-time incremental updates of models. When network topology is restructured on a large scale, distributed federated collaboration enables secure aggregation of multi-node experience and model synchronization, endowing the network with the dual capabilities of dynamic self-healing and covert collaboration. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. In the drawings:

[0037] Figure 1 The figure is a schematic diagram of the steps of a distributed large-scale reverse tracing elastic network intelligent routing method in one embodiment of the present invention.

[0038] Figure 2 Schematic diagram of the various component units of the GAT module of the network intelligent body in one embodiment of the present invention.

[0039] Figure 3 Schematic diagram of the network structure of the QMIX model in one embodiment of the present invention.

[0040] Figure 4 Schematic diagram of a training scenario under the CTDE framework in one embodiment of the present invention.

[0041] Figure 5 This is a diagram of the model architecture of centralized training in one embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0043] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0044] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0045] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0046] In order to solve the problems of low transmission efficiency, insufficient anonymity and imbalanced resource allocation caused by node scale expansion, path-related attack threats and suboptimal local decision-making in dynamic topology network environments, the present invention provides a distributed large-scale reverse tracing elastic network intelligent routing method and system, such as Figure 1 As shown, the method includes the following steps S101 to S104:

[0047] Step S101: Construct a distributed elastic network model based on a tightly coupled proxy architecture to achieve information isolation and dynamic expansion.

[0048] Step S102: Construct a message communication mechanism based on GAT, integrate multi-dimensional feature encoding and dynamic neighbor sampling strategy, and achieve covert collaborative decision-making with low communication overhead between agents through attention weight aggregation.

[0049] Step S103: Introduce the QMIX algorithm to realize the decomposition of multi-agent value function, and perform nonlinear combination and global optimization of local Q values.

[0050] Step S104: Design a global optimization mechanism for the CTDE framework, and jointly train local and hybrid network parameters through TD error to achieve distributed back-tracing elastic routing decisions.

[0051] In step S101, in order to build a distributed network model based on a tightly coupled agent architecture, the network environment is first modeled as a graph structure. in Represents the set of anti-traceability terminal devices. is the set of communication links between devices, It is a set of proxies that are strictly mapped one-to-one with terminal devices.

[0052] Furthermore, each proxy only obtains local observation state through its bound terminal devices and independently selects the next-hop routing path within its local action space. This design fundamentally blocks path-association-based tracing attacks through information isolation and topology hiding, while also supporting network elastic expansion and dynamic adaptation.

[0053] Specifically, each agent Only the corresponding device can be accessed Real-time status information, including the queue depth of data to be forwarded by the device, the bandwidth utilization of adjacent links, etc. Agents do not share global topology or routing path information, and only interact with each other through local observation status and neighbor messages.

[0054] Global state space It is formed by the Cartesian product of the local state spaces of all agents However, the agent cannot directly observe the global state and can only observe the state based on local information sv And the implicit features of neighboring nodes make decisions. Global action space It is defined as the Cartesian product of each agent's local action space Each agent's local action a v ∈A v Indicates the next-hop path selection starting from the current device.

[0055] By constructing this network model, a dual isolation mechanism is implemented: on the one hand, the agent only relies on local information to generate actions, avoiding the exposure of the global topology. Even if some nodes are controlled by attackers, the data source cannot be inferred by path backtracking; on the other hand, the Cartesian product form of the global state and action space disperses the decision-making process to each agent for independent execution, supporting the dynamic joining or exit of nodes.

[0056] In step S102, a lightweight message communication mechanism based on graph attention network is designed to achieve efficient collaborative decision-making between adjacent agents. The internally deployed GAT module consists of three functional units: lightweight feature encoding, attention weight aggregation, and dynamic neighbor sampling. Its core goal is to extract network topology features with low communication overhead and high computing efficiency in a large-scale dynamic network environment, supporting the flexibility and confidentiality of distributed routing decisions.

[0057] like Figure 2 The figure shows a schematic diagram of the components of the GAT module of agent v.

[0058] The lightweight feature encoding unit is responsible for converting the local network state into a vector form that can be processed by GAT. v Three types of indicators are collected: queue load reflects the queue depth of the device's data packets to be forwarded, directly representing the instantaneous congestion level of the node; the adjacency table records the basic attributes of the neighboring nodes directly connected to the device and the corresponding links, which is used to build local topological relationships; the real-time link status dynamically captures network indicators such as bandwidth utilization, latency, and packet loss rate, and characterizes the real-time communication quality of the link. These raw indicators are normalized and nonlinearly mapped through the embedding layer, and finally encoded into a feature vector h of fixed dimension d. v ∈R d , as the input basis of GAT.

[0059] The attention weight aggregation unit is the core of the GAT module to achieve intelligent collaboration. Its design goal is to dynamically perceive the importance of neighboring nodes and fuse local states with global context to support efficient and adaptive routing decisions. In specific implementation, each agent v generates its own feature vector h through a lightweight feature encoding unit. v ∈Rd , and obtain the feature vector {h u |u∈N(v)}. After normalization to eliminate dimensional differences, the feature vectors of the agent itself and its neighbors are input into the shared weight matrix W∈R d ' ×d , project it into a unified latent space: h′ v =Wh v ,h u ′=Wh u Extract high-level features related to routing decisions and reduce model complexity through parameter sharing.

[0060] Then, for each pair of nodes (v,u), the projection features [h′] of the agent itself and its neighbors are concatenated. v ||h u ′], and use the attention parameter vector a∈R 2d′ Calculate the original attention coefficient: e vu =LeakyReLU(a T [h′ v ||h u ′]), where LeakyReLU is used to introduce nonlinearity to avoid gradient disappearance. Attention coefficient e vu Reflects the potential importance of neighbor u to the decision of agent v.

[0061] Next, the attention coefficient of the neighbor set N(v) is normalized by the Softmax function to obtain an interpretable attention weight:

[0062] Finally, the aggregated features h of agent v v ″Generated by weighted summation: h v ″=σ(∑ u∈N(v) α vu ·h′ u ). Among them, σ is the ELU activation function, which is used to enhance the nonlinear expression ability.

[0063] This process enables the agent to adaptively focus on high-value neighbors without relying on preset rules. When the bandwidth utilization of a link suddenly increases, the attention weight of its corresponding neighbors will automatically decrease, thus guiding the routing decision to avoid congested areas. The aggregated feature h v It contains both local observation information and implicit global collaborative context, providing input for subsequent Q-value calculation.

[0064] The dynamic neighbor sampling unit is the key to the GAT module's adaptation to the dynamics of the elastic network. In specific implementation, each agent dynamically selects some high-value neighbor nodes (nodes connected by links with low latency, sufficient bandwidth, and light load) based on the real-time status of its bound terminal devices, and only performs feature aggregation on these nodes. In specific implementation, the agent periodically calculates the comprehensive score s of the adjacent links. uv =α·U -1 +β·D -1 +γ·LB (α, β, and γ are adjustable weight coefficients), where U represents bandwidth utilization, D represents latency, and LB represents load balancing. Based on the overall score, the top k neighbors with the highest scores are selected for calculation. This strategy significantly reduces resource consumption among distributed nodes while ensuring that the characteristics of key links are captured first.

[0065] In step S103, collaborative learning is completed through the QMIX algorithm to achieve multi-agent value function decomposition and global strategy optimization. Figure 3 Figure 2 shows a schematic diagram of the network structure of the QMIX model.

[0066] During execution, each agent v∈V proxy Deploy an independent local Q network v (s v ,a v θ v ), the input consists of three parts: one is the hidden state h generated by the GAT module v (K) aggregates the dynamic features of neighbor nodes through a multi-layer graph attention network to form a contextual representation containing cross-agent collaboration information; second, the local observation state s v , local network indicators collected in real time by the terminal device bound to the agent, including device node d v The queue depth, link bandwidth utilization, current link load rate, etc.; the third is optional action a v ∈A v , the set of next-hop routing path actions that the agent can choose. The output of the local Q network is each optional action a in the current state. v The local Q value represents the expected value of the future cumulative reward that can be obtained by the next hop forwarding after selecting this action in the current local network state.

[0067] Furthermore, in order to achieve global collaborative optimization, QMIX uses a hybrid network Q tot (s,a;ψ) the local Q value Q v The nonlinear combination is the global Q value. The super network is introduced to generate the dynamic weight matrix W(s;ψ w ) and the bias vector b(s;ψ b ), and adjust the combination strategy based on the global graph state representation s constructed by GAT.

[0068] Among them, the calculation form of the hybrid network is: Q tot =W(s)·σ([Q1,Q2,...,Q |V| ])+b(s), weight matrix W(s)∈R 1×|V| And the bias b(s)∈R is dynamically generated through the global state s. This design satisfies the monotonicity constraint That is, the global Q value Q tot With any local Q value Q v It increases monotonically with the increase of , thus ensuring that the local optimization direction is consistent with the global goal.

[0069] In step S104, the training is divided into three steps: distributed experience collection, centralized training, and distributed execution. The TD error is used to jointly train local and hybrid network parameters to achieve distributed back-tracing elastic routing decisions. Figure 4 The following is a schematic diagram of the training scenario under the CTDE framework.

[0070] The distributed experience collection described above breaks through the static limitations of traditional experience sampling through an adaptive annealing strategy driven by action entropy, and provides global collaborative data with high information density for subsequent centralized training.

[0071] Specifically, each agent v calculates the information entropy of its historical action distribution in real time Quantifying the diversity of action selection.

[0072] In some embodiments, if the agent chooses the same link to forward data packets for a long time, its action distribution approaches a unimodal distribution, and the entropy value Significantly reduced, at this time the system dynamically adjusts the exploration rate Slowing the decay of ε forces agents to try other links to discover potential low-latency paths. Conversely, if agent actions are evenly distributed, the exploration rate is rapidly reduced, prioritizing the use of verified high-Q paths. This mechanism, through an entropy feedback loop, achieves an autonomous balance between exploration and exploitation, making it particularly suitable for dynamic network environments with frequently fluctuating link states.

[0073] Agent v performs action a v After that, through the local immediate reward function r v =w1·(1-U)+w2·(1-Q)-w3·Δt to quantify the decision effect, where U is the link utilization, Q is the queue occupancy, and Δt is the link delay. At the same time, the agent exchanges hidden state differences with its neighbors through the GAT message passing mechanism at a period of T. Building Enhanced Local State Computing global graph feature vectors based on aggregated neighbor messages in, is the node hidden state after K-layer graph attention aggregation, encoding the dynamic characteristics of neighbor links such as load and delay; The impact of link quality mutation on state representation is explicitly captured by weighted delay information Δt.

[0074] The experience tuple additionally records the hidden state difference Δh of neighbor nodes in each layer of GAT b (k) is used to perceive bandwidth utilization mutations and delay jitter caused by link quality fluctuations, dynamic topology changes such as neighbor node offline or new proxy nodes, and changes in neighbor queue load trends caused by traffic pattern changes.

[0075] Furthermore, the regional coordinator adopts an asynchronous parallel gradient update mechanism to receive the extended experience tuples (s v ,a v ,r v ,S′,Δh b (k),b∈B(v)), build a global experience pool D=(S,a v ,R tot ,S′). Global reward Comprehensive local reward r v , global throughput Θ global and end-to-end delay variance Force the policy model to prioritize path stability during training rather than simply pursuing local low latency.

[0076] In the centralized training, each agent v periodically collects local experience data, including the current state s v , perform action a v 、Instant Reward v and the new state s after the transfer v ′, and the quaternion (s v ,a v ,r v ,s v ′) is uploaded to the regional coordinator. The coordinator calculates the comprehensive reward r based on the global performance index tot =α·Δt+β·Θ-γ·L, where Δt is the end-to-end delay, Θ is the network throughput, L is the packet loss rate, and the weight coefficients α, β, and γ are used to balance the priorities of different optimization objectives. Figure 5 The following figure shows the model architecture diagram for centralized training.

[0077] After the global reward calculation is completed, the coordinator updates the policy model through the TD error. The TD error δ is defined as: δ = r tot +γ·(max a′ Q tot (s′,a′;ψ target)-Q tot (s,a;ψ)), where γ is the discount factor, ψ target is the target network parameter, Q tot (s,a;ψ) is the global Q value predicted by the current hybrid network based on the global state s and the joint action a, and max a′ Q tot (s′,a′;ψ target ) is the maximum expected Q value of the target network for the next state s′. By minimizing the mean square loss of TD error L = E[δ 2 ], the coordinator uses the back propagation algorithm to synchronously update the local Q network parameters θ of each agent v and the hybrid network parameter ψ.

[0078] Furthermore, the loss function is minimized through a double-delayed deep deterministic policy gradient framework:

[0079] y=R tot +γ·max a′ Q tot (S′,a′;ψ target ) is the time series difference target, and β is the regularization coefficient, which is used to suppress the deviation of local Q value. When the Q value of an agent is abnormally high due to local observation deviation, the regularization term By penalizing the squared magnitude of the Q-value, the model is forced to reduce its reliance on unreliable local estimates, thereby preventing the global policy from being misled by noisy data from individual agents.

[0080] Then, the parameter delay synchronization strategy is adopted, and each K θ Step 1: Update the local Q network parameters θ v Synchronizing all agents reduces the overhead of frequent communication while ensuring consistency across all agents during distributed execution. In dynamic topology scenarios, the delayed synchronization mechanism allows for multiple iterations of fine-tuning locally before synchronizing globally once parameters stabilize, thus balancing convergence speed with model stability.

[0081] Finally, to improve training stability, the target network parameter ψ target Keep frozen during training, and only soft-update (ψ target ←τψ+(1-τ)ψ target ,τ<<1). This mechanism effectively alleviates the moving target problem and avoids policy oscillation. At the same time, the experience data uploaded by the agent is sampled in the global experience pool according to priority (Prioritized Experience Replay), focusing on samples with high TD error to accelerate model convergence and improve the model's adaptability to dynamic scenarios.

[0082] In the distributed execution described above, each agent v deploys a local Q network and GAT module with frozen parameters, and only relies on local observations s during execution. v Neighbor messages aggregated with GAT, through greedy strategy Generate routing decisions. The routing sequence of the service flow is generated by an iterative path construction algorithm. The source node agent selects the optimal next hop node based on the Q value. The relay node agent recursively executes the same decision until it reaches the destination node, forming a complete routing path P = v0 → v1 → ... → v d , building a low-latency, high-reliability routing sequence hop by hop.

[0083] Furthermore, the system introduces a path feedback mechanism to cope with network dynamics. When the performance indicators of the routing path exceed the preset threshold, the relevant indicators are back-propagated to the path nodes, triggering the incremental fine-tuning of the local Q network. Its agent updates the local Q network parameters θ online based on the feedback indicator data. v , adjusting the Q value of the node selected as the next hop without interrupting service or retraining the global model. This mechanism achieves local policy adaptation through lightweight gradient adjustment, significantly improving the system's responsiveness in dynamic environments.

[0084] In some embodiments, when the network topology changes dramatically, for example, when the number of newly added or removed nodes exceeds a preset threshold ΔN, the system starts the federated learning protocol to adapt to the new network environment. In the specific implementation steps, all proxy nodes participating in the federated learning initialize the local model parameters θ v , ensuring the consistency of initial parameters.

[0085] Then, each agent node calculates the gradient based on local experience data During the computation process, differential privacy technology is used to inject noise into the gradient information to ensure data privacy. The gradient information is then encrypted to prevent malicious nodes from stealing or tampering.

[0086] Encrypted gradient The data is transmitted to a trusted aggregation node through a secure channel. The aggregation node uses secure multi-party computing (MPC) technology to complete the gradient addition operation without decryption, thereby obtaining the global gradient. Subsequently, the aggregation node decrypts the global gradient and adjusts the global model parameter θ according to the preset learning rate global .

[0087] Updated global model parameters θ globalThe model parameters are transmitted to all proxy nodes through a secure channel to ensure that each proxy node can obtain the latest global model parameters. After receiving the global model parameters, each proxy node updates the local model parameters θ v , ensuring the consistency of model parameters of all nodes.

[0088] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a distributed large-scale reverse tracing elastic network intelligent routing method.

[0089] In summary, the present invention provides a distributed large-scale anti-traceability elastic network intelligent routing method and system, including: constructing a multi-agent distributed anti-traceability elastic network model according to the node scale and security requirements in a dynamic topology network environment, mapping nodes into independent agents and limiting local information visibility to block path-related traceability risks; designing a neighbor sampling and message aggregation module based on a lightweight graph attention mechanism, dynamically screening key link status information through local topology perception, and adapting to large-scale network scenarios; introducing the QMIX multi-agent collaborative learning framework, combining local Q networks with global hybrid networks, realizing nonlinear optimization of distributed routing decisions, and simultaneously improving latency, throughput and load balancing performance; designing an action entropy-driven adaptive exploration mechanism to support rapid policy fine-tuning when the topology changes drastically, and ensuring routing service continuity; constructing a CTDE collaborative architecture to jointly optimize local and global policy parameters through TD errors to generate highly concealed and highly flexible routing sequences. The method provided by the present invention deeply integrates distributed agent decision-making, graph attention mechanism and collaborative reinforcement learning to solve the problems of path-related attack threats and suboptimal local decision-making in large-scale dynamic networks, effectively improve transmission efficiency, anonymity and resource allocation balance, and ensure the reliable operation of intelligent routing in complex anti-tracing scenarios.

[0090] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.

[0091] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0092] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.

[0093] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A distributed large-scale reverse tracing elastic network intelligent routing method, characterized in that: The method comprises the following steps: The distributed large-scale anti-traceability elastic network is a transmission network with a large number of nodes, a wide node distribution span, a complex network topology, and elastic expansion and contraction and anti-traceability functions. The distributed large-scale anti-traceability elastic network intelligent routing method is a distributed large-scale anti-traceability elastic network intelligent routing method based on the neighbor sampling graph attention network NSGAT and Q-value hybrid multi-agent reinforcement learning QMARL. It uses graph neural networks to learn hidden information of topology and is suitable for elastic network environments. Neighbor sampling is used to reduce the computational burden caused by large-scale networks. Nodes only determine the next hop forwarding of data based on local information, without exposing the source of the data. The specific steps are as follows: Step 1: Build a multi-agent distributed anti-traceability elastic network model. Set up an agent for each network node. Agents communicate with each other through directly adjacent agents and exchange network status information. Distributed agents only know the previous and next hop information of the data and have anti-traceability capabilities. Step 2: Message communication between adjacent agents is implemented using the Graph Attention Network (GAT). Through neighbor sampling, full-graph sampling is optimized to small-batch neighbor sampling centered on some nodes. The attention mechanism is used to distinguish the importance of neighbors. GAT is adaptable to large-scale network scenarios. Step 3: Introduce the QMIX algorithm to implement collaborative learning of agent intelligent routing. QMIX improves on the deep Q network DQN and processes the joint action value function in a multi-agent environment through the value function decomposition method. Each network node acts as an independent agent and calculates a local Q value based on its own state and neighbor messages obtained by GAT through a local Q network to represent its local utility for routing decisions. QMIX's hybrid network dynamically integrates the local Q values ​​of all nodes through a supernetwork. Combined with the global graph state representation constructed by GAT, it nonlinearly generates a global Q value to optimize network objectives such as end-to-end latency and throughput. Step 4: Training is divided into three steps: distributed experience collection, centralized training, and distributed execution. Each node selects an action based on the current epsilon-greedy strategy, executes routing decisions, and observes the utilization of the local reward link. The GAT periodically exchanges messages to construct an approximation of the local state s and the global state S. The hybrid network calculates the global Q value based on the global state, and optimizes the parameters of the local Q network and the hybrid network through the TD error. The goal is to minimize the prediction error of the global Q value. After training, each node selects an action based solely on the local Q value and neighbor messages obtained by the GAT, without the participation of the hybrid network. The distributed reverse-source elastic network intelligent routing gives the best next hop for the network node, and after combining them, the routing sequence of the service flow is obtained.

2. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The distributed elastic network model is constructed in step 1, specifically: A one-to-one mapping tightly coupled agent architecture is proposed to model the distributed large-scale reverse traceability elastic network environment as a graph. Represents the set of anti-traceable terminal devices and represents the terminal link set, Represents the agent set, and the device set One-to-one mapping, the agent can only obtain local status through the bound device, avoiding global topology exposure and blocking tracing attacks based on path association. and Represent the global state and action space respectively. Assume that there are N agents in the system and the local observation state of each agent v is represented by s v , then the global state space can be defined as the Cartesian product of these local states, if the local observation state s of each agent v is v From their respective local state space S v , define the global state space is the Cartesian product of the local state space of each agent The local observation state s v Contains device node d v queue depth, link bandwidth utilization, and global action space is the local action a of each agent v ∈A v The Cartesian product of The local action a v Select the next routing path.

3. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The message communication between adjacent agents in step 2 realizes cross-agent collaborative decision-making through the GAT functional component. Each agent The internally deployed GAT module includes the following core functional units: Dynamic neighbor sampling and feature extraction: Each agent extracts local information based on its observations. Each agent then shares these features with neighboring agents in the graph and combines the received neighbor features with its own observations to enhance decision-making capabilities. A dynamic neighbor random sampling strategy is proposed to avoid the traditional GAT, which requires all nodes in the graph to participate in the calculation. Combining the link delay changes, bandwidth utilization, and load balancing of the network topology, only features of some high-value neighbors are aggregated. Each agent only extracts some high-value neighbor nodes from the distributed large-scale reverse traceability elastic network instead of processing the entire graph data, reducing the computational load of distributed nodes and adapting to the dynamic nature of large-scale networks. Through the aforementioned lightweight feature extraction unit, the agent encodes the queue load, adjacency table, bandwidth, and delay into feature vectors as the input basis of the GAT. By combining the attention mechanism with a distributed, large-scale, and resilient back-tracing network through attention weight aggregation, the agent distinguishes the importance of different neighbors. GAT assigns higher attention weights to neighboring nodes with high link quality and low load. This process does not require preset rules, but is dynamically adjusted through learnable parameters. The agent aggregates the feature vectors of neighboring nodes according to the attention weights. The aggregated information contains the context required for global collaboration. The above design enables the GAT component to extract network topology features while reducing communication overhead. The attention mechanism enables the system to flexibly respond to the dynamic changes of distributed large-scale back-traceable elastic networks, while the modular functional units facilitate distributed deployment and expansion, ultimately supporting an efficient, hidden, and highly available elastic network.

4. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The collaborative learning mechanism of distributed elastic routing in step 3 uses the QMIX algorithm to implement multi-agent value function decomposition and global strategy optimization, specifically including the following technical features: Distributed local Q network architecture, for each agent Deploy a local Q network v (s v ,a v θ v ), whose input is the hidden state h generated by the GAT module containing aggregated neighbor messages v (K) and local state representation of link indicators s v With optional action a v ∈A v , the output is the local Q value of the corresponding action, which represents the expected future reward of forwarding to the next hop in the current state of the local network; Design the QMIX network to give different contribution ratios to different agents, integrate the hybrid network with each local agent network, and the hybrid network Q tot (s,a;ψ) generates a dynamic weight matrix W(s;ψ w ) and the bias vector b(s;ψ b ), the local Q value Q of each agent v The nonlinear combination is the global Q value, and each agent’s local Q value Q v Satisfy the monotonicity constraint The specific form is Q tot =W(s)·σ([Q1,Q2,...,Q |V| ])+b(s), where σ is the Exponential Linear Unit (ELU) activation function, W(s)∈R 1×|V| and b(s)∈R is generated by the global graph state representation s constructed based on GAT, where s is the node-level aggregate feature; Distributed training mechanism, adopts centralized training-distributed execution framework, during the training phase, each agent v regularly uploads local experience data (s v ,a v ,r v ,s v ′ ) to the regional coordinator, which is based on the global reward r tot =α·Δt+β·Θ-γ·L, calculate the timing difference error δ=r tot +γ·(max a 'Q tot (s ′ ,a ′ ψ target )-Q tot (s, a; ψ)), where Δt represents the end-to-end delay, Θ represents the throughput, and L represents the packet loss rate. Finally, the local Q network parameters θ are synchronously updated through backpropagation. v and the mixed network parameters ψ, while freezing the target network parameters ψ target To train for stability.

5. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The distributed experience collection described in step 4 includes the following steps: Design an action entropy driven adaptive annealing mechanism to calculate the information entropy of the historical action distribution of each agent v Each agent Based on the exploration rate ε v Dynamic annealing adjustment based on node historical action entropy If the action exploration of agent v is sufficient and the entropy is high, ε is quickly reduced to give priority to the high reward path; if the action exploration is insufficient and the entropy is low, the decay rate of ε is slowed down to encourage diverse exploration; Agent v performs the selected action a v After that, the state is transferred to s v ′ , by collecting local immediate rewards r v =w1·(1-U)+w2·(1-Q)-w3·Δt, where U represents link utilization, Q represents queue occupancy, and Δt represents link delay. Using the GAT message passing mechanism, with T as the period, the hidden state differential is exchanged with the neighboring agent. Message, build local state Each agent v calculates the global graph feature vector based on the aggregated neighbor messages Traditional empirical data cannot explicitly reflect the real-time state changes of neighbor nodes such as bandwidth drops and proxy nodes joining / exiting. Therefore, it is considered to add the hidden state difference Δh of neighbor nodes in GAT multi-layer propagation to the empirical tuple. b (k) To perceive the sudden change in bandwidth utilization and delay jitter caused by link quality fluctuations, dynamic changes in topology such as neighbor nodes going offline or new proxy nodes being added, and changes in neighbor queue load trends caused by traffic pattern changes, the regional coordinator adopts an asynchronous parallel gradient update mechanism to receive the experience tuples uploaded by each proxy (s v ,a v ,r v ,S′,Δh b (k),b∈B(v)), build a global experience pool D=(S,a v ,R tot ,S ′ ), where the global reward r v represents the local immediate reward of agent v, Θ global represents the global throughput, Represents the variance of end-to-end delay.

6. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The centralized training described in step 4 includes the following steps: During the training and learning process, the QMIX hybrid network nonlinearly merges the local Q functions of the single agents. The hybrid weights are dynamically generated by the super network and assisted by the global state information. The hybrid network is conditioned on the global state S and generates a dynamic weight matrix W(S)∈R through the super network. d×|V| With the bias vector b(S)∈R d , the QMIX algorithm converts the local Q value Q v (s v ,a v )The nonlinear combination is the global Q value Q tot =W(S)·tanh([Q1,Q2,...,Q |V| ])+b(S), and calculate the time series difference target y=R tot +γ·max a 'Q tot (S ′ ,a ′ ψ target ), where ψ target is the target hybrid network parameter; Minimizing the loss function via a double-delayed deep deterministic policy gradient optimizer Where β is the regularization coefficient, which suppresses the overestimation of local Q value. At the same time, the parameter delay synchronization strategy is adopted. θ Step synchronization of the local Q network parameters θ v To all agents.

7. The distributed large-scale reverse tracing elastic network intelligent routing method according to claim 1 is characterized in that: The distributed execution in step 4 includes the following steps: Each agent v deploys a local Q network and GAT module with frozen parameters, and only relies on local observations s during execution v Neighbor messages aggregated with GAT, through greedy strategy Generate routing decisions. The routing sequence of the service flow is generated by an iterative path construction algorithm. The source node agent selects the optimal next hop node based on the Q value. The relay node agent recursively executes the same decision until it reaches the destination node, forming a complete routing path P = v0 → v1 → ... → v d , traditional methods require global retraining when the performance degrades due to topology changes. This method uses a path feedback mechanism to backpropagate routing delay and packet loss rate indicators to path nodes, triggering incremental fine-tuning of the local Q network, updating the model without interrupting service, and supporting online adaptive updates. When it is detected that the number of added / removed nodes exceeds ΔN, resulting in a drastic change in the topology structure, the federated learning protocol is started, and each agent calculates the gradient based on local experience Update the global model after multi-party aggregation.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Communication method for distributed training of hybrid expert model and distributed system

    CN121547506A