An intelligent routing optimization method and device, computer equipment and storage medium
By constructing a quantum network graph model and using deep reinforcement learning, the problems of accurate representation and resource assessment of quantum network routing technology in dynamic environments were solved, realizing intelligent adaptive routing decision-making and improving the stability and efficiency of quantum networks.
Patent Information
- Application Number
- CN202610386768.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-27
- Publication Date
- 2026-08-25
AI Technical Summary
Existing quantum network routing technologies lack accurate dynamic topology representation capabilities and comprehensive and non-real-time quantum resource state assessment when dealing with complex dynamic environments. This results in insufficient basis for routing decisions, low intelligence, and difficulty in achieving optimal routing.
A quantum network graph model is constructed to measure and estimate the state of quantum resources in real time. An intelligent routing decision engine is built through deep reinforcement learning. The performance is evaluated and benchmarked on a quantum network simulation and evaluation platform, and finally deployed to quantum hardware and control systems.
It enables accurate dynamic topology characterization and real-time resource status assessment of quantum networks, improves the scientific nature and efficiency of routing decisions, adapts to different scenario requirements, reduces deployment risks and costs, and promotes the practical application of quantum networks.
Smart Images

Figure CN122640337A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of quantum communication technology, and in particular to intelligent routing optimization methods, devices, computer equipment, and storage media. Background Technology
[0002] Existing quantum network routing technologies have significant shortcomings when dealing with complex and dynamic environments. First, they lack accurate dynamic topology representation capabilities. Traditional methods are mostly based on static or simple dynamic models, making it difficult to capture rapid changes in quantum network topology in real time. This leads to a disconnect between routing decisions and the actual network state, failing to adapt to the dynamic addition and removal of quantum nodes and link instability. Second, quantum resource state assessment is incomplete and not real-time. Existing technologies struggle to measure and estimate the states of various quantum resources affecting routing decisions in quantum networks in real time and accurately, such as the coherence time and entanglement generation rate of qubits. This results in insufficient basis for routing decisions, easily leading to resource waste or routing failures. Third, the intelligence level of routing decisions is low. Traditional routing algorithms are mostly rule-based fixed patterns, unable to adaptively adjust routing strategies according to the real-time network state, making it difficult to achieve optimal routing in complex and ever-changing quantum network environments. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention provides an intelligent routing optimization method, employing the following technical solution, including the following steps: Construct a quantum network graph model to perform dynamic topological characterization of the quantum network; Based on the quantum network graph model, the states of various quantum resources that affect routing decisions in the quantum network are measured and estimated in real time, and the states of quantum resources are transformed into feature vectors or state representations. Based on the feature vector or state representation, an intelligent routing decision engine is constructed through deep reinforcement learning, and a routing strategy is output through the intelligent routing decision engine. Translate the routing policy into executable network instructions; Construct a quantum network simulation and evaluation platform to conduct performance evaluation and benchmark testing of quantum networks; The quantum network, after performance evaluation and benchmarking, is deployed to quantum hardware and control systems.
[0004] Preferably, the step of constructing a quantum network graph model and performing dynamic topological characterization of the quantum network specifically includes: Construct a quantum network graph model, abstracting the physical quantum network nodes and channels into a mathematical graph structure; Formalize the quality indicators of quantum services; Modeling the dynamics and uncertainties of quantum networks to characterize their time-varying properties.
[0005] Preferably, the step of measuring and estimating the states of various quantum resources affecting routing decisions in the quantum network in real time based on the quantum network graph model, and converting the quantum resource states into feature vectors or state representations specifically includes: Assess the quantum state transmission quality of current and future links; Perform link throughput and capacity modeling to assess the ability of each link and potential path to serve entangled requests; Assess the health status and service capabilities of relay nodes to determine their suitability as routing relay stations.
[0006] Preferably, the step of constructing an intelligent routing decision engine based on the feature vector or state representation through deep reinforcement learning, and outputting a routing strategy through the intelligent routing decision engine, specifically includes: Design reinforcement learning states, actions, and rewards; Set up a distributed near-end strategy optimization training framework; Transfer learning is performed based on graph structure to output routing strategies.
[0007] Preferably, the step of translating the routing policy into executable network instructions specifically includes: Generate multiple paths based on constrained integer programming; Coordinate and lock in the necessary quantum resources for the selected path; Configure local fast recovery and rerouting mechanisms.
[0008] Preferably, the steps of constructing a quantum network simulation and evaluation platform to perform performance evaluation and benchmark testing of quantum networks specifically include: Concurrent and asynchronous events in quasi-quantum networks; evaluating the performance of quantum networks in complex dynamic scenarios. Conduct multi-dimensional performance evaluation and benchmark testing; Perform automatic hyperparameter tuning and strategy iteration.
[0009] Preferably, the step of deploying the quantum network, after performance evaluation and benchmarking, onto quantum hardware and control systems specifically includes: Design a hardware abstraction layer to provide a unified control interface for different quantum devices; Integrate quantum routing protocols into existing classical network infrastructure and control planes to enable collaborative operation; After actual deployment, continuously monitor the network operation status, evaluate the effectiveness of routing strategies, and make dynamic adjustments and optimizations.
[0010] To address the aforementioned technical problems, the present invention also provides an intelligent routing optimization device, which employs the following technical solution, including: The characterization module is used to construct quantum network graph models and perform dynamic topological characterization of quantum networks. The conversion module is used to measure and estimate the quantum resource states that affect routing decisions in the quantum network in real time based on the quantum network graph model, and convert the quantum resource states into feature vectors or state representations. A construction module is used to build an intelligent routing decision engine based on the feature vector or state representation through deep reinforcement learning, and to output a routing strategy through the intelligent routing decision engine. The translation module is used to translate the routing policy into executable network instructions; The evaluation module is used to build a quantum network simulation and evaluation platform to perform performance evaluation and benchmark testing on quantum networks. The deployment module is used to deploy quantum networks, after performance evaluation and benchmarking, onto quantum hardware and control systems.
[0011] To address the aforementioned technical problems, the present invention also provides a computer device that employs the technical solution described below, comprising a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the aforementioned intelligent routing optimization method.
[0012] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium, which employs the technical solution described below. The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the aforementioned intelligent routing optimization method.
[0013] Compared with the prior art, the present invention has the following main advantages: (1) By constructing a quantum network graph model and dynamically representing the topology, we can accurately grasp the structural characteristics and change patterns of the quantum network, provide a solid foundation for subsequent routing decisions, effectively address the challenges of highly dynamic quantum network topologies, and ensure that the routing strategy adapts to the real-time state of the network. (2) By measuring and estimating the state of quantum resources in real time and converting it into feature vectors, various types of quantum resources can be quantified comprehensively and accurately, enabling the intelligent routing decision engine to make decisions based on rich and accurate information, avoiding routing errors caused by inaccurate resource information, and improving the scientificity and rationality of routing decisions. (3) By leveraging deep reinforcement learning to build an intelligent routing decision engine, it can learn and optimize routing strategies autonomously, adapt to different scenarios and needs, and improve the efficiency and flexibility of routing decisions without frequent manual intervention, thereby achieving intelligent and adaptive routing. (4) By constructing a quantum network simulation and evaluation platform, a comprehensive performance evaluation and benchmark test of the quantum network can be carried out before deployment, potential problems can be identified and optimized in advance, and deployment risks and costs can be reduced. After full verification, deployment to quantum hardware and control systems can ensure the stable and efficient operation of the quantum network, improve overall performance and reliability, and promote the practical application of quantum network technology. Attached Figure Description
[0014] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0015] Figure 1 This is a flowchart of an embodiment of the intelligent routing optimization method of the present invention; Figure 2 This is a schematic diagram of a structure of an embodiment of the intelligent routing optimization device of the present invention; Figure 3 This is a schematic diagram of the structure of an embodiment of the computer device of the present invention. Detailed Implementation
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings are used to distinguish different objects and not to describe a particular order.
[0017] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0018] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.
[0019] It should be noted that the intelligent routing optimization method provided in the embodiments of the present invention is generally executed by a server / terminal device, and correspondingly, the intelligent routing optimization device is generally set in the server / terminal device.
[0020] It should be understood that the number of terminal devices, networks, and servers is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used. Example
[0021] Please refer to Figure 1 The diagram illustrates a flowchart of an embodiment of the intelligent routing optimization method of the present invention. The intelligent routing optimization method includes the following steps: Step S1: Construct a quantum network graph model and perform dynamic topological characterization of the quantum network.
[0022] In this embodiment, the intelligent routing optimization method runs on electronic devices (e.g., Figure 1 The server / terminal device shown can receive intelligent routing optimization requests via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future wireless connection methods.
[0023] In this embodiment, step S1, constructing a quantum network graph model and performing dynamic topological characterization of the quantum network, may specifically include the following steps: S11, constructing a quantum network graph model, abstracting physical quantum network nodes and channels into mathematical graph structures.
[0024] Weighted directed hypergraph modeling uses G = (V, E, W) to represent the network. Here, V is the set of quantum nodes (such as quantum memories and transmitters); E is the set of directed hyperedges, and a hyperedge eT ij Connecting nodes i and j can represent multiple parallel quantum channels (such as free space and fiber optic channels); W is a set of weight functions that dynamically assign multidimensional weights to each edge e and each node v, such as Wi. f (e) represents the channel fidelity weight, W c (e) represents the cost (energy, time) of establishing entanglement, W q (v) represents the quantum memory capacity and coherence time of node v.
[0025] On top of the hypergraph, construct the logical layer L. logical and physical layer L physicalThe physical layer describes the actual photonic transmission link; the logical layer describes the end-to-end entangled connection. The two layers are mapped through an entanglement swap operation, which is represented in the diagram as stitching two adjacent physical edges into a longer logical edge at an intermediate node.
[0026] To handle timing and queuing, a time dimension is introduced, expanding the static graph into a spatiotemporal graph G. ST = (V×T, E) ST T is a set of discrete time slots. Node (v, t) represents node v at time t; edge ((v1, t1), (v2, t2)) represents a quantum state emitted from v1 at time t1 and arriving at v2 at time t2. This allows the algorithm to schedule entanglement requests to be satisfied in future time slots.
[0027] The purpose of step S11 is to abstract the physical quantum network nodes and channels into a mathematical graph structure, which is the underlying data structure for all routing algorithms.
[0028] S12 provides a formal definition of quantum service quality indicators.
[0029] Defined as the generated entangled state ρ actual With the ideal target state ρ ideal Uhlmann fidelity F between (e.g., Bell states) e2e :F e2e =(tr√(√(ρ ideal ) ρ actual √(ρ ideal ))) 2 , where ρ actual It is the mixed-state density matrix after multiple entanglement swaps and channel transmissions. Fidelity F e2e The closer the value is to 1, the higher the quality of quantum information transmission. It is one of the core constraints for routing path selection (F... e2e ≥F threshold ).
[0030] Quantum throughput is defined as the number of successfully distributed, high-fidelity end-to-end entangled pairs per unit time. Let R... success (t,Δt) represents the number of successful requests within the time window Δt, then the throughput T put =lim Δt→∞ R success (t,Δt) / Δt. This measures the overall service capacity of the network.
[0031] Entanglement establishment delay includes queuing delay (the time it takes for a request to wait for resources in a buffer), transmission delay (the time it takes for a photon to cross the channel), and operation delay (the time consumed by quantum operations such as entanglement generation and Bell state measurement). The total delay D total =D queue +Dtrans +D op Low latency is crucial for many quantum applications.
[0032] The purpose of step S12 is to clarify the criteria for measuring the quality of routes and to provide a mathematical definition for the optimization objective.
[0033] S13 models the dynamics and uncertainties of quantum networks, characterizing their time-varying properties.
[0034] Stochastic process modeling of channel attenuation and noise: The depolarization noise and attenuation coefficient of a quantum channel are not constant. Hidden Markov Models (HMMs) or Autoregressive Integral Moving Average (ARIMA) models are used for modeling. For example, the channel fidelity F... channel (t) is modeled as: F channel (t)=F base The equation is: ⋅η(t)+(1-η(t))⋅Noise(t), where η(t) is a time-varying attenuation factor (following a certain distribution), and Noise(t) is a random noise process. By training the model parameters using historical data, short-term channel quality can be predicted.
[0035] Dynamic model of node resources: Entangled pairs in a node quantum memory decay over time due to decoherence. The fidelity F of the stored entangled states... memory Exponential decay over time t: F memory (t) = F0⋅exp(-t / τ), where τ is the storage coherence time. Meanwhile, the memory occupancy state (idle / occupied) is a discrete process that changes over time and can be modeled using a continuous-time Markov chain (CTMC).
[0036] Request arrival and load model: The arrival of quantum entangled requests is usually assumed to be a Poisson process, i.e., P(N(t+Δt)-N(t)=k)=(λΔt) k The formula is exp(-λΔt) / k!, where λ is the average arrival rate. The source-destination pair distribution, required fidelity, priority, and other attributes of the requests are described by a joint probability distribution. This model is used to generate training data and evaluate system load performance.
[0037] The purpose of step S13 is to characterize the time-varying properties of the quantum network, making the model more realistic and preparing intelligent algorithms to cope with uncertainties.
[0038] The purpose of step S1 is to abstract a computable and analyzable mathematical model from a specific problem and to provide a unified descriptive framework and analytical benchmark for all subsequent steps.
[0039] Step S2: Based on the quantum network graph model, measure and estimate the states of various quantum resources that affect routing decisions in the quantum network in real time, and convert the quantum resource states into feature vectors or state representations.
[0040] In this embodiment, step S2, based on the quantum network graph model, involves real-time measurement and estimation of various quantum resource states affecting routing decisions in the quantum network, and the conversion of these quantum resource states into feature vectors or state representations. Specifically, this may include the following steps: S21, assess the quantum state transmission quality of current and future links.
[0041] By inserting probe states (such as weakly coherent states) into the data channel and measuring their state changes after passing through the channel, the quantum process of the channel (Choi matrix X) can be reconstructed using maximum likelihood estimation (MLE) or Bayesian estimation. The fidelity of the channel to the Bell state |Φ+> can be calculated as F. ch = ⟨Φ+|X|Φ+>. This method is accurate but expensive and should not be performed frequently.
[0042] Establish an empirical / theoretical mapping function F between classical physical parameters (such as fiber length L, attenuation coefficient α, ambient temperature T, and bit error rate BER) and quantum fidelity. ch =g(L,α,T,BER,…). By monitoring readily available classical parameters, quantum fidelity can be calculated in real time, enabling lightweight sensing.
[0043] Temporal models such as Long Short-Term Memory (LSTM) networks or Gated Recurrent Units (GRUs) are used. The model is trained to predict the fidelity at future time Δt using the historical fidelity sequence {F(t-τ),…,F(t-1)} and related environmental parameter sequences as input. (t+Δt). The model continuously adapts to changes such as channel degradation or repair through online learning.
[0044] The purpose of step S21 is to accurately assess the quantum state transmission quality of the current and future links, which is the primary basis for determining the feasibility of the path.
[0045] S22 performs link throughput and capacity modeling to assess the ability of each link and potential path to handle entangled requests.
[0046] Capacity calculation based on entanglement generation probability: For an entangled source based on parametric downtransformation, the success probability of generating an entangled pair in a single attempt is p. gen If N attempts are made within period T, the expected number of successful attempts is N⋅p. gen Considering the symmetry of bidirectional transmission, the entanglement generation rate (original capacity) of link (i, j) is C. raw (i,j)=min(p geni ,pgenj ) / T cycle Further considering the channel loss η, the effective bidirectional entanglement distribution rate is C. eff =η 2 ⋅C raw .
[0047] Buffer queue status monitoring: A priority queue is maintained at each output link interface of each node to buffer quantum states to be sent or entanglement swap requests to be executed. The length Q of each queue is monitored in real time. len The oldest group's age Q age And the expected waiting time Q delay These metrics directly reflect the real-time congestion level of the link.
[0048] Entangled flow computation considers not only the physical capacity of the link itself but also the pre-occupied logical entangled flows. The remaining available entangled flow A(i,j) = C for link (i,j) is calculated. eff (i,j)-∑ flow f flow (i,j), where f flow (i,j) represents the rate of the established logical flow through this link. This requires the network to maintain a global or distributed view of the traffic matrix.
[0049] The purpose of step S22 is to assess the ability of each link and potential path to handle entangled requests and avoid routing requests to congested paths.
[0050] S23 assesses the health status and service capabilities of relay nodes to determine whether they are suitable as routing relay stations.
[0051] Each quantum node maintains a resource table that records the entangled state information currently stored in each quantum storage unit: target node ID, storage start time t. start The initial fidelity F0, and according to the formula F cur =F0⋅exp(-(t now -t start The current fidelity is calculated as ) / τ). The number of free storage units M is counted in real time. free and average remaining coherence time.
[0052] Local quantum operations on nodes (such as Bell state measurements and single-qubit gates) are not perfect. By periodically self-calibrating, the process fidelity F of these operations is measured and recorded. BSM F gate These values will be used to calculate the overall fidelity of the end-to-end path. (Simplified model).
[0053] Nodes periodically broadcast digests of their key resources (such as M) to their immediate neighbors.free F BSM_avg The "heartbeat" message (load factor). By receiving heartbeats from neighbors, nodes can build a local resource view and understand the available relay options within their "one-hop" range without global communication, reducing overhead.
[0054] Step S2, based on the model from step S1, measures and estimates in real time the states of various key quantum resources in the network that influence routing decisions, and transforms this information into feature vectors or state representations usable by subsequent intelligent decision-making algorithms. This is a crucial transformation from the physical world to the information world.
[0055] Step S3: Based on the feature vector or state representation, construct an intelligent routing decision engine through deep reinforcement learning, and output a routing strategy through the intelligent routing decision engine.
[0056] In this embodiment, step S3, which involves constructing an intelligent routing decision engine based on the feature vector or state representation using deep reinforcement learning, and outputting a routing strategy through the intelligent routing decision engine, may specifically include the following steps: S31, design reinforcement learning states, actions, and rewards.
[0057] High-dimensional state-space design: state s t It is a combined feature vector, including: Global topological features: Features of nodes and edges encoded by graph neural networks (GNNs) (fidelity, capacity, latency).
[0058] Local request context: the source node s, destination node d, and minimum required fidelity F of the current request to be routed. req Priority p.
[0059] Dynamic resource snapshots: M of critical nodes (such as nodes on candidate paths) free and F BSM .
[0060] Historical performance metrics: average success rate and latency within the sliding window.
[0061] Layered motion space design: Action a t It's not simply about the next jump, but about layered decision-making: Level 1: Path selection: Choose one from K pre-selected paths (generated by step S41), or choose to wait / reject.
[0062] Level 2: Resource reservation parameters: Specify the amount of memory to reserve and the channel slots to be used (in the spatiotemporal graph) for the selected path.
[0063] Level 3: Recovery Action: When a path failure is predicted, a partial recovery mechanism is triggered (such as switching to a backup path).
[0064] Multi-objective reward function design: reward r t Guiding an agent to balance multiple QoS objectives: r t =w1⋅II success +w2⋅(F achieved -F threshol) )-w3⋅D normalized -w4⋅C resource +w5⋅II fairness Among them, II succes : 1 for a successful request, 0 for otherwise. (F achieved -F threshold ): Encourage exceeding fidelity requirements. (D) normalized Normalized end-to-end latency, penalizing slow routes. (C) resource The weighted sum of entanglement, memory, and other resources consumed encourages resource conservation. II fairness : Items that promote fairness in services between different source-destination pairs. w1…w5 are adjustable weights to balance the importance of different objectives.
[0065] The purpose of step S31 is to formalize the complex quantum routing problem into a reinforcement learning (RL) problem, which is a prerequisite for the successful application of DRL.
[0066] S32, set up the Distributed Proximity Policy Optimization (DPPO) training framework.
[0067] Actor-Critic Structure: Actor Network (Strategy Network π) θ ): Input state s t Output action a t probability distribution π θ (a t | s t The parameter θ is updated via policy gradient.
[0068] Critics Network (Value Network V) ϕ ): Input state s t Output the estimated expected cumulative reward V for this state. ϕ (s t The parameter ϕ is updated by minimizing the temporal difference (TD) error.
[0069] The core of the Proximal Policy Optimization (PPO) algorithm is to ensure training stability by limiting the magnitude of policy updates. Its loss function is: , It is the probability ratio between the old and new strategies. It is the estimation of the advantage function. =δ t +(γλ) δ t+1 +⋯, where δ t =r t +γV(s) t+1 )-V(s t ) is the TD error. The clip function will reduce r t (θ) is restricted to [1-ϵ, 1+ϵ] to prevent excessively large single updates. ϵ is a hyperparameter (e.g., 0.2).
[0070] Loss function L CLIP The meaning of (θ) is: to encourage the increase of what brings positive advantages. The probability of actions greater than 0 is calculated, but the update range is limited to avoid drastic policy fluctuations.
[0071] Distributed parallel training: Multiple environment workers are deployed to simulate the network environment in parallel on a CPU cluster (based on the simulator in step S5). Each worker independently collects experience and sends it to the central parameter server. The parameter server aggregates gradients, updates global θ and ϕ, and then synchronizes them to all workers. This greatly accelerates the data collection and training process.
[0072] The purpose of step S32 is to train the routing policy network in a distributed environment using a stable and efficient policy gradient algorithm to cope with large-scale networks and high-dimensional states.
[0073] S33 performs transfer learning based on graph structure and outputs routing strategies.
[0074] Learning through progressively more difficult lessons: Phase 1 (Simple Topology): Training is conducted in a small linear or star topology with low request arrival rates and stable channel quality. The agent learns basic pathfinding and fidelity requirements first.
[0075] Phase Two (Increasing Complexity): Gradually expand the network size to grid and random graph, improve request arrival rate, and introduce random fluctuations in channel quality.
[0076] Phase 3 (Introducing Faults and Dynamics): Simulating link interruptions, node failures, and sudden traffic surges to train the agent's robustness and recovery capabilities. Each phase builds upon the model trained in the previous phase.
[0077] Graph-based transfer learning: When the network topology changes (such as adding nodes), it is not necessary to train from scratch.
[0078] Graph Isomorphic Networks (GINs) are used as feature extractors for the policy network. GINs exhibit good generalization ability to small changes in graph structure.
[0079] The weights of the pre-trained GIN encoder in the original network are fixed, and only the subsequent fully connected decision layer is fine-tuned to adapt to the change in the number of nodes / edges in the new topology. This leverages the prior knowledge that the underlying rules of quantum routing between different topologies (such as fidelity decay and resource contention) are similar.
[0080] Domain adaptation from simulation to reality: To alleviate the gap between simulation (Sim) and real (Real) networks (Sim2Real Gap), in the later stages of training: Introduce a more random noise model and more realistic non-ideal device characteristics into the simulation.
[0081] Collect a small amount of runtime data from real networks or high-fidelity test platforms (such as in step 5.1), and use this data to calibrate the simulator parameters or fine-tune the policy network.
[0082] The purpose of step S33 is to solve the problems of difficulty in exploration and slow convergence in the early stage of DRL training, and to achieve rapid adaptation of policies to networks of different sizes or configurations.
[0083] The purpose of step S3 is to receive the multidimensional resource quantization information from step S2 as state input, and under the constraints of the model and QoS objectives defined in step S1, learn and output the optimal or near-optimal routing strategy through continuous interaction with the environment (simulated or real network). It transforms the high-dimensional, dynamic, and non-convex quantum routing optimization problem into a learnable sequential decision problem.
[0084] Step S4: Translate the routing policy into executable network instructions.
[0085] In this embodiment, step S4, translating the routing policy into executable network instructions, may specifically include the following steps: S41 generates multiple paths based on constrained integer programming.
[0086] Modeling the Multi-Constraint Shortest Path Problem (MCSP): For each request, solve a constrained optimization problem. The objective is to minimize the cost C(p) (e.g., number of hops, total decay) while satisfying: F path (p)≥F req (Fidelity constraint), Bottleneck_Capacity(p)≥B req (Bandwidth constraint), D path (p)≤D max(Delay constraint). This is typically an NP-hard problem.
[0087] Heuristic and metaheuristic algorithms for solving problems: An improved Dijkstra's algorithm operates on an expanded state space. Each state records the lowest path fidelity guaranteed when reaching the node. The algorithm prioritizes expanding paths with high fidelity and returns once the destination node is found and all constraints are satisfied.
[0088] Ant Colony Optimization (ACO): Simulates ants crawling on a graph. The pheromone concentration along a path is positively correlated with historical success rate and fidelity. Ants choose their next hop based on pheromone levels and heuristic information (such as the geometric distance to the target). After multiple iterations, several high-quality paths converge. The pheromone update rule is: τ ij =(1-ρ)τ ij +∑ antk Δτ ij k Where ρ is the volatility coefficient, and τ ij k F is the path obtained by ant k using edge (i, j). path / D path Proportional.
[0089] The K Shortest Paths (KSP) algorithm uses Yen's algorithm. After finding the first shortest path, it systematically deviates from some nodes or edges to find the 2nd, 3rd...Kth shortest paths. Then, it performs a posterior filtering on these K paths, discarding paths that do not meet fidelity constraints, ultimately obtaining a candidate path set {P1, P2, ..., P...}. m (m≤K).
[0090] The purpose of step S41 is to quickly generate a set of high-quality candidate paths based on the current network status and request requirements, so that the DRL engine can choose from them and narrow down its search space.
[0091] S42 coordinates and locks in the required quantum resources for the selected path.
[0092] Two-phase confirmation reservation agreement: Phase 1: Exploration and Reservation Requests: The source node sends a "probe" control message along the selected path to the destination node. The message carries request details and desired resources (such as memory slots that each relay node needs to reserve). Each intermediate node checks its local resources; if available, it temporarily locks these resources and forwards the probe; otherwise, it returns a negative acknowledgment (NACK).
[0093] Phase Two: Confirmation and Execution: If the destination node receives the probe, it sends a confirmation message along the reverse path. Upon receiving the confirmation, the intermediate node converts the temporary lock to a formal reservation and begins executing quantum operations (such as entanglement generation). If the source node receives any NACK or a timeout, it sends a "cancel" message to release all temporary locks and may trigger a retry or notify the DRL engine of a decision failure.
[0094] Priority-based preemptive scheduling: Priorities are assigned to different requests (e.g., based on application type, wait time). When a high-priority request arrives and resources are insufficient, it is allowed to preempt a portion of the resources reserved for lower-priority requests that have not yet begun quantum operations. The preempted request is returned to the waiting queue, its consumed resources are compensated, and it may receive an upgraded priority.
[0095] Slotted operation scheduling: Time is divided into fixed-length slots. The initiation of all quantum operations (photon emission, BSM measurement) is aligned to slot boundaries. The scheduler assigns a specific slot number to each operation requested. This is similar to Time Division Multiplexing (TDM) in classical networks, simplifying synchronization and collision avoidance problems. The scheduling problem can be formalized as a resource-constrained project scheduling problem (RCPSP), solvable online using greedy or genetic algorithms.
[0096] The purpose of step S42 is to coordinate and lock the required quantum resources (memory, channel time slots) for the selected path, avoid conflicts, and ensure that the request can be successfully executed.
[0097] S43, configure local fast recovery and rerouting mechanism.
[0098] Backup next-hop switching: Each node maintains a backup neighbor list for each of its output links. After multiple failed attempts to establish entanglement with the preferred next hop N1, the node automatically switches to the backup next hop N2 and attempts to continue towards the destination node through N2. The backup list is updated periodically based on historical success rates and link quality.
[0099] Segmented re-establishment: The end-to-end path is divided into several logical segments. If a segment (e.g., from node A to node C) fails to establish, retrying or local rerouting (e.g., attempting ADC) is performed only within that segment (ABC), instead of backtracking to the source node. This requires nodes along the way to buffer the intermediate entanglements of the established segment (e.g., AB, BC) until the end-to-end connection is complete.
[0100] Proactive Fault Prediction and Avoidance: Integrating the fidelity predictor from step S21. When it is predicted that the fidelity of a link (i, j) may fall below the threshold in the future time Δt, the scheduler proactively avoids allocating the link to new requests and initiates preventative rerouting for ongoing requests that have already used the link before a failure occurs, migrating their traffic to a more stable alternative path.
[0101] The purpose of step S43 is to quickly repair the local area when a certain link in the path (such as a single entanglement generation) fails, without discarding the entire request or waiting for a global re-decision, thereby improving robustness and resource utilization.
[0102] The purpose of step S4 is to translate and refine the high-level strategy (such as path selection A) output by the DRL engine in step S3 into specific, executable network instructions, including precise path calculation, resource scheduling timing, and handling unexpected failures during execution. It connects intelligent decision-making and physical network operations.
[0103] Step S5: Construct a quantum network simulation and evaluation platform to evaluate the performance and benchmark the quantum network.
[0104] In this embodiment, step S5, constructing a quantum network simulation and evaluation platform to perform performance evaluation and benchmark testing on the quantum network, may specifically include the following steps: S51, concurrent and asynchronous events in quasi-quantum networks, evaluating the performance of quantum networks in complex dynamic scenarios.
[0105] Core event queue and scheduling engine: All future events are managed using a priority queue. Each event includes: trigger time t. event The simulation engine processes events in chronological order, calling callback functions and potentially inserting new events into the queue. This is the standard approach for Discrete Event Simulation (DES). The event type (e.g., PHOTON_ARRIVAL), associated entities (e.g., links, request IDs), and callback functions are all considered.
[0106] The device non-ideal and noise module includes a channel module, a detector module, and a memory module. The channel module implements various noise models, such as a depolarized channel: ε(ρ) = p × I ÷ 2 + (1 - p) × ρ, where p is the error rate, I is the identity matrix, and ρ is the input state; and an attenuation channel (photon loss): passing through the state with probability η and projecting onto the vacuum state with probability 1 - η. The detector module simulates finite efficiency η. d Dark count rate λ dark Dead time τ deadThe Monte Carlo method is used to randomly determine whether an arriving photon is detected. The memory module is used to implement decoherence models, such as amplitude-damped or phase-damped channels, and dynamically updates the density matrix of the stored states.
[0107] Protocol Stack and Controller Simulation: Implement a classic control channel, simulating message exchange, processing latency, and potential message loss of the reserved protocol in step S42. Integrate the intelligent routing decision engine (step S3) as a controller plugin into the simulator, enabling it to respond to simulated events and make decisions.
[0108] The purpose of step S51 is to accurately simulate concurrent and asynchronous events (such as photon arrival, operation completion, and request arrival) in the quantum network and evaluate the performance of the algorithm in complex dynamic scenarios.
[0109] S52 undergoes multi-dimensional performance evaluation and benchmark testing.
[0110] The evaluation metrics system is used to systematically measure and record the following metrics: core QoS metrics, resource efficiency metrics, and scalability and robustness metrics, etc.
[0111] Core QoS metrics: average entanglement establishment success rate, average end-to-end fidelity (and compliance rate), average establishment latency (and latency jitter), and network throughput.
[0112] Resource efficiency metrics: average utilization of quantum memory, channel utilization, and average resource consumption per successful request.
[0113] Scalability and robustness metrics: performance degradation curve as the network scales up, service survival rate under random node / link failures, and performance under burst traffic.
[0114] Benchmark Algorithm Implementation and Comparison: Shortest path (SP): Dijkstra's algorithm that aims to minimize the number of hops.
[0115] Highest Fidelity Path (HFP): Select the path with the largest fidelity product.
[0116] Routing based on classic Q-learning: using tabular Q-learning, with the state being the current node and the destination node, and the reward being whether it is successful immediately.
[0117] Fixed multipath load balancing: such as using equal cost multipath (ECMP).
[0118] Statistical significance testing: Each experiment (specific load, topology, algorithm) should be run independently a sufficient number of times (e.g., 30 times). Use a t-test or Mann-Whitney U test to compare the mean differences between the intelligent method and the traditional method on each metric, and calculate the p-value to determine whether the difference is statistically significant (usually p < 0.05 is considered significant). Use box plots and cumulative distribution function (CDF) plots to visualize the performance distribution.
[0119] The purpose of step S52 is to quantitatively compare the advantages and disadvantages of this intelligent method and traditional methods in different scenarios, and to clarify its gains and applicable scope.
[0120] S53 performs automatic hyperparameter tuning and strategy iteration.
[0121] Bayesian Optimization (BO) framework: used for tuning hyperparameters (such as learning rate lr, discount factor γ, entropy coefficient β, number of neural network layers, etc.) in DRL. BO treats the mapping of hyperparameter combination x to performance target y (such as the average success rate on the validation set) as a black-box function f(x). It uses a Gaussian process (GP) as a surrogate model to model the prior distribution of f(x) and defines a sampling function (such as the desired improvement in EI) to determine the next evaluation point x. next In the formula, EI is defined as: EI(x) = [max(f(x) - f(x)] + ),0)], where f(x + The current observed performance is 0. BO selects the x that maximizes EI(x) for the next round of simulation evaluation, efficiently finding the optimal value in the high-dimensional parameter space.
[0122] Offline policy evaluation and selection: During training, the current policy is periodically run in an independent validation environment to evaluate its performance. A policy buffer is maintained to store several policies that have historically performed best. After training, the policy that performs best on the validation set from the buffer is selected as the final deployment policy to avoid overfitting the training environment.
[0123] The simulation-to-policy feedback loop feeds policy weaknesses (such as poor performance in a certain failure mode) identified during simulation evaluation back into the training process. More training data containing these weaknesses can be generated in a targeted manner, or the reward function can be adjusted (e.g., increasing the penalty for specific failures) to guide the DRL agent to focus on overcoming these weaknesses in the next training iteration.
[0124] The purpose of step S53 is to automatically find the optimal configuration parameters for the intelligent routing algorithm (especially the DRL part) and continuously improve the strategy based on simulation feedback.
[0125] The purpose of step S5 is to provide a high-fidelity, configurable simulation environment for the DRL training in step S3, and to conduct a comprehensive and repeatable performance evaluation and benchmarking after the algorithm design is completed. It connects theoretical design with practical effect evaluation and is the cornerstone of algorithm iterative improvement.
[0126] Step S6 involves deploying the quantum network, after performance evaluation and benchmarking, onto quantum hardware and control systems.
[0127] In this embodiment, step S6, deploying the quantum network, after performance evaluation and benchmarking, onto the quantum hardware and control system, may specifically include the following steps: S61 is designed as a hardware abstraction layer that provides a unified control interface for different quantum devices.
[0128] The API includes: generate_entanglement(neighbor_id, attempts), perform_bsm(memory_id1, memory_id2), measure_qubit(memory_id, basis), get_memory_status(), etc. The routing controller issues commands through the HAL, without needing to concern itself with the underlying hardware differences.
[0129] The real-time control system architecture employs a soft real-time control loop. The control computer runs routing algorithms and protocol stacks, interacting with the quantum device via an FPGA or high-speed DAC / ADC card. The control tasks are divided into: Real-time critical tasks: timing control and rapid feedback (nanosecond-microsecond level) for photon emission / detection, implemented by FPGA. Near real-time tasks: routing decisions and protocol processing (millisecond-second level), implemented by software on the CPU. Both communicate via shared memory or a high-speed bus.
[0130] Pruning (removing unimportant neuron connections) and quantization (converting 32-bit floating-point weights to 8-bit integers) are performed on the trained DRL policy network to reduce model size and computational latency.
[0131] In large networks, a portion of the policy network (such as a feature extraction GNN) can be deployed on regional aggregation nodes for distributed state awareness and preliminary decision-making, while only the summary information is uploaded to the central controller for final coordination, thereby reducing the central load and communication latency.
[0132] The purpose of step S61 is to ensure that the intelligent routing algorithm can effectively interact with diverse quantum hardware platforms and meet real-time requirements.
[0133] S62 integrates quantum routing protocols into existing classical network infrastructure and control planes to enable collaborative operation.
[0134] SDN-based architecture: Employing the Software-Defined Networking (SDN) concept, a Quantum Network Operating System (QNOS) is designed as the controller. QNOS communicates with quantum switches / nodes via a southbound interface (such as an extended OpenFlow protocol), collecting resource status and issuing flow tables (containing routing decisions and operational instructions). Intelligent routing algorithms run as an application on QNOS.
[0135] In-band and out-of-band signaling: In-band signaling: Classical control bits and quantum bits are multiplexed in the same optical fiber (e.g., using different wavelengths) to achieve tight synchronization, but may be affected by quantum channel loss. Out-of-band signaling: Control messages are transmitted using independent, reliable classical IP networks, offering greater flexibility and robustness. A hybrid approach is typically used, with critical synchronization signals in-band and complex protocol interactions out-of-band.
[0136] Service Level Agreements (SLAs) and Queue Management: Define SLA templates for different quantum applications (such as quantum key distribution and distributed quantum computing), specifying their required fidelity, latency, and bandwidth. Routers and schedulers place requests into different priority queues based on their SLA category and apply corresponding resource allocation and routing policies.
[0137] Security and Authentication Mechanisms: Ensure the security of routing control signaling. Use classical authentication and encryption protocols (such as TLS) to protect control channels between controllers and nodes, and between nodes. Design quantum-safe mechanisms to prevent attacks against routing protocols (such as false link-state announcements and resource exhaustion attacks). For example, critical resource reservation requests may require credentials generated by quantum digital signatures (such as those based on quantum inadvertent transmission).
[0138] The purpose of step S62 is to seamlessly integrate the quantum routing protocol into the existing classical network infrastructure and control plane to achieve collaborative operation.
[0139] After actual deployment, S63 continuously monitors the network operation status, evaluates the effectiveness of routing policies, and makes dynamic adjustments and optimizations.
[0140] A lightweight agent is deployed on each quantum node to collect local performance counters: the number of successful / failed entanglement attempts for each link, historical memory usage, operation fidelity measurements, request processing latency, etc. The agent periodically (or event-triggered) pushes the aggregated data to a central monitoring database (such as the time-series database InfluxDB).
[0141] Based on the collected data, visual dashboards are built using tools such as Grafana or Kibana. The dashboards dynamically display: network topology diagrams (node colors represent load, edge thickness represents traffic), real-time throughput and success rate dashboards, historical performance trend curves, heatmaps, and fault alarms. This helps operations personnel intuitively grasp the overall status.
[0142] The deployed network environment may differ from the training environment. A secure online learning mode is set up: a "shadow" DRL agent runs in parallel in the background. This agent receives real-time network conditions, makes routing suggestions, but does not actually execute them. The results of the "shadow" decisions are compared with those of the actually executed decisions (potentially from a stable version of the policy). If, within a certain time window, the "shadow" policy consistently outperforms the simulation and has undergone sufficient safety verification (e.g., it will not cause network instability), the new policy can be gradually rolled out to a small number of nodes, replacing the old policy. This achieves continuous adaptation and evolution of the routing policy.
[0143] The purpose of step S63 is to continuously monitor the network operation status after actual deployment, evaluate the effectiveness of routing strategies, and make dynamic adjustments and optimizations when necessary.
[0144] The purpose of step S6 is to consider how to adapt the algorithms and protocols designed in the first 5 steps to real quantum hardware and control systems, solve specific challenges in engineering implementation, and plan deployment and operation and maintenance schemes.
[0145] The beneficial effects of implementing this embodiment are: (1) By constructing a quantum network graph model and dynamically representing the topology, we can accurately grasp the structural characteristics and change patterns of the quantum network, provide a solid foundation for subsequent routing decisions, effectively address the challenges of highly dynamic quantum network topologies, and ensure that the routing strategy adapts to the real-time state of the network. (2) By measuring and estimating the state of quantum resources in real time and converting it into feature vectors, various types of quantum resources can be quantified comprehensively and accurately, enabling the intelligent routing decision engine to make decisions based on rich and accurate information, avoiding routing errors caused by inaccurate resource information, and improving the scientificity and rationality of routing decisions. (3) By leveraging deep reinforcement learning to build an intelligent routing decision engine, it can learn and optimize routing strategies autonomously, adapt to different scenarios and needs, and improve the efficiency and flexibility of routing decisions without frequent manual intervention, thereby achieving intelligent and adaptive routing. (4) By constructing a quantum network simulation and evaluation platform, a comprehensive performance evaluation and benchmark test of the quantum network can be carried out before deployment, potential problems can be identified and optimized in advance, and deployment risks and costs can be reduced. After full verification, deployment to quantum hardware and control systems can ensure the stable and efficient operation of the quantum network, improve overall performance and reliability, and promote the practical application of quantum network technology.
[0146] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0147] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0148] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps. Example
[0149] Further reference Figure 2As a response to the above Figure 1 The present invention provides an embodiment of an intelligent routing optimization device, which is similar to the method shown. Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0150] like Figure 2 As shown, the intelligent routing optimization device 70 in this embodiment includes: a characterization module 71, a transformation module 72, a construction module 73, a translation module 74, an evaluation module 75, and a deployment module 76. Wherein: Characterization module 71 is used to construct a quantum network graph model and perform dynamic topological characterization of the quantum network; The conversion module 72 is used to measure and estimate the quantum resource states that affect routing decisions in the quantum network in real time based on the quantum network graph model, and convert the quantum resource states into feature vectors or state representations. Module 73 is used to build an intelligent routing decision engine based on the feature vector or state representation through deep reinforcement learning, and output routing strategies through the intelligent routing decision engine; Translation module 74 is used to translate the routing policy into executable network instructions; Evaluation module 75 is used to build a quantum network simulation and evaluation platform to perform performance evaluation and benchmark testing of quantum networks; Deployment module 76 is used to deploy quantum networks, after performance evaluation and benchmarking, onto quantum hardware and control systems.
[0151] The beneficial effects of implementing this embodiment are: (1) By constructing a quantum network graph model and dynamically representing the topology, we can accurately grasp the structural characteristics and change patterns of the quantum network, provide a solid foundation for subsequent routing decisions, effectively address the challenges of highly dynamic quantum network topologies, and ensure that the routing strategy adapts to the real-time state of the network. (2) By measuring and estimating the state of quantum resources in real time and converting it into feature vectors, various types of quantum resources can be quantified comprehensively and accurately, enabling the intelligent routing decision engine to make decisions based on rich and accurate information, avoiding routing errors caused by inaccurate resource information, and improving the scientificity and rationality of routing decisions. (3) By leveraging deep reinforcement learning to build an intelligent routing decision engine, it can learn and optimize routing strategies autonomously, adapt to different scenarios and needs, and improve the efficiency and flexibility of routing decisions without frequent manual intervention, thereby achieving intelligent and adaptive routing. (4) By constructing a quantum network simulation and evaluation platform, a comprehensive performance evaluation and benchmark test of the quantum network can be carried out before deployment, potential problems can be identified and optimized in advance, and deployment risks and costs can be reduced. After full verification, deployment to quantum hardware and control systems can ensure the stable and efficient operation of the quantum network, improve overall performance and reliability, and promote the practical application of quantum network technology. Example
[0152] To address the aforementioned technical problems, embodiments of the present invention also provide a computer device. Please refer to [link / reference needed]. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0153] The aforementioned computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that only the computer device 8 with components 81, 82, and 83 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0154] The aforementioned computer devices can be desktop computers, laptops, handheld computers, and cloud servers, among other computing devices. These devices can facilitate human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.
[0155] The aforementioned memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the aforementioned memory 81 may be an internal storage unit of the aforementioned computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the aforementioned memory 81 may also be an external storage device of the aforementioned computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the aforementioned memory 81 may also include both the internal storage unit and its external storage device of the aforementioned computer device 8. In this embodiment, the aforementioned memory 81 is typically used to store the operating system and various application software installed on the aforementioned computer device 8, such as computer-readable instructions for intelligent routing optimization methods. In addition, the aforementioned memory 81 can also be used to temporarily store various types of data that have been output or will be output.
[0156] In some embodiments, the processor 82 described above may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to execute computer-readable instructions stored in the memory 81 or to process data, such as executing computer-readable instructions for the intelligent routing optimization method described above.
[0157] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.
[0158] The beneficial effects of implementing this embodiment are as follows: It can accurately grasp the structural characteristics and changing patterns of quantum networks, providing a solid foundation for subsequent routing decisions, effectively addressing the challenges of highly dynamic quantum network topologies, and ensuring that routing strategies adapt to the real-time state of the network; it can comprehensively and accurately quantify various quantum resources, enabling the intelligent routing decision engine to make decisions based on rich and precise information, avoiding routing errors caused by inaccurate resource information, and improving the scientific and rational nature of routing decisions; it can autonomously learn and optimize routing strategies, adapting to different scenarios and needs without frequent manual intervention, improving the efficiency and flexibility of routing decisions, and achieving intelligent and adaptive routing; it allows for comprehensive performance evaluation and benchmark testing of the quantum network before deployment, identifying and optimizing potential problems in advance, reducing deployment risks and costs. After thorough verification, deployment to quantum hardware and control systems ensures stable and efficient operation of the quantum network, improves overall performance and reliability, and promotes the practical application of quantum network technology. Example
[0159] The present invention also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the intelligent routing optimization method as described above.
[0160] The beneficial effects of implementing this embodiment are as follows: It can accurately grasp the structural characteristics and changing patterns of quantum networks, providing a solid foundation for subsequent routing decisions, effectively addressing the challenges of highly dynamic quantum network topologies, and ensuring that routing strategies adapt to the real-time state of the network; it can comprehensively and accurately quantify various quantum resources, enabling the intelligent routing decision engine to make decisions based on rich and precise information, avoiding routing errors caused by inaccurate resource information, and improving the scientific and rational nature of routing decisions; it can autonomously learn and optimize routing strategies, adapting to different scenarios and needs without frequent manual intervention, improving the efficiency and flexibility of routing decisions, and achieving intelligent and adaptive routing; it allows for comprehensive performance evaluation and benchmark testing of the quantum network before deployment, identifying and optimizing potential problems in advance, reducing deployment risks and costs. After thorough verification, deployment to quantum hardware and control systems ensures stable and efficient operation of the quantum network, improves overall performance and reliability, and promotes the practical application of quantum network technology.
[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0162] Obviously, the embodiments described above are merely some embodiments of the present invention, not all embodiments. The accompanying drawings show preferred embodiments of the present invention, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the patent protection scope of this invention.
Claims
1. A smart routing optimization method, characterized in that, Includes the following steps: Construct a quantum network graph model to perform dynamic topological characterization of the quantum network; Based on the quantum network graph model, the states of various quantum resources that affect routing decisions in the quantum network are measured and estimated in real time, and the states of quantum resources are transformed into feature vectors or state representations. Based on the feature vector or state representation, an intelligent routing decision engine is constructed through deep reinforcement learning, and a routing strategy is output through the intelligent routing decision engine. Translate the routing policy into executable network instructions; Construct a quantum network simulation and evaluation platform to conduct performance evaluation and benchmark testing of quantum networks; The quantum network, after performance evaluation and benchmarking, is deployed to quantum hardware and control systems.
2. The intelligent routing optimization method according to claim 1, characterized in that, The steps for constructing a quantum network graph model and performing dynamic topological representation of the quantum network specifically include: Construct a quantum network graph model, abstracting the physical quantum network nodes and channels into a mathematical graph structure; Formalize the quality indicators of quantum services; Modeling the dynamics and uncertainties of quantum networks to characterize their time-varying properties.
3. The intelligent routing optimization method according to claim 1, characterized in that, The step of measuring and estimating the states of various quantum resources affecting routing decisions in the quantum network in real time based on the quantum network graph model, and converting the quantum resource states into feature vectors or state representations, specifically includes: Assess the quantum state transmission quality of current and future links; Perform link throughput and capacity modeling to assess the ability of each link and potential path to serve entangled requests; Assess the health status and service capabilities of relay nodes to determine their suitability as routing relay stations.
4. The intelligent routing optimization method according to claim 1, characterized in that, The steps of constructing an intelligent routing decision engine based on the feature vector or state representation through deep reinforcement learning, and outputting a routing strategy through the intelligent routing decision engine, specifically include: Design reinforcement learning states, actions, and rewards; Set up a distributed near-end strategy optimization training framework; Transfer learning is performed based on graph structure to output routing strategies.
5. The intelligent routing optimization method according to claim 1, characterized in that, The step of translating the routing policy into executable network instructions specifically includes: Generate multiple paths based on constrained integer programming; Coordinate and lock in the necessary quantum resources for the selected path; Configure local fast recovery and rerouting mechanisms.
6. The intelligent routing optimization method according to claim 1, characterized in that, The steps for constructing a quantum network simulation and evaluation platform, and for conducting performance evaluation and benchmark testing of quantum networks, specifically include: Concurrent and asynchronous events in quasi-quantum networks; evaluating the performance of quantum networks in complex dynamic scenarios. Conduct multi-dimensional performance evaluation and benchmark testing; Perform automatic hyperparameter tuning and strategy iteration.
7. The intelligent routing optimization method according to any one of claims 1 to 6, characterized in that, The steps for deploying the quantum network, after performance evaluation and benchmarking, onto quantum hardware and control systems specifically include: Design a hardware abstraction layer to provide a unified control interface for different quantum devices; Integrate quantum routing protocols into existing classical network infrastructure and control planes to enable collaborative operation; After actual deployment, continuously monitor the network operation status, evaluate the effectiveness of routing strategies, and make dynamic adjustments and optimizations.
8. A smart routing optimization device, characterized in that, include: The characterization module is used to construct quantum network graph models and perform dynamic topological characterization of quantum networks. The conversion module is used to measure and estimate the quantum resource states that affect routing decisions in the quantum network in real time based on the quantum network graph model, and convert the quantum resource states into feature vectors or state representations. A construction module is used to build an intelligent routing decision engine based on the feature vector or state representation through deep reinforcement learning, and to output a routing strategy through the intelligent routing decision engine. The translation module is used to translate the routing policy into executable network instructions; The evaluation module is used to build a quantum network simulation and evaluation platform to perform performance evaluation and benchmark testing on quantum networks. The deployment module is used to deploy quantum networks, after performance evaluation and benchmarking, onto quantum hardware and control systems.
9. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the intelligent routing optimization method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the intelligent routing optimization method as described in any one of claims 1 to 7.