Network routing and resource allocation method and device based on quantum key distribution

By using the routing policy model to process quantum key path requests in the quantum key distribution network, an optimized routing strategy is generated, which solves the problems of insufficient cost, complexity and flexibility of the routing solution in the prior art, and achieves optimization of network performance and stability improvement.

CN120165857APending Publication Date: 2025-06-17CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371079.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

Existing quantum key distribution network routing solutions have problems such as high cost and system complexity, insufficient flexibility and weak adaptability to dynamic network environments.

Method used

The network routing and resource allocation method based on quantum key distribution is adopted. By obtaining the routing strategy model trained by quantum key path requests and using the near-end strategy optimization algorithm, the routing strategy is optimized to generate a routing strategy including path information, wavelength and time slot.

Benefits of technology

It improves the optimization of network performance, ensures the stability and transmission efficiency of the network, and enhances the adaptability to dynamic network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120165857A_ABST
    Figure CN120165857A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of routing optimization, in particular to a network routing and resource allocation method and device based on quantum key distribution, which are used for improving the rationality of network routing and resource allocation. The method comprises the following steps: a first device obtains a quantum key path request; wherein the quantum key path request is used for requesting a path between the source node and the destination node. And the first device processes the quantum key path request by using a routing policy model to obtain a routing policy. Wherein the routing policy model is obtained by performing deep reinforcement learning on a plurality of historical training data based on a near-end policy optimization algorithm, and the historical training data comprises a historical source node, a historical destination node, historical network state information, a historical routing policy and network feedback information; the network feedback information is feedback information generated by executing the historical routing strategy. And the first device outputs the routing policy. Wherein the routing strategy comprises path information, wavelength and time slot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of routing optimization, and in particular, to a network routing and resource allocation method and apparatus based on quantum key distribution. Background Art

[0002] Quantum key distribution (QKD) technology is a technology that uses the characteristics of quantum mechanics to ensure communication security. It enables two communicating parties to generate and share a random and secure key to encrypt and decrypt messages. Quantum keys not only achieve the secure distribution of keys but also ensure that the key exchange between the two communicating parties cannot be stolen or tampered with by a third party.

[0003] In existing QKD solutions, due to complex system integration, high hardware costs, and lack of intelligent dynamic adjustment capabilities, network performance is limited, resources are wasted, and the blocking rate is relatively high. At the same time, the load balancing and resource allocation strategies in existing QKD solutions rely on static information and lack an intelligent prediction mechanism, making it difficult to cope with frequently changing network states.

[0004] Therefore, the QKD network routing solutions in the prior art have problems of high cost, insufficient system complexity and flexibility, and weak adaptability to dynamic network environments, and need to be improved. Summary of the Invention

[0005] Embodiments of this application provide a network routing and resource allocation method and apparatus based on quantum key distribution, which are used to improve the rationality of network routing and resource allocation.

[0006] In a first aspect, this application provides a network routing and resource allocation method based on quantum key distribution. The method includes: obtaining a quantum key path request for requesting a path between a source node and a destination node; processing the quantum key path request using a routing policy model to obtain a routing policy. The routing policy model is obtained by training multiple historical training data using the proximal policy optimization algorithm. The historical training data includes historical source nodes, historical destination nodes, historical network state information, historical routing policies, and network feedback information, and the network feedback information is feedback information generated by executing the historical routing policy; outputting the routing policy, where the routing policy includes path information, wavelength, and time slot.

[0007] With this method, since the training data of the routing policy model includes historical source nodes, historical destination nodes, historical network status information, historical routing policies, and network feedback information, and the network feedback information is the feedback information generated by executing the historical routing policies. Therefore, the routing policy generated by using this routing policy model can optimize network performance and ensure network stability and transmission efficiency. In addition, the training data of the routing policy model includes the feedback information corresponding to the routing policy to feedback the network status of the routing policy, improving the learning efficiency of the model.

[0008] In a possible embodiment, processing the quantum key path request by using a routing policy model to obtain a routing policy includes: processing the quantum key path request by using the routing policy model to obtain a current routing policy; comparing the current routing policy with the historical routing policy corresponding to the quantum key path request to determine the deviation value between the current routing policy and the historical routing policy; when the deviation value meets a preset threshold, using the current routing policy as the routing policy.

[0009] Optionally, when the deviation value does not meet the preset threshold, the method includes: determining the historical routing policy with the smallest deviation value from the current routing policy among multiple historical routing policies, and using the historical routing policy as the routing policy.

[0010] Based on this embodiment, it is possible to avoid excessive routing policy updates and ensure the stability of the network system.

[0011] In a possible embodiment, the network feedback information includes at least one of path status information, path hop count information, resource utilization information, key update information, delay information, and a preset factor.

[0012] In a possible embodiment, detecting the network operating state, where the network operating state includes network load information; updating the routing policy by using the routing policy model based on the network operating state.

[0013] Based on this embodiment, updating the routing policy based on the network operating state can ensure the stability and transmission efficiency of the network system.

[0014] Second aspect, the present application provides a network routing and resource allocation device based on quantum key distribution. The device includes: a communication module, configured to obtain a quantum key path request for requesting a path between a source node and a destination node; a processing module, configured to process the quantum key path request by using a routing policy model to obtain a routing policy. The routing policy model is obtained by training multiple historical training data by using a proximal policy optimization algorithm. The historical training data includes a historical source node, a historical destination node, historical network state information, a historical routing policy, and network feedback information. The network feedback information is feedback information generated by executing the historical routing policy; the communication module is further configured to output the routing policy, where the routing policy includes path information, wavelength, and time slot.

[0015] In a possible embodiment, the communication module processes the quantum key path request by using the routing policy model to obtain a routing policy. Specifically, the processing module is configured to: process the quantum key path request by using the routing policy model to obtain a current routing policy; compare the current routing policy with the historical routing policy corresponding to the quantum key path request to determine a deviation value between the current routing policy and the historical routing policy. When the deviation value meets a preset threshold, use the current routing policy as the routing policy.

[0016] In a possible embodiment, when the deviation value does not meet the preset threshold, the processing module is specifically configured to: determine a historical routing policy with the smallest deviation value from the current routing policy among multiple historical routing policies, and use the historical routing policy as the routing policy.

[0017] In a possible embodiment, the network feedback information includes at least one of path state information, path hop count information, resource utilization information, key update information, delay information, and a preset factor.

[0018] In a possible embodiment, the processing module is further configured to: detect a network operation state, where the network operation state includes network load information; and update the routing policy by using the routing policy model based on the network operation state.

[0019] Third aspect, the present application provides an electronic device, including:

[0020] a memory, configured to store program instructions;

[0021] a processor, configured to call the program instructions stored in the memory and execute the steps included in the method according to any one of the first aspect according to the obtained program instructions.

[0022] Fourthly, the present application provides a computer-readable storage medium storing a computer program including program instructions, which when executed by a computer, cause the computer to execute the method according to any one of the first aspect.

[0023] Fifthly, the present application provides a computer program product, which includes computer program code that, when running on a computer, causes the computer to execute the method according to any one of the first aspect.

[0024] For the technical effects brought by the second aspect to the fifth aspect and any of their designs, reference may be made to the technical effects brought by the corresponding designs in the first aspect, which will not be elaborated herein. Description of the Drawings

[0025] Figure 1 It is a schematic structural diagram of a quantum key distribution network provided by an embodiment of the present application;

[0026] Figure 2 It is a schematic flowchart of a network routing and resource allocation method based on quantum key distribution provided by an embodiment of the present application;

[0027] Figure 3 It is a schematic structural diagram of routing deep reinforcement learning provided by an embodiment of the present application;

[0028] Figure 4 It is a schematic diagram of the test result of the blocking probability of a routing policy provided by an embodiment of the present application;

[0029] Figure 5 It is a schematic diagram of the test result of the average reward of a routing policy provided by an embodiment of the present application;

[0030] Figure 6 It is a schematic diagram of the test result of the blocking probability and the average traffic arrival rate provided by an embodiment of the present application;

[0031] Figure 7 It is a schematic diagram of the test result of the resource utilization rate and the average traffic arrival rate provided by an embodiment of the present application;

[0032] Figure 8 It is a schematic structural diagram of a network routing and resource allocation device based on quantum key distribution provided by an embodiment of the present application;

[0033] Figure 9 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0034] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the following will describe the technical solutions in the embodiments of this application clearly and completely in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part rather than all of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application. Without conflict, the embodiments in this application and the features in the embodiments can be combined arbitrarily with each other. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0035] In the description and claims of this application and the above accompanying drawings, the terms "first" and "second" are used to distinguish different objects rather than to describe a specific order. In addition, the term "including" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices. "Multiple" in this application can mean at least two, for example, it can be two, three, or more, and there is no limitation in the embodiments of this application.

[0036] In the technical solutions of this application, the collection, dissemination, use, etc. of data all comply with the requirements of relevant national laws and regulations.

[0037] For ease of understanding, some terms in this application will be explained first to facilitate understanding by those skilled in the art.

[0038] 1) Quantum key distribution. Quantum key distribution is a method of transmitting encryption keys using the principles of quantum mechanics. It securely shares keys by sending quantum bits and can detect eavesdropping based on the interference of quantum states, thus ensuring the security of key transmission.

[0039] 2) Heisenberg uncertainty principle. The Heisenberg uncertainty principle is a core concept in quantum mechanics, used to indicate that certain basic physical properties (such as the position and momentum of a particle) cannot be precisely measured simultaneously. That is, the more precisely one physical quantity is measured, the greater the uncertainty of another related physical quantity.

[0040] 3) Quantum Lightpath Request (QLPR). A quantum lightpath request is a resource request in a quantum key distribution secure optical network, used to establish a quantum lightpath in the network for secure communication. This request includes quantum communication between the source node and the destination node and requires available resources in the network (such as wavelengths, time slots, etc.) to establish this quantum link.

[0041] The application scenarios involved in this application are introduced in detail below.

[0042] Figure 1 It is a schematic structural diagram of a quantum key distribution network provided by an embodiment of this application. As Figure 1 shown, the core architecture of the quantum key distribution network includes the following three types of communication channels:

[0043] 1) Quantum Signal Channel (QSCh). The QSCh can be used to transmit qubits. Among them, qubits are generated and distributed through quantum key distribution technology. The QSCh is responsible for generating secure keys. Quantum signals can transmit quantum data according to the principles of quantum mechanics (such as the uncertainty principle, the no-cloning theorem), thereby ensuring the security of data transmission.

[0044] 2) Public Interaction Channel (PICh). The PICh is used to transmit the verification key information between the sender and the receiver, that is, to transmit key confirmation and exchange information. The PICh can be used to ensure the integrity of quantum key distribution and ensure that the generated keys can be correctly used and verified.

[0045] 3) Traditional Data Channel (TDCh). The TDCh is used to transmit actual encrypted data.

[0046] The quantum transmitter (i.e., node A) includes a quantum state source, a random number generator, and a polarization modulator. Among them, the quantum state source is used to generate qubits, which is the basic channel required for qubits through quantum key distribution. The random number generator is used to generate random numbers for encoding during the transmission of qubits. The polarization modulator is used to determine the random numbers of specific polarization states, thereby ensuring that the quantum signal has a suitable quantum state during transmission.

[0047] The quantum receiver (i.e., Node B) includes a random number generator, a quantum bit measurer, and a quantum dot (QD). Among them, the quantum bit generator can be used to receive the quantum bit signal from Node A and perform quantum measurement on the received signal. The random number generated by the random number generator can be used to decode and measure the received signal.

[0048] To address the defects of high cost, poor system complexity and flexibility, and insufficient intelligence in the existing quantum key distribution network solutions, this application provides a network routing and resource allocation method based on quantum key distribution to improve the rationality of network routing and resource allocation.

[0049] The method provided in this application includes: The first device obtains a quantum key path request. Among them, the quantum key path request is used to request the path between the source node and the destination node. The first device processes the quantum key path request using a routing policy model to obtain a routing policy. Among them, the routing policy model is obtained by training multiple historical training data using the proximal policy optimization algorithm. The historical training data includes historical source nodes, historical destination nodes, historical network state information, historical routing policies, and network feedback information. The network feedback information is the feedback information generated by executing the historical routing policy. The first device outputs the routing policy. Among them, the routing policy includes path information, wavelength, and time slot.

[0050] Using this method, since the training data of the routing policy model includes historical source nodes, historical destination nodes, historical network state information, historical routing policies, and network feedback information, and the network feedback information is the feedback information generated by executing the historical routing policy. Therefore, the routing policy generated using this routing policy model can optimize network performance and ensure network stability and transmission efficiency. In addition, the feedback information corresponding to the routing policy in the training data of the routing policy model reflects the network state of this routing policy, improving the learning efficiency of the model.

[0051] Among them, the first device can be a management device for allocating network routing and resources. The first device can be included in a computer system for executing the method shown in this application, or can be a processing device in the computer system for executing the method shown in this application, such as a processor or a processing module, etc. This application does not specifically limit.

[0052] In addition, based on the above application scenario, the exemplary embodiments of this application will be described in more detail below. It should be noted that the above application scenario is only shown for the convenience of understanding the spirit and principle of this application, and the implementation manner of this application is not limited by any of them here. On the contrary, the implementation manner of this application can be applied to any applicable scenario.

[0053] Figure 2Schematic flowchart of a network routing and resource allocation method based on quantum key distribution provided by an embodiment of the present invention. Taking the first device as the execution entity as an example, the process may include the following steps:

[0054] S201, the first device obtains a quantum key path request. Among them, the quantum key path request is used to request a path between the source node and the destination node.

[0055] In one or more embodiments, the quantum key path request may be triggered by a user according to service requirements. For example, the user may trigger an option to access the destination node on the front-end page of the system through an input device (such as a mouse, keyboard, part of the touch screen) of the terminal device, or trigger an option to access the destination node in the application program. When the system detects an operation of the option to access the destination node, it may generate a corresponding quantum key path request for requesting a path between the source node and the destination node.

[0056] Optionally, the quantum key path request may carry information such as the source node, the destination node, the request arrival time, the key update period, and the resource requirements. For example, different bytes in the message of the quantum key path request are used to represent different information identifiers.

[0057] Among them, the source node may refer to the quantum node that sends the key request. The source node may be, for example, Figure 1 node A in

[0058] The destination node may refer to the quantum node that receives the key. The destination node may be, for example, Figure 1 node B in

[0059] The request arrival time may refer to the time when the quantum key path request arrives at the network.

[0060] The key update period may refer to the time interval for key update. The key update period may be preset according to the security requirements of the key to meet the security requirements of the network.

[0061] The resource requirements may refer to the wavelength and time slot resources required by the quantum key path request. Among them, the resource requirements may also be used to create and / or refine the key.

[0062] In one or more embodiments, the quantum key path request may also be generated by the first device according to preset rules. For example, the first device may periodically update the routing path between the source node and the destination node to avoid network congestion, improve resource utilization, and system flexibility. The first device may generate a corresponding quantum key path request before updating the routing path.

[0063] Or, the first device may receive a quantum key path request from other devices.

[0064] S202. The first device processes the quantum key path request using a routing policy model to obtain a routing policy.

[0065] Among them, the routing policy model is obtained by training multiple historical training data using the proximal policy optimization algorithm. The historical training data includes historical source nodes, historical destination nodes, historical network state information, historical routing policies, and network feedback information. The network feedback information is the feedback information generated by executing the historical routing policy.

[0066] In one or more embodiments, the routing policy model can be a pre-trained model. The routing policy model can be used to process the quantum key path request. That is, taking the quantum key path request as the input data of the routing policy model, the routing policy model can process the quantum key path request and output the corresponding routing policy.

[0067] Exemplarily, the routing policy model is a model that can be based on deep reinforcement learning (DRL) and can be obtained by training the model training data using the proximal policy optimization algorithm.

[0068] S203. The first device outputs the routing policy. Among them, the routing policy includes path information, wavelength, and time slot.

[0069] In one or more embodiments, after the first device determines the routing policy corresponding to the quantum key path request, it can send the routing policy to the source node and / or the destination node. Correspondingly, the source node and / or the destination node receive the routing policy from the first device and establish a path between the source node and the destination node based on the routing policy.

[0070] The following is a detailed description of the routing policy model in step S202 based on Figure 3 the content shown. Figure 3 It is a schematic structural diagram of a routing deep reinforcement learning provided for the implementation of this application.

[0071] As Figure 3 shown, the routing policy model is deployed in the DRL agent. Correspondingly, the DRL agent can use the routing policy model to output the corresponding routing policy.

[0072] The software defined networking controller (SDN) can collect the network state information in the network environment and send the network state information to the DRL agent. Among them, the network state can be represented as (state, S t ).

[0073] Correspondingly, the DRL agent receives network status information from the SDN controller and processes the network status information using a routing policy model to output an optimal routing policy. Among them, the routing policy can also be referred to as an action, which can be expressed as (action, a t ).

[0074] After the routing policy is executed, the SDN controller can collect network feedback information output by the network environment and send the network feedback information to the DRL agent. Correspondingly, the DRL agent can generate corresponding rewards based on the network feedback information. Among them, the rewards can be used to indicate network performance metrics. Exemplarily, the network performance metrics include the success rate of quantum key path requests, the utilization rate of network resources, the network blocking rate, etc. When the quantum key path request is successful, the DRL agent generates a positive reward. When the quantum key path request fails, the DRL agent generates a negative reward.

[0075] The DRL agent can store the network status information, routing policy, and corresponding rewards in an experience buffer so that the DRL agent can train the routing decision model based on this information to optimize the routing decision model.

[0076] The following is a detailed introduction to the network status information:

[0077] In one or more embodiments, the network status information may include current network topology information, in-service quantum key path request information, link resource utilization information, candidate path information, and status information in the time dimension.

[0078] Among them, the network topology information may include quantum nodes, quantum links, and the connection relationships of quantum nodes. Exemplarily, in a quantum key distribution network, different quantum nodes may represent different quantum devices in the network, and the chain links may be represented as fiber connections between quantum devices. Different links may also include information on available resources (such as wavelengths, time slots, etc.), and the information on available resources can be used to indicate the current resources available for quantum key distribution.

[0079] The in-service quantum key path request information refers to the quantum key requests currently being processed.

[0080] The link resource utilization information may include resource information such as the number of available wavelengths and the number of time slots on the link. It can be understood that based on the link resource utilization information, the congestion degree of each link can be determined, and a path with sufficient resources can be selected based on the congestion degree to reduce network blocking.

[0081] Candidate path information may include multiple path information between a source node and a destination node. The path information may include the number of hops from the source node to the destination node and the resource status of the link. Additionally, the path information may further include the status information of the path. For example, when the available resources of the path are insufficient or occupied, the status information of the path is unavailable. It is understandable that based on the candidate path information, the corresponding path can be determined quickly and accurately, avoiding the selection of congested paths.

[0082] The status information in the time dimension may include the time information when the quantum key request arrives at the network, the key update time information, and the validity period of the key. It is understandable that when determining the routing decision, the long-term utilization information of the path resources can be determined based on the status information in the time dimension, so that when a new quantum key request appears, the resources can be adjusted to achieve dynamic resource allocation, thereby ensuring that the key can be updated in a timely manner and the resources can be utilized reasonably.

[0083] Exemplarily, the source node can be represented as o t , the destination node can be represented as d t , the time slot resource can be represented as The wavelength resource can be represented as The number of time slots can be represented as K T , the number of wavelengths can be represented as W T , the key update period can be represented as T, the number of hops can be represented as H t , the network status information can be represented as s t , then the network status information, source node, destination node, time slot resource, wavelength resource, number of time slots, number of wavelengths, key update period, and number of hops satisfy:

[0084]

[0085] That is to say, the network status information can be represented as a vector including multiple dimensions, and the network status vector can be used to indicate the current resource utilization information and link status information of the network.

[0086] The process of generating the corresponding reward based on the network feedback information is introduced below:

[0087] In one or more embodiments, the network feedback information includes at least one of path status information, path hop count information, resource utilization information, key update information, delay information, and a preset factor. Among them, the path status information can be used to indicate the path status of the quantum key path request. The path hop count information can be used to indicate the number of hops of the path. The resource utilization information can be used to indicate the overall resource utilization rate of the network. The key update information can be used to indicate the resource information occupied by updating the key. The delay information can be used to indicate the delay time of the path. The preset factor can be set in advance according to requirements.

[0088] Optionally, the DRL agent can generate corresponding rewards based on the reward functions corresponding to different network feedback information.

[0089] Exemplarily, the reward functions corresponding to the path status information, path hop count information, resource utilization information, key update information, latency information, and preset factors are introduced below, respectively.

[0090] 1), Path status information:

[0091] When the path status information indicates that resources have been successfully allocated for the quantum key path request and the path has been established, the system generates a positive reward +R s . When the path status information indicates that resource allocation has failed, resulting in the QLR request being blocked or failed, the system generates a negative reward -R f .

[0092] The reward function corresponding to the path status information satisfies:

[0093]

[0094] where r t represents the reward corresponding to the network feedback information.

[0095] 2), Path hop count information:

[0096] When the path hop count information indicates that a path with fewer hops is selected, the system generates a positive reward +R h . When the path hop count information indicates that a path with more hops is selected, the system generates a negative reward -R h .

[0097] The reward function corresponding to the path hop count information satisfies:

[0098]

[0099] where r t represents the reward corresponding to the network feedback information.

[0100] 3), Resource utilization information:

[0101] When the resource utilization information indicates that a path and time slot with higher resource utilization are selected, the system generates a positive reward +R u . When the resource utilization information indicates that a path and time slot with lower resource utilization are selected, the system generates a negative reward -R u .

[0102] The reward function corresponding to the resource utilization information satisfies:

[0103]

[0104] Among them, r t represents the reward corresponding to the network feedback information.

[0105] 4), Key update information:

[0106] When the remaining resources meet the resources required for key update indicated by the key update information, the system generates a positive reward +R k . When the remaining resources do not meet the resources required for key update indicated by the key update information, the system generates a negative reward -R k .

[0107] The reward function corresponding to the key update information satisfies:

[0108]

[0109] Among them, r t represents the reward corresponding to the network feedback information.

[0110] 5), Delay information:

[0111] When the delay information indicates to select a low-delay path, the system generates a positive reward +R d . When the delay information indicates to select a high-delay path, the system generates a negative reward -R d .

[0112] The reward function corresponding to the delay information satisfies:

[0113]

[0114] Among them, r t represents the reward corresponding to the network feedback information.

[0115] 6), Preset factor:

[0116] The preset factor can also be called the discount factor, which can be expressed as γ. The preset factor can be used to balance the current reward and the future reward. The value range of the preset factor belongs to [0, 1]. Among them, the larger the value of the preset factor, the greater the proportion of the cumulative reward. The goal of the long-term reward is to maximize the cumulative return G t starting from the current time t:

[0117]

[0118] Among them, r t+j is the reward at the (t + j)-th moment, and γ can be used to indicate the discount degree of the future reward.

[0119] Further, when the network feedback information is path status information, path hop count information, resource utilization information, key update information, delay information, and a preset factor, the reward function corresponding to the network feedback information satisfies:

[0120] r t

[0121] = R s ·II(QLR success) - R f ·II(QLR failure) + R h ·II(hop count low) - R h ·II(hop count high) + R u ·II(high resource utilization) - R u ·II(low resource utilization) + R k ·II(key update success) - R k ·II(key update failure) + R d ·II(low delay) - R d ·II(high delay)

[0122] where II() is an indicator function used to determine whether a certain condition holds. If the condition holds, the return value is 1, otherwise it is 0.

[0123] It can be understood that the reward function corresponding to the network feedback information combines various factors such as path status information, path hop count information, resource utilization information, key update information, delay information, and a preset factor, so as to achieve comprehensive optimization of the routing selection and resource allocation strategies.

[0124] The process of determining the routing strategy using the routing decision model is introduced as follows:

[0125] The routing strategy includes the routing path and the wavelengths and time slots allocated to the links on this path. Therefore, the determination of the routing strategy can be introduced from four aspects: path selection, wavelength allocation, time slot allocation, and combined routing strategy. In addition, the routing strategy can also be an action.

[0126] 1). Path selection:

[0127] After obtaining the quantum key path request, a candidate path set containing multiple paths can be generated based on the routing decision model, and an optimal path can be selected from this candidate path set. Among them, each path in the candidate path set at different times is from the source node o t to the destination node dt Communication paths. Each path in the candidate path set can be represented as P, and the candidate path set satisfies:

[0128] Candidate path set = {P1, P2, …, P K};

[0129] The paths in the selected path set satisfy:

[0130] P k = {n1, n2, …, n h};

[0131] where n i represents the nodes in the path, and h is the hop count of the path.

[0132] Selecting an optimal path from the path sets at different times can be represented as Then the optimal path satisfies:

[0133]

[0134] 2), Wavelength allocation:

[0135] After determining the optimal path, a wavelength needs to be allocated to each link on this path. Among them, there are M available wavelengths on each link, and the wavelengths allocated to each link on the entire path are the same. The selection space of wavelengths can be represented as W, and the selection space of wavelengths satisfies:

[0136] W = {w1, w2, …, w M};

[0137] where w i represents the i-th wavelength on the link.

[0138] The wavelengths allocated at different times can be represented as Then the wavelengths allocated at different times satisfy:

[0139]

[0140] 3), Time slot allocation:

[0141] After determining the optimal path, time slots need to be allocated to each link on this path, and the time slots allocated to each link on the entire path are the same. There are N time slots available for allocation on each link, and the selection space of time slots can be represented as K, and the selection space of time slots satisfies:

[0142] K = {k1, k2, …, k N};

[0143] where k j represents the j-th time slot on the link.

[0144] 4) Composite action space:

[0145] Actions at different times are obtained by combining the path selection, wavelength allocation, and time slot allocation at that time. Among them, the action space can be represented as the Cartesian product of paths, wavelengths, and time slots. The routing policy can be represented as a t , then the routing policy, path, wavelength, and time slot satisfy:

[0146]

[0147] The size of the action space is:

[0148] |A| = K × M × N;

[0149] where K is the number of candidate paths, M is the number of wavelengths, and N is the number of time slots.

[0150] It can be understood that from the above formula, the size of the action space increases exponentially with the increase in the number of candidate paths, wavelengths, and time slots.

[0151] Based on this method, the action space can be reasonably defined, thereby improving the learning efficiency of the routing policy model.

[0152] Next, the algorithm in the deep reinforcement learning process of the routing policy model will be introduced.

[0153] In one or more embodiments, the first device processes the quantum key path request using the routing policy model to obtain a routing policy, including: The first device processes the quantum key path request using the routing policy model to obtain the current routing policy. The first device compares the current routing policy with the historical routing policy corresponding to the quantum key path request to determine the deviation value between the current routing policy and the historical routing policy. When the deviation value meets the preset threshold, the current routing policy can be used as the routing policy.

[0154] Optionally, when the deviation value does not meet the preset threshold, the first device can determine the historical routing policy with the smallest deviation value from the current routing policy among multiple historical routing policies and use this historical routing policy as the routing policy.

[0155] Exemplarily, the algorithm of deep reinforcement learning can be the proximal policy optimization (PPO) algorithm.

[0156] Based on the PPO algorithm, the routing policy can be optimized by clipping the objective function while ensuring the stability of the routing policy.

[0157] The form of the objective function can be expressed as:

[0158]

[0159] Among them, r t (θ) is the policy ratio, representing the current routing policy πθ(a t |s t ) relative to the historical routing policy . That is, the policy ratio can be a form of manifestation of the deviation value between two routing policies.

[0160] The policy ratio satisfies:

[0161]

[0162] is the advantage function, which can be used to indicate the performance of the current action relative to the average behavior.

[0163] The advantage function satisfies:

[0164]

[0165] Among them, Q(s t , a t ) is the action value function, and V(s t ) is the state value function.

[0166] Clipping mechanism: When the policy ratio r t (θ) does not belong to (1 - ∈, 1 + ∈), it is directly clipped to the boundary values of this interval to prevent excessive policy updates and ensure stability. That is, (1 - ∈, 1 + ∈) is the preset threshold.

[0167] Furthermore, the PPO algorithm can optimize the routing policy based on the policy network and the value network.

[0168] Among them, policy network: The PPO algorithm can input the state of the current network into the policy network, and the policy network outputs the probability distribution corresponding to this state, so that the optimal routing path and resource allocation scheme can be selected based on this probability distribution. Among them, the network state can be expressed as s t , and the probability distribution can be expressed as π θ (a t |s t ), where a t is the routing policy.

[0169] Value network: The PPO algorithm can determine the value of the current network state based on the value network, so as to provide a reference for the long-term reward of the update of the routing policy. Among them, the value of the network state output by the value network can be expressed as V(s t)。It is understandable that according to the above formula, the value of the network state can be used to determine the advantage function.

[0170] In addition, the PPO algorithm can also store information such as the network state, routing policy, and reward corresponding to each output decision in an experience replay buffer for subsequent training based on the data stored in the experience replay buffer. It is understandable that training based on the data in the experience replay buffer can break the temporal correlation and improve the utilization rate of training data. For example, the experience replay buffer can store information such as network states, routing policies, and rewards at multiple moments. When training the routing policy model, training data can be randomly sampled from the experience replay buffer for training.

[0171] In one or more embodiments, the first device can detect the operating state of the network. The network operating state includes network load information. The first device can update the routing policy using a routing policy model based on the network operating state.

[0172] Exemplarily, taking the above DRL agent as an example, the DRL agent can satisfy the quantum key path request by selecting the optimal routing path and allocating wavelengths and time slots, and dynamically adjust the routing policy based on changes in network load.

[0173] The DRL agent dynamically adjusts the routing policy based on changes in network load, including the following steps:

[0174] Step 1: Determine the initial routing policy.

[0175] When the DRL agent receives a quantum key path request, it first selects the optimal path for routing and allocates wavelengths and time slots on this path.

[0176] Among them, the resource allocation follows the following rules:

[0177] Path selection: Based on the available resources on each path, select the optimal path P K from the candidate path set {P1, P2, …, P k}:

[0178]

[0179] where h i represents the hop count of path P i .

[0180] Wavelength allocation: Allocate wavelengths w that are consistent and continuous for each link on the optimal path.

[0181] The wavelength satisfies:

[0182] W = {w1, w2, …, wM}

[0183]

[0184] Among them, W represents the set of wavelengths, and L k represents the number of links on path P k .

[0185] Time slot allocation: Allocate the time slot k with the minimum delay and the highest availability to each link on the optimal path. The time slots satisfy:

[0186] K = {k1, k2, …, k N};

[0187]

[0188] Among them, K represents the set of time slots, and delay(k j ) represents the transmission delay corresponding to time slot k j .

[0189] Step 2: Dynamically adjust the routing strategy based on the network load.

[0190] Among them, as the network state changes, the DRL agent can detect the network state in real time and adjust the routing strategy based on the current network state to improve the network performance. Among them, the adjustment methods include load balancing and path reallocation.

[0191] Load balancing means that when the load on a link is overloaded, the traffic is allocated to a path with sufficient resources, thereby reducing the blocking rate.

[0192] Path reallocation means that when a link fails or there is a shortage of resources, the path can be dynamically adjusted to ensure the completion of the QLR request.

[0193] It can be understood that when dynamically adjusting the routing strategy, the blocking rate can be minimized by balancing the network load and maximizing the number of successfully allocated QLRs; the utilization efficiency of wavelengths and time slots can be improved to avoid resource waste, thereby maximizing the resource utilization rate; ensure that the key update operation is carried out on time and reserve sufficient resources to meet future needs, thereby ensuring key update.

[0194] Optionally, the routing strategies corresponding to quantum key path requests at different times can be represented as a t , the optimal path can be represented as P k , the wavelength can be represented as w, and the time slot can be represented as k. Then the routing strategy, optimal path, wavelength, and time slot satisfy:

[0195] a t = (P k , w, k);

[0196] Among them, the routing policy can be determined based on the corresponding reward function, that is, the routing policy with the largest reward value is determined from multiple routing policies. As can be seen from the above, the reward function is composed of factors such as resource utilization rate, blocking rate, and key update.

[0197] The beneficial effects of the method provided by this application are described below through model training and testing.

[0198] Example 1, during the training process of the routing policy model, the blocking probability is an important indicator to measure the effectiveness of the routing policy. Therefore, the effectiveness of the routing policy can be indicated by the blocking probability.

[0199] Figure 4 It is a schematic diagram of the comparative analysis of the blocking probability between the routing policy based on deep reinforcement learning and the benchmark algorithms of the prior art in different network environments. Among them, the network environments include NSFNET, UBN24, European Backbone, and AsiaNet. NSFNET corresponds to Figure 4 in (a), UBN24 corresponds to Figure 4 in (b), European Backbone corresponds to Figure 4 in (c), and AsiaNet corresponds to Figure 4 in (d).

[0200] (a) NSFNET network: In the NSFNET network, the initial blocking probability is relatively high, especially the RF (Random Fit) and FF (First Fit) algorithms show relatively large blocking probabilities. The DRL-base RRA scheme has a rapid decline in the blocking probability as the number of training iterations increases, and reaches a relatively low stable level after 40,000 iterations, significantly outperforming other benchmark methods.

[0201] (b) UBN24 network: In the UBN24 network, the DRL-base RRA scheme also shows strong learning ability. The blocking probability is relatively high at the beginning of training, but it is gradually optimized as the number of iterations increases, and finally outperforms all benchmark routing methods. Especially when approaching 50,000 iterations, the blocking probability of DRL drops to the lowest level.

[0202] (c) European Backbone: In the European Backbone network, the blocking probabilities of other methods such as SP (Shortest Path) and HC (Hop Count) algorithms are relatively stable. However, in contrast, the blocking probability of DRL-base RRA decreases significantly as the training process progresses, and always maintains the best performance.

[0203] (d) AsiaNet Network: In the AsiaNet network environment, the DRL agent also performs excellently. Although the initial blocking probability is relatively high, as the number of iterations increases, DRL-base RRA gradually learns how to optimally allocate resources, and finally the blocking probability drops to the lowest, outperforming DQN and other benchmark algorithms.

[0204] It can be understood that based on Figure 4 it can be seen that in the learning and optimization process of the model provided by this application in different network environments, as the training iterations increase, the blocking probability corresponding to the network policy gradually decreases.

[0205] Example 2, Average Reward (AR) is a key metric for measuring the learning effect of the model during model training. Therefore, the learning result of the model can be indicated by analyzing the change of the average reward with the training iterations in different network environments.

[0206] Figure 5 It is a schematic diagram comparing the change results of the average rewards of the DRL-based routing policy and the DQN method with the number of training iterations in different network environments. Among them, the network environments include NSFNET, UBN24, European Backbone, and AsiaNet networks. NSFNET corresponds to Figure 5 the (a) in Figure 5 UBN24 corresponds to Figure 5 the (b) in Figure 5 European Backbone corresponds to

[0207] (a) NSFNET Network: In the NSFNET network, the average reward of the DRL-base RRA scheme is significantly better than that of DQN. As the training iterations proceed, the average reward of DRL increases rapidly, reaches a relatively high level after 10,000 iterations, and stabilizes at 40,000 iterations. In contrast, the learning curve of DQN is relatively flat, and the final average reward value is also lower than that of the DRL scheme.

[0208] (b) UBN24 Network: In the UBN24 network, DRL-base RRA also shows a faster learning speed and higher reward value. Although DQN performs well in the initial stage of iteration, in the subsequent training process, the DRL agent quickly surpasses DQN, showing stronger resource allocation and path optimization capabilities.

[0209] (c) European Backbone: In the European Backbone network, the average reward of DRL-base RRA increases significantly and exceeds that of DQN in the earlier training stage. As the training progresses, the average reward of DRL gradually stabilizes and remains at a higher level, indicating its strong adaptability under complex network conditions.

[0210] (d) AsiaNet network: In the AsiaNet network, the average reward of DRL-base RRA also increases faster than that of DQN.

[0211] Based on Figure 5 it can be seen that the model training provided by this application can learn network states faster and more effectively, optimize resource allocation and path selection, and ultimately achieve higher cumulative rewards.

[0212] Example 3, the adaptability of the routing policy model can be determined by analyzing the impact of the traffic arrival rate on the blocking probability of the routing policy.

[0213] Figure 6 The relationship between the blocking probability (BP) and the average traffic arrival rate in different network environments. The test network environments include NSFNET, UBN24, European Backbone, and AsiaNet. Under different traffic arrival rates, the performance of the DRL-base RRA method is compared with several other benchmark routing algorithms, including random fit (RF), first fit (FF), shortest path (SP), hop count (HC), and DQN.

[0214] (a) NSFNET network: In the NSFNET network, as the traffic arrival rate gradually increases from low to high, the blocking probability of all algorithms also increases accordingly. However, the DRL-base RRA solution always maintains a lower blocking probability, especially in the case of high load (>0.7), and its performance is particularly outstanding compared with other algorithms.

[0215] (b) UBN24 network: In the UBN24 network, as the traffic arrival rate increases, the blocking probabilities of the RF and FF algorithms rise rapidly, especially the blocking rate is significantly increased under high load. The DRL-base RRA solution can significantly reduce the blocking probability and always maintain better performance than other methods.

[0216] (c) European Backbone: In the European Backbone network, the blocking probabilities of various algorithms also increase with the increase in the traffic arrival rate. The DRL-base RRA demonstrates strong resource optimization capabilities at medium and high loads, and its blocking probability is much lower than that of traditional routing methods such as SP and HC.

[0217] (d) AsiaNet network: In the AsiaNet network, when the traffic arrival rate increases, the blocking probability of DRL-base RRA has the smallest increase, and it can still maintain a low blocking probability under high load conditions, outperforming DQN and other benchmark algorithms.

[0218] Based on Figure 6 it can be seen that the model training scheme provided by this application can effectively reduce the blocking rate through an intelligent resource allocation strategy under high load conditions.

[0219] Example 4 can determine the optimization ability of the routing policy model in resource utilization by analyzing the impact of the traffic arrival rate on the resource utilization rate of the routing policy.

[0220] Figure 7 To test the relationship between the resource utilization rate (RU) and the average traffic arrival rate in different network environments. The test network environments include NSFNET, UBN24, European Backbone, and AsiaNet networks. The figure compares the resource utilization performance of the DRL-base RRA method and several benchmark algorithms at different traffic arrival rates.

[0221] (a) NSFNET network: In the NSFNET network, the DRL-base RRA scheme always maintains a high resource utilization rate, especially outstanding under high traffic loads, significantly outperforming traditional methods such as RF and FF, demonstrating its advantages in resource optimization.

[0222] (b) UBN24 network: In the UBN24 network environment, as the traffic arrival rate increases, the resource utilization rate of all algorithms increases. However, the DRL-base RRA scheme has a higher resource utilization rate than the other several algorithms both under low and high traffic loads, especially showing its strong resource optimization ability under high load.

[0223] (c) European Backbone Network: In the European Backbone Network, the resource utilization rate of each algorithm gradually increases as the traffic arrival rate increases. The DRL agent has a better resource utilization rate than other algorithms under various load conditions, especially in medium and high traffic loads, indicating its advantage in resource scheduling under complex network conditions.

[0224] (d) AsiaNet Network: In the AsiaNet Network, as the traffic load increases, the DRL-base RRA scheme shows obvious advantages in terms of resource utilization rate. Especially under high load conditions, its resource utilization rate is always higher than other benchmark algorithms, proving its excellent adaptability under high traffic loads.

[0225] Based on Figure 7 it can be seen that the model training scheme provided by this application can maximize the utilization rate of network resources through an intelligent resource management strategy under the condition of high traffic arrival rate.

[0226] Based on Examples 1 - 4, it can be seen that the scheme provided by this application can not only effectively reduce the blocking rate, but also significantly improve the resource usage efficiency of the network, fully demonstrating its adaptability and optimization performance in the actual network environment.

[0227] Based on the same inventive concept, an embodiment of this application provides a network routing and resource allocation device based on quantum key distribution. Figure 8 The structural schematic diagram of a network routing and resource allocation device based on quantum key distribution provided by an embodiment of this application is shown. As Figure 8 shown, the device includes a communication module 801 and a processing module 802.

[0228] The communication module 801 is used to obtain a quantum key path request, and the quantum key path request is used to request a path between a source node and a destination node; the processing module 802 is used to process the quantum key path request by using a routing policy model to obtain a routing policy. The routing policy model is obtained by training multiple historical training data by using the proximal policy optimization algorithm. The historical training data includes historical source nodes, historical destination nodes, historical network state information, historical routing policies, and network feedback information. The network feedback information is the feedback information generated by executing the historical routing policy; the communication module 801 is further used to output the routing policy, and the routing policy includes path information, wavelength, and time slot.

[0229] In a possible embodiment, the communication module 801 processes the quantum key path request by using a routing policy model to obtain a routing policy. Specifically, the processing module 802 is configured to: process the quantum key path request by using the routing policy model to obtain a current routing policy; compare the current routing policy with the historical routing policy corresponding to the quantum key path request to determine a deviation value between the current routing policy and the historical routing policy. When the deviation value meets a preset threshold, use the current routing policy as the routing policy.

[0230] In a possible embodiment, when the deviation value does not meet the preset threshold, the processing module 802 is specifically configured to: determine, from multiple historical routing policies, a historical routing policy with the smallest deviation value from the current routing policy, and use the historical routing policy as the routing policy.

[0231] In a possible embodiment, the network feedback information includes at least one of path status information, path hop count information, resource utilization information, key update information, delay information, and a preset factor.

[0232] In a possible embodiment, the processing module 802 is further configured to: detect a network operating state, where the network operating state includes network load information; and update the routing policy by using the routing policy model based on the network operating state.

[0233] Based on the same inventive concept, an embodiment of the present application provides an electronic device, which can implement the functions of the device described above. Figure 9 FIG. shows a schematic structural diagram of an electronic device provided by an embodiment of the present application.

[0234] The electronic device in the embodiment of the present application may include a processor 901. The processor 901 is the control center of the device, and can connect various parts of the device by using various interfaces and lines, and run or execute instructions stored in the memory 903 and call data stored in the memory 903. Optionally, the processor 901 may include one or more processing units. The processor 901 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 901. In some embodiments, the processor 901 and the memory 903 may be implemented on the same chip, and in some embodiments, they may also be implemented on separate chips independently.

[0235] The processor 901 may be a general-purpose processor, such as a Central Processing Unit (CPU), a digital signal processor, an application specific integrated circuit, a field programmable gate array or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The method steps disclosed in combination with the embodiments of the present application may be directly executed by the hardware processor, or executed by a combination of hardware and software modules in the processor.

[0236] In the embodiments of the present application, the memory 903 stores instructions executable by at least one processor 901, and the at least one processor 901 can be used to execute the method steps disclosed in the embodiments of the present application by executing the instructions stored in the memory 903.

[0237] As a non-volatile computer-readable storage medium, the memory 903 can be used to store non-volatile software programs, non-volatile computer-executable programs and modules. The memory 903 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, Random Access Memory (RAM), Static Random Access Memory (SRAM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. The memory 903 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 903 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0238] In the embodiments of the present application, the device may further include a communication interface 902, and the electronic device can transmit data through the communication interface 902.

[0239] Optionally, it can be implemented by Figure 9 the shown processor 901 (or the processor 901 and the communication interface 902) Figure 8The processing module 802 and / or the communication module 801 shown, that is to say, the actions of the processing module 802 and / or the communication module 801 can be executed by the processor 901 (or the processor 901 and the communication interface 902).

[0240] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, in which instructions can be stored. When the instructions are run on a computer, the computer is made to execute the operation steps provided in the above method embodiment. The computer-readable storage medium can be Figure 9 the memory 903 shown.

[0241] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] The present application is described with reference to the flowcharts and / or block diagrams of the method, device (system), and computer program product according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0243] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0244] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocksFigure 1 Steps of functions specified in one or more boxes.

[0245] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

Claims

1. A network routing and resource allocation method based on quantum key distribution, characterized in that: The method comprises: Obtaining a quantum key path request, where the quantum key path request is used to request a path between a source node and a destination node; A routing strategy model is used to process the quantum key path request to obtain a routing strategy, wherein the routing strategy model is obtained by training a plurality of historical training data using a proximal strategy optimization algorithm, wherein the historical training data includes a historical source node, a historical destination node, historical network status information, a historical routing strategy, and network feedback information, and the network feedback information is feedback information generated by executing the historical routing strategy; The routing strategy is output, where the routing strategy includes path information, wavelength, and time slot.

2. The method according to claim 1, characterized in that The routing strategy model is used to process the quantum key path request to obtain a routing strategy, including: Processing the quantum key path request using the routing strategy model to obtain a current routing strategy; Comparing the current routing strategy with the historical routing strategy corresponding to the quantum key path request, and determining a deviation value between the current routing strategy and the historical routing strategy; When the deviation value meets a preset threshold, the current routing strategy is used as the routing strategy.

3. The method according to claim 2, characterized in that The deviation value does not meet a preset threshold, and the method includes: A historical routing strategy with the smallest deviation value from the current routing strategy is determined from multiple historical routing strategies, and the historical routing strategy is used as the routing strategy.

4. The method according to any one of claims 1 to 3, characterized in that The network feedback information includes at least one of path status information, path hop count information, resource utilization information, key update information, delay information and a preset factor.

5. The method according to claim 1, characterized in that The method further comprises: Detecting a network operation status, wherein the network operation status includes network load information; The routing policy model is used to update the routing policy based on the network operation status.

6. A network routing and resource allocation device based on quantum key distribution, characterized in that: The device comprises: A communication module, used to obtain a quantum key path request, where the quantum key path request is used to request a path between a source node and a destination node; A processing module, configured to process the quantum key path request using a routing strategy model to obtain a routing strategy, wherein the routing strategy model is obtained by training a plurality of historical training data using a proximal strategy optimization algorithm, wherein the historical training data includes historical source nodes, historical destination nodes, historical network status information, historical routing strategies, and network feedback information, wherein the network feedback information is feedback information generated by executing the historical routing strategy; The communication module is further used to output the routing strategy, which includes path information, wavelength and time slot.

7. The device according to claim 6, characterized in that The routing strategy model is used to process the quantum key path request to obtain a routing strategy, and the processing module is specifically used to: Processing the quantum key path request using the routing strategy model to obtain a current routing strategy; Comparing the current routing strategy with the historical routing strategy corresponding to the quantum key path request, and determining a deviation value between the current routing strategy and the historical routing strategy; When the deviation value meets a preset threshold, the current routing strategy is used as the routing strategy.

8. The device according to claim 7, characterized in that The deviation value does not meet the preset threshold value, and the processing module is specifically used to: A historical routing strategy with the smallest deviation value from the current routing strategy is determined from multiple historical routing strategies, and the historical routing strategy is used as the routing strategy.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method according to any one of claims 1 to 5.

10. A computer program product, characterized in that The computer program product comprises: a computer program code, and when the computer program code is run on a computer, the computer is enabled to execute the method according to any one of claims 1 to 5.