A communication coordination system for distributed quantum computing

By constructing a dynamic heterogeneous computation graph and a deep reinforcement learning model, combined with a high-fidelity digital twin model, the problem of low resource utilization in existing technologies is solved, and efficient coordination and adaptive scheduling of distributed quantum computing systems are achieved, improving task success rate and system stability.

CN122433935APending Publication Date: 2026-07-21BEIJING XIDIAN QUANTUM TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING XIDIAN QUANTUM TECHNOLOGY CO LTD
Filing Date
2026-04-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing quantum computing task scheduling and management strategies fail to deeply perceive the heterogeneous characteristics of hardware and monitor the performance status of nodes in real time, resulting in low resource utilization, extended task execution time, and even quantum state decoherence, which seriously restricts the efficiency and feasibility of distributed quantum computing.

Method used

By constructing a dynamic heterogeneous computing graph through a monitoring and perception module, an intelligent decision-making module, a prediction and inference module, and an execution and learning module, scheduling decisions are generated using a deep reinforcement learning model. Combined with a high-fidelity digital twin model, forward-looking inferences are performed to identify performance degradation and network congestion risks, generate remedial plans, and achieve elastic coordination and adaptive scheduling.

Benefits of technology

It improved resource utilization efficiency, increased task success rate and overall coordination intelligence, continuously optimized resource allocation and task execution, and reduced task waiting and decoherence risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433935A_ABST
    Figure CN122433935A_ABST
Patent Text Reader

Abstract

The application relates to a communication coordination system for distributed quantum computing, and in particular to the field of distributed quantum computing, wherein a dynamic computing graph is constructed through real-time sensing and unified representation of the states of heterogeneous quantum processors and networks, and a deep reinforcement learning model is used to generate an optimized scheduling decision; a high-fidelity digital twin can perform forward-looking deduction on the decision scheme, identify performance degradation and network congestion risks, and generate an evaluated remedial plan; during execution, the baseline scheme and the optimized plan can be fused to generate an elastic instruction sequence, and dynamic rescheduling can be triggered according to the deviation between the actual and predicted values, so that adaptive coordination is achieved; the entire scheme can also use execution feedback data to periodically perform incremental learning on the scheduling model and the digital twin model, thereby continuously improving resource utilization efficiency, task success rate and overall coordination intelligence level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed quantum computing, and more specifically, to a communication coordination system for distributed quantum computing. Background Technology

[0002] As quantum computing technology advances towards practical application, heterogeneous distributed computing clusters integrating quantum processors from multiple physical systems have become an important architecture for overcoming the limitations of single-system scale and performance. These clusters are typically composed of quantum chips employing different technological approaches, such as superconducting quantum processors and ion trap quantum processors, aiming to integrate their respective advantages in terms of qubit count, coherence time, gate operation fidelity, and connectivity to execute complex algorithms such as large-scale quantum optimization and quantum simulation. In practical operating environments, these algorithms are often decomposed into subtasks with different resource requirements; for example, some modules may require high-fidelity two-qubit gates. Some rely on long-range entanglement or a large number of qubits. Meanwhile, the performance of quantum processors is significantly affected by multiple external factors such as temperature fluctuations and electromagnetic interference in the laboratory environment. This causes dynamic changes in the stability of key components such as superconducting qubits based on Josephson junctions or their equivalent nonlinear oscillators. Consequently, core parameters such as the number of available qubits, gate error rate, and quantum state lifetime of each computing node are not constant, and the bandwidth and latency of the communication link also fluctuate. This inherent heterogeneity of hardware, the diversity of task requirements, and the real-time nature of environmental disturbances intertwine to form a highly complex and dynamic computing scenario.

[0003] Currently, most scheduling and management strategies for quantum computing tasks are designed for homogeneous or small-scale systems, and generally adopt polling or static allocation schemes based on fixed priorities. These schemes fail to fully consider the fundamental differences between heterogeneous processors in terms of native quantum gate sets, control interfaces, and performance indicators, and cannot respond to real-time changes in the state of computing nodes and networks. When multiple subtasks are executed concurrently and require frequent exchange of quantum states or classical data, static scheduling strategies cannot dynamically coordinate the allocation of communication resources, causing limited interconnect bandwidth to become a system bottleneck, leading to serious resource conflicts and task blocking. The fundamental problem is the lack of an intelligent management system that can deeply perceive the heterogeneous characteristics of hardware, monitor the performance status of nodes and network load in real time, and dynamically map and coordinate tasks accordingly. This results in low utilization of expensive quantum computing resources, significantly extended execution time of complex algorithms, and even quantum state decoherence due to task waiting timeouts, leading to overall computation failure, which seriously restricts the efficiency and feasibility of distributed quantum computing in solving practical problems. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a communication coordination system for distributed quantum computing. This system utilizes a monitoring and sensing module, an intelligent decision-making module, a prediction and inference module, and an execution and learning module to solve the problems mentioned in the background.

[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: Specifically, it includes: a monitoring and sensing module, an intelligent decision-making module, a prediction and inference module, and an execution and learning module connected in sequence, wherein; Monitoring and Sensing Module: Deployed in each quantum processor node and network exchange node in the heterogeneous quantum computing cluster, it concurrently collects quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and performs spatiotemporal alignment and standardization processing on the collected quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and outputs a unified heterogeneous computing graph that uses dynamic attribute graph nodes to represent quantum processor nodes and directed edges to represent dependencies and connections. Intelligent decision-making module: Receives heterogeneous computing graph as input, generates scheduling decisions based on a preset deep reinforcement learning model. This deep reinforcement learning model evaluates and optimizes the mapping relationship of quantum computing tasks on quantum processor nodes, the allocation of communication links and the execution timing through a preset reward function, and outputs a preliminary scheduling scheme containing task mapping, link allocation and timestamps of quantum computing tasks. Prediction and simulation module: Maintains a high-fidelity digital twin model, receives preliminary scheduling schemes and performs simulations, monitors and predicts the performance degradation inflection point and network link congestion risk of each quantum processor node during the simulation, and generates a set of remedial plans that include subset migration or communication rerouting of quantum computing tasks when the performance degradation inflection point or congestion risk is predicted, and evaluates the effect of the remedial plan set using the high-fidelity digital twin model to generate estimated effect data; The execution and learning module integrates the initial scheduling scheme and the set of remedial plans to generate a flexible coordination instruction sequence with time stamps, which is then issued to each quantum processor node and network exchange node for execution. During execution, it collects standardized quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data obtained in real time from the monitoring and sensing module as real-time running data. The difference between the real-time running data and the predicted effect data at the corresponding time point generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module is used to trigger real-time rescheduling, and the deep reinforcement learning model and the high-fidelity digital twin model are periodically and incrementally updated. In a preferred embodiment, the specific process of concurrently collecting quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data in the monitoring and sensing module is as follows: The monitoring and sensing module periodically collects the readout fidelity, relaxation time, and dephase time of each qubit on the quantum processor node through local agents deployed on the quantum processor node, as well as the average single-quantum gate fidelity and double-quantum gate fidelity statistically analyzed in the most recent calibration period. These data are used as quantum hardware telemetry parameters. Through probes deployed on the quantum processor node and network exchange node, the module collects the end-to-end latency, available bandwidth, and message loss rate of the communication link between the nodes. These data are used as network performance indicators. The module receives quantum computing task description data in the form of a directed acyclic graph through the task submission interface. The vertices of the directed acyclic graph represent subtasks and include resource requirement metadata, while the edges represent dependencies between subtasks and are marked with estimated values ​​of the amount of data to be exchanged.

[0006] In a preferred embodiment, the specific operation of spatiotemporally aligning the collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data is as follows: The monitoring and sensing module establishes a global synchronization clock to mark all collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data with timestamped batch numbers. A preset fixed duration is used as the coordination period, and each coordination period is defined as a time window. Spatiotemporal alignment is performed on all quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data arriving within a time window. For quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that are sampled multiple times within a time window, the sampled value with the timestamp closest to the end of the time window is selected as the representative value of this type of data in the current time window; for quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that do not have new sampled values ​​arriving within a time window, the representative value determined in the most recent time window is used; through this spatiotemporal alignment operation, system state snapshot data corresponding to the end of each time window is generated.

[0007] In a preferred embodiment, the specific process of outputting the heterogeneous computation graph is as follows: The monitoring and sensing module calculates a dynamic normalized confidence level for each quantum processor node based on system state snapshot data; Next, a comprehensive state attribute vector is constructed for each quantum processor node. This vector includes at least a performance sub-vector aggregated from key quantum hardware telemetry parameters, a network sub-vector composed of network performance metrics from the quantum processor node to other quantum processor nodes, dynamically normalized confidence, and static metadata describing the quantum processor node type and the total number of physical qubits. Finally, based on the task subgraph defined by the quantum computing task description data, the task subgraph is mapped onto the physical topology composed of quantum processor nodes and network exchange nodes: each quantum processor node is mapped to a vertex in the heterogeneous computing graph, and the comprehensive state attribute vector constructed for the quantum processor node is assigned to the corresponding vertex. Each vertex representing a subtask in the task subgraph is assigned a specific attribute based on the associated metadata. The source requirement metadata is associated with one or more quantum processor nodes capable of executing the subtask. This association information serves as an additional attribute of the corresponding vertex in the heterogeneous computing graph. The network connections between quantum processor nodes and network exchange nodes, as well as the computational or communication dependencies between vertices in the task subgraph, together constitute the directed edges in the heterogeneous computing graph. The weights of the directed edges formed by the network connections between quantum processor nodes and network exchange nodes are initialized based on the reciprocal of the end-to-end latency or available bandwidth of the network connection in the current time window. The weights of the directed edges formed by the dependencies between vertices in the task subgraph are initialized based on the estimated data volume of the dependency. This generates a heterogeneous computing graph containing a vertex set, an edge set, a vertex attribute set, and an edge attribute set, with a timestamp of the end time of the time window.

[0008] In a preferred embodiment, the specific process of generating scheduling decisions based on a preset deep reinforcement learning model in the intelligent decision-making module is as follows: The intelligent decision-making module first performs hierarchical embedding encoding on the received heterogeneous computing graph. It decomposes the comprehensive state attribute vector of each vertex representing a quantum processor node in the heterogeneous computing graph. The numerical attribute components of the comprehensive state attribute vector, including dynamically normalized confidence and end-to-end latency and available bandwidth in network performance metrics, are directly normalized. The categorical attribute components of the comprehensive state attribute vector, including the quantum processor node type, are one-hot encoded. Then, a multi-head attention graph neural network is used to perform message passing and feature aggregation on the encoded heterogeneous computing graph. Simultaneously, the association information between vertices representing quantum computing subtasks and physical vertices representing quantum processor nodes in the task subgraph is encoded into attention edge weights that span the heterogeneous computing graph and the task subgraph. A global graph context embedding vector capturing the cluster's global load and health status is generated, along with a refined vertex embedding vector for each vertex containing its own and its multi-hop neighbors' topology and state information. Next, the intelligent decision-making module models the scheduling problem as a sequential decision-making process. It takes the global graph context embedding vector, the features of the quantum computing subtasks to be scheduled, and the encoded partial scheduling schemes already decided as input. Through the policy network of a deep reinforcement learning model, it sequentially selects the quantum processor node to execute each quantum computing subtask and allocates communication link paths and transmission time slots for each data dependency between quantum computing subtasks. The merits of each decision action output by the policy network are evaluated by a preset multi-objective reward function. The value of this multi-objective reward function is a negative sum, which is composed of the first computational depth cost term, the second communication time cost term, the third reliability penalty cost term, and the fourth conflict penalty cost term. The policy network of a deep reinforcement learning model is obtained through training with the environment based on the evaluation of a multi-objective reward function, or through forward inference of a trained model, to generate a complete sequence of decision actions, which is the scheduling decision.

[0009] In a preferred embodiment, the specific process of outputting the preliminary scheduling scheme is as follows: Set up a meta-learner that takes the global graph context embedding vector and the overall features of all quantum computing tasks in the current batch as input. The overall features include the ratio of the number of computationally intensive quantum computing subtasks to the number of communication-intensive quantum computing subtasks. The meta-learner analyzes the global graph context embedding vector and overall features, and outputs a set of dynamic weight coefficients. This set of dynamic weight coefficients includes the first dynamic weight coefficient, the second dynamic weight coefficient, the third dynamic weight coefficient, and the fourth dynamic weight coefficient, which are used to dynamically adjust the relative weights between the first computational depth cost term, the second communication time cost term, the third reliability penalty cost term, and the fourth conflict penalty cost term in the multi-objective reward function, in order to replace the preset fixed weight coefficients. The policy network of the deep reinforcement learning model optimizes decisions under the guidance of a reward function defined by the dynamic weight coefficients output by the meta-learner. Finally, the intelligent decision-making module decodes the optimized decision action sequence into a preliminary scheduling scheme. The preliminary scheduling scheme is a structured list in which each entry explicitly includes the identifier of the quantum computing subtask, the identifier of the quantum processor node to which it is mapped, a list of communication link identifiers allocated to each output dependency, and the precise timestamp of the start of execution of the quantum computing subtask and the time slot arrangement for communication transmission.

[0010] In a preferred embodiment, the specific process of using a high-fidelity digital twin model to simulate and extrapolate the preliminary scheduling scheme in the prediction and extrapolation module is as follows: Before initiating the simulation, the prediction and extrapolation module first initializes a high-fidelity digital twin model, which includes a quantum processor performance evolution sub-model and a network behavior simulation sub-model. The quantum processor performance evolution sub-model establishes a mathematical evolution model for each quantum processor node and each qubit on it. The mathematical evolution model is based on physical degradation laws and describes the evolution of the qubit's dephase time with virtual time through a differential equation. The initial state parameters of this mathematical evolution model, including the initial dephase time of each qubit and the dynamic normalized confidence level characterizing the overall performance state of the quantum processor node, are obtained from the latest time window system state snapshot data output by the monitoring and sensing module. The network behavior simulation sub-model establishes a simulation model for each communication link, including queuing, transmission, and packet loss behavior. Its initial parameters are obtained from network performance indicators in system state snapshot data, including end-to-end latency, available bandwidth, and packet loss rate. The simulation engine uses the start timestamps of all quantum computing subtasks defined in the initial scheduling scheme and the communication transmission time slot arrangements as an event sequence to drive the high-fidelity digital twin model to perform discrete execution on the virtual timeline. At each virtual moment of the simulation, the simulation engine performs the following operations: A1. Update the computational resource occupancy status and communication link bandwidth occupancy status of the quantum processor nodes in the high-fidelity digital twin model according to the event sequence; A2. Calculate and monitor two forward-looking risk indicators in parallel. The first risk indicator is the node coherence time health index, which is calculated for each quantum processor node. For a quantum processor node currently calculating its node coherence time health index, identify all qubits involved in all quantum computing subtasks running on that node at the current virtual moment. For each identified qubit, calculate a ratio. The numerator of this ratio is the out-of-phase time of the qubit calculated by the quantum processor performance evolution sub-model at the current virtual moment minus the theoretical minimum coherence time estimate required for the quantum computing subtask to complete its remaining operations. The denominator of this ratio is the initial out-of-phase time of the qubit at the start of the simulation. Take the minimum value among all these ratios as the node coherence time health index of the quantum processor node at the current virtual moment. The second risk indicator is the link queue pressure index, which is calculated for each communication link. The specific calculation process is as follows: Obtain the queue length and nominal bandwidth of the communication link calculated by the network behavior simulation sub-model at the current virtual moment; divide the queue length by the nominal bandwidth to obtain the instantaneous pressure term; obtain the instantaneous rate of change of the queue length and a predefined sensitivity coefficient at the current virtual moment; multiply the instantaneous rate of change by the sensitivity coefficient to obtain the trend pressure term; add the instantaneous pressure term and the trend pressure term to obtain the link queue pressure index of the communication link at the current virtual moment.

[0011] In a preferred embodiment, the specific process of generating a set of remedial contingency plans for subset migration or communication rerouting of quantum computing tasks, and generating estimated effect data, is as follows: The prediction and simulation module presets a coherence time health threshold and a queue pressure threshold. During the simulation process, it continuously compares the calculated node coherence time health index of each quantum processor node and the link queue pressure index of each communication link with the corresponding thresholds. When the node coherence time health index of any quantum processor node is detected to be lower than the coherence time health threshold, it is determined that the quantum processor node has the risk of performance degradation inflection point. The predicted time of occurrence of the risk is the current virtual time plus the estimated time extrapolated based on the downward trend of the index. When the link queue pressure index of any communication link is found to be higher than the queue pressure threshold and its rate of change is continuously positive, it is determined that the communication link has a congestion risk. The predicted time of risk occurrence is the current virtual time plus the estimated time extrapolated based on the growth trend of the index. After determining the risk, record the risk type, the specific quantum processor node or communication link where the risk occurred, the predicted time of the risk occurrence, and the identifiers of all quantum computing subtasks affected by the risk. Once a risk is detected, the simulation at the current time point is immediately paused, and multiple remedial plans are generated for each risk. These remedial plans are then combined to form a remedial plan set. For the risk of performance degradation inflection point, the generated remedial plan includes: before the predicted risk occurs, migrating part or all of the affected quantum computing subtasks to one or more backup quantum processor nodes with the lightest current load and the highest node coherence time health index. For the risk of communication link congestion, the generated remedial plan includes: before the predicted time of risk occurrence, pre-calculating and allocating one or more alternative communication routing paths that do not pass through the risk link for the data flow passing through the risk link; For each generated remedial plan, the simulation engine creates a new simulation branch from the currently paused virtual time point, injects the operations defined in the remedial plan into the simulation branch, and then continues to execute the simulation until all quantum computing subtasks are simulated and completed. The differences between each plan branch and the original plan in terms of the estimated overall task completion time and the estimated success rate of quantum computing subtasks are recorded and compared. Structured estimated effect data is generated. The estimated effect data includes at least the plan identifier, the specific operation description of the plan, the difference between the estimated overall task completion time after adopting the plan and the original plan, and the estimated additional computing or communication resource overhead required to execute the plan.

[0012] In a preferred embodiment, the specific process of generating a time-stamped elastic coordination instruction sequence by integrating the preliminary scheduling scheme and the set of remedial contingency plans in the execution and learning module is as follows: First, the estimated effect data of each remedial plan is quantified. The execution and learning module calculates a plan fusion benefit index for each remedial plan. The calculation process of the plan fusion benefit index is as follows: after the prediction and simulation module simulates the preliminary scheduling plan, the estimated overall task completion time and the estimated success rate of the quantum computing subtask are obtained. The difference between the estimated overall task completion time of the preliminary scheduling plan and the estimated overall task completion time of the remedial plan is calculated. This difference is divided by the estimated overall task completion time of the preliminary scheduling plan to obtain the relative time benefit ratio. The difference between the estimated success rate of quantum computing subtasks in the remedial plan and the estimated success rate of quantum computing subtasks in the preliminary scheduling plan is calculated to obtain the absolute difference in reliability benefits. The estimated additional computing or communication resource overhead of the remedial plan is divided by a preset resource overhead normalization benchmark to obtain the relative ratio of resource costs. The relative ratios of time benefits, absolute differences in reliability benefits, and relative ratios of resource costs are weighted using preset time benefit weighting coefficients, reliability benefit weighting coefficients, and resource cost weighting coefficients, respectively. The weighted relative ratios of time benefits and absolute differences in reliability benefits are added together, and then the weighted relative ratios of resource costs are subtracted to obtain the plan fusion benefit index of the remedial plan. Next, the execution and learning module sets an activation threshold and compares it with the plan fusion benefit index of each plan to make a final decision: if the plan fusion benefit index of all remedial plans is not greater than the activation threshold, it is determined that the expected comprehensive benefit of all remedial plans is not significant or the cost is too high, and the decision is to directly execute the preliminary scheduling plan; if there is a remedial plan with a plan fusion benefit index greater than the activation threshold, the one with the highest plan fusion benefit index among these plans is selected as the optimal remedial plan. Then, using the preliminary scheduling scheme as the baseline instruction sequence, before the time when the risk predicted by the optimal remedial plan occurs, the quantum computing subtask migration operation or communication rerouting operation defined in the optimal remedial plan is inserted into the corresponding position in the baseline instruction sequence, and the original related instruction fragments at that position are replaced, thereby generating a flexible coordination instruction sequence containing precise time stamps that integrates the preliminary scheduling scheme and the optimal remedial plan operations.

[0013] In a preferred embodiment, the specific process of using the difference between real-time running data and estimated performance data to trigger real-time rescheduling and periodically incrementally updating the deep reinforcement learning model and the high-fidelity digital twin model is as follows: During the execution of the elastic coordination command sequence, the execution and learning module continuously acquires real-time operational data from the monitoring and sensing module, including the actual values ​​of quantum hardware telemetry parameters and network performance indicators; at the same time, it acquires the predicted values ​​of the corresponding parameters contained in the predicted effect data generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module, which corresponds to the same time point. For each monitored quantum hardware telemetry parameter and key parameter in the network performance index, a normalized real-time deviation value is calculated. The calculation method for the real-time deviation value is as follows: First, calculate the absolute value of the difference between the actual value in the real-time running data of the key parameter at that moment and the predicted value of the corresponding parameter in the estimated effect data. Then, divide this absolute value by the standard deviation of the key parameter calculated on the historical running data sequence. A deviation tolerance threshold is preset for each key parameter. When the real-time deviation value of any key parameter exceeds its corresponding deviation tolerance threshold, or when the norm of the vector formed by the real-time deviation values ​​of all key parameters exceeds a global deviation threshold, real-time rescheduling is immediately triggered, and a rescheduling request is sent to the intelligent decision-making module. This request carries the latest real-time running data to trigger the intelligent decision-making module to generate a new preliminary scheduling scheme based on the latest system state. After the task execution cycle ends, the execution and learning module collects complete actual execution data, including the final actual task completion time, actual resource consumption, and the real-time running data sequence and its corresponding prediction data sequence throughout the entire execution process. The deep reinforcement learning model is incrementally updated periodically using these actual execution data.

[0014] The beneficial effects of this invention are as follows: By real-time perception and unified representation of the heterogeneous quantum processor and network status, a dynamic computation graph is constructed, and a deep reinforcement learning model is used to generate optimized scheduling decisions. The high-fidelity digital twin can perform forward-looking deduction of decision schemes, identify performance degradation and network congestion risks, and generate evaluated remedial plans. During execution, the baseline scheme and the optimized plan can be integrated to generate a flexible instruction sequence, and rescheduling can be dynamically triggered according to the deviation between the actual and the prediction to achieve adaptive coordination. The entire scheme can also use execution feedback data to periodically perform incremental learning on the scheduling model and the digital twin model, thereby continuously improving resource utilization efficiency, task success rate and overall coordination intelligence level. Attached Figure Description

[0015] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a block diagram of the system structure of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] In the description of this application, the terms "" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined with "" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.

[0018] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application. Example

[0019] This embodiment provides, for example Figure 1-2 The diagram illustrates a communication coordination system for distributed quantum computing, specifically comprising: a monitoring and sensing module, an intelligent decision-making module, a prediction and inference module, and an execution and learning module connected sequentially, wherein; Monitoring and Sensing Module: Deployed in each quantum processor node and network exchange node in the heterogeneous quantum computing cluster, it concurrently collects quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and performs spatiotemporal alignment and standardization processing on the collected quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and outputs a unified heterogeneous computing graph that uses dynamic attribute graph nodes to represent quantum processor nodes and directed edges to represent dependencies and connections. Intelligent decision-making module: Receives heterogeneous computing graph as input, generates scheduling decisions based on a preset deep reinforcement learning model. This deep reinforcement learning model evaluates and optimizes the mapping relationship of quantum computing tasks on quantum processor nodes, the allocation of communication links and the execution timing through a preset reward function, and outputs a preliminary scheduling scheme containing task mapping, link allocation and timestamps of quantum computing tasks. Prediction and simulation module: Maintains a high-fidelity digital twin model, receives preliminary scheduling schemes and performs simulations, monitors and predicts the performance degradation inflection point and network link congestion risk of each quantum processor node during the simulation, and generates a set of remedial plans that include subset migration or communication rerouting of quantum computing tasks when the performance degradation inflection point or congestion risk is predicted, and evaluates the effect of the remedial plan set using the high-fidelity digital twin model to generate estimated effect data; The execution and learning module integrates the initial scheduling scheme and the set of remedial contingency plans to generate a flexible coordination instruction sequence with time stamps, which is then issued to each quantum processor node and network exchange node for execution. During execution, it collects standardized quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data obtained in real time from the monitoring and sensing module as real-time running data. The difference between the real-time running data and the predicted effect data at the corresponding time point generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module is used to trigger real-time rescheduling, and the deep reinforcement learning model and the high-fidelity digital twin model are periodically and incrementally updated.

[0020] In this embodiment, the specific process of concurrently collecting quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data in the monitoring and sensing module is as follows: The monitoring and sensing module periodically collects the readout fidelity, relaxation time, and decoherence time of each qubit on the quantum processor node through local agents deployed on the node, as well as the average single-qubit fidelity and two-qubit fidelity statistically analyzed within the most recent calibration period. This data is used as telemetry parameters for the quantum hardware. The collection period of the local agents can be dynamically adjusted based on the decoherence time of the qubit and the system load. For example, for superconducting qubits with short decoherence times, the collection period can be set to the second level to capture their rapid performance fluctuations; for systems with longer coherence times, such as ion traps, the collection period can be appropriately extended. The collected data is encapsulated in a structured format (such as JSON or Protocol Buffers) and uploaded through a dedicated control network. Probes deployed on the quantum processor node and network exchange node also collect data. The end-to-end latency, available bandwidth, and message loss rate of the inter-communication link are used as network performance metrics. The probe evaluates the link quality by sending fixed-length probe data packets and measuring round-trip time and throughput. This can be implemented based on a lightweight protocol similar to TWAMP (Two-Way Active Measurement Protocol) to minimize the impact of measurement overhead on the computation task. The task submission interface receives quantum computing task description data in the form of a directed acyclic graph. The vertices of the directed acyclic graph represent subtasks and include resource requirement metadata, while the edges represent dependencies between subtasks and are marked with estimated amounts of data to be exchanged. The resource requirement metadata may include, but is not limited to: the minimum and expected number of physical qubits required, the upper limit of the tolerable gate error rate, the required set of quantum gates (such as the Clifford+T gate set or specific native gates), and the maximum allowed execution time of the subtask. The specific steps for spatiotemporal alignment of the collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data are as follows: The monitoring and sensing module establishes a global synchronization clock to mark all collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data with timestamped batch numbers. A preset fixed duration is used as the coordination period, and each coordination period is defined as a time window. The duration of the coordination period is a key system parameter, and its setting needs to strike a balance between sensing real-time performance and system overhead. A typical range is 100 milliseconds to 5 seconds. Shorter periods can respond to state changes more quickly but increase system coordination overhead; longer periods have the opposite effect. Spatiotemporal alignment is performed on all quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data arriving within a time window. For quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that are sampled multiple times within a time window, the sampled value with the timestamp closest to the end of the time window is selected as the representative value of this type of data in the current time window. For quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that do not have new sampled values ​​arriving within a time window, the representative value determined in the most recent time window is used. This alignment strategy of "prioritizing the most recent value and using historical values ​​as a fallback" ensures that the system state snapshot contains the latest dynamic information and maintains continuity when some data is temporarily missing, avoiding the invalidation of the entire snapshot due to a single data packet loss or delay. Through this spatiotemporal alignment operation, system state snapshot data corresponding to the end of each time window is generated. The specific process of outputting the heterogeneous computation graph is as follows: The monitoring and sensing module calculates a dynamically normalized confidence score for each quantum processor node based on system state snapshot data. This confidence score is a scalar between 0 and 1, designed to quantitatively integrate multiple dimensions reflecting the instantaneous reliability of the quantum processor node. The calculation process is as follows: For a quantum processor node, operations are performed on all qubits on the node and summed. The operations performed on a single qubit include: obtaining the single-gate fidelity of the qubit within the current time window, calculating the ratio of the single-gate fidelity to the historical average single-gate fidelity to obtain the first operation term. The first operation term reflects the degree to which the current gate operation accuracy of the qubit is maintained relative to its normal level; the closer the ratio is to or greater than 1, the more stable or better-than-average the performance. Next, the module obtains the out-of-phase time of the qubit within the current time window, calculates the ratio of the coordination period duration to the out-of-phase time, uses the negative of this ratio as an exponent to calculate the natural exponential function value, and subtracts this natural exponential function value from one to obtain the second operation term. The second operation term is a coherence time health index based on an exponential decay model, and its value increases with the decrease in out-of-phase time. The shorter and smaller value intuitively reflects the relative risk of spontaneous decay of the quantum state in the next coordination period; the measured value of the two-quantum gate fidelity involved in the qubit within the current time window is obtained, the difference between this measured value and the historical average two-quantum gate fidelity is calculated and the absolute value is taken, and this absolute value is subtracted from one to obtain the third operation term. The third operation term is used to capture sudden fluctuations in the performance of the two-quantum gate. Its value is close to 1 when the measured value is close to the historical average, and decreases when significant drift occurs, which helps to identify correlation errors or calibration failures; using hardware for quantum processor nodes The first, second, and third harmonic coefficients are pre-calibrated based on the device characteristics. The sum of the first, second, and third harmonic coefficients is one. These coefficients are used to weight the first, second, and third operational terms, respectively. These harmonic coefficients define the importance of different reliability dimensions. For example, for a processor where gate errors are the main noise source, a larger value (e.g., 0.5) can be assigned to the first harmonic coefficient; for a processor where decoherence is dominant, a larger value (e.g., 0.5) can be assigned to the second harmonic coefficient; and the third harmonic coefficient is usually set to a smaller value (e.g., 0.1-0).2) Used to sensitively capture abnormal fluctuations, the coefficients can be initialized and calibrated through analysis of historical fault data or expert experience. The weighted first, second, and third operational terms are added together to obtain the contribution value of the qubit. After summing the contribution values ​​of all qubits on the quantum processor node, the result is divided by a normalization factor to ensure that the final dynamic normalized confidence value is between zero and one. The normalization factor can usually be taken as the total number of physical qubits on the node, thus obtaining the average value of the contribution of all qubits. The dynamic normalized confidence value is used to characterize the performance stability and reliability of the quantum processor node in the current time window. The closer the confidence value is to 1, the more reliable the node is, and the more suitable it is for executing critical subtasks with high fidelity and coherence time requirements. When the confidence value is below a certain threshold (e.g., 0.7), it may indicate that the node has a risk of performance degradation. Next, a comprehensive state attribute vector is constructed for each quantum processor node. This vector includes at least a performance sub-vector aggregated from key quantum hardware telemetry parameters, a network sub-vector consisting of network performance metrics from the quantum processor node to other quantum processor nodes, dynamically normalized confidence, and static metadata describing the quantum processor node type and the total number of physical qubits. The performance sub-vector may include, for example, statistics such as the minimum and median readout fidelity of all qubits, the minimum and median single / dual quantum gate fidelity, and the median relaxation time and dephase time. The network sub-vector is a list recording the distance from the current node to every other quantum processor in the cluster. The current latency, available bandwidth, and packet loss rate of a node are static metadata, which are fixed and unchanging identifying information. Finally, based on the task subgraph defined by the quantum computing task description data, the task subgraph is mapped to the physical topology composed of quantum processor nodes and network switching nodes: each quantum processor node is mapped to a vertex of the heterogeneous computing graph, and the comprehensive state attribute vector constructed for the quantum processor node is assigned to the corresponding vertex; each vertex representing a subtask in the task subgraph is associated with one or more quantum processor nodes capable of executing that subtask, based on the resource requirement metadata attached to that vertex. This association information serves as an additional attribute of the corresponding vertex in the heterogeneous computing graph. The process is implemented through a matching algorithm: traversing all quantum processor nodes, selecting nodes whose static metadata (such as gate sets) satisfies the subtask requirements and whose dynamic performance (such as confidence and gate fidelity) exceeds the subtask's specified threshold, forming the "feasible execution node set" for that subtask; the network connections between quantum processor nodes and network exchange nodes, as well as the computational or communication dependencies between vertices in the task subgraph, together constitute the directed edges in the heterogeneous computation graph; among them, the weights of the directed edges formed by the network connections between quantum processor nodes and network exchange nodes are initialized based on the reciprocal of the end-to-end latency or available bandwidth of the network connection in the current time window. The specific initialization formula for the edge weights can be... The weight of a physical network edge equals the end-to-end latency; or the weight equals 1 / available bandwidth. The weight of a task-dependent edge equals the estimated data volume. These weights directly quantify the cost of communication or data exchange. The weights of directed edges formed by dependencies between vertices in the task subgraph are initialized based on the estimated data volume labeled by that dependency. This generates a heterogeneous computation graph containing a vertex set, edge set, vertex attribute set, and edge attribute set, along with a timestamp of the end of the time window. This heterogeneous computation graph is a dynamic and semantically rich data structure that uniformly represents the state of physical resources, the constraints of logical tasks, and the mapping relationship between them, providing a unique and complete input view for subsequent intelligent scheduling decisions.

[0021] In this embodiment, it is necessary to specifically explain the process by which the scheduling decision is generated based on a preset deep reinforcement learning model in the intelligent decision-making module: The intelligent decision-making module first performs hierarchical embedding encoding on the received heterogeneous computing graph. It decomposes the comprehensive state attribute vector of each vertex representing a quantum processor node in the heterogeneous computing graph. For the numerical attribute portion of the comprehensive state attribute vector, including dynamically normalized confidence and end-to-end latency and available bandwidth in network performance metrics, direct normalization is performed. Normalization can employ methods such as max-min normalization or Z-score standardization to map values ​​of different dimensions and ranges to a unified interval, such as [0,1] or a standard distribution with zero mean and unit variance, to eliminate the impact of numerical scale differences on model training. For the categorical attribute portion of the comprehensive state attribute vector, including the quantum processor node type, one-hot encoding is performed. For example, the processor type "superconducting" can be encoded as [1,0], and "ion trap" can be encoded as [0,1]. Then, a multi-head attention graph neural network is used to perform message passing and feature aggregation on the encoded heterogeneous computing graph. Through the multi-layer message passing mechanism of this multi-head attention graph neural network, the final embedding vector obtained by each vertex in the heterogeneous computing graph not only contains the attribute information of that vertex itself. It also encodes the topology and state information of neighboring vertices with multi-hop connections to the vertex. The network can be composed of multiple stacked graph attention layers. In each layer, each vertex aggregates the information of its direct neighbors through learnable attention weights. After multiple layers of transmission, the information can propagate multi-hop in the graph. The attention mechanism enables the network to dynamically focus on more important neighboring nodes. At the same time, the association information between the vertices representing quantum computing subtasks and the physical vertices representing quantum processor nodes in the task subgraph is encoded as an attention edge weight that spans the heterogeneous computing graph and the task subgraph. This can be achieved by introducing a cross-graph attention mechanism, calculating the "matching degree" of the task subtask vertex to each physical processor node as an attention score. This score can be calculated based on the similarity between task requirements and node attributes. A global graph context embedding vector that captures the global load and health status of the cluster is generated, as well as a refined vertex embedding vector for each vertex that contains its own and its multi-hop neighbors' topology and state information. The global graph context embedding vector can be obtained through a graph readout operation, such as averaging or summing all vertex embeddings, or generated through a dedicated pooling network. Next, the intelligent decision-making module models the scheduling problem as a sequential decision-making process. It takes the global graph context embedding vector, the features of the quantum computing subtasks to be scheduled, and the encoded partial scheduling schemes of the already decided options as input. Through a policy network of a deep reinforcement learning model, it sequentially selects the quantum processor node for execution for each quantum computing subtask and allocates communication link paths and transmission time slots for each data dependency between quantum computing subtasks. The order of sequential decisions can be determined based on the topological sorting of the task subgraph to ensure priority scheduling of prerequisite tasks. The merits of each decision action output by the policy network are evaluated by a pre-defined multi-objective reward function. The value of this multi-objective reward function is a negative sum, which is determined by the first... The cost is composed of four components: a first cost term for computational depth, a second cost term for communication time, a third cost term for reliability penalty, and a fourth cost term for conflict penalty. A negative reward value indicates a penalty. The training objective of the agent is to maximize the cumulative reward (i.e., minimize the cumulative penalty). The first cost term for computational depth is the sum of the estimated quantum circuit depths required for all quantum computing subtasks to execute on their respective assigned quantum processor nodes, weighted by a first dynamic weighting coefficient. The estimated quantum circuit depth, considering the current average gate error rate of the target node, is an estimate of the number of physical quantum gates (including error correction operations) required to implement a specific quantum algorithm; it is a key indicator for measuring computational time overhead. The second cost term for communication time is the sum of all existing data... The first term, representing the sum of the data to be transmitted between dependent quantum computing subtask pairs, is calculated by dividing the amount of data by the quotient obtained from the end-to-end available bandwidth of the path allocated for this communication, and then weighted by the second dynamic weighting coefficient. This quotient approximately represents the theoretical minimum time required for data transmission and serves as an indicator of communication overhead. The third reliability penalty cost is calculated for each quantum processor node by subtracting the node's dynamic normalized confidence level from the given value, and then multiplying this value by the sum of the computational load weights of all quantum computing subtasks assigned to that node. The computational load weights are determined based on the estimated quantum circuit depth and the number of qubits required for the quantum computing subtasks. The penalty values ​​for all quantum processor nodes are weighted by the third dynamic weighting coefficient. The design of this cost item is innovative: it not only penalizes assigning tasks to low-confidence (unreliable) nodes, but also the strength of the penalty is proportional to the total load borne by the node. This encourages the synergy between load balancing and reliability assurance. A specific implementation of calculating the load weight is: the calculated load weight of the subtask = the normalized estimated line depth * the normalized required number of bits; the fourth conflict penalty cost item is the sum of the indication values ​​of communication link resource conflicts occurring in all scheduling time slots, weighted by the fourth dynamic weight coefficient. The criterion for determining resource conflicts is: in the same time slot, the same physical communication link is assigned to two or more different data transmission tasks. The indication value is 1 in the time slot with conflict and 0 in the time slot without conflict. The policy network of a deep reinforcement learning model is obtained through training with the environment based on the evaluation of a multi-objective reward function, or through forward inference of a trained model, to generate a complete sequence of decision actions, which is the scheduling decision. The specific process of outputting the preliminary scheduling plan is as follows: A meta-learner is set up, which takes the global graph context embedding vector and the overall features of all quantum computing tasks in the current batch as input. The overall features are a statistical description of all quantum computing tasks in the current batch, including the ratio of the number of computationally intensive quantum computing subtasks to the number of communication-intensive quantum computing subtasks. The threshold for distinguishing between computationally intensive and communication-intensive tasks can be set empirically. For example, subtasks with an estimated computation depth / average computation depth of subtasks greater than 1.5 can be regarded as computationally intensive, and subtasks with an output data volume / average data volume of subtasks greater than 1.5 can be regarded as communication-intensive. The meta-learner analyzes the global graph context embedding vector and overall features, outputting a set of dynamic weight coefficients. This set includes four dynamic weight coefficients—the first, second, third, and fourth—used to dynamically adjust the relative weights among the first computational depth cost term, the second communication time cost term, the third reliability penalty cost term, and the fourth conflict penalty cost term in the multi-objective reward function. These coefficients replace preset fixed weight coefficients. The meta-learner itself can be a small fully connected neural network. Its input is the concatenated global graph embedding and overall feature vector, and its output is four dynamic weight coefficients. These coefficients are processed by a Softmax function before output to ensure their sum is one, thus giving them a clear and consistent weight distribution. Explanation of relative importance: For example, when the global graph context embedding vector reflects a low overall dynamic normalization confidence of the cluster, and the overall features reflect a high proportion of computationally intensive quantum computing subtasks in the current batch of quantum computing tasks, the meta-learner outputs the first and third dynamic weight coefficients after adjustment. This means that when the overall system reliability is poor and the computational tasks are heavy, the scheduling strategy will focus more on reducing computational depth (optimizing computational efficiency) and avoiding the use of unreliable nodes (ensuring task success rate). When the global graph context embedding vector reflects severe network congestion, the meta-learner outputs the second and fourth dynamic weight coefficients after adjustment. This means that when the network becomes a bottleneck, the scheduling strategy will focus more on reducing communication volume and avoiding communication conflicts. The policy network of the deep reinforcement learning model optimizes decisions under the guidance of a reward function defined by the dynamic weight coefficients output by the meta-learner. During the training phase, the meta-learner and the policy network can be trained collaboratively using a two-layer optimization strategy. In the inference phase of actual deployment, the meta-learner first generates dynamic weights based on the current state and task characteristics, and the policy network then makes decisions based on the expected reward defined by these weights. Finally, the intelligent decision-making module decodes the optimized decision action sequence into a preliminary scheduling scheme. The preliminary scheduling scheme is a structured list, in which each entry explicitly includes the identifier of the quantum computing subtask, the identifier of the mapped quantum processor node, a list of communication link identifiers allocated to each output dependency, and the precision of the quantum computing subtask's execution start time. The scheduling scheme precisely defines the timestamps and communication transmission time slots. The communication link identifier list can be a sequence of node IDs, defining the complete routing path from the source node to the destination node. The timestamp and time slot arrangement must satisfy the dependency constraints between tasks, that is, a subtask can only start after all its predecessor tasks have been completed and the necessary data transmissions have been completed. This preliminary scheduling scheme fully defines the mapping relationship and execution sequence of quantum computing subtasks on quantum processor nodes, as well as the communication link paths and timing on which data transmission between quantum computing subtasks depends. This scheme is the baseline input for subsequent prediction and inference modules to perform sandbox simulations and performance risk assessments. Its precise timing arrangement enables the digital twin to perform discrete event simulations and accurately predict task completion times and resource conflicts.

[0022] In this embodiment, it is necessary to specifically explain the process of using a high-fidelity digital twin model to simulate and extrapolate the preliminary scheduling scheme in the prediction and extrapolation module: Before initiating the prediction and extrapolation module, a high-fidelity digital twin model is first initialized. This model includes a quantum processor performance evolution sub-model and a network behavior simulation sub-model. The quantum processor performance evolution sub-model establishes a mathematical evolution model for each quantum processor node and each qubit on it. Based on physical degradation laws, this mathematical evolution model describes the evolution of the qubit's dephase time over virtual time using a differential equation. The differential equation states that the rate of change of the dephase time equals a negative decay coefficient associated with the quantum processor node multiplied by the current dephase time, plus a... The stochastic process term simulating environmental noise has an attenuation coefficient that can be obtained by exponentially fitting the historical decoherence data of the target quantum processor node during long-term operation, reflecting the unique performance degradation rate of the node. The stochastic process term can be modeled as zero-mean Gaussian white noise, and its variance can be dynamically adjusted according to the environmental electromagnetic noise level collected by the monitoring and sensing module. The initial state parameters of this mathematical evolution model, including the initial dephase time of each quantum bit and the dynamic normalized confidence level characterizing the overall performance state of the quantum processor node, are obtained from the latest time window system state snapshot data output by the monitoring and sensing module. The network behavior simulation sub-model establishes a simulation model for each communication link, including queuing, transmission, and packet loss behavior. Its initial parameters are obtained from network performance indicators in system state snapshot data, including end-to-end latency, available bandwidth, and packet loss rate. The simulation engine uses the start timestamps of all quantum computing subtasks defined in the initial scheduling scheme and the communication transmission time slot arrangements as event sequences to drive the high-fidelity digital twin model to execute discretely on the virtual timeline. The simulation engine adopts a discrete event simulation strategy, with the virtual timeline stepping at a fixed and sufficiently small time granularity (e.g., 1 millisecond) to ensure accurate simulation of the dynamic processes of quantum gate operations (nanosecond to microsecond level) and network packet transmission. At each virtual moment of the simulation, the simulation engine performs the following operations: A1. Update the computational resource occupancy status and communication link bandwidth occupancy status of quantum processor nodes in the high-fidelity digital twin model according to the event sequence. The update operation follows the precise timing in the initial scheduling scheme. For example, when the start timestamp of a quantum computing subtask arrives, the engine marks the quantum processor node and bit resources it occupies as "busy" and begins simulating the execution process of its quantum circuit. A2. Compute and monitor two forward-looking risk indicators in parallel. The first risk indicator is the node coherence time health index. This index is calculated for each quantum processor node. For a quantum processor node currently calculating its coherence time health index, identify all qubits involved in all quantum computing subtasks running on that node at the current virtual moment. For each identified qubit, calculate a ratio. The first risk indicator is the dephase time of the qubit calculated by the quantum processor performance evolution sub-model at the current virtual moment, minus the estimated theoretical minimum coherence time required for the quantum computing subtask to complete its remaining operations. This estimated theoretical minimum coherence time is calculated based on the remaining quantum circuit depth of the subtask and the typical execution time of quantum gates on the quantum processor node. The denominator of this ratio is the initial dephase time of the qubit at the start of the simulation. The minimum of all these ratios is taken as the node coherence time health index of the quantum processor node at the current virtual moment. The closer the node coherence time health index is to zero, the more likely it is that at least one quantum computing subtask on the quantum processor node is about to fail due to quantum state decoherence. The second risk indicator is the link queue pressure index, which is calculated for each communication link. The specific calculation process is as follows: Obtain the queue length and nominal bandwidth of the communication link calculated by the network behavior simulation sub-model at the current virtual moment; divide the queue length by the nominal bandwidth to obtain the instantaneous pressure term; obtain the instantaneous rate of change of the queue length at the current virtual moment and a predefined sensitivity coefficient. The instantaneous rate of change can be approximated by comparing the difference between the current queue length and the queue length at the previous virtual moment. The sensitivity coefficient is used to amplify or reduce the impact of the trend, and its typical value can be set empirically between 0.1 and 1.0 to adjust the system's sensitivity to congestion trends; multiply the instantaneous rate of change by the sensitivity coefficient to obtain the trend pressure term; add the instantaneous pressure term and the trend pressure term to obtain the link queue pressure index of the communication link at the current virtual moment; the larger the value of the link queue pressure index and the more continuously its growth trend is positive, the more likely the communication link is to experience congestion in subsequent moments; The specific process for generating a set of remedial contingency plans for subset migrations or communication rerouting involving quantum computing tasks, and for generating data to estimate the expected effects, is as follows: The prediction and simulation module presets a coherence time health threshold and a queue pressure threshold. A typical value for the coherence time health threshold is 0.1. This threshold means that when the decoherence safety margin of a node is only 10% of its initial capacity, the system determines that it is in a high-risk state and needs intervention. The queue pressure threshold is set according to the historical network load data. For example, it can be set to 1.5 times the historical average queue pressure value, or set to the pressure value corresponding to 80% utilization of the nominal bandwidth of the link. During the simulation, the calculated node coherence time health index of each quantum processor node and the link queue pressure index of each communication link are continuously compared with the corresponding thresholds. When the node coherence time health index of any quantum processor node is detected to be lower than the coherence time health threshold, it is determined that the quantum processor node has the risk of performance degradation inflection point. The predicted time of risk occurrence is the current virtual time plus the estimated time extrapolated based on the downward trend of the index. The extrapolated estimated time can be achieved by performing linear regression on the node coherence time health index of the past few virtual times to predict the time when its value drops to zero, which is a conservative estimate of the risk occurrence. When the link queue pressure index of any communication link is found to be higher than the queue pressure threshold and its rate of change is consistently positive, it is determined that the communication link has a congestion risk. The predicted time of risk occurrence is the current virtual time plus the estimated time extrapolated based on the growth trend of the index. In addition, the rate of change "consistently positive" can be defined as the derivative of which is greater than zero within several consecutive virtual times (e.g., 3-5 times). After determining the risk, record the risk type, the specific quantum processor node or communication link where the risk occurred, the predicted time of the risk occurrence, and the identifiers of all quantum computing subtasks affected by the risk. Once a risk is detected, the simulation at the current time point is immediately paused, and multiple remedial plans are generated for each risk. These remedial plans are then combined to form a remedial plan set. For the risk of performance degradation inflection point, the generated remedial plan includes: before the predicted risk occurs, migrating part or all of the affected quantum computing subtasks to one or more backup quantum processor nodes with the lightest current load and the highest node coherence time health index. The migration strategy needs to consider the resource capacity of the target node, the network connection quality with the source node, and the overhead and time introduced by the migration operation itself (quantum state transmission, context switching). For the risk of communication link congestion, the generated remedial plan includes: before the predicted risk occurs, pre-calculating and allocating one or more backup communication routes that do not pass through the risk link for the data flow passing through the risk link. The backup routes can be calculated using the k-shortest path algorithm, and priority should be given to paths with low overall latency and large available bandwidth. For each generated remedial plan, the simulation engine creates a new simulation branch from the currently paused virtual time point, injects the operations defined in the remedial plan into this branch, and then continues to execute the simulation until all quantum computing subtasks are simulated and completed. The differences between each plan branch and the original plan in terms of the estimated overall task completion time and the estimated success rate of the quantum computing subtasks are recorded and compared. The estimated success rate of the quantum computing subtasks can be calculated based on the final fidelity of the digital twin model simulation, combined with the task success determination threshold, to generate structured estimated effect data. The estimated effect data includes at least the plan identifier, the specific operation description of the plan, the difference between the estimated overall task completion time after adopting the plan and the original plan, and the estimated additional computing or communication resource overhead required to execute the plan. This estimated effect data provides a quantitative decision basis for the subsequent execution and learning modules to select which plan to execute, realizing a closed loop from "prediction" to "decision".

[0023] In this embodiment, the specific process of generating a time-stamped elastic coordination instruction sequence by integrating the preliminary scheduling scheme and the remedial contingency plan set in the execution and learning module is as follows: The execution and learning module receives a preliminary scheduling plan from the intelligent decision-making module, and a set of remedial plans containing multiple remedial plans and the estimated effect data corresponding to each remedial plan from the prediction and inference module. First, the estimated effect data of each remedial plan is quantified. The execution and learning module calculates a plan fusion benefit index for each remedial plan. The calculation process of the plan fusion benefit index is as follows: after the prediction and inference module simulates the preliminary scheduling plan, the estimated overall task completion time and the estimated success rate of the quantum computing subtask are obtained. The difference between the estimated overall task completion time of the preliminary scheduling plan and the estimated overall task completion time of the remedial plan is calculated. This difference is divided by the estimated overall task completion time of the preliminary scheduling plan to obtain the relative time benefit ratio. The difference between the estimated success rate of the quantum computing subtasks in the remedial plan and the estimated success rate in the initial scheduling plan is calculated to obtain the absolute difference in reliability gains. The estimated additional computing or communication resource overhead in the remedial plan is divided by a preset resource overhead normalization benchmark to obtain the relative ratio of resource costs. The resource overhead normalization benchmark can be set as the typical maximum additional overhead allowed for a single task execution by the cluster; for example, its value can be set to five percent of the total computing resources of the cluster (such as total computing power). Preset time gain weighting coefficients, reliability gain weighting coefficients, and resource cost weighting coefficients are used to add to the relative ratio of time gains, the absolute difference in reliability gains, and the relative ratio of resource costs, respectively. The weighting coefficients for time benefit, reliability benefit, and resource cost can be dynamically set according to the Service Level Agreement (SLA) of the task. For example, for delay-sensitive tasks, the time benefit weighting coefficient can be set to 0.6, the reliability benefit weighting coefficient to 0.3, and the resource cost weighting coefficient to 0.1; for reliability-priority tasks, the reliability benefit weighting coefficient can be set to 0.6. The weighted relative ratio of time benefit is added to the absolute difference of reliability benefit, and then the weighted relative ratio of resource cost is subtracted to obtain the plan integration benefit index of the remedial plan. This index is used to comprehensively quantify the expected net benefit of adopting the remedial plan after considering the time improvement, reliability improvement, and resource consumption costs. Next, the execution and learning module sets an activation threshold to filter out contingency plans with insignificant benefits. A typical value is 0.05, which means that a contingency plan will only be considered for adoption if its expected overall net benefit exceeds 5% of the baseline plan. The final decision is made by comparing this activation threshold with the contingency plan fusion benefit index of each contingency plan: if the contingency plan fusion benefit index of all remedial contingency plans is not greater than the activation threshold, it is determined that the expected overall benefit of all remedial contingency plans is insignificant or the cost is too high, and the decision is to directly execute the preliminary scheduling plan; if there are remedial contingency plans with a contingency plan fusion benefit index greater than the activation threshold, the one with the highest contingency plan fusion benefit index among these contingency plans is selected as the optimal remedial contingency plan. Then, using the preliminary scheduling scheme as the baseline instruction sequence, before the time when the risk predicted by the optimal remedial plan occurs, the quantum computing subtask migration operation or communication rerouting operation defined in the optimal remedial plan is inserted into the corresponding position in the baseline instruction sequence, and the original related instruction fragments at that position are replaced, thereby generating a flexible coordination instruction sequence containing precise time stamps that integrates the preliminary scheduling scheme and the optimal remedial plan operations. The instruction insertion process must ensure the correctness of the timing logic. For example, the migration operation must be issued only after the source task is suspended and the target node resources are ready, and sufficient time must be reserved for quantum state transmission and context reconstruction. The specific process of using the difference between real-time running data and predicted performance data to trigger real-time rescheduling, and periodically incrementally updating the deep reinforcement learning model and the high-fidelity digital twin model is as follows: During the execution of the elastic coordination command sequence, the execution and learning module continuously acquires real-time operational data from the monitoring and sensing module, including the actual values ​​of quantum hardware telemetry parameters and network performance indicators; at the same time, it acquires the predicted values ​​of the corresponding parameters contained in the predicted effect data generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module, which corresponds to the same time point. For each monitored quantum hardware telemetry parameter and key parameter in the network performance indicators, a normalized real-time deviation value is calculated. The real-time deviation value is calculated as follows: First, the absolute value of the difference between the actual value in the real-time running data of the key parameter at that moment and the predicted value of the corresponding parameter in the estimated effect data is calculated. Then, this absolute value is divided by the standard deviation of the key parameter calculated on the historical running data sequence. The standard deviation of the historical running data sequence can be calculated based on data from a past time window (e.g., the past 1 hour), which reflects the degree of dispersion of the parameter within the recent normal fluctuation range. A deviation tolerance threshold is preset for each key parameter. The deviation tolerance threshold can be set according to the importance of the parameter. For example, for quantum gate fidelity, the deviation tolerance threshold can be set to 1.5 (i.e., the actual value deviates from the predicted value by 1.5 historical standard deviations); for network latency, it can be set to 2.0. When the real-time deviation value of any key parameter exceeds its corresponding deviation tolerance threshold, or when the norm of the vector formed by the real-time deviation values ​​of all key parameters exceeds a global deviation threshold, the global deviation threshold is used to monitor the overall deviation. For example, it can be set to the square root of the sum of the squares of the deviation tolerance thresholds of all key parameters (i.e., the L2 norm threshold). Real-time rescheduling is immediately triggered, and a rescheduling request is sent to the intelligent decision module. This request carries the latest real-time running data to trigger the intelligent decision module to generate a new preliminary scheduling scheme based on the latest system state. After the task execution cycle ends, the execution and learning module collects complete actual execution data, including the final actual task completion time, actual resource consumption, and the real-time running data sequence and its corresponding predicted data sequence throughout the entire execution process. The deep reinforcement learning model is incrementally updated periodically using these actual execution data. The update process involves: constructing a complete task execution as an enhanced experience trajectory and storing it in an experience replay buffer. This enhanced experience trajectory includes the system state reflected by real-time running data during task execution, the actual sequence of decision actions taken, and the actual cumulative reward calculated based on the actual completion time and resource consumption; periodically sampling experience from the experience replay buffer and fine-tuning the network parameters of the deep reinforcement learning model using gradient descent; when calculating the loss function, assigning lower weights to the system state and actual decision actions reflected by real-time running data corresponding to moments with larger real-time deviation values; specifically, multiplying the temporal difference error of this experience by a weighting factor inversely proportional to the real-time deviation value; simultaneously, using the time-varying quantum hardware telemetry parameter sequence and network performance index sequence collected by the monitoring and sensing module, and comparing it with high... The corresponding parameter prediction sequences generated by the high-fidelity digital twin model are compared, and an online learning method based on adaptive adjustment of the learning rate according to prediction bias is used to fine-tune the internal parameters of the quantum processor performance evolution sub-model and the network behavior simulation sub-model in the high-fidelity digital twin model. The specific principle of this online learning method is: for model parameters with persistently large prediction bias, a higher learning rate is used for updating to quickly correct them; for example, the base learning rate can be multiplied by a factor (such as 2-5 times) for updating. For model parameters with small prediction bias, a lower learning rate is used for updating to maintain model stability; for example, the base learning rate can be multiplied by a factor less than 1 (such as 0.5) for updating. This achieves a confidence-weighted incremental update of the deep reinforcement learning model and the high-fidelity digital twin model. This closed-loop learning mechanism enables the entire system to become more and more accurate in its scheduling decisions and more and more reliable in digital twin prediction as the running time increases, forming a self-reinforcing and continuously optimizing intelligent coordinator.

[0024] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0025] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0026] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0027] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0028] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0029] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0030] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A communication coordination system for distributed quantum computing, characterized in that, Specifically, it includes: The monitoring and sensing module, intelligent decision-making module, prediction and inference module, and execution and learning module are connected in sequence, wherein; Monitoring and Sensing Module: Deployed in each quantum processor node and network exchange node in the heterogeneous quantum computing cluster, it concurrently collects quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and performs spatiotemporal alignment and standardization processing on the collected quantum hardware telemetry parameters, network performance indicators and quantum computing task description data, and outputs a unified heterogeneous computing graph that uses dynamic attribute graph nodes to represent quantum processor nodes and directed edges to represent dependencies and connections. Intelligent decision-making module: Receives heterogeneous computing graph as input, generates scheduling decisions based on a preset deep reinforcement learning model. This deep reinforcement learning model evaluates and optimizes the mapping relationship of quantum computing tasks on quantum processor nodes, the allocation of communication links and the execution timing through a preset reward function, and outputs a preliminary scheduling scheme containing task mapping, link allocation and timestamps of quantum computing tasks. Prediction and simulation module: Maintains a high-fidelity digital twin model, receives preliminary scheduling schemes and performs simulations, monitors and predicts the performance degradation inflection point and network link congestion risk of each quantum processor node during the simulation, and generates a set of remedial plans that include subset migration or communication rerouting of quantum computing tasks when the performance degradation inflection point or congestion risk is predicted, and evaluates the effect of the remedial plan set using the high-fidelity digital twin model to generate estimated effect data; The execution and learning module integrates the initial scheduling scheme and the set of remedial contingency plans to generate a flexible coordination instruction sequence with time stamps, which is then issued to each quantum processor node and network exchange node for execution. During execution, it collects standardized quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data obtained in real time from the monitoring and sensing module as real-time running data. The difference between the real-time running data and the predicted effect data at the corresponding time point generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module is used to trigger real-time rescheduling, and the deep reinforcement learning model and the high-fidelity digital twin model are periodically and incrementally updated.

2. The communication coordination system for distributed quantum computing according to claim 1, characterized in that: The specific process of concurrently collecting quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data in the monitoring and sensing module is as follows: The monitoring and sensing module periodically collects the readout fidelity, relaxation time, and dephase time of each qubit on the quantum processor node through local agents deployed on the quantum processor node, as well as the average single-quantum gate fidelity and double-quantum gate fidelity statistically analyzed in the most recent calibration period. These data are used as quantum hardware telemetry parameters. The module also collects the end-to-end delay, available bandwidth, and message loss rate of the communication link between the nodes through probes deployed on the quantum processor node and the network switching node. These data are used as network performance indicators. The task submission interface receives quantum computing task description data in the form of a directed acyclic graph. The vertices of the directed acyclic graph represent subtasks and are accompanied by resource requirement metadata. The edges represent the dependencies between subtasks and are marked with the estimated amount of data to be exchanged.

3. The communication coordination system for distributed quantum computing according to claim 2, characterized in that: The specific operation of spatiotemporally aligning the collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data is as follows: The monitoring and sensing module establishes a global synchronization clock to mark all collected quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data with timestamped batch numbers. A preset fixed duration is used as the coordination period, and each coordination period is defined as a time window. Spatiotemporal alignment is performed on all quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data arriving within a time window. For quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that are sampled multiple times within a time window, the sampled value whose timestamp is closest to the end of the time window is selected as the representative value of this type of data in the current time window; for quantum hardware telemetry parameters, network performance indicators, and quantum computing task description data that do not have new sampled values ​​arriving within a time window, the representative value determined in the most recent time window is used. This spatiotemporal alignment operation generates system state snapshot data corresponding to the end time of each time window.

4. The communication coordination system for distributed quantum computing according to claim 3, characterized in that: The specific process of outputting the heterogeneous computation graph is as follows: The monitoring and sensing module calculates a dynamic normalized confidence level for each quantum processor node based on system state snapshot data; Next, a comprehensive state attribute vector is constructed for each quantum processor node. The comprehensive state attribute vector includes at least a performance sub-vector aggregated from key quantum hardware telemetry parameters, a network sub-vector composed of network performance indicators from the quantum processor node to other quantum processor nodes, a dynamically normalized confidence score, and static metadata describing the type of quantum processor node and the total number of physical qubits. Finally, based on the task subgraph defined by the quantum computing task description data, the task subgraph is mapped to the physical topology composed of quantum processor nodes and network exchange nodes: each quantum processor node is mapped to a vertex of the heterogeneous computing graph, and the comprehensive state attribute vector constructed for the quantum processor node is assigned to the corresponding vertex. Each vertex representing a subtask in the task subgraph is associated with one or more quantum processor nodes capable of executing the subtask, based on the resource requirement metadata attached to that vertex. This association information serves as an additional attribute of the corresponding vertex in the heterogeneous computing graph. The network connections between quantum processor nodes and network exchange nodes, as well as the computational or communication dependencies between vertices in the task subgraph, together constitute the directed edges in the heterogeneous computation graph. The weights of the directed edges formed by the network connections between quantum processor nodes and network exchange nodes are initialized based on the reciprocal of the end-to-end latency or available bandwidth of the network connection in the current time window. The weights of the directed edges formed by the dependencies between vertices in the task subgraph are initialized based on the estimated data volume of the dependency. This generates a heterogeneous computation graph containing a vertex set, an edge set, a vertex attribute set, and an edge attribute set, with a timestamp of the end time of the time window.

5. A communication coordination system for distributed quantum computing according to claim 4, characterized in that: In the intelligent decision-making module, the specific process of generating scheduling decisions based on a preset deep reinforcement learning model is as follows: The intelligent decision-making module first performs hierarchical embedding encoding on the received heterogeneous computing graph, and splits the comprehensive state attribute vector of each vertex representing a quantum processor node in the heterogeneous computing graph. The numerical attribute part of the comprehensive state attribute vector, including the dynamic normalized confidence, end-to-end latency and available bandwidth in the network performance indicators, is directly normalized. One-hot encoding is performed on the categorical attribute portion of the comprehensive state attribute vector, including the quantum processor node type; Then, a multi-head attention graph neural network is used to perform message passing and feature aggregation on the encoded heterogeneous computation graph; Meanwhile, the association information between the vertices representing quantum computing subtasks and the physical vertices representing quantum processor nodes in the task subgraph is encoded into an attention edge weight that spans the heterogeneous computing graph and the task subgraph. Generate a global graph context embedding vector that captures the global load and health status of the cluster, as well as a refined vertex embedding vector for each vertex that includes its own topology and state information and that of its multi-hop neighbors. Next, the intelligent decision-making module models the scheduling problem as a sequential decision-making process. It takes the global graph context embedding vector, the features of the quantum computing subtasks to be scheduled, and the encoded partial scheduling schemes already decided as input. Through the policy network of a deep reinforcement learning model, it sequentially selects the quantum processor node to execute each quantum computing subtask and allocates communication link paths and transmission time slots for each data dependency between quantum computing subtasks. The merits of each decision action output by the policy network are evaluated by a preset multi-objective reward function. The value of this multi-objective reward function is a negative sum, which is composed of the first computational depth cost term, the second communication time cost term, the third reliability penalty cost term, and the fourth conflict penalty cost term. The policy network of a deep reinforcement learning model is obtained through training with the environment based on the evaluation of a multi-objective reward function, or through forward inference of a trained model, to generate a complete sequence of decision actions, which is the scheduling decision.

6. A communication coordination system for distributed quantum computing according to claim 5, characterized in that: The specific process of outputting the preliminary scheduling scheme is as follows: Set up a meta-learner that takes the global graph context embedding vector and the overall features of all quantum computing tasks in the current batch as input. The overall features include the ratio of the number of computationally intensive quantum computing subtasks to the number of communication-intensive quantum computing subtasks. The meta-learner analyzes the global graph context embedding vector and overall features, and outputs a set of dynamic weight coefficients. This set of dynamic weight coefficients includes the first dynamic weight coefficient, the second dynamic weight coefficient, the third dynamic weight coefficient, and the fourth dynamic weight coefficient, which are used to dynamically adjust the relative weights between the first computational depth cost term, the second communication time cost term, the third reliability penalty cost term, and the fourth conflict penalty cost term in the multi-objective reward function, in order to replace the preset fixed weight coefficients. The policy network of the deep reinforcement learning model optimizes decisions under the guidance of a reward function defined by the dynamic weight coefficients output by the meta-learner. Finally, the intelligent decision-making module decodes the optimized decision action sequence into a preliminary scheduling scheme. The preliminary scheduling scheme is a structured list in which each entry explicitly includes the identifier of the quantum computing subtask, the identifier of the quantum processor node to which it is mapped, a list of communication link identifiers allocated to each output dependency, and the precise timestamp of the start of execution of the quantum computing subtask and the time slot arrangement for communication transmission.

7. A communication coordination system for distributed quantum computing according to claim 6, characterized in that: In the prediction and simulation module, the specific process of using a high-fidelity digital twin model to simulate and extrapolate the preliminary scheduling scheme is as follows: Before initiating the simulation, the prediction and extrapolation module first initializes a high-fidelity digital twin model, which includes a quantum processor performance evolution sub-model and a network behavior simulation sub-model. The quantum processor performance evolution sub-model establishes a mathematical evolution model for each quantum processor node and each qubit on it. The mathematical evolution model is based on physical degradation laws and describes the evolution of the qubit's dephase time with virtual time through a differential equation. The initial state parameters of this mathematical evolution model, including the initial dephase time of each qubit and the dynamic normalized confidence level characterizing the overall performance state of the quantum processor node, are obtained from the latest time window system state snapshot data output by the monitoring and sensing module. The network behavior simulation sub-model establishes a simulation model for each communication link, including queuing, transmission, and packet loss behavior. Its initial parameters are obtained from network performance indicators in system state snapshot data, including end-to-end latency, available bandwidth, and packet loss rate. The simulation engine uses the start timestamps of all quantum computing subtasks defined in the initial scheduling scheme and the communication transmission time slot arrangements as an event sequence to drive the high-fidelity digital twin model to perform discrete execution on the virtual timeline. At each virtual moment of the simulation, the simulation engine performs the following operations: A1. Update the computing resource occupancy status and communication link bandwidth occupancy status of the quantum processor node in the high-fidelity digital twin model according to the event sequence; A2. Parallel computation and monitoring of two forward-looking risk indicators. The first risk indicator is the node coherence time health index, which is calculated for each quantum processor node. For a quantum processor node currently calculating its node coherence time health index, identify all qubits involved in all quantum computing subtasks running on that node at the current virtual moment. For each identified qubit, calculate a ratio. The numerator of this ratio is the dephase time of the qubit calculated by the quantum processor performance evolution sub-model at the current virtual moment, minus the theoretical minimum coherence time estimate required for the quantum computing subtask to complete its remaining operations. The denominator of this ratio is the initial dephase time of the qubit at the start of the derivation. Take the minimum of all these ratios as the node coherence time health index of the quantum processor node at the current virtual moment. The second risk indicator is the link queue pressure index, which is calculated for each communication link. The specific calculation process is as follows: Obtain the queue length and nominal bandwidth of the communication link calculated by the network behavior simulation sub-model at the current virtual moment; divide the queue length by the nominal bandwidth to obtain the instantaneous pressure term; obtain the instantaneous rate of change of the queue length and a predefined sensitivity coefficient at the current virtual moment; multiply the instantaneous rate of change by the sensitivity coefficient to obtain the trend pressure term; add the instantaneous pressure term and the trend pressure term to obtain the link queue pressure index of the communication link at the current virtual moment.

8. A communication coordination system for distributed quantum computing according to claim 7, characterized in that: The specific process of generating a set of remedial contingency plans for subset migration or communication rerouting of quantum computing tasks, and generating estimated effect data, is as follows: The prediction and simulation module presets a coherence time health threshold and a queue pressure threshold. During the simulation process, it continuously compares the calculated node coherence time health index of each quantum processor node and the link queue pressure index of each communication link with the corresponding thresholds. When the node coherence time health index of any quantum processor node is detected to be lower than the coherence time health threshold, it is determined that the quantum processor node has the risk of performance degradation inflection point. The predicted time of occurrence of the risk is the current virtual time plus the estimated time extrapolated based on the downward trend of the index. When the link queue pressure index of any communication link is found to be higher than the queue pressure threshold and its rate of change is continuously positive, it is determined that the communication link has a congestion risk. The predicted time of risk occurrence is the current virtual time plus the estimated time extrapolated based on the growth trend of the index. After determining the risk, record the risk type, the specific quantum processor node or communication link where the risk occurred, the predicted time of the risk occurrence, and the identifiers of all quantum computing subtasks affected by the risk. Once a risk is detected, the simulation at the current time point is immediately paused, and multiple remedial plans are generated for each risk. These remedial plans are then combined to form a remedial plan set. For the risk of performance degradation inflection point, the generated remedial plan includes: before the predicted risk occurs, migrating part or all of the affected quantum computing subtasks to one or more backup quantum processor nodes with the lightest current load and the highest node coherence time health index. For the risk of communication link congestion, the generated remedial plan includes: before the predicted time of risk occurrence, pre-calculating and allocating one or more alternative communication routing paths that do not pass through the risk link for the data flow passing through the risk link; For each generated remedial plan, the simulation engine creates a new simulation branch from the currently paused virtual time point, injects the operations defined in the remedial plan into the simulation branch, and then continues to execute the simulation until all quantum computing subtasks are simulated and completed. The differences between each plan branch and the original plan in terms of the estimated overall task completion time and the estimated success rate of quantum computing subtasks are recorded and compared. Structured estimated effect data is generated. The estimated effect data includes at least the plan identifier, the specific operation description of the plan, the difference between the estimated overall task completion time after adopting the plan and the original plan, and the estimated additional computing or communication resource overhead required to execute the plan.

9. A communication coordination system for distributed quantum computing according to claim 8, characterized in that: In the execution and learning module, the specific process of integrating the preliminary scheduling scheme and the set of remedial contingency plans to generate a flexible coordination instruction sequence containing time stamps is as follows: First, the estimated effect data of each remedial plan is quantified. The execution and learning module calculates a plan fusion benefit index for each remedial plan. The calculation process of the plan fusion benefit index is as follows: after the prediction and simulation module simulates the preliminary scheduling plan, the estimated overall task completion time and the estimated success rate of the quantum computing subtask are obtained. The difference between the estimated overall task completion time of the preliminary scheduling plan and the estimated overall task completion time of the remedial plan is calculated. This difference is divided by the estimated overall task completion time of the preliminary scheduling plan to obtain the relative time benefit ratio. The absolute difference in reliability gains is obtained by calculating the difference between the estimated success rate of the quantum computing subtasks in the remedial plan and the estimated success rate of the quantum computing subtasks in the initial scheduling plan. The estimated additional computational or communication resource overhead of the remedial plan is calculated and divided by a preset normalized baseline value of resource overhead to obtain the relative resource cost ratio. The relative time cost ratio, the absolute difference in reliability cost and the relative resource cost ratio are weighted by preset time benefit weighting coefficient, reliability benefit weighting coefficient and resource cost weighting coefficient respectively. The weighted relative time cost ratio is added to the absolute difference in reliability cost and then the weighted relative resource cost ratio is subtracted to obtain the plan integration benefit index of the remedial plan. Next, the execution and learning module sets an activation threshold and compares it with the plan fusion benefit index of each plan to make a final decision: if the plan fusion benefit index of all remedial plans is not greater than the activation threshold, it is determined that the expected comprehensive benefit of all remedial plans is not significant or the cost is too high, and the decision is to directly execute the preliminary scheduling plan; if there is a remedial plan with a plan fusion benefit index greater than the activation threshold, the one with the highest plan fusion benefit index among these plans is selected as the optimal remedial plan. Then, using the preliminary scheduling scheme as the baseline instruction sequence, before the time when the risk predicted by the optimal remedial plan occurs, the quantum computing subtask migration operation or communication rerouting operation defined in the optimal remedial plan is inserted into the corresponding position in the baseline instruction sequence, and the original related instruction fragments at that position are replaced, thereby generating a flexible coordination instruction sequence containing precise time stamps that integrates the preliminary scheduling scheme and the optimal remedial plan operations.

10. A communication coordination system for distributed quantum computing according to claim 9, characterized in that: The specific process of using the difference between real-time running data and predicted performance data to trigger real-time rescheduling and periodically incrementally updating the deep reinforcement learning model and the high-fidelity digital twin model is as follows: During the execution of the elastic coordination command sequence, the execution and learning module continuously acquires real-time operational data from the monitoring and sensing module, including the actual values ​​of quantum hardware telemetry parameters and network performance indicators; at the same time, it acquires the predicted values ​​of the corresponding parameters contained in the predicted effect data generated by the high-fidelity digital twin model during the simulation process of the prediction and inference module, which corresponds to the same time point. For each monitored quantum hardware telemetry parameter and key parameter in the network performance index, a normalized real-time deviation value is calculated. The calculation method for the real-time deviation value is as follows: First, calculate the absolute value of the difference between the actual value in the real-time running data of the key parameter at that moment and the predicted value of the corresponding parameter in the estimated effect data. Then, divide this absolute value by the standard deviation of the key parameter calculated on the historical running data sequence. A deviation tolerance threshold is preset for each key parameter. When the real-time deviation value of any key parameter exceeds its corresponding deviation tolerance threshold, or when the norm of the vector formed by the real-time deviation values ​​of all key parameters exceeds a global deviation threshold, real-time rescheduling is immediately triggered, and a rescheduling request is sent to the intelligent decision-making module. This request carries the latest real-time running data to trigger the intelligent decision-making module to generate a new preliminary scheduling scheme based on the latest system state. After the task execution cycle ends, the execution and learning module collects complete actual execution data, including the final actual task completion time, actual resource consumption, and the real-time running data sequence and its corresponding prediction data sequence throughout the entire execution process. The deep reinforcement learning model is incrementally updated periodically using these actual execution data.