Method and system for unified scheduling of cross-domain computing power
Patent Information
- Application Number
- CN202610923984.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-25
AI Technical Summary
[0005]鉴于现有技术中的上述缺陷或不足,本申请旨在提供一种量超智跨域算力统一调度方法与系统,以解决相关技术中跨量子域、超算域以及智算域的资源协调性差的问题,提高跨域资源调度成功率
[0019]综上所述,本申请提出一种量超智跨域算力统一调度方法,该方法响应于接收到混合作业请求,生成混合作业对应的跨域协同对象,进而从该混合作业涉及到的各目标域读取资源状态信息,根据资源状态信息确定目标资源组合,并发起对目标资源组合的预留,通过各目标域的资源状态信息确定各目标域的单域就绪度,并确定各目标域之间的跨域启动对齐度以及取消预留目标资源组合的回滚代价,进而通过各单域就绪度、跨域启动对齐度与回滚代价,判断目标资源组合是否满足联合启动条件,得到一致性许可判断结果,最后根据一致性许可判断结果,将跨域协同对象与目标资源组合进行绑定并提交到各目标域,或者,维持当前预留并进入等待控制,或者,对跨域协同对象进行回滚处理,实现量子域、超算域以及智算域的跨域资源统一调度。该方法通过各个目标域的单域就绪度进行跨域全局一致性的判定,能够保证在所有目标域预留稳定充足的情况下绑定资源,避免仅在单个目标域稳定的情况下进行绑定,从而避免出现部分资源占用而部分资源无法及时获得的场景,有助于提升跨域联合调度的成功率,并减少重复排队、无效等待及重复回滚造成的资源浪费,此外,该方法通过跨域启动对齐度进行跨域全局一致性的判定,能够保证各个目标域启动时间的一致性,减少因启动时机不一致导致的智算节点、超算节点或量子设备长时间空转、等待或占位的问题,并且,该方法通过回滚代价进行跨域全局一致性的判定,能够在继续等待与立即回滚之间进行量化决策,避免在不可恢复或收益较低的情况下长期维持无效预留,从而降低资源碎片化和无效占位的现象。
Smart Images

Figure CN122450688B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of cross-domain resource scheduling technology, specifically to a unified scheduling method and system for cross-domain computing power. Background Technology
[0002] In the evolution of converged computing platforms, high performance computing (HPC) and artificial intelligence computing (AI) have long had differences in usage patterns and scheduling.
[0003] Specifically, when the same platform simultaneously supports online services, offline training, and multi-node HPC jobs, using multiple entry points and independent scheduling systems can easily lead to problems such as resource silos, policy fragmentation, auditing difficulties, and challenges in unifying the orchestration of cross-domain workflows. Furthermore, quantum hardware (Quantum Processing Unit, QPU) typically provides services in queues or sessions, while hybrid quantum and classical algorithms require classical computing resources and quantum execution resources to operate collaboratively.
[0004] Existing solutions still lack a consistent resource coordination mechanism across HPC, AI, and quantum scheduling domains, which can easily lead to partial scheduling success and partial scheduling failure of cross-domain resources, resulting in a low success rate of unified scheduling. Summary of the Invention
[0005] In view of the above-mentioned defects or deficiencies in the prior art, this application aims to provide a unified scheduling method and system for cross-domain computing power of quantum, supercomputing and intelligent computing, so as to solve the problem of poor resource coordination across quantum domain, supercomputing domain and intelligent computing domain in related technologies, and improve the success rate of cross-domain resource scheduling.
[0006] This application provides a method for unified scheduling of cross-domain computing power in a super-intelligent computing system, the method comprising:
[0007] In response to receiving a hybrid job request, a cross-domain collaborative object corresponding to the hybrid job is generated, wherein the hybrid job includes computing tasks in at least two target domains, and the target domains are quantum domain, supercomputing domain, or intelligent computing domain;
[0008] Read resource status information from each of the target domains, determine the target resource combination based on the resource status information, and initiate the reservation of the target resource combination;
[0009] The single-domain readiness of each target domain is determined based on the resource status information of each target domain, and the cross-domain launch alignment between each target domain and the rollback cost of canceling the reserved target resource combination are determined.
[0010] Based on the single-domain readiness, the cross-domain startup alignment, and the rollback cost, it is determined whether the target resource combination meets the joint startup conditions, and a consistency permission judgment result is obtained.
[0011] Based on the consistency permission judgment result, the cross-domain collaboration object is bound to the target resource combination and submitted to each of the target domains, or the current reservation is maintained and entered into waiting control, or the cross-domain collaboration object is rolled back.
[0012] This application also provides a unified scheduling system for cross-domain computing power of a super-intelligent computing system, the system including a unified scheduling kernel module and a multi-backend adapter module;
[0013] The unified scheduling kernel module is used to execute the intelligent cross-domain computing power unified scheduling method provided in any embodiment of this application;
[0014] The multi-backend adapter module is used to receive the scheduling request sent by the unified scheduling kernel module, determine the target adapter according to the scheduling request, and submit the corresponding job in the mixed job to the corresponding target domain through the target adapter.
[0015] This application embodiment also provides an electronic device, the electronic device comprising:
[0016] Processor and memory;
[0017] The processor executes the steps of the unified scheduling method for cross-domain computing power provided in any embodiment of this application by calling the program or instructions stored in the memory.
[0018] This application also provides a computer-readable storage medium that stores a program or instructions that cause a computer to execute the steps of the unified scheduling method for cross-domain computing power provided in any embodiment of this application.
[0019] In summary, this application proposes a unified scheduling method for cross-domain computing power in quantum, supercomputing, and intelligent computing domains. Upon receiving a mixed job request, this method generates a cross-domain collaborative object corresponding to the mixed job. It then reads resource status information from each target domain involved in the mixed job, determines the target resource combination based on the resource status information, and initiates a reservation for the target resource combination. It determines the single-domain readiness of each target domain through the resource status information of each target domain, and determines the cross-domain startup alignment between target domains and the rollback cost of canceling the reserved target resource combination. Furthermore, it judges whether the target resource combination meets the joint startup conditions based on the single-domain readiness, cross-domain startup alignment, and rollback cost, obtaining a consistency permission judgment result. Finally, based on the consistency permission judgment result, it binds the cross-domain collaborative object to the target resource combination and submits it to each target domain; or, it maintains the current reservation and enters waiting control; or, it rolls back the cross-domain collaborative object, thereby achieving unified scheduling of cross-domain resources in the quantum domain, supercomputing domain, and intelligent computing domain. This method determines cross-domain global consistency by assessing the readiness of each target domain. This ensures that resources are bound when all target domains have stable and sufficient reservations, avoiding binding only when a single target domain is stable. This prevents scenarios where some resources are occupied while others cannot be obtained in time, thus improving the success rate of cross-domain joint scheduling and reducing resource waste caused by repeated queuing, invalid waiting, and repeated rollbacks. Furthermore, this method determines cross-domain global consistency by assessing cross-domain startup alignment, ensuring the consistency of startup time for each target domain. This reduces the problem of long-term idling, waiting, or occupation of intelligent computing nodes, supercomputing nodes, or quantum devices due to inconsistent startup timing. Moreover, this method determines cross-domain global consistency by assessing rollback costs, enabling quantitative decision-making between continuing to wait and immediate rollback. This avoids maintaining invalid reservations for a long time in situations where recovery is impossible or the benefits are low, thereby reducing resource fragmentation and invalid occupation. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the specific embodiments of this application or the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a unified scheduling method for cross-domain computing power provided in an embodiment of this application;
[0022] Figure 2 This is a schematic diagram of the overall system architecture and cross-domain consistency scheduling closed loop provided in an embodiment of this application;
[0023] Figure 3This is a schematic diagram of CDR object relationships provided in an embodiment of this application;
[0024] Figure 4 This is a flowchart of a unified scheduling method provided in an embodiment of this application;
[0025] Figure 5 This is a timing diagram of hybrid quantum-classical cooperative reservation and session licensing provided in an embodiment of this application;
[0026] Figure 6 This is a placeholder operation execution path diagram for an HPC adapter interfacing with Slurm, provided in an embodiment of this application.
[0027] Figure 7 This is a schematic diagram of a quantum-classical hybrid iterative workflow provided in an embodiment of this application;
[0028] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0029] The present application will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings.
[0030] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0031] Before introducing the method provided in the embodiments of this application, the technical problem solved by the method will be explained first. Existing solutions still lack a consistent resource coordination mechanism across HPC, AI and quantum domains, which can easily lead to problems such as resource silos, policy fragmentation, auditing difficulties and difficulties in unified orchestration of cross-domain workflows. Furthermore, it can easily lead to situations where some resources are scheduled successfully while others fail to be scheduled.
[0032] Unlike existing schemes that make scheduling decisions only for a single scheduling domain (such as Kubernetes node resources), this application addresses multiple independent scheduling domains: for example, HPC scheduling domains such as Slurm, intelligent computing scheduling domains of Kubernetes native or batch scheduling systems, and quantum service scheduling domains that provide capabilities via queues / sessions. These scheduling domains differ in resource state representation, availability determination, job submission semantics, and failure recovery methods, making it difficult for traditional schedulers to directly guarantee consistency in cross-domain resource allocation. This application addresses this by designing a mechanism for extending the scheduling lifecycle of the Kubernetes Scheduling Framework, enabling the Kubernetes scheduler to not only perform node selection but also act as a cross-scheduling domain resource coordinator.
[0033] The consistency scheduling goal to be achieved by the embodiments of this application is not simply a unified entry point submission, but rather, after the unified entry point, further ensuring that the following technical conditions are met:
[0034] 1. Cross-domain resource coordination and reachability: The classical and quantum execution resources required for the same hybrid operation must be coordinated within the same scheduling cycle to avoid "invalid occupation caused by only partial resource reachability";
[0035] 2. Atomicity of cross-domain resource allocation: For the same cross-domain collaborative object, the resource allocation result should satisfy the atomicity constraint of all successes or all reversals, rather than allowing one scheduling domain to succeed first and another scheduling domain to fail later;
[0036] 3. Cross-domain failure recovery consistency: When any scheduling domain fails during the reservation, permission, or commit phase, reservation actions already completed in other scheduling domains should be able to be rolled back uniformly to avoid issues such as resource hanging, inconsistent states, or inconsistent ledgers.
[0037] 4. Scalability of the scheduling process: After introducing quantum resources, it should not be implemented by forking the Kubernetes core scheduler, but should be based on the extension points of the Scheduling Framework to complete the mechanism injection, so as to extend more heterogeneous resource types and policy plugins in the future.
[0038] The embodiments of this application construct a cross-heterogeneous computing power consistency scheduling mechanism under the above-mentioned objective constraints. Figure 1 This is a flowchart illustrating a unified scheduling method for cross-domain computing power in a super-intelligent computing system, as provided in an embodiment of this application. See also... Figure 1 The unified scheduling method for super-intelligent cross-domain computing power specifically includes:
[0039] S110. In response to receiving a hybrid job request, generate a cross-domain collaboration object corresponding to the hybrid job.
[0040] Hybrid jobs include computational tasks in at least two target domains: quantum, supercomputing, or intelligent computing. The intelligent computing domain can be a scheduling domain based on the Kubernetes scheduler, and the supercomputing domain can be a scheduling domain based on the HPC scheduler.
[0041] In this embodiment of the application, a job request sent by a user terminal or an upper-layer platform interface can be received, the request parameters can be validated and standardized, and the standardized task description can be written into a unified task object (such as UnifiedComputeTask).
[0042] The number of unified task objects corresponds to the number of tasks involved in the job. A unified task object describes the resource requests, scheduling strategies, operating parameters, and unified status of a single intelligent computing task, supercomputing task, or quantum computing task. The unified task object serves as the main task object in this embodiment, carrying the core task description after the unified entry point is submitted. This object is preferably used to uniformly express the three types of tasks and their hybrid forms, and serves as the main input object for the unified scheduling kernel.
[0043] For example, a unified task object is used to describe at least the following: task type (HPC, AI, Quantum, or Hybrid), classic resource request and quantum resource request, scheduling constraints and policy attributes, runtime parameters (image, command, script, data location, etc.), lifecycle state and backend job identifier, and association with cross-domain collaboration objects or cross-domain reserved objects. The state of a unified task object may also include task state, identifier of the module processing the task, task start time, and task end time.
[0044] The following is an example description of the field structure of a unified task object:
[0045] 1) Metadata fields, including name, namespace, tags, annotations, etc., are used for tenant isolation, policy matching, tracking and auditing;
[0046] 2) Task specification field (spec), which may include, but is not limited to:
[0047] a. Task type field
[0048] spec.workloadType: Used to identify the task type, such as HPC (Hypercomputing Domain), AI (Intelligent Computing Domain), Quantum (Quantum Domain), Hybrid (Hybrid).
[0049] b. Resource Request Fields (Unified Resource Abstraction)
[0050] spec.resources.classic: Classic resource requests, which must include at least one or a combination of CPU (Central Processing Unit), memory, GPU (Graphics Processing Unit), number of nodes, topology hints, wall clock time, etc.
[0051] spec.resources.quantum: Quantum resource request, which includes at least one or a combination of QPU selection criteria, shots, execution mode (batch / session / hybrid), session attributes, latency / window preferences, etc.
[0052] c. Scheduling Constraints and Policy Fields
[0053] spec.scheduling includes priority, queue / partition, QoS (Quality of Service), deadline, preemption policy, backfill policy, failure handling policy, retry policy, etc.
[0054] spec.profileRef (optional): Used to bind a specific Profile or policy set in the unified scheduling kernel;
[0055] d. Runtime fields
[0056] pec.runtime: includes container images, commands, parameters, script references, environment variables, input / output data locations, data preprocessing or distribution parameters, etc.
[0057] e. Coordinated scheduling related fields
[0058] spec.coScheduleGroupRef: Associated cross-domain collaboration object or cross-domain reserved object;
[0059] spec.reservationPolicy: Defines whether cross-domain joint reservation is required, whether partial waiting is allowed, timeout policy, etc.
[0060] 3) Status field, which may include, but is not limited to:
[0061] status.phase: Unifies lifecycle states (such as Pending, Scheduled, Reserved, Permitting, Submitted, Running, Succeeded, Failed, RolledBack, etc.).
[0062] status.conditions: A set of conditional states used to reflect the results of scheduling, reservation, permission, execution, and other phases in a fine-grained manner;
[0063] status.backendRefs: Records the backend job identifier, session identifier, or request identifier for HPC / AI / Quantum;
[0064] status.reservationRef: Associated cross-domain reservation object;
[0065] status.startTime / endTime: Task start and end times;
[0066] status.accounting: Resource usage statistics, wait time, execution time, failure reason code, number of rollbacks, etc.
[0067] status.lastError: The most recent error message or error code.
[0068] To support the consistency permission determination process in the subsequent Permit phase, the status fields may further include: status.readinessScores (records the single-domain readiness of each scheduling domain), status.alignmentScore (records the cross-domain startup alignment), status.rollbackCost (records the rollback cost), status.globalConsistencyScore (records the global consistency determination value), status.permitDecision (records the allow, wait, or rollback result output by the Permit phase), and status.permitThresholds (records the threshold configuration or threshold version identifier used in the Permit phase).
[0069] The aforementioned fields can be used to record the consistency permission judgment inputs, calculation results, and decision outputs during the Permit phase, thereby supporting status write-back, policy review, audit trail, retry judgment, and fault replay. By generating a unified task object, heterogeneous jobs can be abstracted into a single scheduling object after a unified entry point. The unified scheduling kernel module can then use this object to complete cross-domain resource arbitration, consistency control, and backend commits, thus avoiding the problem of multiple interface objects and multiple status definitions coexisting in traditional solutions.
[0070] Furthermore, in the hybrid quantum-classical job scenario, the "classical resource set" and the "quantum execution window / session resource" are logically constructed into a unified task object, rather than treating quantum calls as an adjunct step after classical job execution. The classical resource set is used to carry classical computational processes such as optimizer calculations, parameter updates, data processing, and orchestration control, while the quantum execution window / session resource is used to carry quantum circuit execution and measurement processes. Its availability is typically manifested as queue status, session permissions, or time slot reachability. Through this unified construction mechanism, the scheduler can simultaneously process "spatial resources" (such as CPU / GPU / nodes) and "time window resources" (such as quantum sessions) within the same lifecycle, thereby achieving coordinated startup and consistency control of hybrid jobs. This mechanism overcomes the limitations of existing solutions that separate "classical computation scheduling" and "quantum service calls."
[0071] Specifically, job requests sent by the user terminal or upper-layer platform interface can be hybrid job requests or independent job requests. A hybrid job consists of computational tasks from at least two different target domains, while an independent job request consists of computational tasks from a single target domain. For example, a hybrid job may include intelligent computing domain computational tasks and supercomputing domain computational tasks, or a hybrid job may include intelligent computing domain computational tasks, supercomputing domain computational tasks, and quantum domain computational tasks.
[0072] In this embodiment, for independent job requests, only a unified task object for the object can be generated. For mixed job requests, in addition to generating unified task objects for each task involved, a cross-domain collaboration object for the entire mixed job also needs to be generated. This cross-domain collaboration object is used to describe the DAG (Directed Acyclic Graph), iterative cycles, or group-level collaborative initiation constraints between multiple tasks in the mixed job.
[0073] Specifically, a cross-domain collaborative object (CoScheduleGroup) can describe the collaborative relationships, dependencies, and group-level consistency control requirements among multiple unified task objects. It is suitable for gang scheduling, hybrid quantum-classical (intelligent computing or supercomputing) iterative workflows, and cross-domain pipeline scenarios. Specifically, the cross-domain collaborative object can be used in the following scenarios: group-level joint scheduling where multiple tasks need to "simultaneously satisfy resource requirements before starting"; cross-domain task orchestration with DAG dependencies; hybrid quantum-classical algorithms with iterative relationships (such as classical optimizers repeatedly calling quantum execution); and group-level failure rollback, retry, or degradation control. For example, a cross-domain collaborative object may include the following fields:
[0074] 1) Group Member Field
[0075] spec.members: A list of references to the associated Uniform Task objects, or select group members using a tag selector;
[0076] 2) Graph structure fields
[0077] spec.graph: Used to describe DAGs, phase partitions, or cycle relationships;
[0078] spec.loopPolicy (optional): Used to describe the number of iterations, termination condition, convergence condition, etc.
[0079] 3) Collaborative Reserved Fields
[0080] spec.coReserve: Indicates whether cross-domain synchronization reservation / window alignment is required;
[0081] spec.groupPermitPolicy: Group-level permit waiting policy, timeout policy, partial success handling policy, etc.;
[0082] 4) Failure handling field
[0083] spec.failurePolicy: Group-level rollback, retry, skip, and degradation (e.g., switching to simulator) policies;
[0084] 5) Status field
[0085] status.phase: Group-level status (Pending, Reserving, Permitting, Running, Failed, Completed, etc.);
[0086] status.memberStatusSummary: Summary of member task status;
[0087] status.lastFailureReason: Group-level failure reason;
[0088] status.iterationState: The round status and intermediate metrics of an iterative task.
[0089] By generating cross-domain collaborative objects corresponding to hybrid jobs, single-task scheduling can be extended to cross-domain group-level collaborative scheduling, enabling the unified scheduling kernel module to perform group-level waiting and consistency permission judgment during the Permit phase, and to perform group-level rollback or retry control in abnormal scenarios.
[0090] In addition to generating the aforementioned cross-domain collaboration object, in this embodiment of the application, in order to carry the description of the consistency permission judgment in the subsequent Permit stage, a cross-domain reserved object corresponding to the cross-domain collaboration object can also be generated, so that the consistency scheduling mechanism has the feasibility, observability and auditability.
[0091] In some implementations, after generating the cross-domain collaboration object corresponding to the hybrid job, the method further includes:
[0092] Generate a cross-domain reserved object corresponding to the cross-domain collaboration object; the status of the cross-domain reserved object includes resource waiting for reservation, resource reservation in progress, partial reservation successful, reservation completed, waiting for joint condition verification, license obtained, binding submission completed, waiting for rollback, resource release completed, retry refused, and job completed.
[0093] Among them, the cross-domain reserved object is used to describe the intermediate state, reservation identifier and rollback information of the cross-domain collaborative object in the resource reservation process, consistency permission judgment process, waiting process and rollback process.
[0094] In this embodiment, the cross-domain reservation object includes at least the following states: resource waiting for reservation (Pending: entered the unified scheduling process, but cross-domain resource reservation has not yet been completed), resource reservation in progress (Reserving: initiating resource reservation requests to one or more scheduling domains), partially reserved successfully (PartiallyReserved: some scheduling domains have successfully reserved, while others are still waiting or have failed), reservation completed (Reserved: reservation on the classic resource side is completed, meeting the conditions for entering the consistency permission judgment), waiting for joint condition verification (Permitting: consistency permission judgment is in progress), permission obtained (Permitted: unified permission has been obtained, and the binding and submission stage can be entered), binding and submission completed (Committed: binding and backend job submission have been completed, entering unified lifecycle management), waiting for rollback (Rollingback: rollback is executed because any scheduling domain condition is not met), resource release completed (RolledBack: unified resource release and intermediate state cleanup have been completed), retry refused (Failed: scheduling or execution failed and no further retrying is allowed), and job completed (Completed: job execution is completed and resource reclamation is completed).
[0095] It should be noted that by defining the state of the aforementioned cross-domain reserved objects, the phased states of resources in different scheduling domains during the joint scheduling process can be recorded, making the entire cross-domain scheduling mechanism feasible and auditable.
[0096] To ensure the reliability of state transitions for cross-domain reserved objects, the following constraints can be designed: 1. Unidirectional progression constraint (the state progresses from resource waiting for reservation to reservation completion, license acquisition, binding submission completion, etc., and no reverse jumps are allowed except for rollback paths to avoid ambiguity in state rollback under concurrent conditions); 2. Unified submission precondition constraint (tasks that have not reached the license acquisition state cannot enter the binding submission completion state; tasks that have not completed classical resource reservation and quantum license verification cannot execute final binding and submission); 3. Partial success visibility constraint (when in the partial reservation success state, the system regards the task as "not completed joint scheduling", and it must not be exposed as a successfully scheduled state, and must continue to wait or enter rollback when the policy conditions are met); 4. Failure rollback closed-loop constraint (when any scheduling domain fails, times out, or is rejected in the Reserve or Permit phase, the state enters the waiting rollback state, and enters the resource release completion state after the reserved resources are released, to ensure the consistency of the cross-domain resource state closed loop); 5. Audit trail constraints (recording transition time, triggering reason, associated backend identifier, and failure reason code during state transitions for unified billing, auditing, and fault analysis). By defining the above state machine and state transition rules for cross-domain reserved objects, the cross-heterogeneous computing power consistency scheduling mechanism can have clear implementation boundaries, observability, and verifiability, avoiding keeping consistency control at the level of abstract description.
[0097] In this embodiment, a unified task object can be associated with a cross-domain collaborative object or a cross-domain reserved object in the following ways: 1. Explicit reference field association (e.g., a unified task object can directly reference a cross-domain collaborative object or a cross-domain reserved object); 2. Tag / annotation association (used for policy selection, resource matching, and audit tracing); 3. OwnerReference association (used for object lifecycle linkage and automatic cleanup); 4. Controller coordination association (triggered by the unified scheduling kernel module based on object state changes, including reservation, permission, commit, rollback, and state writeback). Through the above association methods, a collaborative closed loop of "task object—resource directory—cross-domain reserved object—cross-domain collaborative object" can be formed in the Kubernetes declarative object system.
[0098] In addition to the objects mentioned above, policy binding objects (PolicyBinding / SchedulingProfileBinding) can also be constructed for hybrid jobs. To avoid embedding all policy fields into a unified task object, a policy binding object can be further introduced to bind tenants, projects, task types, or business lines to configurations such as policy templates, scheduling profiles, permit thresholds, and rollback policies. Specifically, this policy binding object is used to achieve: decoupling policy configuration from task objects, differentiated scheduling policies for different tenants / projects, canary releases of new scheduling policies, and unified policy changes without modifying the structure of historical task objects. It is worth noting that this policy binding object is optional and does not constitute a necessary condition for implementing a unified scheduling mechanism.
[0099] S120. Read resource status information from each target domain, determine the target resource combination based on the resource status information, and initiate the reservation of the target resource combination.
[0100] In this embodiment, the cross-domain consistency control process based on the Kubernetes scheduling lifecycle can be divided into three stages: 1. Reserve stage: cross-domain resource reservation; 2. Permit stage: consistency permission judgment; 3. Unreserve stage: unified failure rollback.
[0101] In this embodiment, after receiving the hybrid job request and generating the corresponding cross-domain collaboration object, a cross-domain resource reservation phase can be entered to perform reservation operations on the various resources required by the hybrid job. The classical computing resources (intelligent computing or supercomputing resources) required by the hybrid job include, but are not limited to, CPU, memory, GPU, HPC nodes, node topology constraints, runtime constraint-related resources, etc., while the quantum computing resources required by the hybrid job include Session, QueueSlot Session, or execution window.
[0102] Specifically, during the cross-domain resource reservation phase, resource status information can first be read from each target domain. For example, resource status information can be synchronized or read from the HPC supercomputing domain, the Kubernetes intelligent computing domain, and the quantum service scheduling domain. Based on this read resource status information, a unified resource view is constructed, and a candidate resource set is formed according to the resources and constraints required for the hybrid job. The resource status information must include at least one of the following:
[0103] 1) HPC queues / partitions, node availability, QoS / account constraints;
[0104] 2) Kubernetes nodes, GPU extended resources, and tag / taint / affinity attributes;
[0105] 3) QPU device availability, queue status, session capability, and window reachability.
[0106] For example, to describe the view of schedulable resources in different scheduling domains, a unified resource view can be generated by constructing a unified computing resource catalog object. This unified computing resource catalog object is used to maintain a unified view of schedulable resources such as HPC partitions, Kubernetes node pools, and quantum endpoints.
[0107] For example, a unified computing resource catalog object can be used to extract the following types of resources: HPC partitions / queues and their node sets, Kubernetes node pools or resource pools with specific tags, quantum device endpoints (QPUEndpoint) or quantum service capability units, simulator resource endpoints (Quantum Simulator), and other extended heterogeneous resource pools (such as FPGAs, dedicated inference devices, etc.). A unified computing resource catalog object can include the following fields:
[0108] 1) Type field
[0109] spec.kind: Resource directory item type, such as HPCPartition, K8sNodePool, QPUEndpoint, etc.
[0110] 2) Capacity field
[0111] spec.capacity / spec.allocatable: Total and allocatable resources, supporting representation in a uniform unit or a mapped unit;
[0112] 3) Ability tag field
[0113] spec.labels: Used to describe GPU model, interconnect type, region, QPU provider, device type, gate set limitations, network characteristics, etc.
[0114] 4) Scheduling attribute field
[0115] spec.schedulerAttributes: includes queueing policies, QoS restrictions, session support capabilities, window characteristics, reserved capability flags, etc.
[0116] 5) Status field
[0117] status.health: Health status; status.availability: Availability status; status.queueDepth: Queue depth or congestion metric (applicable to HPC / Quantum resources); status.lastSyncTime: Last synchronization time; status.version: Resource view version number or snapshot number.
[0118] By constructing the aforementioned unified computing power resource catalog object, the heterogeneous resource states of different scheduling domains can be mapped to unified and readable resource catalog entries, enabling the unified scheduling kernel module to perform policy judgments in the filtering, scoring, and consistency permission judgment stages without directly relying on the original backend format.
[0119] In this embodiment, the provided unified task object, cross-domain collaboration object, cross-domain reservation object, and unified computing resource catalog object are preferably constructed using Kubernetes Custom Resource Definitions (CRDs) to build a data model for unified tasks, unified resource catalogs, and cross-domain collaboration reservation objects. This approach allows heterogeneous resource requests, scheduling constraints, lifecycle states, and collaboration relationships of HPC, AI, and quantum operations to be incorporated into the Kubernetes API management system in the form of declarative objects. Kubernetes' native support for CRDs enables this embodiment to obtain versioned APIs, declarative object management, controller tuning, and state write-back capabilities, thereby providing a standardized platform for unified entry, unified scheduling, and unified auditing.
[0120] It should be noted that the above objects are merely exemplary implementations used to illustrate the implementation path of the unified abstraction and scheduling mechanism. The specific fields, naming conventions, splitting granularity, and version evolution methods of the objects can be tailored, merged, or extended according to platform engineering requirements. The system implementation is not limited to a specific quantum service provider, a specific HPC scheduling system, or a specific AI batch scheduling system. As long as the underlying execution domain can provide resource status query, job submission / cancellation, and status feedback capabilities, it can be connected to the unified scheduling and service orchestration system of this application embodiment through a backend adapter.
[0121] The object design in this application prioritizes the following modeling objectives:
[0122] 1) Unified expression ability
[0123] It can express classical computing resources (CPU, memory, GPU, nodes, topology, wall clock time, etc.) and quantum resources (QPU selection, shots, session / batch processing mode, queue / window constraints, etc.) within the same object system.
[0124] 2) Scheduling Consumability
[0125] The field structure should facilitate direct reading and processing by the kernel during the Filter, Score, Reserve, Permit, and Bind stages.
[0126] 3) Lifecycle observability
[0127] Supports writing back the status field to reflect the status of stages such as reservation, permission, commit, execution, rollback, and completion;
[0128] 4) Cross-domain collaboration and relevance
[0129] Supports association and state synchronization between unified task objects and cross-domain reserved objects, cross-domain collaborative objects, and unified computing resource catalog objects;
[0130] 5) Scalability and version compatibility
[0131] It supports the subsequent integration of new heterogeneous resource types, scheduling backends and policy fields, and maintains compatibility through CRD (Custom Resource Definition, Kubernetes) version evolution.
[0132] In terms of modeling principles, this application preferably adopts a declarative pattern where "spec describes the desired state and status describes the actual state," and introduces "intermediate state objects (such as Reservation / Lease class objects)" to carry out cross-domain consistency control processes when necessary. A unified computing power abstraction model is established through Kubernetes CRDs to uniformly describe HPC, AI, and quantum resources, and a pluggable policy mechanism is used to achieve: cross-domain fairness and priority control; unified quota and billing; preemption, backfilling, and cost optimization; data and network proximity policies; and end-to-end observability and auditing capabilities are achieved through a unified job ledger.
[0133] In this embodiment, after forming a candidate resource set based on the resources and constraints required for the hybrid job, the candidate resource set can be further filtered according to the resources and constraints required for the hybrid job. Then, based on strategies such as priority, fairness, cost, congestion level, and data proximity, the remaining resources in the candidate resource set are scored, and sorted in descending order of score to select the preferred resource combination, thus obtaining the target resource combination. Specifically, the target resource combination consists of cross-domain computing resources, which may include classical resources (CPU / memory / GPU / HPC nodes, etc.) and quantum resources (QPU / session / window, etc.).
[0134] After determining the target resource combination, a further step is to initiate a cross-domain reservation operation for the target resource combination, so that the hybrid operation enters an intermediate state of "reserved but not submitted". It should be noted that reserving the target resource combination does not mean that the quantum resources have been finally licensed.
[0135] Specifically, unlike traditional Kubernetes scheduling scenarios where resource occupancy marking is only applied to a single node, the Reserve phase in this application embodiment can perform reservation actions for multiple backend scheduling domains, including: if the target resource combination includes intelligent computing domain resources, then the reservation status of resources such as GPUs, nodes, and affinity is recorded on the Kubernetes or intelligent computing side; if the target resource combination includes supercomputing domain resources, then the reservation application or placeholder reservation of node or queue resources is triggered on the HPC or Slurm side; if the target resource combination includes quantum domain resources, then the pre-application, soft placeholder, queuing context attachment, or candidate window registration of candidate sessions, queue slots, or execution windows is initiated on the quantum domain side, and the corresponding placeholder token, globally unique request ID, queue ticket, or equivalent associated identifier is obtained.
[0136] In addition to performing the above reservation operations, the corresponding reservation identifier, reservation validity period, reservation status and associated task identifier can also be recorded in the cross-domain reservation object.
[0137] By reserving the target resource combination, an intermediate state of resource locking but not yet final submission can be established, providing a basis for subsequent global consistency permission judgment and avoiding inconsistent resource visibility and duplicate occupation caused by multiple jobs competing for resources concurrently.
[0138] S130. Determine the single-domain readiness of each target domain based on the resource status information of each target domain, and determine the cross-domain startup alignment between each target domain and the rollback cost of canceling the reserved target resource combination.
[0139] Specifically, after reserving the target resource combination, the consistency permission judgment phase (i.e., the Permit phase) can be entered. In this phase, a global consistency judgment can be made on whether the cross-domain target resource combination meets the joint startup conditions. Unlike the conventional Permit usage, which only allows rapid release based on node resource availability, this application embodiment extends the Permit phase to a "unified permission gate" across heterogeneous computing power, used to handle the special scheduling semantics of quantum execution resources.
[0140] Quantum resources are typically provided in the form of Sessions, Queue Slots, or execution windows. They inherently possess distinct time window attributes and service-side priority semantics, and cannot be simply equated to static capacity resources. Therefore, this application's embodiments can introduce the following control logic in the Permit phase: 1. Quantum Session Availability Judgment: Determine whether the target QPU device or quantum service has an available session, an acceptable queuing state, or an execution window that meets the policy threshold; 2. Joint Condition Verification: Verify whether the classical resource reservation state, quantum session state, and task dependencies in cross-domain collaborative objects are simultaneously satisfied; 3. Waiting and Timeout Control: When classical resources are reserved but the quantum session is temporarily unavailable, the Permit phase can enter a waiting state and execute timeout exit, retry, or downgrade according to a preset policy; 4. Allow / Reject Decision: When the joint startup conditions are met, a unified allow permission is issued; when the conditions are not met and cannot be recovered, a rejection signal is issued and a rollback process is initiated.
[0141] Through the above mechanism, the Permit phase is no longer just the starting permit point for a single Pod, but becomes a global commit gate for cross-domain resource consistency coordination, thereby ensuring the consistency of the startup timing of hybrid jobs.
[0142] Specifically, the readiness of each target domain can be assessed by evaluating the resource status information of each target domain. Furthermore, the cross-domain startup alignment between target domains and the rollback cost of canceling reserved target resource combinations can be evaluated to achieve a global consistency permission judgment for target resource combinations.
[0143] The single-domain readiness level reflects the degree to which a target domain can allow resources from the corresponding domain in the target resource combination. In this embodiment, the likelihood of a target domain allowing resources from the corresponding domain in the target resource combination can be assessed based on whether the target domain's capacity is satisfied, whether the policy is satisfied, whether the target domain's resource status information is reliable, and the stability of maintaining effective reservations, thus obtaining the single-domain readiness level.
[0144] It should be noted that since the target resource portfolio contains multiple resources across domains, the readiness level of each target domain involved in the target resource portfolio needs to be determined separately.
[0145] In one specific implementation, determining the single-domain readiness of each target domain based on the resource status information of each target domain includes the following steps:
[0146] Step 21: For each target domain, determine the effective allocable quantity for each resource dimension based on the original allocable quantity, resource fragmentation correction coefficient, topology matching correction coefficient, and duration matching correction coefficient for each resource dimension of the target domain.
[0147] Step 22: Based on the effective allocable quantity and request quantity under each resource dimension, determine the single-dimensional resource satisfaction rate corresponding to each resource dimension, and determine the capacity satisfaction of the target domain based on the single-dimensional resource satisfaction rate;
[0148] Step 23: Determine whether the target resource combination satisfies the hard constraints of the target domain, obtain the hard constraint satisfaction factor, and determine the soft constraint score based on the degree to which the target resource combination satisfies the soft constraints of the target domain. Based on the hard constraint satisfaction factor and the soft constraint score, determine the policy satisfaction of the target domain.
[0149] Step 24: Based on the target domain's state change rate, queue fluctuation degree, historical synchronization error, and interface reliability, determine the freshness decay parameter of the target domain, and determine the state freshness of the target domain according to the target domain's resource state synchronization time and freshness decay parameter.
[0150] Step 25: Determine the predicted preemption probability, predicted expiration probability, predicted conflict probability, and state fluctuation risk value of the target resource combination, and determine the reserved stability of the target domain based on the predicted preemption probability, predicted expiration probability, predicted conflict probability, and state fluctuation risk value.
[0151] Step 26: Determine the single-domain readiness of the target domain based on capacity satisfaction, policy satisfaction, state freshness, and reservation stability.
[0152] Specifically, the readiness of a single domain can be obtained by determining the capacity satisfaction, policy satisfaction, state freshness, and reservation stability of the target domain, thereby comprehensively assessing the likelihood of releasing corresponding resources in the target resource combination.
[0153] Capacity satisfaction can be used to describe the degree to which the currently allocable resources of the target domain can satisfy the corresponding resources in the target resource combination. In the embodiments of this application, the capacity satisfaction of the target domain can be determined by evaluating the single-dimensional resource satisfaction rate of the target domain under different resource dimensions.
[0154] Specifically, in step 21, for each target domain, considering that different target domains involve different resource dimensions, the effective allocable quantity under each resource dimension can be determined first based on the original allocable quantity, resource fragmentation correction coefficient, topology matching correction coefficient, and duration matching correction coefficient under each resource dimension.
[0155] For example, for the Kubernetes intelligent computing domain, dimensions may include CPU, memory, GPU, number of nodes, affinity-related schedulable capacity, etc.; for the HPC scheduling domain, dimensions may include number of nodes, number of CPU cores, memory, GPU, remaining capacity of partitions / queues, wall clock time constraints, etc.; for the quantum resource domain, dimensions may include number of available sessions, number of queue slots, available execution window length, shot capacity (the ability to process multiple instance / request / task instances simultaneously), device service concurrency, etc.
[0156] For example, for any resource dimension, the effective allocatable quantity under that resource dimension can be calculated using the following formula:
[0157] ;
[0158] In the formula, Representing resource dimensions The effective allocatable quantity below, Representing resource dimensions The original allocatable quantity below, Representing resource dimensions Resource fragmentation correction coefficient, Representing resource dimensions The topology matching correction coefficient (or affinity matching correction coefficient) can be used. Representing resource dimensions The duration matching correction factor (which can be understood as the time window or retention duration matching correction factor).
[0159] After obtaining the effective allocable quantity for each resource dimension, further, in step 22, the single-dimensional resource satisfaction rate corresponding to each resource dimension can be determined based on the effective allocable quantity and the request quantity for each resource dimension, as shown in the following formula:
[0160] ;
[0161] In the formula, Representative resource dimension The corresponding single-dimensional resource satisfaction rate, This represents a hybrid operation in this resource dimension. The number of requests. Furthermore, the capacity satisfaction of the target domain can be determined by the satisfaction rate of all single-dimensional resources, thus comprehensively considering the bottleneck resource constraints and the overall resource satisfaction level, as shown in the following formula:
[0162] ;
[0163] In the formula, Represents the target domain Single-domain readiness, Represents the target domain All resource dimensions, For the resource sensitivity coefficient of the short board, , Represents the target domain Lower resource dimension The weights must satisfy the target domain. The sum of the weights of all resource dimensions equals 1. The above capacity satisfaction not only reflects whether the overall resources of the target domain are sufficient, but also reflects the inhibitory effect of insufficient single key resource dimension on the joint start-up conditions.
[0164] In this embodiment, policy satisfaction can be used to characterize whether the target domain's queue, QoS, quota, affinity, device constraints, permission constraints, and fairness constraints are met. Specifically, in step 23, policy constraints can be divided into two types: hard constraints and soft constraints. Based on whether the target resource combination meets the target domain's hard constraints and the degree to which the target resource combination meets the target domain's soft constraints, a hard constraint satisfaction factor and a soft constraint score are determined, thereby obtaining the policy satisfaction.
[0165] For example, hard constraints may include: whether tenant or project quotas are allowed, whether the target partition is available, whether device types are compatible, whether quantum gate sets are satisfied, whether regional or compliance policies are allowed, and whether strong affinity or strong repulsion between tasks and nodes (or devices) is satisfied. Assume the result of each hard constraint is... , The value of is 0 or 1 (0 represents not satisfied, 1 represents satisfied), and the hard constraint satisfaction factor can be calculated using the following formula:
[0166] ;
[0167] In the formula, For the target domain Hard constraint satisfaction factor, Represents the target domain The number of hard constraints.
[0168] For example, soft constraints may include QoS matching degree, priority compatibility degree, affinity satisfaction degree, fairness deviation, congestion tolerance, energy consumption or cost preference matching degree, etc. For each soft constraint, a soft constraint score can be determined based on the degree to which the target resource combination satisfies the soft constraint, and then a comprehensive soft constraint score can be obtained based on all soft constraint scores.
[0169] For example, suppose the normalized soft constraint scores are respectively used as (QoS matching degree) (Priority compatibility) (Degree of satisfaction with affinity) (Fairness bias) (Congestion tolerance) represents the overall soft constraint score, which can then be expressed as:
[0170] ;
[0171] In the formula, The weights are non-negative and satisfy the following conditions: .
[0172] After obtaining the hard constraint satisfaction factor and the soft constraint score, the hard constraint satisfaction factor and the soft constraint score can be multiplied to obtain the policy satisfaction degree of the target domain, as shown in the following formula:
[0173] ;
[0174] In the formula, For the target domain The strategy satisfaction level is determined by the soft constraint score. Therefore, when any hard constraint is not met, the strategy satisfaction level can be directly reduced to 0; when all hard constraints are met, the soft constraint score reflects the quality of the target resource combination.
[0175] In this embodiment, state freshness is used to describe the reliability of the resource state information of the target domain at the current moment. Specifically, in step 24, the freshness decay parameter of the resource state of the target domain can be dynamically evaluated by the state change rate of the target domain, the queue fluctuation degree, the historical synchronization error, and the interface reliability, as shown in the following formula:
[0176] ;
[0177] In the formula, For the target domain Freshness decay parameter For the target domain The basic decay time constant (can be pre-calibrated). Represents the target domain The frequency of resource state changes per unit time (i.e., the rate of state change). Represents the target domain The degree of fluctuation in queue depth or available capacity (i.e., queue variability). Represents the target domain The historical synchronization error rate or state drift rate (i.e., historical synchronization error). Represents the target domain The backend interface response error rate or instability (i.e., interface reliability). ~ This is the adjustment coefficient. Through the above method, scheduling domains with high volatility, high drift, and high anomaly rates can undergo faster state decay, while relatively stable scheduling domains retain a longer freshness validity period.
[0178] After obtaining the freshness decay parameter, we can further determine the time of the most recent synchronization of resource state information in the target domain, i.e., the resource state synchronization time, and combine it with the freshness decay parameter to determine the state freshness of the target domain, as shown in the following formula:
[0179] ;
[0180] In the formula, For the target domain Freshness of the state For the current time, The time of the last synchronization of resource status information for the target domain. For the target domain The freshness decay parameter.
[0181] In this embodiment, reservation stability describes the probability or degree of stability of a target domain in maintaining a valid reservation before the reservation is completed and a consensus clearance is pending. For example, reservation stability can be modeled as the survival probability related to the risk of reservation failure.
[0182] Specifically, in step 25, the predicted preemption probability, predicted expiration probability, predicted conflict probability, and state fluctuation risk value of the target resource combination can be estimated based on the target domain's historical scheduling logs, preemption statistics, reservation validity period, queue fluctuation, equipment failure rate, or current resource contention intensity. The predicted preemption probability is the probability that a reserved resource will be preempted by a higher-priority task; the predicted expiration probability is the probability of expiration and invalidation during the period from reservation completion to waiting for consistency permission determination; the predicted conflict probability is the probability that a reservation will become invalid due to subsequent resource conflicts; and the state fluctuation risk value reflects the state fluctuation risk of the target domain.
[0183] Furthermore, the reserved stability of the target domain can be determined based on the predicted preemption probability, predicted expiration probability, predicted conflict probability, and state fluctuation risk value, as shown in the following formula:
[0184] ;
[0185] In the formula, For the target domain Reserved stability, For the target domain The predicted preemption probability, For the target domain The predicted probability of maturity, For the target domain The predicted probability of conflict, For the target domain State fluctuation risk value, This represents the remaining waiting time from the current moment until the estimated time when the consensus permission judgment is completed. ~ This represents the weight of the risk item.
[0186] After calculating the capacity satisfaction, policy satisfaction, state freshness, and reservation stability, further, in step 26, the single-domain readiness of the target domain can be determined using the capacity satisfaction, policy satisfaction, state freshness, and reservation stability, as shown in the following formula:
[0187] ;
[0188] In the formula, For the target domain Single-domain readiness, for Representing the Kubernetes intelligent computing domain, for Indicates the HPC / Slurm scheduling domain. Represents the quantum field. , , , The target domains are respectively The capacity satisfaction, strategy satisfaction, state freshness, and reservation stability. ~ The weights for each evaluation item are specified. The capacity satisfaction, policy satisfaction, state freshness, and reservation stability used in the above formulas can all be normalized and mapped to a unified numerical range to facilitate weighted combination.
[0189] Steps 21-26 above, by determining the capacity satisfaction, policy satisfaction, state freshness, and reservation stability of each target domain, can be combined with the capacity, policy, freshness, and stability of the scheduling domain to evaluate the single-domain readiness of each target domain, ensuring the comprehensiveness of the evaluation, thereby further ensuring that the single-domain resources can meet the user's needs in the subsequent consistency permission judgment process.
[0190] In this embodiment of the application, cross-domain startup alignment can be used to measure the time consistency of each target domain reaching the joint startup condition. For example, the time for each target domain to reach the committable state can be estimated, and then the cross-domain startup alignment between each target domain can be determined by combining the estimated time for all target domains to reach the committable state.
[0191] In one specific implementation, determining the cross-domain launch alignment between target domains includes the following steps:
[0192] Step 31: For each target domain, determine the waiting time for resources in the target domain to meet the conditions for joint submission, the first time required for resource reservation or session status confirmation, the second time for resource status to remain stable or conflict to be resolved, and the handover processing time before resources enter the binding process.
[0193] Step 32: Based on the waiting time, the first time, the second time, and the handover processing time, determine the estimated time for the target domain to reach the submittable state;
[0194] Step 33: Determine the submission time variance between each target domain based on the estimated time corresponding to each target domain, and determine the cross-domain launch alignment between each target domain based on the submission time variance.
[0195] In step 31, for each target domain, the waiting time for resources in the target domain to reach the joint commit condition can be determined. For example, for the Kubernetes intelligent computing domain, it can be determined based on group scheduling conditions (gang conditions), GPU availability prediction, or node release prediction; for the HPC scheduling domain, it can be determined based on the expected queue waiting time, backfill window prediction, and reserved confirmation time; for the quantum domain, it can be determined based on the session allocation waiting time, queue slot queuing time, execution window open time, or quantum device congestion.
[0196] Furthermore, for each target domain, the first time required for resource reservation or session status confirmation of the target domain can be determined, as well as the second time required for the stable and continuous resource status or conflict resolution of the target domain, and the handover processing time before resources in the target domain enter the binding process can be determined.
[0197] Furthermore, in step 32, the estimated time for the target domain to reach a submittable state can be determined based on the waiting time, the first time, the second time, and the handover processing time, as shown in the following formula:
[0198] ;
[0199] In the formula, , , , These represent the waiting time, the first time, the second time, and the handover processing time, respectively. ~ As the weight of the time component, For the target domain The estimated time to reach the submittable state. It should be noted that the waiting time, first time, second time, and handover processing time used in the above formula can all be the result after normalization and mapping to a uniform numerical range, so as to facilitate weighted combination.
[0200] Furthermore, in step 33, the submission time variance between target domains can be determined based on the estimated time corresponding to each target domain, and the cross-domain launch alignment between target domains can be determined based on the submission time variance. Taking the target domains including the intelligent computing domain, supercomputing domain, and quantum domain as an example, the cross-domain launch alignment is shown in the following formula:
[0201] ;
[0202] In the formula, These represent the estimated time for the intelligent computing domain, supercomputing domain, and quantum domain, respectively. The variance of submission time between target domains. The variance reference constant is the preset value. Alignment for cross-domain startup.
[0203] Steps 31-33 above determine the estimated time for each target domain to reach the committable state, thereby obtaining the cross-domain startup alignment between target domains. This enables the determination of the degree of coordination between different scheduling domains in the time dimension, ensuring that the startup timing of each scheduling domain is close, and solving the problem of GPU, HPC nodes or QPU idling, waiting or occupying space for a long time due to inconsistent startup timing.
[0204] In this embodiment of the application, in order to quantify the trade-off between waiting and rolling back, the rollback cost of canceling the reserved target resource combination can be determined. The rollback cost can consist of the cost required for each target domain to cancel the corresponding reserved resources in the target resource combination.
[0205] In one specific implementation, determining the rollback cost of canceling the reserved target resource combination includes the following steps:
[0206] Step 41: For each target domain, determine the action time required to cancel the reserved corresponding resources in the target resource combination, the waiting loss after cancellation and re-enqueueing, the resource fragmentation loss caused by cancellation, the opportunity loss caused by lost resources, and the restoration cost after cancellation.
[0207] Step 42: Based on action time, waiting loss, resource fragmentation loss, opportunity loss, and restoration cost, determine the cancellation reservation cost corresponding to the target domain;
[0208] Step 43: Determine the rollback cost of canceling the reservation of the target resource combination based on the cancellation cost corresponding to each target domain.
[0209] In step 41, for each target domain, the action time required to cancel the reserved corresponding resource in the target resource combination can be determined. This action time can be understood as the time or overhead required to perform the release or cancellation action itself. Furthermore, for each target domain, the waiting loss after canceling the reserved corresponding resource and re-queuing, as well as the resource fragmentation loss caused by canceling the reserved corresponding resource in the target resource combination, are determined. This resource fragmentation loss can be understood as the resource fragmentation loss (such as port discontinuity) caused by releasing discontinuous resources.
[0210] Furthermore, for each target domain, determine the opportunity loss due to the loss of the corresponding resource after the reservation in the target resource combination is cancelled, and the restoration overhead required to restore the resource. This restoration overhead can be understood as the overhead required for state restoration, consistency restoration or compensation operations.
[0211] For example, in the quantum domain, opportunity loss can be further considered as scarce execution window loss, priority window loss, or session reconstruction loss; in the supercomputing scheduling domain, restoration overhead can be further considered as lost backfilling opportunities and queue reordering losses; in the intelligent computing scheduling domain, resource fragmentation loss can be further considered as fragmentation caused by the partial release of GPUs or nodes and the increased difficulty of subsequent group scheduling.
[0212] Furthermore, in step 42, for each target domain, the cancellation reservation cost corresponding to the target domain can be determined based on the action time, waiting loss, resource fragmentation loss, opportunity loss, and restoration overhead, as shown in the following formula:
[0213] ;
[0214] In the formula, For the target domain The corresponding cost of canceling the reservation, , , , , The target domains are respectively Action time, waiting loss, resource fragmentation loss, opportunity loss, and recovery cost. ~ , where represents the weighting coefficient. It should be noted that the action time, waiting loss, resource fragmentation loss, opportunity loss, and restoration cost used in the above formula can all be normalized and mapped to a uniform numerical range to facilitate weighted combination.
[0215] Furthermore, in step 43, the rollback cost of canceling the reservation of target resources can be determined based on the cancellation cost corresponding to each target domain. Taking the target domains including the quantum domain, supercomputing domain, and intelligent computing domain as an example, the rollback cost is shown in the following formula:
[0216] ;
[0217] In the formula, Indicates the cost of rollback. , , These represent the cancellation costs for the intelligent computing domain, supercomputing domain, and quantum domain, respectively. Specifically, This can be understood as the cost of releasing reserved GPUs or nodes in the intelligent computing domain. This can be understood as the cost of canceling reserved queues or nodes in the supercomputing domain. This can be understood as the cost of revoking a reserved quantum session, placeholder, or queuing context in the quantum field. ~ These represent the cost weights corresponding to each target domain. It should be noted that the cancellation cost used in the above formula can be the result after normalization and mapping to a uniform numerical range, to facilitate weighted combination.
[0218] Steps 41-43 above determine the cancellation reservation cost for the corresponding resources in the target resource portfolio by considering the action time, waiting loss, resource fragmentation loss, opportunity loss, and restoration cost of each target domain. Then, by combining the cancellation reservation costs of all target domains, the rollback cost of the entire target resource portfolio is obtained. This allows for subsequent decisions on whether to continue waiting or roll back immediately based on the rollback cost, avoiding the long-term maintenance of invalid reservations in situations that are unrecoverable or have low returns, thereby reducing resource fragmentation and zombie occupation phenomena.
[0219] S140. Based on the readiness of each single domain, cross-domain startup alignment, and rollback cost, determine whether the target resource combination meets the joint startup conditions and obtain the consistency permission judgment result.
[0220] Specifically, after determining the single-domain readiness of each target domain, the cross-domain startup alignment between target domains, and the rollback cost of canceling the target resource combination, the single-domain readiness, cross-domain startup alignment, and rollback cost can be combined to determine whether the target resource combination meets the joint startup conditions, that is, to determine whether each resource in the target resource combination can be started simultaneously, so as to realize the verification of cross-domain global scheduling consistency and obtain the consistency permission judgment result.
[0221] For example, it can be determined whether the readiness of each single domain exceeds a preset readiness threshold, whether the cross-domain startup alignment is higher than a preset alignment threshold, and whether the rollback cost is less than a preset cost threshold. If the readiness of all single domains exceeds the preset readiness threshold, the cross-domain startup alignment is higher than the preset alignment threshold, and the rollback cost is less than the preset cost threshold, then it can be determined that the target resource combination meets the joint startup conditions, and the target resource combination can be allowed to proceed; that is, the consistency permission determination result is approval.
[0222] In one specific implementation, based on the readiness of each single domain, cross-domain startup alignment, and rollback cost, it is determined whether the target resource combination meets the joint startup conditions to obtain a consistency permission determination result, including the following steps:
[0223] Step 51: Determine the minimum readiness level based on the minimum readiness level of each single domain, and determine the global consistency judgment value between each target domain based on the minimum readiness level, cross-domain startup alignment, and rollback cost.
[0224] Step 52: If the minimum readiness is not less than the preset readiness threshold and the global consistency judgment value is not less than the preset first judgment threshold, then the target resource combination is determined to meet the joint startup conditions, and the consistency permission judgment result is to allow.
[0225] Step 53: If the minimum readiness is less than the preset readiness threshold, it is determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is rollback;
[0226] Step 54: If the minimum readiness is not less than the preset readiness threshold, and the global consistency judgment value is less than the preset first judgment threshold and not less than the preset second judgment threshold, then it is determined that the target resource combination does not meet the joint start condition, and the consistency permission judgment result is waiting.
[0227] Step 55: If the minimum readiness is not less than the preset readiness threshold, and the global consistency judgment value is less than the preset second judgment threshold, then it is determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is rollback.
[0228] Specifically, in step 51, the minimum value of the readiness of all single domains can be queried first to obtain the minimum readiness. Then, the minimum readiness, cross-domain startup alignment, and rollback cost are combined to calculate the global consistency judgment value between each target domain. This global consistency judgment value is used to reflect the feasibility of all target domains starting the corresponding resources in the target resource combination at the same time, that is, to describe the consistency of cross-domain global resource scheduling.
[0229] Taking the target domain, which includes the intelligent computing domain, the supercomputing domain, and the quantum domain, as an example, the global consistency determination value can be calculated using the following formula:
[0230] ;
[0231] In the formula, This represents the global consistency judgment value. These represent the single-domain readiness of the intelligent computing domain, the supercomputing domain, and the quantum domain, respectively. For minimum readiness, For cross-domain startup alignment, To the cost of rollback, ~ These represent the weights corresponding to readiness, alignment, and cost, respectively. It should be noted that the single-domain readiness, cross-domain startup alignment, and rollback cost used in the above formulas can all be the results after normalization and mapping to a unified numerical range, in order to facilitate weighted combination and threshold comparison.
[0232] Furthermore, the weighting coefficients involved in the formulas provided in the embodiments of this application, such as , , , , , , , These parameters can be pre-set based on historical scheduling statistics, strategy configuration, business priorities, or platform experience. Alternatively, they can be dynamically adjusted through offline data training, online feedback optimization, or rule engine.
[0233] After obtaining the global consistency determination value, it is further possible to determine whether the target resource combination meets the joint startup conditions based on the minimum readiness and the global consistency determination value.
[0234] Specifically, in step 52, if the minimum readiness is not less than the preset readiness threshold and the global consistency judgment value is not less than the preset first judgment threshold, then it can be determined that the target resource combination meets the joint startup conditions, and the consistency permission judgment result is to allow.
[0235] In step 53, if the minimum readiness is less than the preset readiness threshold, it can be determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is rollback.
[0236] In step 54, if the minimum readiness is not less than the preset readiness threshold, and the global consistency judgment value is less than the preset first judgment threshold and not less than the preset second judgment threshold, then it can be determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is waiting, wherein the preset first judgment threshold is greater than the preset second judgment threshold.
[0237] In step 55, if the minimum readiness is not less than the preset readiness threshold and the global consistency judgment value is less than the preset second judgment threshold, it can be determined that the benefits of continuing to wait are insufficient, the rollback cost is high, or the cross-domain startup alignment is insufficient, the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is rollback.
[0238] Steps 51-55 above determine whether the target resource combination meets the joint startup conditions by considering the single-domain readiness, cross-domain startup alignment, and rollback cost of each target domain. Unlike methods that only allow startup based on single-domain resource availability or reservation status, this implementation can quantitatively judge the joint startup of cross-domain resources, reducing problems such as partial success, startup mismatch, long-term invalid waiting, and zombie reservations. Furthermore, this implementation also allows the Permit phase to no longer be based solely on the existence or reservation of resources, but rather on a globally consistent judgment value to control unified startup, waiting, or rollback, thereby helping to improve the success rate of cross-domain joint scheduling and reduce resource waste caused by repeated queuing, invalid waiting, and repeated rollbacks.
[0239] S150. Based on the consistency permission judgment result, bind the cross-domain collaboration object with the target resource combination and submit it to each target domain, or maintain the current reservation and enter the waiting control, or roll back the cross-domain collaboration object.
[0240] After obtaining the consistency permission judgment result, it can be determined whether to enter the unified failure rollback phase (Unreserve phase) based on the result. This phase can be understood as triggering the unified failure rollback process of the target resource combination.
[0241] Specifically, if the consistency permission judgment result is rollback, the unified failure rollback stage can be entered to roll back the cross-domain collaborative objects and release the reservation of the target resource combination. The rollback process can include at least one of the following: releasing the node or GPU reservation that has been occupied or marked in the intelligent computing domain; canceling the node / queue reservation that has been applied for in the supercomputing domain; terminating or revoking the quantum pre-session application, candidate execution window, placeholder request or queuing context.
[0242] In addition, rollback processing may include cleaning up intermediate states, temporary tokens, queuing identifiers, and associated mappings in cross-domain reserved objects. After rollback is complete, it may proceed to any of the following processing paths: retry scheduling, downgrade to an alternative device or simulator, re-queue for waiting, or terminate and mark as failed.
[0243] If the consistency permission judgment result is pending, the Permit phase can continue, the current reservation can be maintained and the system can enter the waiting control phase. The consistency permission judgment result can be re-determined by recalculating the single-domain readiness of each target domain, the cross-domain startup alignment between target domains, and the rollback cost of canceling the reservation. Then, based on the new consistency permission judgment result, the system can proceed with the release, waiting, or rollback.
[0244] If the consistency permission judgment result is "allow," then the cross-domain collaboration object can be bound to the target resource combination and submitted to each target domain. Specifically, binding the cross-domain collaboration object to the target resource combination maps the unified scheduling decision to the actual execution actions in each target domain of the backend. For example, this can be executed through multiple backend adapters, such as submitting batch jobs to the HPC scheduling domain, creating Job / Deployment or batch job objects to the Kubernetes intelligent computing domain, and submitting Hybrid Job / Session / Batch requests to subdomains.
[0245] Through the above mechanism, at least the following technical effects can be achieved: 1. Avoid cross-domain resource dangling: Avoid partial success scenarios where classical resources are occupied but quantum resources are unavailable, or quantum resources are approved but classical resources are insufficient; 2. Improve resource utilization: Reduce invalid occupation and duplicate queuing, reduce resource fragmentation, and improve the overall utilization efficiency of HPC, GPU, and QPU; 3. Improve scheduling stability and recoverability: In failure scenarios, resource states are restored through unified rollback, reducing the probability of cross-system state mismatch; 4. Enhance platform governance capabilities: Provide a consistent state transition basis for unified billing, unified auditing, unified policies, and observability; 5. Support subsequent heterogeneous resource expansion: The consistency mechanism is not limited to QPU and can be extended to FPGA, dedicated inference chips, or other resource types with independent scheduling semantics.
[0246] In one specific implementation, maintaining the current reservation and entering waiting control includes:
[0247] Maintain the reservation of the target resource combination, and return to redetermine the single-domain readiness of each target domain, determine the cross-domain launch alignment between each target domain and the rollback cost of canceling the reserved target resource combination, so as to determine whether the target resource combination meets the joint launch conditions based on the new single-domain readiness, cross-domain launch alignment and rollback cost, and obtain the consistency permission judgment result.
[0248] If, during the reservation process, the waiting time exceeds the preset duration, or the freshness of the state of any target domain is lower than the preset freshness threshold, or the global consistency judgment value between target domains is lower than the minimum judgment value, then the cross-domain collaborative object will be rolled back.
[0249] Specifically, if the consistency permission judgment result is "waiting", the reservation of the target resource combination can be maintained, and the single-domain readiness, cross-domain startup alignment and rollback cost can be returned to re-determine whether the target resource combination meets the joint startup conditions and obtain a new consistency permission judgment result.
[0250] During this process, if the waiting time exceeds a preset duration, or if the freshness of the state of any target domain falls below a preset freshness threshold, or if the global consistency judgment value among the target domains falls below the minimum judgment value, a unified failure rollback process can be triggered to roll back the target resource combination. It should be noted that if the minimum readiness level falls below a preset readiness threshold during this process, a unified failure rollback process can also be triggered to roll back the target resource combination.
[0251] The above implementation method can avoid long waiting times for target resource combinations, thereby avoiding the problem of low resource utilization caused by long-term resource reservation.
[0252] In this embodiment of the application, for the unified failure rollback phase, a strategy-based rollback can be executed based on the current joint scheduling status, the benefits of continuing to wait, the risk of cross-domain state mismatch, and the resource occupation cost of each scheduling domain.
[0253] In one specific implementation, rolling back cross-domain collaborative objects includes the following steps:
[0254] Step 61: Determine the rollback urgency of the target resource combination based on the single-domain readiness, state freshness, current reservation duration, and allowed waiting time of cross-domain reservation objects for each target domain.
[0255] Step 62: If the rollback urgency is greater than the preset urgency threshold, then for each target domain, determine the rollback priority of the target domain based on the undo reservation cost, maintenance resource loss, maintenance resource risk and resource retry difficulty of the target domain.
[0256] Step 63: Based on the rollback priority of each target domain, release the resources corresponding to each target domain in the target resource combination in sequence.
[0257] Specifically, in step 61, the urgency of the target resource portfolio entering the rollback process can be assessed based on the single-domain readiness, state freshness, current reservation duration, and allowed waiting time for cross-domain reservation objects of each target domain, thus obtaining the rollback urgency. The current reservation duration can refer to the time the target domain has maintained reservations for the corresponding resources in the target resource portfolio.
[0258] For example, the rollback urgency can be calculated using the following formula:
[0259] ;
[0260] In the formula, This indicates the urgency of the rollback. To preset the second judgment threshold, This is the value for global consistency determination. For minimum readiness, These represent the state freshness of the intelligent computing domain, the supercomputing domain, and the quantum domain, respectively. The current reserved duration, Allowed waiting time (which can be understood as the validity period of a cross-domain reserved object). ~ These are the weighting coefficients. .
[0261] After calculating the rollback urgency, in step 62, it can be determined whether the rollback urgency is greater than the preset urgency threshold. If so, it means that the benefits of continuing to wait are insufficient or the mismatch risk is too high, and the unified failure rollback process can be entered immediately. If not, the waiting can continue or a partial retry can be performed.
[0262] Specifically, if the rollback urgency exceeds a preset urgency threshold, the rollback priority for each target domain is evaluated based on the undo reservation cost, maintenance resource loss, maintenance resource risk, and resource retry difficulty, as shown in the following formula:
[0263] ;
[0264] In the formula, For the target domain Rollback priority, For the target domain The cost of canceling the reserved space, For the target domain The loss of maintenance resources, that is, the continued maintenance of the target domain. The degree of opportunity loss or waste of scarce resources resulting from resource occupation. For the target domain The risk of maintaining resources, i.e., not prioritizing the release of the target domain. The risk of inconsistency in the state caused by resources, For the target domain The difficulty of resource retry, i.e., the target domain The difficulty of re-acquiring resources in subsequent retries. ~ , where represents the weighting coefficients. It should be noted that the cancellation reservation cost, resource maintenance loss, resource maintenance risk, and resource retry difficulty used in the above formulas can all be normalized results.
[0265] After calculating the rollback priority of each target domain, in step 63, the resources corresponding to each target domain in the target resource combination can be released sequentially according to their rollback priorities. Specifically, the target domains can be sorted in descending order of rollback priority, and the resources corresponding to each target domain in the target resource combination can be released sequentially according to the sorting result.
[0266] Steps 61-63 above determine the rollback urgency and, when the rollback urgency is high, the rollback priority of each target domain is determined. Then, resources of each target domain are released sequentially according to the rollback priority. This allows the quantum domain to have a higher rollback priority when the quantum execution window is scarce and the waiting window loss is high. When the risk of HPC queue occupation or GPU resource fragmentation is high, the HPC supercomputing domain or Kubernetes intelligent computing domain can be released first, achieving the goal of prioritizing the release of scarce or time-sensitive resources and ensuring resource utilization.
[0267] For example, a unified failure rollback process may include the following steps: 1. Freeze subsequent binding and submission actions of cross-domain collaboration objects, prohibiting new downstream submission requests from continuing to execute; 2. Determine the rollback domain set and rollback order based on the current state of the cross-domain reserved object, the backend return state, and the rollback priority; 3. Perform session cancellation, candidate window release, queuing context cancellation, or request invalidation processing on the quantum domain; 4. Perform node / queue reservation release, placeholder cancellation, or reservation token invalidation processing on the HPC supercomputing domain; 5. Perform GPU, node, affinity, or gang resource reservation release on the Kubernetes intelligent computing domain; 6. Clean up associated identifiers such as Reservation Token, Request ID, Queue Ticket, and BackendCorrelation ID in the cross-domain reserved object; 7. Perform state resynchronization, updating the state of the unified task object and the cross-domain reserved object to RollingBack, RolledBack, or Failed; 8. When any target domain release action fails or the response is inconsistent, enter the compensatory rollback and idempotent retry process.
[0268] In this embodiment, when a target domain rollback is successful while another target domain rollback fails or times out, the system can perform compensatory rollback and idempotent retries based on associated identifiers such as Reservation Token, Request ID, and Backend Correlation ID to avoid duplicate releases, state corruption, or resource hanging. For example, delayed confirmation, status re-query, duplicate release requests, or reverse compensation cleanup can be performed on target domains that have not yet been successfully released. After exceeding a preset number of retries, the mixed job is marked as Failed, while the fault context of the incomplete rollback is retained for manual intervention or background consistency repair.
[0269] Furthermore, in this embodiment of the application, the completion rate of the rollback can also be used to determine whether the unified failure rollback process is complete.
[0270] In some implementations, after releasing the resources corresponding to each target domain in the target resource combination in sequence, the following steps are also included:
[0271] Step 64: For each target domain, determine the proportion of resources that have been successfully released in the target domain, the proportion of relevant information in the target resource combination that has been cleared, and the proportion of the target domain's state synchronization.
[0272] Step 65: Determine the rollback completion rate of the target domain based on the resource ratio, the cleanup ratio, and the state synchronization ratio, and determine the global rollback completion rate based on the rollback completion rate of each target domain.
[0273] Step 66: If the global rollback completion rate is greater than the preset completion rate threshold, then the rollback of the target resource combination is determined to be complete, and the status of the cross-domain reserved object is updated to resource release complete.
[0274] Specifically, in step 64, for each target domain, the proportion of resources successfully released by the target domain, the proportion of relevant information cleared in the target resource combination, and the state synchronization proportion of the target domain can be determined. The proportion of successfully released resources can refer to the proportion of reserved resources successfully released by the target domain; the proportion of cleared relevant information can refer to the proportion of target domain association identifiers, session tickets, or request tokens cleared; and the state synchronization proportion can refer to the proportion of the target domain's local state that is synchronized with the unified state machine state.
[0275] Specifically, in step 65, the rollback completion rate of the target domain can be determined based on the resource ratio, the cleanup ratio, and the state synchronization ratio, as shown in the following formula:
[0276] ;
[0277] In the formula, For the target domain The rollback completion rate, , , These are the resource ratio, the cleaned-up ratio, and the status synchronization ratio, respectively. These are the weighting coefficients.
[0278] Furthermore, the global rollback completion rate can be determined based on the rollback completion rates of all target domains. Taking a target domain that includes the intelligent computing domain, supercomputing domain, and quantum domain as an example, as shown in the following formula:
[0279] ;
[0280] In the formula, To determine the completeness of the global rollback, , , These represent the rollback completion rates for the intelligent computing domain, supercomputing domain, and quantum domain, respectively. , , The weights for the intelligent computing domain, the supercomputing domain, and the quantum domain are respectively, and satisfy the following conditions: .
[0281] After calculating the global rollback completion rate, further in step 66, if the global rollback completion rate is greater than the preset completion rate threshold, it indicates that the unified failure rollback has been basically completed, and the status of the cross-domain reserved object can be advanced to RolledBack, that is, the status of the cross-domain reserved object is updated to resource release completed. If the global rollback completion rate is less than the preset completion rate threshold, the compensatory rollback, idempotent retry, or consistency repair process can continue to be executed.
[0282] Through the above steps, a unified failure rollback mechanism with trigger judgment, sequence control, compensation and repair and completion verification is formed. This can significantly reduce the cross-domain mismatch problems caused by resource hanging, zombie reservation, inconsistent state ledger, duplicate release and partial rollback success in cross-scheduling domain scenarios, and improve the stability, recoverability and maintainability of the unified scheduling mechanism in real engineering environment.
[0283] It is worth mentioning that after binding the target resources together and submitting the mixed job, status write-back, metering records and lifecycle tracking can also be performed. That is, the results of the mixed job submission and the status changes during the execution process are written back to the unified task object and cross-domain collaboration object, and the unified job ledger and audit information are recorded simultaneously to form a unified lifecycle view and auditable status trajectory.
[0284] For example, the write-back content may include: backend job identifier, quantum session identifier, association identifier; state transition time point (such as Reserved, Permitted, Committed, Running, Completed, etc.), waiting time, execution time, resource usage statistics; failure reason code, rollback count, retries count.
[0285] Furthermore, when a task in a hybrid operation completes, is canceled, or fails, unified resource reclamation and final state processing can be performed. Based on a strategy, decisions can be made regarding retrying, rescheduling, or degradation to complete the closed-loop processing of the task's entire lifecycle. For example, this process may include: releasing resources and reserved objects occupied during the runtime phase; updating the status of unified task objects (Succeeded, Failed, Cancelled, etc.); writing to the ledger and audit records; and performing retrying, alternative device switching, or degradation execution based on the failure strategy. It should be noted that the above status write-back and final state processing can be performed continuously in stages.
[0286] In this embodiment, if the hybrid job is a hybrid quantum-classical job, classical resources can be reserved first, and then a consistency permission judgment can be performed on the quantum session / window. If the quantum permission is successful, the job is uniformly released and submitted. If the quantum permission fails or times out, the reserved classical resources are uniformly rolled back, and then retry, requeuing, or downgrading to the quantum simulator is performed according to the strategy. For non-quantum tasks, the quantum permission-related sub-steps can be omitted.
[0287] The cross-scheduling domain transactional consistency joint scheduling mechanism based on the Kubernetes Scheduling Framework provided in this application can be used to uniformly coordinate resource consistency among HPC supercomputing domains, Kubernetes native intelligent computing domains, and quantum computing services. This mechanism achieves atomic joint scheduling of multiple scheduling backend resources by introducing cross-domain resource reservation in the Reserve phase, global consistency permission judgment in the Permit phase, and unified failure rollback in the Unreserve phase throughout the scheduling lifecycle. This allows classical computing resources (including CPU, memory, GPU, HPC nodes, etc.) and quantum execution windows (QPU Session or Queue Slot) to be scheduled and committed as a whole, thereby avoiding resource waste and job inconsistency caused by partial resource success and partial resource failure. Furthermore, this mechanism forms a scheduling process with two-phase consistency characteristics, enabling the Kubernetes scheduler to have cross-system consistency control capabilities.
[0288] Furthermore, the consistent joint scheduling mechanism provided in this application can map quantum domain session states (such as queued, active, reserved, etc.) to scheduling states that are perceptible to the Kubernetes scheduler. This transforms quantum sessions from merely external API call results into participants in the scheduling process. Through this mapping mechanism, the scheduler can determine quantum session availability during the Reserve phase and wait for or verify the quantum execution window during the Permit phase, achieving coordinated control between the quantum execution window and the classical resource startup timing. Essentially, the quantum session participates in the unified scheduling process as a scheduling resource with a time window attribute, thereby transforming quantum computing power from "external platform calls" to "unified scheduling objects."
[0289] Furthermore, this application proposes a unified cross-domain reservation object, which can be used to describe and manage the resource reservation status of different scheduling domains. This object supports HPC node and queue resource reservation, Kubernetes GPU / node resource reservation, and QPU session or execution window reservation, and records the state changes of resources from Pending and Reserved to Committed or Rollback through a unified state machine, thereby realizing unified resource lifecycle management across scheduling systems.
[0290] Furthermore, based on the pluggable architecture of the kube-scheduler Scheduling Framework, unified scheduling logic can be implemented at extension points such as QueueSort, Filter, Score, Reserve, Permit, and Bind, enabling scheduling capabilities to continuously evolve without modifying the Kubernetes core code. Through multi-profile scheduling chain support, it can simultaneously handle various workloads such as AI training, HPC batch jobs, and hybrid quantum-classical tasks.
[0291] The embodiments of this application can also construct a unified backend scheduling adaptation layer, which abstracts different execution systems into a unified interface, including HPC adapters (such as Slurm / Slinky), intelligent batch scheduling adapters (such as Volcano or Kubernetes native workloads), and quantum service adapters (interfacing with cloud quantum task systems). This system enables unified scheduling decisions to be mapped to execution actions of different backends, while maintaining unified lifecycle management and state feedback.
[0292] Furthermore, in this embodiment, considering that in the joint scheduling process across HPC, AI, and quantum domains, each scheduling domain has different state refresh cycles, interface response latency, and failure semantics, this embodiment can further design abnormal scenario handling and consistency recovery strategies to enhance the stability and recoverability of the unified scheduling process in a real-world engineering environment. Abnormal scenarios can include the following types:
[0293] 1) Quantum session unavailable or queue wait timeout
[0294] When classical resources have been reserved, but the scheduler detects that the quantum session is unavailable, the target QPU queue is congested beyond a threshold, or the waiting time exceeds a preset timeout during the Permit phase, the scheduler will perform one of the following actions: trigger the Unreserve phase and release the reserved classical resources; enter the retry process according to the policy and maintain or partially release the reserved resources during the retry; downgrade the task to a quantum simulator or alternative quantum device; or requeue the task and update its priority or candidate resource set.
[0295] 2) Classic resource reservation was successful but subsequent status became invalid.
[0296] When classic resource reservation has been completed in the Reserve phase, but the following situations occur before entering Permit or Bind (e.g., node state change, GPU being preempted by a higher priority task, HPC queue policy change causing reservation failure), the system performs consistency recovery processing, including: marking the cross-domain reservation object state as invalid; aborting Permit; performing unified rollback on other successful target domain reservations; and writing the failure reason to the state machine and audit log for retry or policy adjustment.
[0297] 3) Quantum license successful, but classical resources did not meet the final binding conditions.
[0298] When a quantum session has entered an executable state, but the classical resource no longer meets the binding conditions before final commit (e.g., node unreachable, resource fragmentation causing actual unbinding), the system preferentially performs the following actions: quickly canceling the quantum session reservation or releasing the queue context; rolling back the classical resource reservation; and rescheduling, delaying the commit, or downgrading execution according to the strategy. Through these processes, the invalidation of the quantum resource window is avoided.
[0299] 4) Inconsistent success or response in cross-domain API calls.
[0300] When the backend adapter interacts with different scheduling domains, some interface calls may succeed, some may time out, or the status returns may be inconsistent. This application's embodiment uses intermediate states in cross-domain reserved objects and idempotent identifiers (such as Reservation Token / Request ID / Backend Correlation ID) for association verification to support: idempotent retries, delayed acknowledgments, compensatory rollbacks, and state resynchronization. This reduces state drift issues caused by network jitter, interface latency, or temporary backend unavailability.
[0301] 5) Co-scheduling failure due to unmet job group dependencies.
[0302] For jobs that include cross-domain collaborative objects, DAGs, or iterative loop constraints, when other tasks in the group have not reached the predetermined state (e.g., not completing reservations or obtaining quantum permission), this application embodiment can perform group-level wait control during the Permit phase and uniformly trigger group-level rollback or partial retry after timeout to ensure consistent behavior of tasks in the group.
[0303] In the above-mentioned abnormal scenarios, the embodiments of this application do not only perform local error handling on a single target domain, but also restore the cross-domain resource state in a consistent manner through a unified state machine, a unified rollback entry point and a strategy-based recovery path, thereby significantly improving the scheduling stability, resource utilization and maintainability of hybrid quantum-classical jobs in a real platform environment.
[0304] The quantum computing, supercomputing, and intelligent computing power unified scheduling method provided in this application embodiment responds to a received mixed job request by generating a cross-domain collaborative object corresponding to the mixed job, then reading resource status information from each target domain involved in the mixed job, determining a target resource combination based on the resource status information, and initiating a reservation for the target resource combination. The method determines the single-domain readiness of each target domain through the resource status information of each target domain, and determines the cross-domain startup alignment between each target domain and the rollback cost of canceling the reserved target resource combination. Then, based on the single-domain readiness, cross-domain startup alignment, and rollback cost, it determines whether the target resource combination meets the joint startup conditions, obtaining a consistency permission judgment result. Finally, based on the consistency permission judgment result, the method binds the cross-domain collaborative object to the target resource combination and submits it to each target domain, or maintains the current reservation and enters waiting control, or performs rollback processing on the cross-domain collaborative object, thereby realizing unified scheduling of cross-domain resources in the quantum domain, supercomputing domain, and intelligent computing domain. This method determines cross-domain global consistency by assessing the readiness of each target domain. This ensures that resources are bound when all target domains have stable and sufficient reservations, avoiding binding only when a single target domain is stable. This prevents scenarios where some resources are occupied while others cannot be obtained in time, thus improving the success rate of cross-domain joint scheduling and reducing resource waste caused by repeated queuing, invalid waiting, and repeated rollbacks. Furthermore, this method determines cross-domain global consistency by assessing cross-domain startup alignment, ensuring the consistency of startup time for each target domain. This reduces the problem of long-term idling, waiting, or occupation of intelligent computing nodes, supercomputing nodes, or quantum devices due to inconsistent startup timing. Moreover, this method determines cross-domain global consistency by assessing rollback costs, enabling quantitative decision-making between continuing to wait and immediate rollback. This avoids maintaining invalid reservations for a long time in situations where recovery is impossible or the benefits are low, thereby reducing resource fragmentation and invalid occupation.
[0305] For the same purpose, embodiments of this application also provide a unified scheduling system for cross-domain computing power, which includes a unified scheduling kernel module and a multi-backend adapter module, wherein:
[0306] A unified scheduling kernel module is used to execute the unified scheduling method for cross-domain computing power provided in any embodiment of this application;
[0307] The multi-backend adapter module is used to receive scheduling requests sent by the unified scheduling kernel module, determine the target adapter based on the scheduling request, and submit the corresponding jobs in the mixed jobs to the corresponding target domain through the target adapter.
[0308] Specifically, this application embodiment can construct a unified computing power scheduling and service orchestration system running on Kubernetes (i.e., the Liangchao Intelligent Cross-Domain Unified Computing Power Scheduling System), enabling users (or platform upper-layer APIs) to submit jobs through a single entry point, and for the system to complete parsing, scheduling, execution, status write-back, auditing, and resource reclamation within a unified lifecycle. The jobs submitted by users include one or more of the following types:
[0309] 1. Supercomputing / HPC jobs
[0310] Batch jobs with explicit constraints such as the number of nodes, cores, GPUs, memory, wall clock time, and topology / parallelism can be executed by HPC scheduling systems such as Slurm, or orchestrated and executed in a container environment through Slurm and Kubernetes interoperability components.
[0311] 2. Computing / AI Assignments
[0312] Training, inference, and distributed computing tasks using native Kubernetes workloads or batch job formats can expose hardware resources such as GPUs through device plugins, and can be combined with batch scheduling enhancement systems to improve batch job scheduling capabilities.
[0313] 3. Quantum Homework
[0314] It interfaces with the quantum service side's task models (such as sessions, batch processing, device queues, etc.) and supports hybrid quantum-classical iterative workflows, enabling quantum resource calls to be incorporated into a unified lifecycle and policy control.
[0315] The aforementioned system is not simply positioned to provide a "unified submission portal for three types of jobs," but rather to further achieve collaborative scheduling and platform-level governance closed loop across HPC, AI, and quantum domains through unified resource abstraction, unified scheduling control, and a unified status write-back mechanism.
[0316] It should be noted that the multi-backend adapter module is used to translate the scheduling decisions generated by the unified scheduling kernel into executable commit, cancel, query, and status synchronization actions for specific backend scheduling domains, and to write the backend execution status back to the unified task object, forming a unified lifecycle closed loop. Different backend adapters can follow unified interface semantics, including but not limited to: Submit, Cancel, Status, Logs, Accounting, Reserve, and Release. Reserve / Release can be used in conjunction with the Reserve / Unreserve phases to complete cross-domain consistency control.
[0317] The multi-backend adapter module can include an HPC adapter (such as Slurm / Slinky), an intelligent batch scheduling adapter (Kubernetes native / batch scheduling enhancement system), and a quantum service adapter (QPU Adapter). The HPC adapter interfaces with HPC scheduling systems like Slurm, supporting node / queue constraint mapping, job submission, and status tracking; it can be combined with Slurm and Kubernetes interoperability components (such as slurm-operator / slurm-bridge) to achieve unified orchestration in container environments. The intelligent batch scheduling adapter creates Kubernetes Jobs, Deployments, and other objects, or interfaces with Kubernetes-native batch scheduling enhancement systems (such as systems supporting unified batch job scheduling and gang scheduling) to meet the needs of AI / big data tasks. The quantum service adapter interfaces with the quantum service-side Hybrid Job / Session / Queue model, mapping the quantum service state to a scheduler-aware state and providing a state interface for quantum permission determination during the Permit phase.
[0318] The multi-backend adapter module can write back the job identifier, session identifier, queue identifier, and status information returned by the backend to the unified task object and cross-domain reserved object after submission, and support idempotent retries, compensatory rollbacks, and status resynchronization through associated identifiers (correlation id, reservation token, etc.).
[0319] In some implementations, in addition to a unified scheduling kernel module and multiple backend adapter modules, the system also includes a unified entry point and job parsing module, a unified computing resource catalog module, a policy plugin and rule engine module, a lifecycle state management and ledger module, and an observability and auditing module. These modules can be deployed as one or more controllers, service processes, scheduling plugins, and adaptation components, running in a Kubernetes cluster; alternatively, they can be deployed separately in the control plane and business plane according to platform engineering requirements.
[0320] The unified entry point and job parsing module receives job requests from users, upper-level platform APIs, command-line tools, workflow systems, or automated task systems, and standardizes heterogeneous submission semantics into unified task objects for subsequent unified scheduling processing. For example, input formats can include: API requests / remote call requests, command-line script-based parameter submissions, configuration file-based submissions, task descriptions generated by the workflow engine, and internal service forwarding requests from the upper-level platform. This module performs the following processing on input requests: parameter validation (resource parameters, scheduling parameters, identity information, tenant information, etc.); load type identification (HPC / AI / Quantum / Hybrid); resource request normalization (unified structuring of classic and quantum parameters); scheduling policy parameter completion (priority, quota domain, failure policy, timeout policy, etc.); and generation of unified task objects and writing them to the Kubernetes API.
[0321] The unified computing resource catalog module is used to periodically or event-drivenly synchronize the resource status of different scheduling domains and form a unified schedulable view, which is then used by the unified scheduling kernel to perform cross-domain resource filtering, scoring, reservation, and licensing decisions. The unified resource view includes at least the following information:
[0322] 1. HPC / Slurm Resource View
[0323] Queue / partition information; nodes and node status; available capacity such as CPU / GPU / memory; QoS, account, quota, and priority information; scheduling-related information such as reservation, backfilling, and queue waiting status;
[0324] 2. Kubernetes / Intelligent Computing Resource View
[0325] Scheduling attributes such as Node, Label, Taint, and Affinity; availability of extended resources such as GPUs; namespace quotas and isolation information; running status of Pods / Jobs / batch jobs, etc.
[0326] 3. Quantum Resource View
[0327] QPU / Simulator / Service endpoint identifier; queue status, session availability, window reachability; capability tags (such as device type, gate set limitation, region, etc.); priority, congestion or availability metrics returned by the service side (if available).
[0328] In this embodiment, a unified resource mapping rule can be used to convert the heterogeneous states of each scheduling domain into a data structure that the scheduling kernel can consume, and a resource cache with version number / timestamp can be established to support: reading state snapshots during consistent scheduling; cache invalidation and state refresh; state drift detection; and state write-back and re-verification after scheduling. Through the above-mentioned unified computing resource catalog module, the system can form a platform-level unified resource cognition foundation without requiring the underlying scheduling domains to adopt a unified data format.
[0329] The strategy plugin and rule engine module provides pluggable scheduling and recovery strategies during the execution of the unified scheduling kernel module. This allows the platform to dynamically adjust scheduling behavior according to business objectives without altering the core consistency mechanism. Strategies may include, but are not limited to: cross-domain fairness and priority strategies; quota and tenant isolation strategies; preemption and backfilling strategies; data / storage / network proximity strategies; cost and energy awareness strategies; quantum device selection and queue congestion threshold strategies; Permit wait timeout, retry, and degradation (emulator) strategies; and group-level failure handling and rescheduling strategies. This module provides parameterized or rule-based support for the Reserve, Permit, and Unreserve phases, such as how long to wait in the Permit phase, which scenarios allow partial reservation and retries, which scenarios require immediate rollback, and which scenarios allow degradation to alternative devices or emulators.
[0330] The lifecycle state management and ledger module is used to uniformly manage the entire lifecycle state of a job from submission, scheduling, reservation, licensing, submission, execution, completion / failure to recycling, and records resource usage, state transitions, metering information, and audit information across scheduling domains. The objects of state management include at least a unified task object, a cross-domain collaboration object, a cross-domain reservation object, an optional policy binding object, and backend jobs (mapping the individual states of HPC, AI, and Quantum to a unified state semantic).
[0331] This application embodiment can set up a unified job ledger to record: resource requests and actual allocations; reservation and release times; permit waiting time and reasons for failure; backend metering information (such as job runtime, resource usage statistics, etc.); rollback paths and recovery actions. This ledger provides basic data support for unified billing, unified auditing, problem tracking, and platform operation analysis.
[0332] The observability and auditing module is used to collect metrics, aggregate logs, track events, and output audit results for the unified scheduling and execution process, supporting platform operation and maintenance and policy optimization. Observable content includes, but is not limited to: scheduling time (staged time for Filter / Score / Reserve / Permit / Bind); Permit waiting time and rejection rate; number of rollbacks and distribution of rollback reasons; cross-domain resource utilization (HPC / GPU / QPU); state drift and resynchronization counts; and changes in scheduling success rate and resource utilization under different policy configurations. The observability and auditing module can record key operation logs for unified task objects, cross-domain reserved objects, and backend adapters to meet the needs of platform-level auditing, accountability tracking, and fault review.
[0333] For example, the entire system's operation process may include the following steps:
[0334] 1. Access Request and Object Generation
[0335] The unified entry module receives job requests, parses them, and generates a unified task object (for mixed jobs, a cross-domain collaboration object also needs to be generated).
[0336] 2. Resource view synchronization and candidate construction
[0337] The resource catalog module synchronizes the status of HPC, AI, and quantum resources and generates a unified resource view;
[0338] 3. Unified scheduling decision-making
[0339] The scheduling kernel performs filtering and scoring, and forms candidate and preferred decisions based on the policy plugins;
[0340] 4. Cross-domain consistency control
[0341] The kernel schedules the execution of the Reserve / Permit / Unreserv process to achieve consistent coordination between classical resources and quantum execution windows;
[0342] 5. Backend submission and execution
[0343] Submit job or session requests to HPC, AI, and quantum scheduling domains via multiple backend adapters and track execution status.
[0344] 6. Status write-back and ledger recording
[0345] Write back the backend status, associated identifiers, and metering information to a unified task object and a unified operation ledger;
[0346] 7. Completion / Failure Handling and Resource Recovery
[0347] After a task is completed, canceled, or fails, a unified resource recovery and audit record is executed, and a strategy is used to determine whether to retry or downgrade the execution.
[0348] The system provided in this application has the following technical effects:
[0349] 1. Atomized joint scheduling of resources across scheduling domains avoids resource hanging and partial allocation.
[0350] By treating classic computing resources (CPU, memory, GPU, HPC nodes, etc.) and quantum execution windows / session resources (QPUSession / Queue Slot) as unified scheduling objects, cross-domain consistency control is achieved through the Reserve / Permit / Unreserve mechanism during the Kubernetes scheduling lifecycle: permission judgment only proceeds after successful reservation; if any condition is not met, a unified rollback occurs. This avoids the cross-domain resource hanging problem of "classic resources are occupied but QPU is unavailable" or "QPU is approved but classic resources are insufficient," reduces invalid occupation and duplicate queuing, and improves the overall resource utilization and scheduling stability of the converged platform.
[0351] 2. Establish a unified, observable, and auditable cross-domain state machine closed loop to enhance fault recovery and operational governance capabilities.
[0352] By defining a unified state machine (Pending, Reserving, Reserved, Permitting, Permitted, Committed, RollbackInProgress, RolledBack, etc.) through cross-domain reservation objects, and recording the transition time, triggering reason, backend association identifier, and failure reason code, the cross-domain reservation, permission, submission, and rollback processes become observable and verifiable. Compared to traditional multi-system splicing methods, this significantly reduces operational issues such as state drift, zombie reservations, and inconsistent ledger data, and provides a stable data foundation for unified billing and auditing.
[0353] 3. A truly unified entry point and lifecycle, reducing interface fragmentation and R&D / maintenance costs.
[0354] Users or upper-layer platforms no longer need to differentiate between "submitting a Slurm job / submitting a Kubernetes job / calling a quantum service," but instead submit them uniformly as declarative task objects; the platform side can provide unified submission, query, cancellation, logging, and auditing interfaces. Versioned APIs based on Kubernetes CRDs can form stable backend contracts, reducing interface fragmentation and long-term maintenance costs caused by multiple entry points and protocols.
[0355] 4. Implement continuously evolving scheduling strategies and multi-load parallelism without modifying Kubernetes.
[0356] Based on the extension points and multi-profile mechanism of the kube-scheduler Scheduling Framework, pluggable scheduling links and policy injection are implemented, enabling the platform to continuously evolve its scheduling capabilities without modifying the Kubernetes core code. For example, policies such as group-level coordination, permit waiting, rollback compensation, cost / congestion awareness, and quantum window alignment are gradually introduced according to business stages, and multiple types of workloads such as AI inference, AI training, HPC batch jobs, and hybrid quantum-classical jobs can coexist in parallel.
[0357] 5. Compatible with existing HPC and intelligent computing ecosystems, reducing dual-stack fragmentation and improving resource sharing efficiency.
[0358] By connecting the HPC scheduling domain and the Kubernetes intelligent computing scheduling domain through a multi-backend adapter system, we can reuse both Slurm's mature scheduling capabilities and strict policies, as well as Kubernetes' advantages in container orchestration and service-oriented architecture. On the HPC side, we can combine Slurm and Kubernetes interoperability components to achieve containerized orchestration and resource sharing; on the intelligent computing side, we can reuse Kubernetes' device plugin system and batch scheduling enhancement capabilities, making the resource orchestration of training / inference / distributed jobs more aligned with actual needs, thereby reducing dual-stack fragmentation and improving resource sharing and overall utilization.
[0359] 6. Incorporate quantum session / hybrid semantics into unified scheduling and governance to support the tight coupling trend of quantum-HPC and achieve scalability.
[0360] The temporal resource semantics of the quantum-side "session / execution window" are mapped to a scheduler-aware state, forming a unified and consistent control with the classical resource reservation process. This transforms quantum resources from manual calls outside the platform into a manageable job type within the platform, allowing for waiting / permission / denial and policy-based degradation (such as switching simulators or alternative devices) during the Permit phase. Furthermore, the object model and consistency mechanism provided in this application can be extended to more heterogeneous devices and more scheduling backends (such as FPGAs, dedicated inference chips, network / storage slices, etc.), enabling smooth expansion through a "backend adapter + policy plugin" approach, resulting in lower evolution costs for the future.
[0361] Figure 2 This is a schematic diagram of the overall system architecture and cross-domain consistency scheduling closed loop provided in an embodiment of this application, such as... Figure 2As shown, requests sent by user / platform API clients are parsed into CDR objects through a unified entry point and job and stored in the Kubernetes API Server. The controller and resource catalog synchronize resource status information of the supercomputing domain, intelligent computing domain, and quantum domain and trigger scheduling; the unified scheduling kernel performs filtering and scoring based on the Scheduling Framework, and executes Reserve / Permit / Unreserve and Bind processes, and completes status updates / reservation tokens, status verification, and failure rollback with the support of the CrossDomainReservation state machine; scheduling decisions are distributed to the supercomputing domain, intelligent computing domain, and quantum domain through policy plugins / rule engines and multiple backend adapter layers; execution status and metering information are fed back to the monitoring / audit / metering ledger and CDR status, forming an end-to-end closed loop.
[0362] Figure 3 This is a schematic diagram of CDR object relationships provided in an embodiment of this application. For example... Figure 3 As shown, the UnifiedComputeTask object matches the UnifiedComputeResource object with selectors and constraints to determine the range of schedulable resources. When a task belongs to a collaborative scheduling scenario, it becomes a member of the CoScheduleGroup cross-domain collaborative object through groupRef. The UnifiedComputeTask object associates with the CrossDomainReservation object through reservationRef and receives its state write-back. The CrossDomainReservation object records classical resource reservation information, quantum session / window information, phase state machine, and associated tokens to support the state loop of cross-domain consistent reservation, permission, and rollback.
[0363] Figure 4 This is a flowchart of a unified scheduling method provided in an embodiment of this application, such as... Figure 4 As shown, firstly, a job request is received and a unified task object is generated and written to the Kubernetes CRD; then, the resource view is refreshed, the scheduling cycle is entered, and the filtering and scoring process is executed; subsequently, it is determined whether cross-domain consistency control is enabled (for hybrid jobs). If cross-domain consistency control is not enabled, the multi-backend adapter is called to submit the job / session request and bind the backend identifier (backendRefs), and then unified state write-back, metering recording, and resource reclamation are performed.
[0364] If cross-domain consistency control is enabled, the process enters the Reserve phase, reserving resources and generating a cross-domain reserved state. It then enters the Permit phase, waiting for the consistency permission judgment result. When the consistency permission judgment result is "allow," the multi-backend adapter is invoked to submit the job, ultimately completing unified state write-back, metering records, and resource reclamation, thus forming a consistent scheduling closed loop across heterogeneous scheduling domains. When the consistency permission judgment result is "waiting" or "rollback," the process enters the exception handling procedure, continuing to wait for judgment, or executing a unified rollback to release reserved resources, and then re-enqueuing, downgrading, or terminating the process according to the policy.
[0365] Figure 5 This application provides a timing diagram for a hybrid quantum-classical cooperative reservation and session licensing mechanism, as shown in the embodiments below. Figure 5 As shown, users create unified task objects and cross-domain collaborative objects through the platform interface (API). After the unified scheduling kernel pulls the job from the API, it enters the consistency scheduling process. In the Reserve phase, it reserves computing resources, and the API side updates the state (phase=Reserved) and associated tokens of the cross-domain reserved object. Subsequently, the unified scheduling kernel requests a quantum session / execution window from the quantum domain (QPU). In the Permit phase, it performs a permission judgment on the job based on the consistency permission judgment result returned by the QPU, and updates the state of the cross-domain reserved object to Permitting / Permitted. When the permission is granted, the unified scheduling kernel calls multiple backend adapters in the Bind phase to submit classical job and quantum session requests and bind backend identifiers (backendRefs). The multiple backend adapters trigger the start of classical tasks and submit quantum tasks respectively. During and after the task execution, the execution status and metering information are sent back to the API side for unified status write-back, auditing, and resource reclamation, thus forming a closed-loop time sequence of collaborative reservation, session permission, and submission execution of hybrid quantum-classical jobs.
[0366] Figure 6 This is a placeholder operation execution path diagram for an HPC adapter interfacing with Slurm, provided in an embodiment of this application. Figure 6As shown, the HPC adapter / bridge scheduler is used to achieve interoperability between Kubernetes workloads and the Slurm scheduling domain. After a Kubernetes workload enters the adapter for queuing and management, the adapter creates or queries a placeholder job to represent the corresponding Kubernetes workload in the Slurm queue, and submits or queries the job status and resource allocation results by calling the Slurm REST API. After obtaining the resource allocation results generated by Slurm scheduling, the adapter performs a binding operation, binding the Pod to be run to the target node. Subsequently, the kubelet on the target node starts the container to run, thus forming an interoperability execution path of "placeholder queuing - resource allocation - node binding - Pod startup". Among them, the placeholder job is used to optimize scheduling speed, trigger cluster scaling, buffer traffic peaks, and use low-priority Pods to occupy resource slots, so that resources can be instantly released when the real business Pod arrives, achieving second-level startup.
[0367] Figure 7 This is a schematic diagram of a quantum-classical hybrid iterative workflow provided in an embodiment of this application, such as... Figure 7 As shown, in a hybrid quantum-classical algorithm scenario, the classical computing module executes classical computational processes such as optimizer iteration, training, or simulation, and constructs quantum circuit parameters or task descriptions in each iteration. The classical computing module submits quantum tasks to the quantum execution module (QPU or quantum simulator) through a session / execution window or hybrid job mechanism. The quantum execution module returns feedback information such as measurement results, expected values, or gradients, which the classical computing module uses to update parameters and proceed to the next iteration, thus forming an iterative loop / multiple-call workflow of "classical update—quantum execution—result feedback." The session / window or hybrid job mechanism carries the continuity and scheduling semantics of multiple rounds of quantum calls, enabling quantum execution resources to be collaboratively managed and controlled with classical computing resources within the same lifecycle.
[0368] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 400 includes one or more processors 401 and memory 402.
[0369] The processor 401 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 400 to perform desired functions.
[0370] The memory 402 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 401 may execute the program instructions to implement the unified scheduling method for cross-domain computing power of any embodiment of this application described above, and / or other desired functions. Various contents such as initial extrinsic parameters and thresholds may also be stored in the computer-readable storage medium.
[0371] In one example, the electronic device 400 may further include an input device 403 and an output device 404, these components being interconnected via a bus system and / or other forms of connection mechanisms (not shown). The input device 403 may include, for example, a keyboard, a mouse, etc. The output device 404 may output various information to the outside, including warning messages, braking force, etc. The output device 404 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0372] Of course, for the sake of simplicity, Figure 8 Only some of the components of the electronic device 400 relevant to this application are shown in this illustration; components such as buses, input / output interfaces, etc., are omitted. In addition, the electronic device 400 may include any other suitable components depending on the specific application.
[0373] In addition to the methods and devices described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the unified scheduling method for cross-domain computing power provided in any embodiment of this application.
[0374] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0375] Furthermore, embodiments of this application may also be computer-readable storage media storing computer program instructions, which, when executed by a processor, cause the processor to perform the steps of the unified scheduling method for cross-domain computing power provided in any embodiment of this application.
[0376] The computer-readable storage medium may be any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0377] It should be noted that the terminology used in this application is for the purpose of describing specific embodiments only and is not intended to limit the scope of this application. As shown in the specification and claims of this application, unless the context clearly indicates otherwise, words such as "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, or apparatus. Without further limitations, an element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes said element.
[0378] It should also be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Unless otherwise expressly specified and limited, the terms "installed," "connected," "linked," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication between two elements. For those skilled in the art, the specific meaning of the above terms in this application can be understood according to the specific circumstances.
[0379] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. The above descriptions are only preferred embodiments of this application. It should be noted that due to the limitations of written expression, while there are objectively infinite specific structures, those skilled in the art can make several improvements, modifications, or changes without departing from the principles of this application, and can also combine the above technical features in an appropriate manner. These improvements, modifications, changes, or combinations, or the direct application of the inventive concept and technical solution to other situations without modification, should all be considered within the scope of protection of this application.
Claims
1. A unified scheduling method for cross-domain computing power in a high-performance computing environment, characterized in that, include: In response to receiving a hybrid job request, a cross-domain collaborative object corresponding to the hybrid job is generated, wherein the hybrid job includes computing tasks in at least two target domains, and the target domains are quantum domain, supercomputing domain, or intelligent computing domain; Read resource status information from each of the target domains, determine the target resource combination based on the resource status information, and initiate the reservation of the target resource combination; The single-domain readiness of each target domain is determined based on the resource status information of each target domain, and the cross-domain startup alignment between each target domain and the rollback cost of canceling the reserved target resource combination are determined; the cross-domain startup alignment is used to measure the time consistency of each target domain in reaching the joint startup conditions. Based on the single-domain readiness, the cross-domain startup alignment, and the rollback cost, determine whether the target resource combination meets the joint startup conditions and obtain the consistency permission judgment result. Based on the consistency permission judgment result, the cross-domain collaboration object is bound to the target resource combination and submitted to each of the target domains; or, the current reservation is maintained and the waiting control is entered; or, the cross-domain collaboration object is rolled back. Determining the single-domain readiness of each target domain based on the resource status information of each target domain includes: For each target domain, the effective allocable quantity under each resource dimension is determined based on the original allocable quantity, resource fragmentation correction coefficient, topology matching correction coefficient, and duration matching correction coefficient under each resource dimension of the target domain. Based on the effective allocable quantity and request quantity under each resource dimension, determine the single-dimensional resource satisfaction rate corresponding to each resource dimension, and determine the capacity satisfaction of the target domain based on the single-dimensional resource satisfaction rate. Determine whether the target resource combination satisfies the hard constraints of the target domain to obtain the hard constraint satisfaction factor, and determine the soft constraint score based on the degree to which the target resource combination satisfies the soft constraints of the target domain. Based on the hard constraint satisfaction factor and the soft constraint score, determine the policy satisfaction degree of the target domain. Based on the target domain's state change rate, queue fluctuation degree, historical synchronization error, and interface reliability, the freshness decay parameter of the target domain is determined, and the state freshness of the target domain is determined according to the resource state synchronization time of the target domain and the freshness decay parameter. Determine the predicted preemption probability, predicted expiration probability, predicted conflict probability, and state fluctuation risk value of the target resource combination, and determine the reservation stability of the target domain based on the predicted preemption probability, the predicted expiration probability, the predicted conflict probability, and the state fluctuation risk value. The single-domain readiness of the target domain is determined based on the capacity satisfaction, the policy satisfaction, the state freshness, and the reservation stability.
2. The method for unified scheduling of cross-domain computing power according to claim 1, characterized in that, Determining the cross-domain launch alignment between the target domains includes: For each target domain, determine the waiting time for resources in the target domain to meet the conditions for joint submission, the first time required for resource reservation or session status confirmation, the second time for resource status to remain stable or conflict to be resolved, and the handover processing time before resources enter the binding process. Based on the waiting time, the first time, the second time, and the handover processing time, determine the estimated time for the target domain to reach a submittable state; The submission time variance between each target domain is determined based on the estimated time corresponding to each target domain, and the cross-domain launch alignment between each target domain is determined based on the submission time variance.
3. The method for unified scheduling of cross-domain computing power according to claim 1, characterized in that, Determine the rollback cost of canceling the reservation of the target resource combination, including: For each target domain, determine the action time required to cancel the reserved corresponding resources in the target resource combination, the waiting loss for re-enqueuing after cancellation, the resource fragmentation loss caused by cancellation, the opportunity loss caused by lost resources, and the restoration cost after cancellation. Based on the action time, the waiting loss, the resource fragmentation loss, the opportunity loss, and the restoration cost, the cancellation reservation cost corresponding to the target domain is determined; Based on the cancellation reservation cost corresponding to each of the target domains, determine the rollback cost for canceling the reservation of the target resource combination.
4. The method for unified scheduling of cross-domain computing power according to claim 1, characterized in that, Based on the single-domain readiness, the cross-domain launch alignment, and the rollback cost, it is determined whether the target resource combination meets the joint launch conditions, and a consistency permission determination result is obtained, including: The minimum readiness is determined based on the minimum value among the readiness of each single domain, and a global consistency judgment value among the target domains is determined based on the minimum readiness, the cross-domain startup alignment, and the rollback cost. If the minimum readiness is not less than a preset readiness threshold, and the global consistency judgment value is not less than a preset first judgment threshold, then the target resource combination is determined to meet the joint startup conditions, and the consistency permission judgment result is to allow. If the minimum readiness is less than the preset readiness threshold, then the target resource combination is determined not to meet the joint startup conditions, and the consistency permission judgment result is rollback; If the minimum readiness is not less than a preset readiness threshold, and the global consistency judgment value is less than the preset first judgment threshold and not less than the preset second judgment threshold, then it is determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is waiting. If the minimum readiness is not less than a preset readiness threshold, and the global consistency judgment value is less than the preset second judgment threshold, then it is determined that the target resource combination does not meet the joint startup conditions, and the consistency permission judgment result is rollback. Wherein, the preset first determination threshold is greater than the preset second determination threshold.
5. The method for unified scheduling of cross-domain computing power according to claim 1, characterized in that, Maintain the current reservation and enter the waiting control phase, including: Maintain the reservation of the target resource combination, and return to redetermine the single-domain readiness of each target domain, and determine the cross-domain launch alignment between each target domain and the rollback cost of canceling the reservation of the target resource combination, so as to determine whether the target resource combination meets the joint launch conditions based on the new single-domain readiness, cross-domain launch alignment and rollback cost, and obtain the consistency permission judgment result; If, during the reservation process, the waiting time exceeds a preset duration, or the freshness of the state of any of the target domains is lower than a preset freshness threshold, or the global consistency judgment value among the target domains is lower than the minimum judgment value, then the cross-domain collaborative object will be rolled back.
6. The method for unified scheduling of cross-domain computing power according to claim 1, characterized in that, After generating the cross-domain collaboration object corresponding to the hybrid job, the following is also included: Generate the cross-domain reserved object corresponding to the cross-domain collaboration object; The states of the cross-domain reserved objects include: resource waiting for reservation, resource reservation in progress, partial reservation successful, reservation completed, waiting for joint condition verification, license obtained, binding submission completed, waiting for rollback, resource release completed, retry refused, and job completed.
7. The method for unified scheduling of cross-domain computing power in super-intelligent computing as described in claim 6, characterized in that, Rollback processing of the cross-domain collaboration object includes: The rollback urgency of the target resource combination is determined based on the single-domain readiness, state freshness, current reservation duration, and allowed waiting time of the cross-domain reservation object for each target domain. If the rollback urgency is greater than a preset urgency threshold, then for each target domain, the rollback priority of the target domain is determined based on the undo reservation cost, maintenance resource loss, maintenance resource risk, and resource retry difficulty of the target domain. Based on the rollback priority of each target domain, the resources corresponding to each target domain in the target resource combination are released sequentially.
8. The method for unified scheduling of cross-domain computing power according to claim 7, characterized in that, After releasing the resources corresponding to each target domain in the target resource combination in sequence, the method further includes: For each target domain, determine the proportion of resources that have been successfully released in the target domain, the proportion of relevant information that has been cleared in the target resource combination, and the proportion of the target domain's state synchronization. Based on the resource ratio, the cleanup ratio, and the state synchronization ratio, the rollback completion rate of the target domain is determined, and the global rollback completion rate is determined based on the rollback completion rate of each target domain. If the global rollback completion rate is greater than the preset completion rate threshold, then the rollback of the target resource combination is determined to be complete, and the status of the cross-domain reserved object is updated to resource release complete.
9. A unified scheduling system for cross-domain computing power of super-intelligent computing, characterized in that, The system includes a unified scheduling kernel module and a multi-backend adapter module; The unified scheduling kernel module is used to execute the unified scheduling method for cross-domain computing power as described in any one of claims 1-8; The multi-backend adapter module is used to receive the scheduling request sent by the unified scheduling kernel module, determine the target adapter according to the scheduling request, and submit the corresponding job in the mixed job to the corresponding target domain through the target adapter.
Citation Information
Patent Citations
Classified scheduling method and system for computing power resources, electronic equipment and storage medium
CN115794373A
Heterogeneous computing network resource collaborative scheduling optimization method based on adaptive multi-agent
CN121547418A