Computation power position-based cross-domain data and task scheduling method, device and system

CN122526762BActive Publication Date: 2026-09-18BEIJING HUAHENG SHENGSHI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611034713.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-18
Estimated Expiration
2046-07-13

AI Technical Summary

Technical Problem

第一,以中心调度-区域执行为核心的集中式算力编排架构,其调度决策多围绕算力与网络可达性,对数据实际位置/副本分布及其迁移代价缺乏联合建模,导致跨中心任务易受数据搬迁时延制约

Benefits of technology

第一,实现了算力与数据的双向协同最优决策。通过引入综合代价评分模型,将任务的排队等待时间与跨域数据迁移时间统一映射为时间成本进行联合优化,打破了传统调度中算力与数据割裂的局面。系统能够在“在数据所在地排队等待”与“调度数据至空闲算力”之间进行量化权衡,自动选择总完成时间最短的执行路径,从而显著提升跨域任务调度效率与全局算力利用率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526762B_ABST
    Figure CN122526762B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cross-domain data and task scheduling method, device and system based on computing power position, it is related to computer and information network technical field.The method includes: by generating and maintaining the metadata record containing data identification and physical distribution information;When scheduling, system analyzes computing task request, inquires metadata to obtain data distribution list, and based on the comprehensive cost score model containing predicted queuing waiting time and required data scheduling quantity, the optimal target execution cluster is determined;If target cluster has no target data, automatically trigger cross-domain data scheduling, update metadata and issue task after data is ready.The application solves the problem of splitting computing power and data resources, significantly improves global resource utilization, reduces total task completion time and cross-domain bandwidth cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer and information network technology, and in particular to a method, apparatus and system for cross-domain data and task scheduling based on computing power location. Background Technology

[0002] With the deep integration of cloud computing and edge computing, enterprise computing tasks are increasingly characterized by cross-regional and cross-cluster deployments. In such multi-regional, multi-cluster, and heterogeneous chip computing environments, how to efficiently coordinate computing and data resources to achieve efficient task execution has become a key challenge for the industry. Existing technologies have proposed various solutions around resource-aware scheduling and unified orchestration of computing resources, but they generally suffer from core shortcomings such as the separation of computing and data resources, a single optimization objective, and a lack of closed-loop governance of the data lifecycle.

[0003] Specifically, existing technical solutions mainly include the following directions: First, the centralized computing orchestration architecture with central scheduling and regional execution as its core makes scheduling decisions mostly around computing power and network reachability. It lacks joint modeling of the actual location / replica distribution of data and its migration costs, which makes cross-center tasks susceptible to data migration delays.

[0004] Second, it focuses on the task scheduling of computing resources within a single data center, and strives for fine-grained allocation of computing load and time limits, but does not incorporate data location and cross-domain transmission latency, cost, and energy consumption into a unified cost model, making it difficult to support multi-regional collaboration.

[0005] Third, while security / cost-aware scheduling for cloud environments takes into account the trade-off between security and cost, it lacks collaborative strategies for computing power geographic location and cross-domain topology, and its support for the linkage decision-making between data mobility and computing power location selection is insufficient.

[0006] Fourth, task scheduling within a single cluster based on intelligent optimization algorithms is limited to a single data center and has weak collaborative modeling capabilities for cross-domain networks, data locations, and computing power locations.

[0007] The shortcomings of the aforementioned existing technical solutions can be summarized as follows: First, computing power and data are decoupled and not uniformly modeled in scheduling decisions, which makes the one-way selection from computing power to data or from data to computing power unable to adapt, often leading to inefficient data migration.

[0008] Second, the optimization goals and constraints are too singular, with indicators such as security, cost, bandwidth, and latency acting independently, lacking a unified cost function to accommodate multiple constraints such as energy consumption, cross-domain fees, compliance, and data copy lifecycle.

[0009] Third, the lack of a refined management mechanism for the entire lifecycle of data copies can easily lead to storage redundancy and bandwidth waste.

[0010] Fourth, there is insufficient awareness of the geographical location of cross-domain network topology and computing power / data. During scheduling, the geographical location, link service quality, cross-domain costs, etc. are not jointly evaluated, which limits the effectiveness of multi-domain collaboration.

[0011] The aforementioned defects in existing technologies collectively lead to technical problems such as low efficiency in cross-domain task scheduling, high data migration costs, and uneven global resource utilization. Summary of the Invention

[0012] In view of the aforementioned defects or deficiencies in the prior art, this invention provides a method, apparatus, and system for cross-domain data and task scheduling based on computing power location. This invention aims to improve resource utilization and reduce total task completion time and operating costs by constructing a unified decision-making framework that couples computing power location, data distribution, and cross-domain network, and integrating a closed-loop governance mechanism for the entire data lifecycle.

[0013] One aspect of the present invention provides a cross-domain data and task scheduling method based on computing power location, comprising the following steps: Receive and store business data uploaded by users, and generate metadata records containing globally unique data identifiers and physical distribution information; Parse the computing task request submitted by the user to obtain the target data identifier and the required computing resources that the computing task request depends on; Query the metadata record based on the target data identifier to obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; Based on the computing power resources required by the computing task request, at least one edge computing cluster is screened to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined based on the comprehensive cost score. If the target execution cluster does not store the business data corresponding to the target data identifier locally, a data scheduling instruction is generated according to the data distribution list to schedule the business data corresponding to the target data identifier from the data source location to the target execution cluster, and the data scheduling instruction is sent to the target execution cluster. In response to the confirmation message sent by the target execution cluster that the cross-domain data scheduling is complete, the physical distribution information of the business data corresponding to the target data identifier is updated in the metadata record, and a computing task is sent to the target execution cluster. Receive and archive the execution results of the computing tasks returned by the target execution cluster.

[0014] A second aspect of the present invention also provides a cross-domain data and task scheduling device based on computing power location, comprising: The data registration module is used to receive and store business data uploaded by users, and generate metadata records containing globally unique data identifiers and physical distribution information; The task parsing module is used to parse the computing task request submitted by the user to obtain the target data identifier and the required computing resources that the computing task request depends on. The data exploration module is used to query the metadata records based on the target data identifier, obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; The cluster screening and decision-making module is used to screen at least one edge computing cluster based on the computing power resources required by the computing task request to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined based on the comprehensive cost score. The scheduling instruction generation module is used to generate a data scheduling instruction based on the data distribution list if the target execution cluster does not store the business data corresponding to the target data identifier locally, and then send the data scheduling instruction to the target execution cluster. The task distribution module is used to update the physical distribution information of the business data corresponding to the target data identifier in the metadata record in response to the confirmation information sent by the target execution cluster that the cross-domain data scheduling is completed, and to distribute the computing task to the target execution cluster. The results archiving module is used to receive and archive the execution results of the computing tasks returned by the target execution cluster.

[0015] A third aspect of the present invention also provides a cross-domain data and task scheduling system based on computing power location, comprising a central management cluster and at least one edge computing cluster, wherein: The central management cluster is used to receive and store user-uploaded business data, generate metadata records containing globally unique data identifiers and physical distribution information; parse user-submitted computing task requests to obtain the target data identifier and required computing resources that the computing task request depends on; query the metadata records based on the target data identifier to obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; filter at least one edge computing cluster based on the computing resources required by the computing task request to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, calculate the overall computing power based on real-time computing power status information and the data distribution list. A comprehensive cost score is calculated, and a target execution cluster is determined based on the comprehensive cost score. If the target execution cluster does not store the business data corresponding to the target data identifier locally, a data scheduling instruction is generated based on the data distribution list to schedule the business data corresponding to the target data identifier from the data source location to the target execution cluster, and the data scheduling instruction is sent to the target execution cluster. In response to the confirmation information sent by the target execution cluster that the cross-domain data scheduling is completed, the physical distribution information of the business data corresponding to the target data identifier is updated in the metadata record, and a computing task is issued to the target execution cluster. The execution results of the computing task returned by the target execution cluster are received and archived. The edge computing cluster is used to schedule business data corresponding to the target data identifier across domains from the data source location according to the data scheduling instructions sent by the central management cluster, and send a confirmation message of cross-domain data scheduling completion to the central management cluster; monitor the local storage space utilization rate, and trigger a cache cleanup process when the local storage space utilization rate exceeds a preset security threshold; and calculate the elimination score of each piece of business data in the local storage of the edge computing cluster that has not been read or used by any computing task according to the following formula. : in, , , As a weighting factor; This represents the access popularity factor, which is the reciprocal of the difference between the current time and the last access time of the data, or a monotonically decreasing function. This represents the replacement cost factor, which is directly proportional to the size of the business data and inversely proportional to the current download bandwidth; The hit frequency factor represents the number of times the business data is accessed within a preset time window; the business data with the lowest elimination score is deleted first until the local storage space utilization rate falls below the preset target threshold.

[0016] The cross-domain data and task scheduling method, apparatus, and system based on computing power location provided by this invention have the following beneficial effects: First, it achieves optimal decision-making through two-way collaboration between computing power and data. By introducing a comprehensive cost scoring model, the queuing time of tasks and the cross-domain data migration time are uniformly mapped to time costs for joint optimization, breaking the traditional separation between computing power and data in scheduling. The system can quantitatively weigh the trade-offs between "queuing at the data's location" and "scheduling data to idle computing power," automatically selecting the execution path with the shortest total completion time, thereby significantly improving the efficiency of cross-domain task scheduling and the utilization rate of global computing power.

[0017] Secondly, it automates and intelligentizes cross-domain data flow. By triggering data migration through computing power scheduling, the data preparation process is automated. The system can automatically schedule data from the central or other edge clusters to the target computing power location based on decisions, eliminating task blockages caused by data silos and significantly improving the utilization rate and task startup speed of heterogeneous computing power clusters.

[0018] Third, it achieves closed-loop governance of data replicas throughout their entire lifecycle. By maintaining metadata records, reference counting, and a value-density-based eviction strategy, a governance system covering data registration, distribution, use, and recycling is constructed. This invention can automatically clean up expired or low-value replicas, reducing storage redundancy and operational costs, while ensuring the traceability of data flow and rollback capability in abnormal situations, thus solving the problems of chaotic replica management and resource waste.

[0019] Fourth, it improves the robustness and cost-effectiveness of cross-domain scheduling. Through dynamic adaptive data transmission strategies and intelligent edge cache eviction policies, the system can effectively cope with WAN instability, optimize transmission efficiency, and improve edge cache hit rate. This together reduces cross-domain bandwidth consumption and task execution latency, enhancing the system's performance and cost-effectiveness in complex network environments. Attached Figure Description

[0020] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the structure of a cross-domain data and task scheduling system based on computing power location provided in one embodiment of this application; Figure 2 This is a flowchart illustrating a cross-domain data and task scheduling method based on computing power location provided in one embodiment of this application; Figure 3 This is a schematic diagram of the structure of a cross-domain data and task scheduling device based on computing power location provided in another embodiment of this application. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] To avoid ambiguity, the key technical terms used in this embodiment are defined as follows: Metadata records refer to a collection of structured descriptive information created for each piece of user business data. This collection is indexed by a globally unique data identifier and includes at least the data size, hash checksum, and the physical location information (i.e., physical distribution information) of one or more replicas of the data across the entire network. Metadata records are the core basis for the system to perform data exploration, scheduling decisions, and state maintenance.

[0023] Data distribution list: refers to a summary list of information such as the physical location and status of all existing copies of a data obtained after querying its metadata record based on the target data identifier at a certain scheduling decision moment.

[0024] Overall Cost Score: A metric used to quantitatively evaluate the estimated total time cost of scheduling a specific computing task to a candidate edge computing cluster for execution.

[0025] Cross-domain data scheduling: refers to the process of transferring business data from a network storage location in one domain to another domain via a wide area network in a computing environment that contains multiple independent geographical or logical regions, based on scheduling decisions.

[0026] Edge computing clusters refer to entities located in geographical locations or network domains different from the central management cluster, possessing independent pools of computing, storage, and network resources, capable of performing computing tasks, and interconnected with the central management cluster via a wide area network.

[0027] Central Management Cluster: Refers to the entity that serves as the global control center of the system, responsible for receiving tasks and data, maintaining global metadata records, executing scheduling decisions, and typically serving as the core data persistent storage and archiving node.

[0028] like Figure 1 As shown, the cross-domain data and task scheduling system based on computing power location in this embodiment of the invention adopts a two-tier distributed architecture of center-edge. The central management cluster and multiple edge computing clusters are interconnected via a wide area network.

[0029] 1. Central Management Cluster The central management cluster comprises the following components: Computing Platform: Provides an API gateway and user portal to receive user data upload requests and computing task submission requests, and to provide feedback on status and results.

[0030] The computing power scheduling system is the core scheduling component. It is responsible for parsing computing task requests, aggregating computing power status information reported by each edge computing cluster in real time, and making decisions on the target execution cluster based on the comprehensive cost scoring model proposed in this invention and the data distribution information obtained by querying metadata records.

[0031] Data Management System: Used to generate, store, query, and update metadata records for all business data. It maintains a global data catalog view, providing accurate and consistent data location and attribute information for scheduling decisions, and serves as the hub connecting data and computing power status information.

[0032] Central storage system: Used to securely store raw business data uploaded by users (as an authoritative data source) and archive the execution results of computing tasks returned by each edge computing cluster.

[0033] 2. Edge computing cluster Each cluster in an edge computing cluster comprises the following components: Edge storage system: Storage resources located locally in the cluster, serving as a high-speed cache layer for data, providing low-latency data access for locally executed computing tasks, and temporarily storing intermediate results.

[0034] Scheduling system probe: A resident process deployed on the edge computing cluster management node, responsible for periodically reporting the real-time computing power status of the cluster (such as CPU / GPU utilization, memory, task queue depth, etc.) to the central management cluster, and receiving and executing computing task start instructions issued by the center.

[0035] Data Management Probe: A service deployed on edge computing nodes, responsible for receiving and executing data scheduling instructions from the center, completing cross-domain data retrieval, verification, and storage from the data source to the local machine, and reporting the results back to the center. Simultaneously, it provides standard file access interfaces for locally cached data to computing tasks and implements advanced features such as streaming access.

[0036] See Figure 2 The cross-domain data and task scheduling method based on computing power location according to this invention includes the following steps: Step S101: Data reception, storage and metadata record generation.

[0037] This step aims to address the initial entry point for data assetization and global manageability. Its technical principle is to establish a global data registration and description system centered on metadata records. When a user uploads raw business data, the system generates a structured metadata record while completing physical storage. The core of this record is a globally unique data identifier, and it also includes at least the data's size, integrity verification value (such as a hash value), and its initial physical distribution information (e.g., the path in central storage).

[0038] This step elevates the original data block into an asset object with a unique identity, measurability, and location, providing a unique and accurate index and descriptive foundation for subsequent intelligent scheduling and governance throughout the entire process.

[0039] Step S102: Compute task request parsing and resource dependency extraction.

[0040] This step aims to transform the user's high-level computing intent into constraints that the system can precisely process. The underlying technology involves semantic parsing and element extraction of structured computing task requests. Specifically, the system parses two rigid constraints from the request: first, the specifications of the computing resources required to execute the task, such as the specific model and quantity of GPUs, CPU cores, and memory capacity; and second, the list of target data identifiers on which the computing task depends. This achieves the crucial transformation from business requirements to machine-executable instructions, generating precise resource requirement lists and data dependency lists, providing clear input targets for subsequent precise resource matching and data collaborative scheduling.

[0041] Step S103: Data distribution exploration based on metadata records.

[0042] This step aims to provide timely and accurate data location context for scheduling decisions. Its technical principle is to utilize metadata records as a global data knowledge base for real-time querying. Based on the target data identifier extracted in step S102, the system initiates a query to the metadata and directory service. This service retrieves the corresponding metadata record and returns information about the size of the data and its replica status across all storage nodes in the network at the current moment—a data distribution list. This list explicitly indicates which edge clusters have already cached the data.

[0043] This step transforms static metadata registration information into dynamic data location snapshots that are directly related to the current scheduling decision. This enables the scheduling system to accurately perceive where the data is and how large it is. This is a key prerequisite for achieving joint optimization of data and computing power, and avoids decisions based on outdated or erroneous information.

[0044] Step S104: Joint decision-making of the target execution cluster based on the comprehensive cost score.

[0045] This step aims to address the technical problem in traditional scheduling where the selection of computing resources and the cost of data migration are considered separately, making it impossible to achieve global optimization. The technical principle is to construct a joint optimization decision model that normalizes the queuing latency on the computing side and the migration and scheduling latency on the data side of candidate clusters to the same dimension for weighted comprehensive evaluation.

[0046] First, all edge computing clusters are initially screened for hardware compliance based on the computing resource specifications required by the task, resulting in a candidate cluster set. For example, the computing task requires two NVIDIA A100 GPUs. The system has three edge clusters: cluster A has four A100 cards, cluster B has eight V100 cards, and cluster C has two A100 cards, but they are currently in use. During the initial screening, the system checks whether the physical hardware inventory of each cluster contains the A100 GPUs required by the task. Cluster B, lacking any A100 cards, is directly filtered out. Clusters A and C, both possessing A100 cards, pass the initial screening and enter the subsequent candidate cluster set.

[0047] Furthermore, by receiving real-time resource load information (including CPU / GPU utilization, memory level, task queue length, and historical average task execution time) periodically and proactively reported by each edge computing cluster, the scheduling decision engine of the central management cluster dynamically estimates any candidate cluster. Estimated queuing time in the task queue (That is, the expected waiting time for a task to be scheduled to the cluster in its local task queue).

[0048] Furthermore, network monitoring modules (or transmission controllers) deployed in the central management cluster periodically send network probe packets to the target edge computing cluster to measure link packet loss rate and round-trip latency, and based on this, estimate the distance from the data source location to the candidate cluster. Real-time available bandwidth Alternatively, statistical predictions can be made by analyzing the successful windows of recent historical data transmissions to obtain the data source location to the candidate cluster. Real-time available bandwidth It represents network transmission capacity and is a key variable affecting data migration latency.

[0049] Furthermore, any candidate cluster is determined based on the data distribution list generated from the query metadata records. Required data scheduling volume If the data distribution list shows that the data is already local to the cluster, then =0; otherwise, This equals the total amount of data that needs to be transferred from the data source location (also determined by the data distribution list), and the data scheduling volume. The data migration overhead associated with choosing this cluster was quantified.

[0050] Next, for each candidate cluster in the set The comprehensive cost score is calculated using the following mathematical model: in, Candidate cluster Overall cost score; The maximum queuing wait time among all candidate clusters. and The weighting coefficients are preset and satisfy the following conditions: These two parameters allow the system to dynamically adjust its optimization bias based on task type (e.g., compute-intensive, data-intensive) or business objectives (latency-sensitive, cost-sensitive). Increase This means that decision-making favors clusters with less computational overhead; increasing... This means that decision-makers prefer clusters with localized data or high network bandwidth.

[0051] Finally, the candidate cluster with the smallest overall cost score was selected as the target execution cluster.

[0052] This step uses a unified cost function to quantitatively compare and automatically select the best strategy between "waiting for computing power at the data location" and "moving the data to the location of idle computing power". This minimizes the overall task completion time and significantly improves the scientific nature and overall efficiency of scheduling in complex cross-domain environments.

[0053] Step S105: Generation and adaptive transmission of cross-domain data scheduling instructions.

[0054] This step aims to translate the decision result of step S104 into specific execution actions and ensure the efficiency and stability of the cross-domain data migration process.

[0055] Specifically, when a target execution cluster is identified and it lacks the required local data, the scheduling instruction generation module of the central management cluster generates a structured data scheduling instruction based on the data distribution list. This instruction specifies the data source location, target location, target data identifier, verification method, etc., and is then sent to the data management probe of the target execution cluster. When the data management probe executes scheduling, it employs a dynamic adaptive data transmission strategy to cope with unstable wide area network environments. Its core idea is to dynamically adjust the transmission chunk size based on the real-time detected packet loss rate of the link from the data source to the cluster. When the network is stable (packet loss rate less than or equal to the first threshold), use larger chunks (e.g., 64MB). Larger chunks reduce protocol overhead and maximize throughput on stable links.

[0056] When the network is unstable (packet loss rate greater than or equal to the second threshold, and the second threshold is greater than the first threshold), smaller data blocks (e.g., 4MB) are used. Smaller blocks require less retransmission of data when packets are lost, thus improving jitter resistance and transmission efficiency.

[0057] When the packet loss rate is between two thresholds, the current block size can be kept unchanged.

[0058] Furthermore, for ultra-large-scale data, a streaming mounting / on-demand reading technique is employed. The data management probe prioritizes fetching the file's metadata records (such as header description information and index information) and declares readiness to the computation task. When the computation task actually reads a portion of the data, the data management probe only fetches the corresponding data block. This mechanism transforms the waiting period for all data into a process of simultaneous transmission and computation, significantly reducing task startup latency.

[0059] Furthermore, this step integrates an exception handling mechanism. If transmission is interrupted or verification fails, the data management probe deletes the incomplete data fragment and reports the failure status. The central management cluster then marks the replica status as non-existent in the metadata record, ensuring state consistency and providing an accurate basis for possible rescheduling.

[0060] Step S106: Metadata update, task assignment and exception handling.

[0061] This step aims to complete the final confirmation of the scheduled execution and ensure that the computation task starts in a data-ready environment.

[0062] Specifically, after completing data scheduling (or metadata record retrieval), the data management probe of the target execution cluster sends a confirmation message to the center confirming the completion of cross-domain data scheduling. Upon receiving the confirmation, the metadata and directory service of the central management cluster immediately updates the physical distribution information corresponding to the target data identifier in the corresponding metadata record, for example, by adding a new replica record in the target execution cluster. This operation ensures the real-time performance and consistency of the global data view.

[0063] Furthermore, before task distribution, the system performs a final health check. If the target execution cluster is found to have insufficient resources or be offline, a dynamic execution domain switch is triggered. The scheduling engine can either re-execute step S104 to select a new target execution cluster, or directly switch to another edge computing cluster with existing data replicas as the new target execution cluster (i.e., computation is performed on the nearest available cluster). This greatly enhances the system's fault tolerance. If everything is normal, the task distribution module then sends the computation task execution instruction to the scheduling system probe of the target execution cluster.

[0064] This step ensures a high degree of reliability and consistency in the task execution environment through rigorous data readiness verification, metadata synchronization updates, and exception handling.

[0065] Step S107: Task execution, result archiving, and copy lifecycle management.

[0066] This step aims to complete the computational loop and achieve sustainable resource utilization. Specifically, the target execution cluster's scheduling system probe initiates computational tasks. Tasks access locally cached data (or read on demand via a streaming interface), and generate results upon completion. The results are then sent back to central storage for archiving by the data management probe, and relevant metadata is recorded.

[0067] Meanwhile, the system background continuously runs an edge cache eviction policy based on value density and a business data replica reclamation mechanism, forming a closed loop for data lifecycle management. The specific implementation process is as follows: (1) Edge cache eviction strategy based on value density When the local storage usage of the edge computing cluster exceeds a preset safety threshold, a cache cleanup process is triggered. Specifically, this cleanup is performed on each business data item that is not locked by a task. Calculate elimination score : in, , , As a weighting factor; This indicates business data that is not locked by the task. The access popularity factor is the reciprocal of the difference between the current time and the last access time of the data, or a monotonically decreasing function. This indicates business data that is not locked by the task. The replacement valence factor is directly proportional to the size of the business data and inversely proportional to the current download bandwidth; This indicates business data that is not locked by the task. The hit frequency factor is the number of times the business data is accessed within a preset time window; Prioritize deleting the business data with the lowest elimination score until the local storage space utilization rate falls below the preset target threshold.

[0068] The above-mentioned eviction strategy tends to retain business data that is costly to retrieve, frequently accessed, or has high access popularity, which can significantly improve caching efficiency and reduce cross-domain bandwidth costs.

[0069] (2) Business data copy recycling mechanism Maintain reference counts for various computing tasks for each business data copy stored in each edge computing cluster; after the current computing task is completed, perform a business data copy reclamation judgment: if the reference count of the business data copy is zero and its storage duration exceeds the preset retention duration, delete the business data copy on the edge computing cluster and update the metadata record; if the reference count of the business data copy is zero and the local storage space utilization of the edge computing cluster exceeds the preset security threshold, prioritize deleting the business data copy with the lowest elimination score among the business data copies with zero reference counts, until the space utilization falls back to below the preset target threshold, and update the metadata record.

[0070] The aforementioned business data copy reclamation mechanism combines reference counting with intelligent scoring to achieve secure, automated reclamation and space optimization of edge cache.

[0071] The following typical application scenario demonstrates the specific execution process of the method of the present invention.

[0072] An AI R&D team needs to train a model on a 500GB training dataset (data identifier: DS-Train-V1). This data has been uploaded and stored in a central management cluster. A user has submitted a computation task request that requires 8 A100 GPUs.

[0073] The specific steps are as follows: Step S1: Data Registration. The DS-Train-V1 data is uploaded to the central storage, and the system generates its metadata record, including the data identifier, size (500GB), hash value, and initial location (central).

[0074] Step S2: Task Submission. The user submits a training task, specifying a computing power requirement of 8 A100 GPUs and data dependency DS-Train-V1.

[0075] Step S3: Data Probing. The scheduling engine of the central management cluster queries the metadata records of DS-Train-V1 and finds that its size is 500GB. The current data distribution list is: replicas are only stored in the central storage.

[0076] Step S4: Cluster Decision. After screening, edge computing clusters... Meets the 8-card A100 requirement and is completely idle (estimated queuing time in the task queue). The candidate set is { }. Calculate the amount of data to be scheduled. Measurements were taken from the central management cluster to the edge computing cluster. bandwidth ,because And it is the only candidate, determining the edge computing cluster. The target cluster is to be executed, but data migration is required.

[0077] Step S5: Trigger data migration. Based on the data distribution list, generate scheduling instructions and send them to the edge computing cluster. Data management probe.

[0078] Step S6: Data transmission. The data management probe detects a network packet loss rate of 0.05% and uses a large block (64MB) transmission mode to pull data from the central management cluster.

[0079] Step S7: State Synchronization and Task Distribution. After data transmission is complete and verification is successful, the edge computing cluster... Send confirmation to the central management cluster. The central management cluster updates the metadata records of DS-Train-V1, adding the edge computing cluster. The copy information. Then, it is sent to the edge computing cluster. Training tasks were issued.

[0080] Step S8: Task Execution and Governance. The training task is performed on the edge computing cluster. Execution is performed locally. Upon completion, the model results are sent back to the central management cluster for archiving. After a period of time, the edge computing cluster... When storage space reaches 85% threshold, cache cleanup is triggered. The system calculates the eviction score of each local business data replica. Assuming that DS-Train-V1 has not been accessed again since training (low H value), but its replacement cost C value (500GB / current bandwidth) is high, its eviction score is higher than those business data replicas with small data volume but also infrequent access, so it is retained, thus optimizing cache effectiveness.

[0081] This embodiment uses metadata-driven intelligent decision-making to schedule data to the optimal computing power location for execution, avoiding computing power queuing and waiting, and retains high-value data through intelligent cache management, thereby achieving global efficiency optimization.

[0082] See Figure 3 Another embodiment of the present invention provides a cross-domain data and task scheduling device 200 based on computing power location, including a data registration module 201, a task parsing module 202, a data exploration module 203, a cluster filtering and decision-making module 204, a scheduling instruction generation module 205, a task issuance module 206, and a result archiving module 207. This cross-domain data and task scheduling device 200 based on computing power location can execute the cross-domain data and task scheduling method based on computing power location in the method embodiment.

[0083] Specifically, the cross-domain data and task scheduling device 200 based on computing power location includes: Data registration module 201 is used to receive and store business data uploaded by users and generate metadata records containing globally unique data identifiers and physical distribution information; The task parsing module 202 is used to parse the computing task request submitted by the user to obtain the target data identifier and the required computing resources that the computing task request depends on. Data exploration module 203 is used to query the metadata record according to the target data identifier, obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; The cluster screening and decision-making module 204 is used to screen at least one edge computing cluster according to the computing power resources required by the computing task request to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined according to the comprehensive cost score. The scheduling instruction generation module 205 is used to generate a data scheduling instruction based on the data distribution list to schedule the business data corresponding to the target data identifier from the data source location to the target execution cluster if the target execution cluster does not store the business data corresponding to the target data identifier locally, and send the data scheduling instruction to the target execution cluster. The task distribution module 206 is used to update the physical distribution information of the business data corresponding to the target data identifier in the metadata record in response to the confirmation information of cross-domain data scheduling completed sent by the target execution cluster, and to distribute the computing task to the target execution cluster. The result archiving module 207 is used to receive and archive the execution results of the computing task returned by the target execution cluster.

[0084] It should be noted that the technical solutions provided in this embodiment, which are based on the cross-domain data and task scheduling device 200 corresponding to the computing power location, can be used to execute various method embodiments. Their implementation principles and technical effects are similar to those of the methods, and will not be repeated here.

[0085] The above description is merely a preferred embodiment of the present invention. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to the specific combination of the above-described technical features, but should also cover other technical solutions formed by any combination of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

Claims

1. A cross-domain data and task scheduling method based on computing power location, characterized in that, Includes the following steps: Receive and store business data uploaded by users, and generate metadata records containing globally unique data identifiers and physical distribution information; Parse the computing task request submitted by the user to obtain the target data identifier and the required computing resources that the computing task request depends on; Query the metadata record based on the target data identifier to obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; Based on the computing resources required by the computing task request, at least one edge computing cluster is screened to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined based on the comprehensive cost score, including: based on the candidate clusters Periodically reported real-time computing power status information is used to estimate candidate clusters. Estimated queuing time in the task queue ; The network monitoring module of the central management cluster sends signals to the candidate clusters Network probe packets are periodically sent to measure link packet loss rate and round-trip time, estimating the distance from the data source location to the candidate cluster. Real-time available bandwidth Alternatively, by statistically analyzing historical data transmission records, the path from the data source location to the candidate cluster can be predicted. Real-time available bandwidth Candidate clusters are determined based on the data distribution list. Required data scheduling volume ; Calculate candidate clusters according to the following formula Overall cost score: in, Candidate cluster Overall cost score; The maximum queuing wait time among all candidate clusters. and The weighting coefficients are preset and satisfy the following conditions: Select the comprehensive cost score. The smallest candidate cluster is selected as the target execution cluster; If the target execution cluster does not store the business data corresponding to the target data identifier locally, a data scheduling instruction is generated according to the data distribution list to schedule the business data corresponding to the target data identifier from the data source location to the target execution cluster, and the data scheduling instruction is sent to the target execution cluster. In response to the confirmation message sent by the target execution cluster that the cross-domain data scheduling is complete, the physical distribution information of the business data corresponding to the target data identifier is updated in the metadata record, and a computing task is sent to the target execution cluster. Receive and archive the execution results of the computing tasks returned by the target execution cluster; Monitor the local storage space utilization of the edge computing cluster. When the local storage space utilization exceeds a preset safety threshold, trigger the cache cleanup process. Calculate the following formula for each piece of business data in the local storage of the edge computing cluster that has not been read or used by any computing task. Elimination score : in, , , As a weighting factor; This indicates business data that has not been read or used by any computing task. The access popularity factor is the reciprocal of the difference between the current time and the last access time of the data, or a monotonically decreasing function. This indicates business data that has not been read or used by any computing task. The replacement valence factor is directly proportional to the size of the business data and inversely proportional to the current download bandwidth; This indicates business data that has not been read or used by any computing task. The hit frequency factor is the number of times the business data is accessed within a preset time window; the business data with the lowest elimination score is deleted first until the local storage space utilization rate falls below the preset target threshold.

2. The method for cross-domain data and task scheduling based on computing power location according to claim 1, characterized in that, Also includes: The network link quality from the data source location to the target execution cluster is detected to obtain packet loss rate and latency parameters. The data transmission chunk size is dynamically determined based on the packet loss rate. When the packet loss rate is less than or equal to a first threshold, a first chunk size is used. When the packet loss rate is greater than or equal to a second threshold, a second chunk size smaller than the first chunk size is used. When the packet loss rate is between the first and second thresholds, the current chunk size remains unchanged. Wherein, the first threshold is less than the second threshold. The target execution cluster is instructed to pull the business data corresponding to the target data identifier from the data source location according to the block size of the data transmission; for business data whose data size exceeds a preset threshold, the header description information and index information of the business data are pulled first to announce in advance that the business data is ready; when the subsequent computing task actually reads the business data, the corresponding data blocks are pulled from the data source location as needed.

3. The method for cross-domain data and task scheduling based on computing power location according to claim 1, characterized in that, Also includes: When the cross-domain data scheduling process is interrupted or verification fails, the target execution cluster is instructed to delete the incomplete data fragments that have been transmitted and report the scheduling failure status to the central management cluster. When the central management cluster receives a scheduling failure status reported by the target execution cluster, it marks the business data corresponding to the target data identifier as non-existent in the replica status of the target execution cluster. If the target execution cluster experiences resource shortages or disconnection before the computing task is issued, the target execution cluster determination step is re-executed, a candidate cluster is selected as the new target execution cluster, and cross-domain data scheduling is re-initiated, or the computing task is switched to be executed by the edge computing cluster where the business data is currently located.

4. The cross-domain data and task scheduling method based on computing power location according to claim 1, characterized in that, Also includes: Maintain reference counts for various computing tasks for each copy of business data stored in each edge computing cluster; After the current computing task is completed, a business data copy reclamation check is performed: If the reference count of a business data copy is zero and its storage duration exceeds the preset retention period, then the business data copy is deleted on the edge computing cluster and the metadata record is updated. If the reference count of a business data copy is zero and the local storage space utilization of the edge computing cluster exceeds a preset security threshold, the business data copy with the lowest elimination score among the business data copies with a reference count of zero will be deleted first until the space utilization falls back below the preset target threshold, and the metadata record will be updated.

5. A cross-domain data and task scheduling device based on computing power location, characterized in that, include: The data registration module is used to receive and store business data uploaded by users, and generate metadata records containing globally unique data identifiers and physical distribution information; The task parsing module is used to parse the computing task request submitted by the user to obtain the target data identifier and the required computing resources that the computing task request depends on. The data exploration module is used to query the metadata records based on the target data identifier, obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and generate a data distribution list; The cluster screening and decision-making module is used to screen at least one edge computing cluster based on the computing resources required by the computing task request, thereby obtaining a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined based on the comprehensive cost score, including: based on the candidate clusters Periodically reported real-time computing power status information is used to estimate candidate clusters. Estimated queuing time in the task queue ; The network monitoring module of the central management cluster sends signals to the candidate clusters Network probe packets are periodically sent to measure link packet loss rate and round-trip time, estimating the distance from the data source location to the candidate cluster. Real-time available bandwidth Alternatively, by statistically analyzing historical data transmission records, the path from the data source location to the candidate cluster can be predicted. Real-time available bandwidth Candidate clusters are determined based on the data distribution list. Required data scheduling volume ; Calculate candidate clusters according to the following formula Overall cost score: in, Candidate cluster Overall cost score; The maximum queuing wait time among all candidate clusters. and The weighting coefficients are preset and satisfy the following conditions: Select the comprehensive cost score. The smallest candidate cluster is selected as the target execution cluster; The scheduling instruction generation module is used to generate a data scheduling instruction based on the data distribution list if the target execution cluster does not store the business data corresponding to the target data identifier locally, and then send the data scheduling instruction to the target execution cluster. The task distribution module is used to update the physical distribution information of the business data corresponding to the target data identifier in the metadata record in response to the confirmation information sent by the target execution cluster that the cross-domain data scheduling is completed, and to distribute the computing task to the target execution cluster. The result archiving module is used to receive and archive the execution results of the computing tasks returned by the target execution cluster; The cache eviction module monitors the local storage space utilization of the edge computing cluster. When the local storage space utilization exceeds a preset safety threshold, it triggers the cache cleanup process. It calculates the usage of each piece of business data in the local storage of the edge computing cluster that has not been read or used by any computing task according to the following formula. Elimination score : in, , , As a weighting factor; This indicates business data that has not been read or used by any computing task. The access popularity factor is the reciprocal of the difference between the current time and the last access time of the data, or a monotonically decreasing function. This indicates business data that has not been read or used by any computing task. The replacement valence factor is directly proportional to the size of the business data and inversely proportional to the current download bandwidth; This indicates business data that has not been read or used by any computing task. The hit frequency factor is the number of times the business data is accessed within a preset time window; the business data with the lowest elimination score is deleted first until the local storage space utilization rate falls below the preset target threshold.

6. A cross-domain data and task scheduling system based on computing power location, comprising a central management cluster and at least one edge computing cluster, characterized in that: The central management cluster is used to receive and store business data uploaded by users, generate metadata records containing globally unique data identifiers and physical distribution information, and parse computing task requests submitted by users to obtain the target data identifiers and required computing resources that the computing task requests depend on. Based on the target data identifier, the metadata records are queried to obtain the size and actual physical distribution of the business data corresponding to the target data identifier, and a data distribution list is generated; based on the computing power resources required by the computing task request, at least one edge computing cluster is screened to obtain a candidate cluster set; for each candidate cluster in the candidate cluster set, a comprehensive cost score is calculated based on real-time computing power status information and the data distribution list, and the target execution cluster is determined based on the comprehensive cost score, including: based on the candidate clusters Periodically reported real-time computing power status information is used to estimate candidate clusters. Estimated queuing time in the task queue ; The network monitoring module of the central management cluster sends signals to the candidate clusters Network probe packets are periodically sent to measure link packet loss rate and round-trip time, estimating the distance from the data source location to the candidate cluster. Real-time available bandwidth Alternatively, by statistically analyzing historical data transmission records, the path from the data source location to the candidate cluster can be predicted. Real-time available bandwidth Candidate clusters are determined based on the data distribution list. Required data scheduling volume ; Calculate candidate clusters according to the following formula Overall cost score: in, Candidate cluster Overall cost score; The maximum queuing wait time among all candidate clusters. and The weighting coefficients are preset and satisfy the following conditions: Select the comprehensive cost score. The smallest candidate cluster is selected as the target execution cluster. If the target execution cluster does not store the business data corresponding to the target data identifier locally, a data scheduling instruction is generated based on the data distribution list to schedule the business data corresponding to the target data identifier from the data source location to the target execution cluster, and the data scheduling instruction is sent to the target execution cluster. In response to the confirmation information sent by the target execution cluster that the cross-domain data scheduling is completed, the physical distribution information of the business data corresponding to the target data identifier is updated in the metadata record, and a computing task is issued to the target execution cluster. The execution results of the computing task returned by the target execution cluster are received and archived. The edge computing cluster is used to schedule business data corresponding to the target data identifier across domains from the data source location according to the data scheduling instructions sent by the central management cluster, and send a confirmation message of cross-domain data scheduling completion to the central management cluster; monitor the local storage space utilization of the edge computing cluster, and trigger a cache cleanup process when the local storage space utilization exceeds a preset safety threshold; calculate each piece of business data stored locally by the edge computing cluster that has not been read or used by any computing task according to the following formula. Elimination score : in, , , As a weighting factor; This indicates business data that has not been read or used by any computing task. The access popularity factor is the reciprocal of the difference between the current time and the last access time of the data, or a monotonically decreasing function. This indicates business data that has not been read or used by any computing task. The replacement valence factor is directly proportional to the size of the business data and inversely proportional to the current download bandwidth; This indicates business data that has not been read or used by any computing task. The hit frequency factor is the number of times the business data is accessed within a preset time window; the business data with the lowest elimination score is deleted first until the local storage space utilization rate falls below the preset target threshold.

Citation Information

Patent Citations

  • Rendering task scheduling method, device and equipment in heterogeneous multi-cluster

    CN118535345A

  • Multi-cluster task scheduling system and method based on scheduling score and increment synchronization

    CN121233248A