A grid coordination type distributed computing method and system
Patent Information
- Application Number
- CN202611293250.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-25
- Publication Date
- 2026-09-25
AI Technical Summary
但是,远程大型固定式数据中心的业务处理模式,边缘计算任务和信息交互的回传成本高,响应时延大幅提升,无法满足边缘协同等实时应用需求,难以提供敏捷高效的计算运用能力
[0026]有益效果:与现有技术相比,本发明的优点在于:本发明支持多样任务灵活扩展,适合于处理数据量大、处理数据与地理位置相关的应用任务,针对领域任务进行分布式协同处理,可以应用于边缘计算等分布式环境。本发明提出了调度算子和应用算子,实现控制和计算节点松耦合,有利于不同任务的灵活实施;采用了基于任务处理网格区域的资源调度方法,与传统的领域任务单节点处理相比,在缩短了任务平均处理时间的同时,提升了分布式节点的资源利用率。
Smart Images

Figure CN122824744A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed computing, and in particular to grid-based collaborative distributed computing methods and systems. Background Technology
[0002] With the application and development of technologies such as cloud computing and big data, and the profound transformation of system construction and application concepts, deploying information systems based on the cloud and supporting software-based management and control, network virtualization, service computing power, and intelligent applications has become an inevitable trend and an important approach. The efficient operation of complex mega-systems requires ubiquitous computing power. Transforming the current centralized computing environment and forming a more agile, flexible, and survivable new computing architecture is a necessary requirement for the formation of system capabilities. However, the business processing model of remote, large-scale fixed data centers suffers from high backhaul costs for edge computing tasks and information interaction, significantly increased response latency, and cannot meet the real-time application needs of edge collaboration, making it difficult to provide agile and efficient computing capabilities.
[0003] Meanwhile, traditional distributed processing models are insufficient to meet the practical needs of complex, specific domain applications in terms of distributed processing models, robust architecture design, and node load scheduling, such as task decomposition, wide-area distributed robustness, and online task adjustment. Summary of the Invention
[0004] Purpose of the invention: The purpose of this invention is to provide a grid-based collaborative distributed computing method and system to improve resource utilization and task execution speed.
[0005] Technical solution: The grid-based collaborative distributed computing method of the present invention includes the following steps:
[0006] Retrieve application computing tasks on any computing node in a distributed computing system;
[0007] The application computing task is decomposed into several collaborative computing tasks;
[0008] Obtain the required resources for each collaborative computing task, and allocate the collaborative computing tasks to several computing nodes for distributed computing based on their required resources;
[0009] The calculation results of each computing node are aggregated into the control node of the distributed computing system to obtain the calculation result of the application computing task.
[0010] Furthermore, the application computing task is decomposed into several collaborative computing tasks, including:
[0011] The number of collaborative computing tasks is: ,in The number of computational tasks applied, The processing power of a single computing node. This is the redundancy processing coefficient for the switching area. This is the system's redundancy increment coefficient for damage resistance; This serves as a reference coefficient based on historical experience.
[0012] Furthermore, obtaining the required resources for each collaborative computing task and allocating the collaborative computing tasks to several computing nodes for distributed computing based on their required resources includes:
[0013] The system filters out several first computing nodes that meet the resource requirements, and then assigns the collaborative computing tasks based on the geographical location and network latency of the first computing nodes.
[0014] Furthermore, obtaining the required resources for each collaborative computing task and allocating the collaborative computing tasks to several computing nodes for distributed computing based on their required resources also includes:
[0015] The system tracks the execution process of distributed computing on each computing node. If a computing node fails, the collaborative computing tasks executed on that node are distributed to other computing nodes.
[0016] Furthermore, when computing nodes perform distributed computing, automatic backup and recovery management is implemented for the computing nodes.
[0017] Furthermore, after obtaining the calculation result of the application computing task, the process also includes:
[0018] The calculation results are compared with known true values or standard data from historical tasks to assess the accuracy of the calculation.
[0019] When an application computing task is obtained again, the collaborative computing task is allocated according to the resources required by the collaborative computing task and the accuracy.
[0020] The present invention discloses a grid-based collaborative distributed computing system, which includes control nodes and computing nodes; the control node includes a scheduling operator, and the computing node includes an application operator.
[0021] The control node is used to obtain the application computing tasks on any computing node in the distributed computing system. The application computing tasks are decomposed into several collaborative computing tasks, and the required resources for each collaborative computing task are obtained. The collaborative computing tasks are allocated to the application operators in several computing nodes for distributed computing according to their required resources through the scheduling operator. The computing results of each computing node are summarized to obtain the computing result of the application computing task.
[0022] The computing nodes are used to perform distributed computing on the collaborative computing tasks and upload the computing results to the control nodes.
[0023] The electronic device of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the grid-based collaborative distributed computing method.
[0024] The computer-readable storage medium of the present invention stores a computer program, which, when executed by a processor, implements the grid-based collaborative distributed computing method.
[0025] The computer program product of the present invention includes a computer program that, when executed by a processor, implements the grid-based collaborative distributed computing method described above.
[0026] Beneficial Effects: Compared with existing technologies, the advantages of this invention are as follows: This invention supports flexible expansion for diverse tasks, is suitable for processing large amounts of data and application tasks related to geographical location, performs distributed collaborative processing for domain tasks, and can be applied to distributed environments such as edge computing. This invention proposes scheduling operators and application operators to achieve loose coupling between control and computing nodes, which is conducive to the flexible implementation of different tasks; it adopts a resource scheduling method based on task processing grid regions, which, compared with traditional single-node processing of domain tasks, shortens the average processing time of tasks while improving the resource utilization of distributed nodes. Attached Figure Description
[0027] Figure 1 This is a flowchart of the distributed computing method according to an embodiment of the present invention.
[0028] Figure 2 This is a distributed computing system architecture diagram according to an embodiment of the present invention. Detailed Implementation
[0029] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0030] like Figure 1 As shown, the grid-based collaborative distributed computing method includes the following steps.
[0031] Step 1, Computing Resource Management: Collect available computing and storage resources; Based on the computing resource management service, realize unified management of distributed computing and storage resources of computing nodes, including node location and bandwidth information.
[0032] Step 2, computational task management, involves decomposing the application's distributed computing tasks, establishing a computational model, and determining the application operator as the smallest computational unit. During task management, the computational tasks requested resources, primarily including runtime resources, data resources, and computing resources.
[0033] Specifically, in step 2, based on the computing task management portal, the distributed computing tasks of the application are decomposed into multiple collaborative computing tasks. The estimation model for the number of collaborative computing tasks is as follows:
[0034] ;
[0035] in, To estimate the number of collaborative computing tasks. The number of computational tasks applied, The processing power of a single computing node. This is the redundancy processing coefficient for the switching area. This is the system's redundancy increment coefficient for damage resistance; This serves as a reference coefficient based on historical experience.
[0036] Step 3, computing power management and scheduling, realizes the allocation and scheduling of computing resources, and is responsible for controlling the deployment and operation of application operators on various computing nodes.
[0037] Specifically, step 3 includes the following steps:
[0038] Step 3-1: Prepare the runtime environment for the application operator, establish relationships such as information flow and data synchronization, and implement the process of running the application operator on multiple computing nodes;
[0039] Step 3-2: Through memory data synchronization of stateful applications, automatic backup and recovery management is performed on multiple nodes to support the continuity of state during application scheduling and switching.
[0040] Step 4, task management and scheduling, adopts a two-stage strategy of pre-selection plus optimal selection, and realizes the scheduling control of application computing tasks and the distribution of application operators based on scheduling operators.
[0041] Specifically, in the pre-selection phase, computing nodes that do not meet the resource requirements are filtered out. In the optimization phase, among the computing nodes filtered in the pre-selection phase, a latency-sensitive scheduling strategy is used, combined with geographical distribution and network latency, to select the optimal computing node that meets the requirements, thereby realizing the scheduling and distribution of application operators.
[0042] Furthermore, in this optimization phase, scheduling and distribution can be performed based on the accuracy feedback from step 5-3. Scheduling decisions are made by combining geographical distribution, network latency, and the accuracy feedback from step 5. This is achieved by recording node latitude and longitude, calculating physical distances, and periodically measuring inter-node latency via RPC, selecting the optimal computing node that meets the criteria, thus realizing the scheduling and distribution of application operators.
[0043] Specifically, in step 4, the process of tracking and providing real-time feedback on distributed computing is executed. By tracking the execution process of distributed tasks, the current computing scheduling status is checked in real time. When the computing node executing the task fails, it is scheduled to other available computing nodes.
[0044] Step 5: Summarize the calculation results to achieve unified aggregation and fusion of the calculation results from each distributed computing node, forming the final collaborative calculation result.
[0045] Specifically, step 5 includes the following steps:
[0046] Step 5-1: Aggregate the analysis results from each computing node to the control node, and summarize them to form the results of the application task.
[0047] Step 5-2: Compare the results with known true values or standard data from historical tasks, and calculate indicators such as precision and accuracy to evaluate the accuracy.
[0048] Step 5-3: Feed back the accuracy to the scheduling operator in the control node of Step 3 to optimize the task scheduling strategy for the next task.
[0049] like Figure 2 As shown, the grid-cooperative distributed computing system includes control nodes and computing nodes; the control node includes a scheduling operator, and the computing node includes an application operator.
[0050] The computing task management portal in the control node is used to obtain the application computing tasks on any computing node in the distributed computing system. The application computing tasks are decomposed into several collaborative computing tasks, and the required resources for each collaborative computing task are obtained. The collaborative computing tasks are allocated to the application operators in several computing nodes for distributed computing according to their required resources through scheduling operators. The computing results of each computing node are summarized to obtain the computing result of the application computing task.
[0051] The computing nodes are used to perform distributed computing on the collaborative computing tasks and upload the computing results to the control nodes.
[0052] The distributed computing method and distributed computing system described in this invention are mainly aimed at distributed collaborative processing of domain tasks, and can be applied to distributed environments such as edge computing. The following experiment verifies the distributed computing method and distributed computing system described in this invention.
[0053] This experiment demonstrates distributed collaborative processing of multi-source images. (Refer to...) Figure 2The experiment was conducted on a local cluster of 20 physical servers. Each server was equipped with two 8-core Intel Xeon E5-2650v2 2.6GHz processors, 256GB of memory, 1.5TB of disk space, and ran CentOS 6.0, Java 1.7.0_55, and Python 3.5. All servers were connected via a high-speed 1.5Gbps local area network. One server served as the control node, and the remaining 19 servers each functioned as a compute node.
[0054] The control node divides the image into 60 slices. The image recognition operator on a single server can process 8 slices. The redundancy processing coefficient for the exchange area is set to 1.2, the system's survivability redundancy increment coefficient is 1.6, and the historical experience reference coefficient is 1. The estimated number of task nodes to be processed is 60 / 8 × 1.2 × 1.6 × 1 = 14.4. Rounding up to 15, this needs to be divided into 15 grid-coordinated computing nodes, resulting in 15 grid-coordinated computing tasks.
[0055] Images are accessed at the nearest computing node and distributed as image processing tasks to the distributed computing service. The distributed computing service receives the tasks, establishes collaborative processing, load balancing, and failover relationships, and the distributed computing scheduling operator decomposes the image slice queue into tasks, distributing the image slices to 15 computing nodes for collaborative processing. It calls the image recognition application operator to perform image slice recognition and monitors the results and progress of image slice recognition. The distributed computing service returns the recognition results to the control node scheduling operator. Data that needs to be synchronized during the processing of the image recognition application operator is synchronized through a Redis cluster. The scheduling operator in the control node integrates the feedback results from each computing point. The distributed collaborative processing environment built by the distributed computing service completes efficient image recognition.
[0056] The electronic device of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it implements the grid-based collaborative distributed computing method.
[0057] The computer-readable storage medium of the present invention stores a computer program, which, when executed by a processor, implements the grid-based collaborative distributed computing method.
[0058] The computer program product of the present invention includes a computer program that, when executed by a processor, implements the grid-based collaborative distributed computing method described above.
[0059] The computer-readable storage medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other media that can be used to store program code in the form of instructions or data structures and is accessible by a computer.
[0060] The processor is used to execute a computer program stored in memory to implement the various steps in the methods described in the above embodiments.
Claims
1. A grid-based collaborative distributed computing method, characterized in that, Includes the following steps: Retrieve application computing tasks on any computing node in a distributed computing system; The application computing task is decomposed into several collaborative computing tasks; Obtain the required resources for each collaborative computing task, and allocate the collaborative computing tasks to several computing nodes for distributed computing based on their required resources; The calculation results of each computing node are aggregated into the control node of the distributed computing system to obtain the calculation result of the application computing task.
2. The grid-based collaborative distributed computing method according to claim 1, characterized in that, The application computing task is decomposed into several collaborative computing tasks, including: The number of collaborative computing tasks is: ,in The number of computational tasks applied, The processing power of a single computing node. This is the redundancy processing coefficient for the switching area. This is the system's redundancy increment coefficient for damage resistance; This serves as a reference coefficient based on historical experience.
3. The grid-based collaborative distributed computing method according to claim 1, characterized in that, Obtaining the required resources for each collaborative computing task, and allocating the collaborative computing tasks to several computing nodes for distributed computing based on their required resources, includes: The system filters out several first computing nodes that meet the resource requirements, and then assigns the collaborative computing tasks based on the geographical location and network latency of the first computing nodes.
4. The grid-based collaborative distributed computing method according to claim 1, characterized in that, Obtaining the required resources for each collaborative computing task and allocating the collaborative computing tasks to several computing nodes for distributed computing based on their required resources also includes: The system tracks the execution process of distributed computing on each computing node. If a computing node fails, the collaborative computing tasks executed on that node are distributed to other computing nodes.
5. The grid-based collaborative distributed computing method according to claim 1, characterized in that, When computing nodes perform distributed computing, automatic backup and recovery management is implemented for the computing nodes.
6. The grid-based collaborative distributed computing method according to claim 1, characterized in that, After obtaining the calculation results of the application computing task, the process also includes: The calculation results are compared with known true values or standard data from historical tasks to assess the accuracy of the calculation. When an application computing task is obtained again, the collaborative computing task is allocated according to the resources required by the collaborative computing task and the accuracy.
7. A grid-based collaborative distributed computing system, characterized in that, The distributed computing system includes a control node and computing nodes; the control node includes a scheduling operator, and the computing node includes an application operator. The control node is used to obtain the application computing tasks on any computing node in the distributed computing system. The application computing tasks are decomposed into several collaborative computing tasks, and the required resources for each collaborative computing task are obtained. The collaborative computing task is distributed to application operators on several computing nodes according to its required resources by the scheduling operator for distributed computing; the computing results of each computing node are aggregated to obtain the computing result of the application computing task. The computing nodes are used to perform distributed computing on the collaborative computing tasks and upload the computing results to the control nodes.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, it implements the grid-based collaborative distributed computing method according to any one of claims 1-6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the grid-based collaborative distributed computing method according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the grid-based collaborative distributed computing method according to any one of claims 1-6.