Computer cluster distributed parallel computing system and method based on message queue

Through a distributed parallel computing system for computer clusters based on message queues, the problems of insufficient computing power and reliability of a single computing node are solved, powerful computing power and efficient data processing are achieved, and good scalability and domestic characteristics are provided.

CN120336044APending Publication Date: 2025-07-18CHINA ELECTRONIC TECH GRP CORP NO 38 RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510403959.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing single computing nodes have insufficient computing power and unsecured computing reliability.

Method used

A computer cluster distributed parallel computing system based on message queue is adopted, including task management nodes, task computing nodes, data storage servers, message queue servers and network switching devices. Through task management nodes, task computing nodes perform parallel computing and store results, data storage server manages task information and results, and network switching devices transmit messages to ensure the reliability and scalability of the system.

Benefits of technology

It effectively solves the problems of insufficient computing power of a single computing node and the inability to guarantee computing reliability, provides strong computing power and efficient data processing capabilities, and has good scalability and domestic independent controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336044A_ABST
    Figure CN120336044A_ABST
Patent Text Reader

Abstract

The invention discloses a computer cluster distributed parallel computing system and method based on a message queue, and belongs to the technical field of computer cluster task processing, and the computer cluster distributed parallel computing system comprises a task management node, N task computing nodes, a data storage server, a message queue server and network switching equipment. According to the computer cluster distributed parallel computing method based on the message queue, the topological relation among the task management node, the task computing node, the database server, the message queue server and the switch is determined, and the respective function requirements of the task management node and the task computing node are clarified; the problems that an existing single computing node is insufficient in computing capacity and computing reliability cannot be guaranteed can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer cluster task processing, and particularly to a computer cluster distributed parallel computing system and method based on a message queue. Background Art

[0002] With the advent of the big data era, the amount of data has grown exponentially. How to quickly analyze and mine valuable information in the data has put forward higher requirements for computing power, computing efficiency, and computing reliability. The number of cores integrated in a single CPU is limited, and its processing power is limited, making it unable to efficiently process this massive amount of data. Therefore, it is very necessary to develop a computer cluster distributed parallel computing method. As an advanced computing architecture, a computer cluster provides powerful computing power and efficient data processing capabilities by integrating the power of multiple computers, and has very good scalability.

[0003] In addition, the safe and stable operation of computing devices is related to national security and social stability. Therefore, the localization, independence, and controllability of computing devices are of great significance. Only by realizing the localization, independent R & D, and design of key software and hardware such as operating systems and CPUs, and comprehensively mastering the core technology of products, can we effectively maintain information security and safeguard national interests.

[0004] The insufficient computing power of existing single computing nodes and the lack of guarantee for computing reliability are problems that need to be solved urgently. For this reason, a computer cluster distributed parallel computing system and method based on a message queue are proposed. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: how to solve the problems of insufficient computing power of existing single computing nodes and the lack of guarantee for computing reliability, and provide a computer cluster distributed parallel computing system based on a message queue.

[0006] The present invention solves the above technical problems through the following technical solutions. The present invention includes a task management node, N task computing nodes, a data storage server, a message queue server, and a network switching device;

[0007] The task management node is used to receive tasks submitted by users, split the tasks into multiple sub-tasks that can be parallelly computed, generate a message queue of tasks to be computed and publish it, and is also used to receive the status of the task computing node fed back by the task computing node and update the task status;

[0008] The task computing node is used to obtain the message queue of tasks to be computed, perform parallel computing by the computing cores, store the task computing results in the data storage server after each computing core finishes computing, and feed back the task computing node status to the task management node; a single task computing node has M computing cores, where N and M are both positive integers;

[0009] The data storage server is used to store the task information table, task information files, and task calculation results;

[0010] The message queue server is used to manage the calculation task queues published by the task management nodes, including the to-be-calculated task queue and the completed calculation task queue;

[0011] The network switching device is used to transmit message packets among the task management nodes, task calculation nodes, data storage servers, and message queue servers.

[0012] Furthermore, the task management nodes are divided into a primary task management node and a standby task management node, which simultaneously manage and monitor the status of N task calculation nodes and the task execution situation. When the primary task management node experiences software failures or hardware outages, the standby management node takes over the work of the primary task management node.

[0013] Furthermore, the status of the task calculation nodes includes the usage status of computing cores, CPU usage status, memory usage status, the number of node cores, and the number of free cores.

[0014] Furthermore, in the data storage server, the task information table and task calculation results are stored in a structured data storage manner, and are stored in the Shentong database that uses the relational data model as the core data model; the task information files are stored in an unstructured data storage manner, and are stored in the distributed file system MongoDB.

[0015] Furthermore, an Ethernet communication method is adopted between the task management nodes and the task calculation nodes. The computing cores of each task calculation node can only access the local memory, and the multiple computing cores of a single task calculation node adopt a shared memory working method. The data storage server and the message queue server can be combined into one when the hardware resources are limited, that is, the data storage service and the message queue service are installed on one server at the same time.

[0016] Furthermore, the task management nodes and the task calculation nodes are both adaptively installed with the domestic Kylin Xin'an operating system, and the CPU model is FT2000.

[0017] The present invention also provides a computer cluster distributed parallel computing method based on a message queue, which is implemented based on the above-mentioned distributed parallel computing system, and includes the following steps:

[0018] S1: The task management node receives the task submitted by the user, splits the task into sub-tasks that can be parallelly calculated, generates a to-be-calculated task message queue and publishes it;

[0019] S2: The task computing node obtains the message queue of tasks to be computed, and parallel computing is performed by the computing cores. After each computing core finishes the computation, the task computing result is stored in the Shentong database, and meanwhile, the status of the task computing node is fed back to the task management node;

[0020] S3: The task management node receives the status of the task computing node fed back by the task computing node. After the task computing node finishes all tasks to be computed, the task management node updates the task status.

[0021] Furthermore, in the step S1, the specific processing procedure is as follows:

[0022] S11: Receive and upload the task information file. The task information file is stored in the distributed file system MongoDB. For each received task information file, a task information record is generated in the task information table of the Shentong database. The fields in the task information table include: task number, task name, and task status. The task status includes three states: waiting for computation, starting computation, and computation completed. The initial value of the task status is set to waiting for computation;

[0023] S12: The task management node monitors the task information table in the Shentong database to obtain the task information, and meanwhile obtains the corresponding task information file from the distributed file system MongoDB;

[0024] S13: The task management node divides the task into multiple sub - tasks that can be computed in parallel according to the task information file and the task information table;

[0025] S14: The task management node generates and publishes the message queue of tasks to be computed, and meanwhile sets the task status of the current task in the task information table to starting computation.

[0026] Furthermore, in the step S2, the specific processing procedure is as follows:

[0027] S21: The task management node determines whether all current computing tasks are completely computed according to the number of generated message queues of tasks to be computed and the number of completed computing task queues. When the number of completed computing task queues is equal to the total number of message queues of tasks to be computed, it indicates that all the computing tasks are completely computed, and then the computation stops. Otherwise, it turns to step S22;

[0028] S22: The task computing node obtains the message queue of tasks to be computed, and according to the maximum number of cores m participating in the task computation set in the configuration file, at most m tasks to be computed are taken from the message queue of tasks to be computed, where m < M;

[0029] S23: Each computing core of the task computing node obtains the task - related information from the shared memory of the task computing node;

[0030] S24: The m computing cores of the task computing node perform parallel computing;

[0031] S25: After each computing core finishes calculating the current computing task, it uploads the task calculation result to the Shentong database, and then turns to step S21 until the message queue of the tasks to be calculated is empty; and in real time, it feeds back the status of the task computing node to the task management node.

[0032] Furthermore, in step S3, the task management node updates the task status, that is, sets the task status in the task information table in the Shentong database to "calculation completed".

[0033] The present invention has the following advantages compared with the prior art: For this computer cluster distributed parallel computing system based on message queue, it determines the topological relationship among the task management node, the task computing node, the database server, the message queue server and the switch, clarifies the respective functional requirements of the task management node and the task computing node, and through the computer cluster distributed parallel computing method based on message queue, it can effectively solve the problems of insufficient computing power of the existing single computing node and unguaranteed computing reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 is a schematic diagram of the domestic computer cluster distributed parallel computing system based on message queue in an embodiment of the present invention;

[0035] Figure 2 is a schematic diagram of the overall process of the domestic computer cluster distributed parallel computing method based on message queue in an embodiment of the present invention;

[0036] Figure 3 is a schematic diagram of the detailed process of the domestic computer cluster distributed parallel computing method based on message queue in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following details the embodiments of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and provides a detailed implementation manner and a specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0038] Embodiment 1

[0039] As Figure 1As shown in the figure, this embodiment provides a domestic computer cluster distributed parallel computing system based on a message queue, including: 1 main task management node, 1 standby task management node, N task computing nodes, 1 data storage server, 1 message queue server, and 1 network switching device. The task management node and the task computing node are both adaptively installed with the domestic Kirin Xin'an operating system, the CPU is the domestic Feiteng 2000 (FT2000), and each task computing node has M computing cores.

[0040] The task management node is used to receive tasks submitted by users, split the tasks into sub-tasks that can be parallelly computed, generate a message queue of tasks to be computed and publish it, and receive the status of the task computing node feedback by the task computing node and update the task status.

[0041] The task management node is divided into a main task management node and a standby task management node, which can manage and monitor the status of N task computing nodes and the task execution situation at the same time. When the main task management node has a software failure or a hardware outage, the standby management node takes over the work of the main task management node.

[0042] The task computing node is used to obtain the message queue of tasks to be computed, perform parallel computing by the computing cores, store the computing results in the data storage server after each computing core finishes computing, and at the same time feedback the status of the task computing node (including: computing core usage status, CPU usage status, memory usage status, number of node cores, number of free cores) to the task management node.

[0043] The data storage server is used to store the task information table, task information files, and task computing results. Among them, the task information table and task computing results adopt a structured data storage method. Specifically, it uses the Shentong database with a relational data model as the core data model, supports the SQL general database query language, provides standard data access interfaces such as ODBC and JDBC, and has the ability to manage massive data and handle large-scale concurrency; the task information files adopt an unstructured data storage method. Specifically, it uses the distributed file system MongoDB for storage.

[0044] The message queue server is used to manage the computing task queue published by the task management node, including the task queue to be computed and the completed computing task queue.

[0045] The network switching device is used to transmit message packets among the task management node, task computing node, data storage server, and message queue server. Specifically, the network switching device can be a network switch, router, or other devices with network switching and distribution functions.

[0046] Specifically, an Ethernet communication method is adopted between the task management node and the task computing nodes. The computing cores of each task computing node can only access the local memory, and multiple computing cores of a single task computing node adopt a shared memory working method; the data storage server and the message queue server can be combined into one when the hardware resources are limited, that is, the data storage service and the message queue service are installed on one server at the same time.

[0047] Embodiment 2

[0048] As Figure 2 shown, this embodiment provides a domestic computer cluster distributed parallel computing method based on a message queue, which is implemented based on the distributed parallel computing system in Embodiment 1, and includes the following steps:

[0049] S1: The task management node receives the task submitted by the user, divides the task into sub-tasks that can be parallelly computed, generates a message queue of tasks to be computed and publishes it;

[0050] S2: The task computing node obtains the message queue of tasks to be computed, and the computing cores perform parallel computing. After each computing core finishes the computation, it stores the task computing result in the Shentong database, and at the same time feeds back the task computing node status to the task management node;

[0051] S3: The task management node receives the task computing node status fed back by the task computing node. After the task computing node completes all tasks to be computed, the task management node updates the task status.

[0052] In this embodiment, in step S1, the task management node receives the task submitted by the user, divides the task into sub-tasks that can be parallelly computed, generates a message queue of tasks to be computed and publishes it, that is, the task management node first receives the task submitted by the user, stores the task information file in the distributed file system MongoDB, and records the task information in the task information table in the Shentong database, and then divides the task into sub-tasks that can be parallelly computed, generates a message queue of tasks to be computed and publishes it;

[0053] In this embodiment, in step S2, the task computing node obtains the message queue of tasks to be computed, and the computing cores perform parallel computing. After each computing core finishes the computation, it stores the computation result in the database, and at the same time feeds back the task computing node status to the task management node, that is, the task computing node obtains the task to be computed from the message queue of tasks to be computed, and the m (m < M) computing cores of the task computing node perform parallel computing. After each computing core finishes the computation, it stores the computation result in the database, and at the same time feeds back the task computing node status (including: computing core usage status, CPU usage status, memory usage status, number of node cores, number of free cores) to the task management node;

[0054] In this embodiment, in step S3, the task management node receives the status of the task computing node fed back by the task computing node. After the task computing node completes all the tasks to be computed, the task management node updates the task status. That is, the task management node determines whether all the current computing tasks have been computed based on the total number of task message queues to be computed generated and the number of completed computing tasks. When the number of completed computing tasks is equal to the total number of task message queues to be computed, it indicates that all the computing tasks have been computed. At this time, the task management node updates the task status in the task information table in the Shentong database.

[0055] As Figure 3 shown, the specific process of the distributed parallel computing method in this embodiment includes the following steps:

[0056] Step 201: Start;

[0057] Step 202: Receive and upload the task information file. The task information file is stored in the distributed file system MongoDB. Among them, for each received task information file, a task information record is generated in the task information table of the Shentong database. The fields in the task information table include: task number, task name, and task status. The task status includes three states: waiting for computing, starting to compute, and computing completed. The initial value of the task status is set to waiting for computing;

[0058] Step 203: The task management node monitors the task information table in the Shentong database and obtains the task information, and at the same time obtains the corresponding task information file from the distributed file system MongoDB;

[0059] Step 204: The task management node divides the task into parallel computable subtasks according to the task information file and the task information table (such as video transcoding, splitting a 3-hour long video into 10-second sub-videos according to the time axis for parallel transcoding);

[0060] Step 205: The task management node generates and publishes a task message queue to be computed, and at the same time sets the task status of the current task in the task information table to starting to compute;

[0061] Step 206: The task management node determines whether all the current computing tasks have been completely computed according to the number of generated task message queues to be computed and the number of completed computing task queues. When the number of completed computing task queues is equal to the total number of task message queues to be computed, it indicates that all the computing tasks have been completely computed, and it turns to step 211; otherwise, it turns to step 207;

[0062] Step 207: The task computing node obtains the message queue of tasks to be computed, and according to the maximum number of cores m participating in task computing set in the configuration file (this number of cores should be less than the maximum number of cores M of the task computing node server), at most m tasks to be computed are taken from the message queue of tasks to be computed.

[0063] Step 208: Each computing core of the task computing node obtains task-related information from the shared memory of this task computing node.

[0064] Step 209: The m computing cores of the task computing node perform parallel computing.

[0065] Step 210: After each computing core finishes computing the current computing task, the computing result is uploaded to the Shentong database, and then it turns to Step 207 until the message queue of tasks to be computed is empty.

[0066] Step 211: The task management node updates the task status and sets the task status in the task information table in the Shentong database to computed.

[0067] Step 212: End.

[0068] In summary, for the distributed parallel computing system of the computer cluster based on message queue in the above embodiments, the topological relationship among the task management node, the task computing node, the database server, the message queue server and the switch is determined, the respective functional requirements of the task management node and the task computing node are clarified, and through the distributed parallel computing method of the computer cluster based on message queue, the problems of insufficient computing power of the existing single computing node and unreliable computing can be effectively solved.

[0069] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A computer cluster distributed parallel computing system based on a message queue, characterized in that Including: A task management node, N task computing nodes, a data storage server, a message queue server, and a network switching device; The task management node is used to receive tasks submitted by users, split the tasks into multiple subtasks that can be computed in parallel, generate a message queue of tasks to be computed and publish it, and is also used to receive the status of the task computing nodes fed back by the task computing nodes and update the task status; The task computing node is used to obtain the message queue of tasks to be computed, perform parallel computing by the computing cores, store the task computing results in the data storage server after each computing core finishes computing, and at the same time feed back the status of the task computing node to the task management node; A single task computing node has M computing cores, where N and M are both positive integers; The data storage server is used to store task information tables, task information files, and task computing results; The message queue server is used to manage the computing task queue published by the task management node, including the task queue to be computed and the completed computing task queue; The network switching device is used to transmit message packets among the task management node, the task computing node, the data storage server, and the message queue server.

2. The computer cluster distributed parallel computing system based on a message queue according to claim 1, wherein The task management node is divided into a primary task management node and a standby task management node, and simultaneously manages and monitors the status of N task computing nodes and the task execution situation. When the primary task management node has a software failure or a hardware outage, the standby management node takes over the work of the primary task management node.

3. The computer cluster distributed parallel computing system based on a message queue according to claim 1, characterized in that The status of the task computing node includes the usage status of the computing cores, the CPU usage status, the memory usage status, the number of node cores, and the number of idle cores.

4. The computer cluster distributed parallel computing system based on a message queue according to claim 3, wherein In the data storage server, the task information table and the task computing results are stored in a structured data storage manner, and are stored in the Shentong database that uses a relational data model as the core data model; the task information files are stored in an unstructured data storage manner and are stored in the distributed file system MongoDB.

5. The computer cluster distributed parallel computing system based on a message queue according to claim 1, characterized in that, The communication method between the task management node and the task computing node is Ethernet. Each computing core of the task computing node can only access the local memory. The multiple computing cores of a single task computing node work in a shared memory manner. The data storage server and the message queue server can be combined into one when the hardware resources are limited, that is, the data storage service and the message queue service are installed on one server at the same time.

6. The computer cluster distributed parallel computing system based on message queue according to claim 4, wherein, Both the task management node and the task computing node are adaptively installed with the domestic Kylin Xin'an operating system, and the CPU model is FT2000.

7. The computer cluster distributed parallel computing method based on a message queue is implemented based on the distributed parallel computing system as described in claim 6, characterized in that, Including the following steps: S1: The task management node receives the task submitted by the user, splits the task into subtasks that can be computed in parallel, generates a message queue of tasks to be computed and publishes it; S2: The task computing node obtains the message queue of tasks to be computed, performs parallel computing by the computing cores, stores the task computing results in the Shentong database after each computing core finishes computing, and at the same time feeds back the status of the task computing node to the task management node; S3: The task management node receives the status of the task computing node fed back by the task computing node. After the task computing node completes all tasks to be computed, the task management node updates the task status.

8. The method for distributed parallel computing of a computer cluster based on a message queue according to claim 7, characterized in that, In the step S1, the specific processing process is as follows: S11: Receive and upload the task information file. The task information file is stored in the distributed file system MongoDB. Among them, for each received task information file, a task information record is generated in the task information table of the Shentong database; the fields in the task information table include: task number, task name, and task status. The task status includes three statuses: waiting for calculation, starting calculation, and calculation completed. The initial value of the task status is set to waiting for calculation; S12: The task management node monitors the task information table in the Shentong database and obtains the task information, and at the same time obtains the corresponding task information file from the distributed file system MongoDB; S13: The task management node divides the task into multiple sub-tasks that can be calculated in parallel according to the task information file and the task information table; S14: The task management node generates and publishes a message queue of tasks to be calculated, and at the same time sets the task status of the current task in the task information table to starting calculation.

9. The computer cluster distributed parallel computing system based on a message queue according to claim 8, characterized in that, In the step S2, the specific processing process is as follows: S21: The task management node judges whether all current calculation tasks are all calculated completed according to the number of message queues of tasks to be calculated generated and the number of message queues of completed calculation tasks. When the number of message queues of completed calculation tasks is equal to the total number of message queues of tasks to be calculated, it indicates that all the calculation tasks are calculated completed, then stop the calculation, otherwise turn to step S22; S22: The task calculation node obtains the message queue of tasks to be calculated, and according to the maximum number of cores m participating in the task calculation set in the configuration file, takes at most m tasks to be calculated from the message queue of tasks to be calculated, where m < M; S23: Each calculation core of the task calculation node obtains the task-related information from the shared memory of the task calculation node; S24: The m calculation cores of the task calculation node perform parallel calculation; S25: After each calculation core completes the calculation of the current calculation task, uploads the task calculation result to the Shentong database, and turns to step S21 until the message queue of tasks to be calculated is empty; And feedback the status of the task calculation node to the task management node in real time.

10. The computer cluster distributed parallel computing system based on a message queue according to claim 9, characterized in that, In the step S3, the task management node updates the task status, that is, sets the task status in the task information table in the Shentong database to calculation completed.