A distributed-based multi-node algorithm and data automatic parallel method

Through the graph neural network intelligently allocates computing tasks and data blocks in a distributed system, the problems of low resource utilization and unbalanced load in the existing technology are solved, and efficient parallel processing and computing efficiency improvement are achieved.

CN119960993BActive Publication Date: 2025-07-18CHENGDU HAIQING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510053921.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-07-18
Estimated Expiration
2045-01-14

AI Technical Summary

Technical Problem

The existing distributed computing systems have problems such as low resource utilization, unbalanced load and high network load in tasks and data allocation, making it difficult to efficiently handle large-scale and complex computing tasks in dynamic environments.

Method used

Through graph neural network (GNN), the computing tasks are intelligently subdivided into subtasks and data blocks. According to the resource status and requirements of the computing nodes and data blocks, their allocation in the distributed system is optimized to achieve parallel processing.

Benefits of technology

It improves the resource utilization and computing efficiency of distributed systems, is suitable for large-scale data processing and complex computing tasks, without the need for complex model training or large amounts of historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119960993B_ABST
    Figure CN119960993B_ABST
Patent Text Reader

Abstract

The present invention discloses a distributed multi-node algorithm and an automatic data parallel method. By subdividing the user's computing tasks into multiple independent subtasks and reasonably partitioning the data set, the subtasks and data blocks are automatically assigned to different computing nodes for parallel processing. After each node completes the processing, the intermediate results are aggregated to the coordination node or the master node for integration and further processing. The method of the present invention, through automated task subdivision and data partitioning, intelligently assigns tasks and data to each node for parallel processing according to the resource status and task requirements of the computing nodes, without the need for complex model training or a large amount of historical data. It has the advantages of simple implementation and wide applicability, can effectively improve the resource utilization rate and computing efficiency of the distributed system, and is applicable to large-scale data processing and complex computing tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed parallel technology, and particularly relates to a multi-node algorithm and an automatic data parallel method based on distribution. Background Art

[0002] In today's information age, with the rapid development of the Internet, Internet of Things, and cloud computing, various applications and services have generated a vast amount of data. The scale and complexity of this data are constantly increasing, posing huge challenges to data storage, processing, and analysis. The existing single-machine computing mode can no longer meet the needs of large-scale data processing and complex computing tasks, and distributed computing systems have thus become an important means to solve this problem.

[0003] Distributed computing improves computing power and efficiency by distributing computing tasks and data across multiple nodes and using parallel processing. However, existing distributed systems usually rely on preset rules or simple scheduling algorithms, such as round-robin, random allocation, or simple resource-based scheduling, for task and data allocation. These methods often prove inadequate when faced with a dynamically changing system environment and diverse task requirements. First, there may be differences in the performance, resource status, and network bandwidth of computing nodes. Existing static allocation strategies cannot fully utilize these heterogeneous resources. Some nodes may be powerful, but due to the lack of an effective allocation strategy, their resource utilization rate is very low; while weaker nodes may become a bottleneck of the system due to overload. Second, different computing tasks and data blocks have different computing complexities and resource requirements. Simple task decomposition and allocation strategies may lead to uneven distribution of tasks and data among nodes, failing to achieve load balancing. This not only reduces the overall performance of the system but may also cause some nodes to be overloaded, leading to system instability or crashes. In addition, data transmission and communication overhead are also important issues to be considered in distributed systems. Existing methods may ignore data locality, resulting in a large amount of data being transmitted between nodes, increasing network load and latency, and reducing system efficiency. To solve the above problems, many improvement methods have been proposed in the industry and academia. For example, dynamic scheduling algorithms are adopted to allocate tasks according to the real-time status of nodes and task requirements; machine learning techniques are introduced to predict task execution time and resource consumption and optimize the allocation strategy. However, these methods are often highly complex, have high implementation costs, or are only applicable to specific scenarios and are difficult to be widely applied. Therefore, how to design a simple and efficient automatic task and data parallel processing method to fully utilize the resources of distributed systems and improve computing efficiency without increasing system complexity remains an urgent technical problem to be solved. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a distributed multi-node algorithm and data automatic parallel method. Through automatic task subdivision and data chunking, according to the resource status and task requirements of computing nodes, tasks and data are intelligently allocated to each node for parallel processing.

[0005] The technical solution adopted by the present invention is as follows: a distributed multi-node algorithm and data automatic parallel method, and the specific steps are as follows:

[0006] S1. Obtain the required data and tasks through the input information of the user, and clarify the task objectives, data types, and processing requirements;

[0007] S2. Preprocess the data obtained in step S1, and perform feature extraction and selection;

[0008] Among them, after the data preprocessing is completed, the calculation task is divided and transferred to step S3, and data chunking is transferred to step S6.

[0009] S3. Divide the overall calculation task into multiple independent subtasks to ensure that each subtask has a clear function and executability;

[0010] S4. Based on step S3, represent the computing nodes and the subtasks to be allocated in the distributed system as a graph structure to provide a structured representation for intelligent allocation;

[0011] S5. Based on step S4, use the graph neural network GNN to intelligently allocate the subtasks to different computing nodes for execution, and optimize the utilization rate of computing resources;

[0012] S6. Reasonably chunk the dataset to be processed obtained in step S2, and represent the computing nodes and data chunks as a graph structure to lay a foundation for the intelligent allocation of data;

[0013] S7. Based on step S6, through the graph neural network GNN, intelligently allocate each data chunk to different nodes for parallel processing to improve the efficiency of data processing;

[0014] S8. After each node completes the subtask and data processing, summarize the intermediate results to the coordination node or the main node for further integration and processing to generate the final result.

[0015] Further, the specific steps of step S1 are as follows:

[0016] S11. Authenticate the user to ensure that they have the permission to execute relevant tasks. At the same time, according to the user's permission, determine the data range and task type that they can access;

[0017] S12. Provide a friendly user interface or API interface that allows users to input or upload task requirements and obtain detailed descriptions of the tasks, including: task objectives, expected outputs, and special requirements;

[0018] S13. Confirm the sources of the data to be processed, including: local upload, database access, or third-party API, and obtain information on the type and format of the data, including: structured data, unstructured data, images, or audio;

[0019] S14. Collect specific requirements from users for data processing, including: algorithms or models to be used, preprocessing requirements, and parameter settings, and confirm the task priority, execution time, and resource limitations.

[0020] Further, the specific steps of step S2 are as follows:

[0021] S21. Clean the data to be processed obtained in step S1, remove noise, errors, and duplicate data therein, handle missing values and outliers, and ensure the quality and integrity of the data;

[0022] S22. Perform format conversion and standardization processing on the data cleaned in step S21, and unify data from different sources and in different formats into a consistent format and unit;

[0023] Among them, format conversion and standardization processing include: converting data types, unifying date and time formats, and standardizing numerical ranges.

[0024] S23. Extract and select features from the data processed in steps S21 and S22 according to the task requirements;

[0025] For structured data, extract statistical features; for unstructured data, namely text or images, use TF-IDF or convolutional neural network CNN to extract high-order semantic features respectively, and screen out key features highly correlated with the target variable through correlation analysis and dimensionality reduction techniques; at the same time, use statistical methods to detect features with no variance and multicollinearity to eliminate redundant features.

[0026] Further, the specific steps of step S3 are as follows:

[0027] S31. Analyze the overall computing task, find out functional modules or steps that can be processed independently, and then identify the core components of the task, and determine the parts that can be executed as independent subtasks;

[0028] S32. Divide the overall computing task into multiple independent subtasks according to the functional modules, and clarify the specific functions and objectives of each subtask to ensure that they can be executed independently and are logically complete;

[0029] S33. Determine the input, output, and processing boundaries for each subtask, as well as the dependencies between other subtasks, and ensure the smooth flow of data between subtasks.

[0030] Further, the specific steps of step S4 are as follows:

[0031] S41. Collect the resource information of each computing node in the distributed system, and represent each computing node as a node in the graph structure, with each computing node carrying its resource attributes and status information;

[0032] Among them, the resource information of each computing node includes: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, and the current load situation.

[0033] S42. Represent the subtasks obtained by dividing step S3 as nodes in the graph structure as well. Each subtask node includes its resource requirements, estimated execution time, task priority, and possible dependencies;

[0034] Among them, the resource requirements of each subtask node include: the required number of CPUs, the required number of GPUs, and the memory size.

[0035] S43. Add edges connecting the computing nodes and subtask nodes in the graph structure to represent the potential allocation relationship that a computing node can execute a subtask, and construct a graph structure G0 = (V0, E0) containing computing nodes and subtasks;

[0036] Among them, the node set V0 = {N1,..., N p , T1,..., T m}, N i represents the i-th computing node, and T j represents the j-th subtask. p and m respectively represent the number of computing nodes and subtasks; the edge set E0 represents the potential allocation relationship between computing nodes and subtasks, as well as the dependencies between subtasks.

[0037] S44. Based on the graph structure of step S3, extract feature vectors for each node in the graph;

[0038] Each node in the graph includes: computing nodes and subtask nodes. For the computing node N i , the extracted feature vector includes: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, current load. Then, for the subtask T j , the extracted feature vector includes: computational complexity, resource requirements, estimated execution time, task priority, data dependency.

[0039] Furthermore, the specific steps of step S5 are as follows:

[0040] S51. Use the graph neural network model GNN to train the graph structure G0 obtained in step S4;

[0041] Among them, during the information transmission process, the hidden state of the node The update expression is as follows:

[0042]

[0043] Among them, k represents the number of iteration layers in the neural network, represents the set of neighbor nodes of node v, W1 and W2 represent trainable weight matrices, b represents the bias term, and σ represents the Sigmoid function.

[0044] Set the loss function L, with the goal of minimizing the task completion time T, balancing the computing node load B, and maximizing the resource utilization rate U. The expression is as follows:

[0045] L = αT + βB - γU

[0046] Among them, α, β, and γ represent weight coefficients, which are used to balance the importance of each index.

[0047] Use backpropagation through the loss function L to update the graph neural network multiple times until the set number of iterations is reached, and complete the training of the graph neural network.

[0048] Among them, the number of iterations is set according to the actual situation. After the iteration ends, all nodes obtain the embedded representation integrating the context information

[0049] S52. Generate the allocation scheme of subtasks to computing nodes;

[0050] First, input the current graph structure G0 and node features into the GNN model trained in step S51 to obtain the fitness score S j assigned to the computing node N i for each subtask T ij . Then, according to the obtained fitness score S ij , select the optimal computing node for each subtask. This fitness score is mapped through a parametric function, and the expression is as follows:

[0051]

[0052] Among them, K represents the total number of iterations, and respectively represent the computing node N i and the data block T jThe final node indicates that Θ represents the trainable parameters of the function f.

[0053] Finally, according to the fitness score S ij generate the allocation scheme of subtasks to computing nodes {(T j , N optimal )}.

[0054] Among them, the optimal computing node N optimal is selected as follows:

[0055]

[0056] S53. Based on step S52, optimize the utilization rate of computing resources;

[0057] First, distribute the subtasks to the corresponding computing nodes N i for execution according to the allocation scheme, and continuously monitor the task progress and system status during the execution, including: task execution time, node load. If an abnormality is detected or the system status changes, re - call the GNN model for inference, adjust the allocation scheme to ensure the stability and efficiency of the system. Iterate until the allocation scheme no longer changes, and complete the optimized deployment of GNN.

[0058] Furthermore, the specific steps of step S6 are as follows:

[0059] S61. Perform strict block processing on the dataset to be processed obtained in step S2;

[0060] Divide the dataset D to be processed into multiple data block sets {D1, D2,..., D n}, where n represents the number of data blocks.

[0061] Among them, the division process is quantitatively analyzed based on data scale, data type, resource requirements, and network topology. Then the optimization objective function expression for data block division is as follows:

[0062]

[0063] Among them, Load(D i ) represents the load of data block D i , Comm(D i ) represents the communication overhead generated by this data block during network transmission, and ε and θ represent weight coefficients.

[0064] S62. Uniformly represent the computing nodes and the data blocks after division as a graph structure;

[0065] Set the computing node set as {M1, M2,..., M q}, by constructing a corresponding set of graph nodes V1 = {M1,..., M q , D1,..., D n} for each computing node and data block. In this graph, the computing node attribute parameters include: processing capacity, storage space, network bandwidth, computing density, and the data block node record information includes: data volume, feature distribution, processing complexity.

[0066] Among them, q represents the number of computing nodes.

[0067] S63. Add connection edges representing potential data block allocation relationships between computing nodes and data block nodes to construct a graph structure G1 = (V1, E1);

[0068] The edge weight function w(M x , D y ) has the following expression:

[0069]

[0070] Among them, Cap(M x ) represents the available computing resources of the computing node M x , Net(M x ) represents the network transmission performance index of this node, and γ and δ represent weight coefficients.

[0071] Furthermore, the specific steps of step S7 are as follows:

[0072] S71. Use the graph neural network model GNN to jointly model the global information of computing nodes and data blocks;

[0073] Based on the graph structure G1 = (V1, E1) constructed in step S6, initialize the feature vector x v for each node in G1.

[0074] Among them, each node in G1 includes: computing nodes and data block nodes; computing node attributes include: processing performance, storage capacity, network bandwidth, resource occupancy, and data block node features include: data volume, processing complexity, expected computing overhead; concatenate the computing node attributes and data block node features to obtain the node initialization feature vector x v .

[0075] S72. Use the graph neural network model GNN to perform multiple iterative trainings on the graph structure G1;

[0076] Set the hidden representation vector of node v′ to be calculated comprehensively based on its own features and the features of its neighbor nodes. The calculation method is as follows:

[0077]

[0078] Among them, k′ represents the number of iterative layers in the neural network, represents the set of neighbor nodes of node v′, W′1 and W′2 represent trainable parameter matrices, b′ represents the bias term, and σ represents the non-linear activation function.

[0079] Among them, the number of iterations is set according to the actual situation. After the iteration ends, all nodes obtain the embedding representation integrating the context information

[0080] S73. Generate the allocation scheme of data block nodes to computing nodes;

[0081] After completing the node representation learning, for each data block node D y calculate its fitness score S′ allocated to computing node M x . This fitness score is mapped through a parameterized function, and the expression is as follows: xy .

[0082]

[0083] Among them, K′ represents the total number of iterations, and respectively represent the final node representations of computing node M x and data block D y , and Θ′ represents the trainable parameter of function f′.

[0084] According to the fitness score S′ xy select the optimal computing node allocation scheme for each data block. By searching for the computing node M′ j that maximizes S′ xy for each data block D optimal to achieve the optimal allocation, the expression is as follows:

[0085]

[0086] Among them, M′ optimal represents the optimal computing node. Allocate data block D y to the computing node corresponding to M′ optimal to realize the intelligent allocation and efficient parallel processing of data blocks in the distributed system.

[0087] S74. Based on step S73, improve the efficiency of data processing;

[0088] During the actual execution process, continuously monitor the task progress and system status. When changes occur in the data block processing process, node load, or network conditions, re - call the GNN model for inference and dynamically adjust the data allocation scheme. Iterate until the allocation scheme no longer changes, and complete the optimized deployment of the GNN.

[0089] Further, the specific steps of step S8 are as follows:

[0090] After each computing node completes the assigned subtasks and data processing, report the intermediate results to the coordination node or the master node in a unified format, and set to collect the intermediate result set {R1, R2,..., R e} from the distributed computing nodes of data blocks or subtasks;

[0091] where each R i includes the intermediate data, statistical metrics, or model parameters after the corresponding subtask processing; e takes p or q according to the data block or subtask corresponding to the distributed computing node respectively.

[0092] Then the coordination node performs integration and fusion operations on the collected intermediate results. According to the actual task requirements and algorithm strategies, use appropriate aggregation functions or integration rules to merge {R i}. For other types of data integration, select the corresponding integration strategy according to the task characteristics and application scenarios.

[0093] After obtaining the final integration result, the coordination node further processes and formats the result according to the output requirements of the task, and performs quality inspection and verification steps to ensure that the final result meets the expected accuracy and performance standards.

[0094] The beneficial effects of the present invention: The method of the present invention divides the user's computing task into multiple independent subtasks, reasonably divides the data set into blocks, automatically allocates subtasks and data blocks to different computing nodes for parallel processing. After each node completes the processing, the intermediate results are summarized to the coordination node or the master node for integration and further processing. The method of the present invention, through automated task subdivision and data chunking, intelligently allocates tasks and data to each node for parallel processing according to the resource status of the computing nodes and task requirements, without complex model training or a large amount of historical data, has the advantages of simple implementation and wide application range, can effectively improve the resource utilization rate and computing efficiency of the distributed system, and is applicable to large - scale data processing and complex computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0095] Figure 1 It is a flowchart of a distributed multi - node algorithm and data automatic parallel method of the present invention.

[0096] Figure 2A schematic diagram of distributed multi-node graph structure allocation in an embodiment of the present invention.

[0097] Figure 3 This is a neural network intelligent allocation decision flow chart in an embodiment of the present invention. DETAILED DESCRIPTION

[0098] The method of the present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0099] like Figure 1 As shown, a flow chart of a distributed multi-node algorithm and data automatic parallel method of the present invention, the specific steps are as follows:

[0100] S1. Obtain the required data and tasks through user input information, and clarify the task objectives, data types and processing requirements;

[0101] S2, preprocessing the data obtained in step S1, and performing feature extraction and selection;

[0102] After the data preprocessing is completed, the calculation task is divided and the process goes to step S3, and the data is divided into blocks and the process goes to step S6.

[0103] S3. Divide the overall computing task into multiple independent subtasks, ensuring that each subtask has clear functions and executability;

[0104] S4. Based on step S3, the computing nodes and the subtasks to be assigned in the distributed system are represented as a graph structure to provide a structured representation for intelligent assignment;

[0105] S5. Based on step S4, use the graph neural network GNN to intelligently allocate subtasks to different computing nodes for execution, thereby optimizing the utilization of computing resources.

[0106] S6, reasonably dividing the data set to be processed obtained in step S2 into blocks, and representing the computing nodes and data blocks as a graph structure, laying a foundation for intelligent allocation of data;

[0107] S7, based on step S6, through the graph neural network GNN, each data block is intelligently allocated to different nodes for parallel processing to improve the efficiency of data processing;

[0108] S8. After each node completes the subtask and data processing, the intermediate results are aggregated to the coordination node or the master node for further integration and processing to generate the final result.

[0109] In this embodiment, the step S1 is specifically as follows:

[0110] S11. Authenticate the user to ensure that they have the permission to perform relevant tasks. At the same time, determine the data scope and task types that the user can access according to the user's permissions.

[0111] S12. Provide a friendly user interface or API interface to allow the user to input or upload task requirements and obtain a detailed description of the task accordingly, including: task objectives, expected outputs, and special requirements.

[0112] S13. Confirm the source of the data to be processed, including: local upload, database access, or third-party API, and obtain information on the data type and format, including: structured data, unstructured data, images, or audio.

[0113] S14. Collect the user's specific requirements for data processing, including: algorithms or models to be used, preprocessing requirements, and parameter settings, and confirm the task priority, execution time, and resource limitations.

[0114] In this embodiment, the specific steps of step S2 are as follows:

[0115] S21. Clean the data to be processed obtained in step S1, remove noise, errors, and duplicate data therein, handle missing values and outliers, and ensure the quality and integrity of the data.

[0116] S22. Perform format conversion and standardization processing on the data cleaned in step S21, unify data from different sources and different formats into a consistent format and unit, and ensure that the data can be correctly recognized and processed in subsequent processing.

[0117] Among them, format conversion and standardization processing include: converting data types, unifying time and date formats, and standardizing numerical ranges.

[0118] S23. According to the task requirements, perform feature extraction and selection on the data processed in steps S21 and S22.

[0119] For structured data, extract statistical features; for unstructured data, namely text or images, respectively use TF-IDF or convolutional neural network (CNN) to extract high-order semantic features, and screen out key features highly correlated with the target variable through correlation analysis and dimensionality reduction techniques; at the same time, use statistical methods to detect features with no variance and multicollinearity to eliminate redundant features.

[0120] In this embodiment, the specific steps of step S3 are as follows:

[0121] S31. Analyze the overall computing task, identify the functional modules or steps that can be independently processed, thereby recognizing the core components of the task, clarifying the parts that can operate independently without relying on the overall context, and determining the parts that can be executed as independent subtasks, thus laying a foundation for subsequent subtask partitioning;

[0122] S32. According to the functional modules, divide the overall computing task into multiple independent subtasks, and clarify the specific functions and objectives of each subtask to ensure that they can be executed independently and are logically complete, so that each subtask can still achieve the expected computing or processing functions when executed independently;

[0123] S33. Determine the input, output, and processing boundaries for each subtask, as well as the dependencies between other subtasks, and ensure the smooth flow of data between subtasks.

[0124] On this basis, establish an effective data transfer and interface interaction mechanism to ensure that each subtask can smoothly obtain the required data when executed in parallel or serially, thereby ensuring the coherence and consistency of the overall task process.

[0125] In this embodiment, the specific steps of step S4 are as follows:

[0126] S41. Collect the resource information of each computing node in the distributed system, and represent these computing nodes as nodes in the graph structure, with each computing node carrying its resource attributes and status information;

[0127] Among them, the resource information of each computing node includes: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, and the current load situation.

[0128] S42. Represent the subtasks obtained by dividing step S3 as nodes in the graph structure as well, with each subtask node including its resource requirements, estimated execution time, task priority, and possible dependencies;

[0129] Among them, the resource requirements of each subtask node include: the number of required CPUs, the number of required GPUs, and the memory size.

[0130] S43. Add edges connecting the computing nodes and subtask nodes in the graph structure to represent the potential allocation relationship that the computing nodes can execute the subtasks, and construct a graph structure G0=(V0, E0) containing the computing nodes and subtasks;

[0131] Among them, the structure diagram of the subtask node graph structure is as Figure 2 shown. The node set V0={N1,..., N p , T1,..., T m}, N irepresents the i-th computing node, T j represents the j-th subtask, where p and m represent the number of computing nodes and subtasks respectively; the edge set E0 represents the potential allocation relationship between computing nodes and subtasks, as well as the dependency relationship between subtasks; the weight of the edge is calculated based on the resource capabilities of the computing nodes and the resource requirements of the subtasks, reflecting the matching degree between the two.

[0132] S44. Based on the graph structure in step S3, extract feature vectors for each node in the graph;

[0133] Each node in the said graph includes: a computing node and a subtask node. For the computing node N i , extract the feature vector including: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, current load. Then, for the subtask T j , extract the feature vector including: computational complexity, resource requirements, estimated execution time, task priority, data dependency.

[0134] In this embodiment, the specific content of step S5 is as follows:

[0135] S51. Adopt the graph neural network model GNN to train the graph structure G0 obtained in step S4;

[0136] Among them, during the information transmission process, the hidden state of the node The update expression is as follows:

[0137]

[0138] Among them, k represents the number of iteration layers in the neural network, represents the set of neighbor nodes of node v, W1 and W2 represent trainable weight matrices, b represents the bias term, and σ represents the Sigmoid function.

[0139] Set the loss function L, with the goal of minimizing the task completion time T, balancing the load B of the computing nodes, and maximizing the resource utilization rate U. The expression is as follows:

[0140] L = αT + βB - γU

[0141] Among them, α, β, and γ represent weight coefficients, which are used to balance the importance of various indicators.

[0142] Use backpropagation through the loss function L to update the graph neural network multiple times until the set number of iterations is reached, and complete the training of the graph neural network.

[0143] Among them, the number of iterations is set according to the actual situation. After the iteration ends, all nodes obtain the embedded representation integrating the context information.

[0144] S52. Generate an allocation scheme for subtasks to computing nodes;

[0145] First, input the current graph structure G0 and node features into the GNN model trained in step S51 to obtain the fitness score S of each subtask T j assigned to the computing node N i . Then, according to the obtained fitness score S ij , select the optimal computing node for each subtask. This fitness score is mapped through a parameterized function (such as a multi-layer perceptron, MLP), and the expression is as follows: ij where K represents the total number of iterations,

[0146]

[0147] where, K represents the total number of iterations, and respectively represent the final node representations of the computing node N i and the data block T j , and Θ represents the trainable parameters of the function f.

[0148] Finally, generate an allocation scheme for subtasks to computing nodes {(T ij , N j , N optimal )} according to the fitness score S.

[0149] Among them, the optimal computing node selection method is as follows:

[0150]

[0151] where, N optimal represents the optimal computing node.

[0152] S53. Optimize the utilization rate of computing resources based on step S52;

[0153] First, issue the subtasks to the corresponding computing nodes N i for execution according to the allocation scheme, and during the execution process, continuously monitor the task progress and system status, including: task execution time, node load. If an abnormality is detected or the system status changes, re-invoke the GNN model for inference and adjust the allocation scheme to ensure the stability and efficiency of the system. Iterate until the allocation scheme no longer changes, and complete the optimized deployment of the GNN.

[0154] In this embodiment, the specific steps of step S6 are as follows:

[0155] S61. Perform strict chunking on the dataset to be processed obtained in step S2;

[0156] Divide the dataset D to be processed into multiple data chunk sets {D1, D2,..., D n}, where n represents the number of data chunks.

[0157] Among them, the partitioning process performs quantitative analysis based on data scale, data type, resource requirements, and network topology to ensure that the volume and complexity of each data chunk meet the requirements of subsequent parallel processing. To maintain a balance between resource utilization and computing performance in the data chunking scheme, the optimization objective function expression for data chunk partitioning is as follows:

[0158]

[0159] Among them, Load(D i ) represents the load of data chunk D i , and Comm(D i ) represents the communication overhead generated by this data chunk during network transmission. ε and θ represent weight coefficients.

[0160] S62. Uniformly represent the computing nodes and the chunked data chunks as a graph structure;

[0161] Set the computing node set as {M1, M2,..., M q}, and construct a corresponding graph node set V1 = {M1,..., M q , D1,..., D n} for each computing node and data chunk. In this graph, the computing node attribute parameters include: processing capacity, storage space, network bandwidth, computing density, and the data chunk node record information includes: data volume, feature distribution, processing complexity, providing rich structured input for the execution of subsequent data allocation strategies.

[0162] Among them, q represents the number of computing nodes.

[0163] S63. Add connection edges representing potential data chunk allocation relationships between the computing nodes and the data chunk nodes to construct the graph structure G1 = (V1, E1);

[0164] The structure diagram of the data chunk node graph structure is as Figure 2 shown. The weights of the edges are strictly defined based on the computing node characteristics and data chunk requirements to measure the fitness of the computing node for a specific data chunk. The edge weight function w(M x , D y ) expression is as follows:

[0165]

[0166] Among them, Cap(M x ) represents the available computing resources of computing node M x , Net(M x ) represents the network transmission performance index of this node, and γ and δ represent weight coefficients used to balance computing and communication factors.

[0167] In this embodiment, step S7 is specifically as follows:

[0168] S71. Use the graph neural network model GNN to jointly model the global information of computing nodes and data blocks;

[0169] Based on the graph structure G1 = (V1, E1) constructed in step S6, initialize the feature vector x v for each node in G1.

[0170] Among them, each node in G1 includes: computing nodes and data block nodes; computing node attributes include: processing performance, storage capacity, network bandwidth, resource occupancy, and data block node features include: data volume, processing complexity, expected computing overhead; concatenate the computing node attributes and data block node features to obtain the node initialization feature vector x v for each node in G1.

[0171] S72. Use the graph neural network model GNN to perform multiple iterative trainings on the graph structure G1;

[0172] Among them, during the information transfer process, the hidden representation vector of node v′ is calculated as follows:

[0173]

[0174] Among them, k′ represents the number of iterative layers in the neural network, represents the set of neighbor nodes of node v′, W′1 and W′2 represent trainable parameter matrices, b′ represents the bias term, and σ represents the non-linear activation function.

[0175] Among them, the number of iterations is set according to the actual situation. After the iteration ends, all nodes obtain the embedding representation integrating the context information

[0176] S73. Generate an allocation plan from data block nodes to computing nodes;

[0177] After completing the node representation learning, for each data block node D y calculate its fitness score S′ x for being assigned to computing node M xyThis adaptation degree score is mapped through a parametric function (such as a multi-layer perceptron, MLP), and the expression is as follows:

[0178]

[0179] Among them, K′ represents the total number of iterations, and respectively represent the final node representations of the computing node M x and the data block D y . Θ′ represents the trainable parameters of the function f′.

[0180] According to the adaptation degree score S′ xy , select the optimal computing node allocation scheme for each data block. By searching for the computing node M′ j that maximizes S′ xy for each data block D optimal , the optimal allocation is achieved, and the expression is as follows:

[0181]

[0182] Among them, M′ optimal represents the optimal computing node. Allocate the data block D y to the computing node corresponding to M′ optimal , so as to achieve intelligent allocation and efficient parallel processing of data blocks in the distributed system.

[0183] S74. Based on step S73, improve the efficiency of data processing;

[0184] During the actual execution process, continuously monitor the task progress and system status. When changes occur in the data block processing process, node load, or network conditions, re-invoke the GNN model for inference and dynamically adjust the data allocation scheme. Iterate until the allocation scheme no longer changes, and complete the optimized deployment of the GNN.

[0185] In this embodiment, the flowchart of the intelligent allocation decision of the graph neural network is as Figure 3 shown.

[0186] In this embodiment, the specific steps of step S8 are as follows:

[0187] After each computing node completes the assigned subtasks and data processing, report the intermediate results to the coordination node or the master node in a unified format, and set the intermediate result set {R1, R2,..., R e} collected from the data block or subtask distributed computing nodes;

[0188] Among them, each R iIncluding the intermediate data, statistical metrics, or model parameters after corresponding subtask processing; e takes p or q respectively according to the data blocks or subtasks corresponding to the distributed computing nodes.

[0189] Then the coordination node performs integration and fusion operations on the collected intermediate results. According to the actual task requirements and algorithm strategies, an appropriate aggregation function or integration rule is used to merge {R i}. For other types of data integration, the corresponding integration strategy is selected according to the task characteristics and application scenarios.

[0190] After obtaining the final integration result, the coordination node further processes and formats the result according to the output requirements of the task, and performs quality inspection and verification steps to ensure that the final result meets the expected accuracy and performance standards.

[0191] In summary, the method of the present invention intelligently allocates tasks and data to each node for parallel processing according to the resource status and task requirements of the computing nodes through automated task subdivision and data chunking, without complex model training or a large amount of historical data. It has the advantages of simple implementation and wide applicability, can effectively improve the resource utilization rate and computing efficiency of the distributed system, and is applicable to large-scale data processing and complex computing tasks.

[0192] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. For those skilled in the art, various changes and modifications can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the scope of the claims of the present invention.

Claims

1. A distributed multi-node algorithm and data automatic parallel method, the specific steps are as follows: S1. Obtain the required data and tasks through the user's input information, and clarify the task objectives, data types, and processing requirements; S2. Preprocess the data obtained in step S1, and perform feature extraction and selection; Among them, After the data preprocessing is completed, if the calculation task is divided, transfer to step S3, and if data chunking is performed, transfer to step S6; S3. Divide the overall calculation task into multiple independent subtasks to ensure that each subtask has a clear function and executability; S4. Based on step S3, represent the computing nodes and the subtasks to be allocated in the distributed system as a graph structure to provide a structured representation for intelligent allocation; S5. Based on step S4, use the graph neural network GNN to intelligently allocate subtasks to different computing nodes for execution, and optimize the utilization rate of computing resources; S6. Reasonably chunk the dataset to be processed obtained in step S2, and represent the computing nodes and data chunks as a graph structure to lay the foundation for the intelligent allocation of data; The specific steps of step S6 are as follows: S61. Perform strict chunking on the dataset to be processed obtained in step S2; Divide the dataset D to be processed into multiple data block sets {D1, x2,..., D n}, where n represents the number of data blocks; Among them, the partitioning process is quantitatively analyzed according to the data scale, data type, resource requirements, and network topology; the optimization objective function expression for data chunk partitioning is as follows: Among them, Load(D i ) represents the load of data block D i , Comm(D i ) represents the communication overhead generated by this data block during network transmission, and ε and θ represent weight coefficients; S62. Uniformly represent the computing nodes and the data chunks after chunking as a graph structure; Set the set of computing nodes as {M1, M2,..., M q}, and construct a corresponding set of graph nodes V1 = {M1,..., M q , D1,..., D n} for each computing node and data block; in this graph, the attribute parameters of the computing nodes include: processing power, storage space, network bandwidth, computing density, and the information recorded by the data block nodes includes: data volume, feature distribution, processing complexity; Among them, q represents the number of computing nodes; S63. Add connection edges representing potential data chunk allocation relationships between the computing nodes and the data chunk nodes to construct the graph structure G1=(V1,E1); Edge weight function w(M x ,D y ) is expressed as follows: Among them, Cap(M x ) represents the available computing resources of computing node M x , Net(M x ) represents the network transmission performance index of this node, and γ and δ represent weight coefficients; S7. Based on step S6, use the graph neural network GNN to intelligently allocate each data chunk to different nodes for parallel processing to improve the efficiency of data processing; S8. After each node completes the subtask and data processing, aggregate the intermediate results to the coordination node or the master node for further integration and processing to generate the final result.

2. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S1 are as follows: S11. Authenticate the user to ensure that they have the permission to execute relevant tasks, and at the same time, according to the user's permissions, determine the data scope and task types that they can access; S12. Provide a friendly user interface or API interface to allow the user to input or upload task requirements and obtain a detailed description of the task accordingly, including: task objectives, expected outputs, and special requirements; S13. Confirm the source of the data to be processed, including: local upload, database access, or third-party API, and obtain the data type and format information, including: structured data, unstructured data, images, or audio; S14. Collect the user's specific requirements for data processing, including: the algorithms or models to be used, preprocessing requirements, and parameter settings, and confirm the priority, execution time, and resource limitations of the task.

3. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that, The specific steps of step S2 are as follows: S21. Clean the data to be processed obtained in step S1, remove the noise, errors, and duplicate data therein, handle the missing values and outliers, and ensure the quality and integrity of the data; S22. Perform format conversion and standardization processing on the data after cleaning in step S21 to unify data from different sources and in different formats into a consistent format and unit. Among them, the format conversion and standardization processing include: converting the data type, unifying the time and date format, and standardizing the numerical range. S23. According to the requirements of the task, perform feature extraction and selection on the data processed in steps S21 and S22. For structured data, extract statistical features; for unstructured data, namely text or images, use TF-IDF or convolutional neural network CNN to extract high-order semantic features respectively, and screen out key features highly correlated with the target variable through correlation analysis and dimensionality reduction techniques; at the same time, use statistical methods to detect features with no variance and multicollinearity to eliminate redundant features.

4. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S3 are as follows: S31. Analyze the overall computing task, find out the functional modules or steps that can be independently processed, and then identify the core components of the task, and determine the parts that can be executed as independent subtasks. S32. According to the functional modules, divide the overall computing task into multiple independent subtasks, and clarify their specific functions and goals for each subtask to ensure that they can be executed independently and are logically complete. S33. Determine the input, output, and processing boundaries for each subtask, as well as the dependencies between other subtasks, and ensure the smooth flow of data between subtasks.

5. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S4 are as follows: S41. Collect the resource information of each computing node in the distributed system, and represent each computing node as a node in the graph structure, and each computing node carries its resource attributes and status information. Among them, the resource information of each computing node includes: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, and the current load situation. S42. Represent the subtasks divided in step S3 as nodes in the graph structure as well, and each subtask node includes its resource requirements, estimated execution time, task priority, and possible dependencies. Among them, the resource requirements of each subtask node include: the number of required CPUs, the number of required GPUs, and the memory size. S43. Add edges connecting the computing nodes and subtask nodes in the graph structure to represent the potential allocation relationship that the computing nodes can execute the subtasks, and construct a graph structure G0=(V0,E0) containing computing nodes and subtasks. Among them, the node set V0 = {n1,..., N p , T1,..., t m}, N i represents the i-th computing node, T j represents the j-th sub-task, p and m respectively represent the number of computing nodes and sub-tasks; the edge set E0 represents the potential allocation relationship between the computing nodes and the sub-tasks, as well as the dependency relationship between the sub-tasks; S44. Based on the graph structure in step S3, extract feature vectors for each node in the graph. Each node in the figure includes: a computing node and a subtask node; for computing node N i , extract the feature vector including: CPU performance, GPU performance, memory capacity, storage space, network bandwidth, current load; then, for subtask T j , extract the feature vector including: computational complexity, resource requirements, estimated execution time, task priority, data dependency.

6. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S5 are as follows: S51. Use the graph neural network model GNN to train the graph structure G0 obtained in step S4. Among them, during the information transmission process, the hidden state of the node The update expression is as follows: where k represents the iteration layer number in the neural network, represents the set of neighbor nodes of node v, W1 and W2 represent trainable weight matrices, b represents the bias term, and σ represents the Sigmoid function; Set the loss function L, with the goal of minimizing the task completion time T, balancing the load B of the computing nodes, and maximizing the resource utilization rate U. The expression is as follows: L = αT + βB - γU Among them, α, β, and γ represent weight coefficients used to balance the importance of each index. Use backpropagation through the loss function L to update the graph neural network multiple times until the set number of iterations is reached to complete the training of the graph neural network. Among them, the number of iterations is set according to the actual situation; after the iteration ends, all nodes obtain the embedded representation integrating the context information S52. Generate an allocation plan for subtasks to computing nodes; First, the current graph structure G0 and node features are input into the GNN model trained in step S51 to obtain each subtask T j Assigned to computing node N i The fitness score S ij ; Then, according to the obtained fitness score S ij , select the best computing node for each subtask, and the fitness score is mapped through a parameterized function, as shown below: where K represents the total number of iterations, and respectively represent the final node representations of computing node N i and data block T j ; Θ represents the trainable parameters of function f; Finally, based on the adaptation score S ij generate an allocation plan of subtasks to computing nodes {(T j , N optimal )}; Among them, the optimal computing node N optimal The selection method is as follows: S53. Optimize the utilization rate of computing resources based on step S52; First, the subtasks are sent to the corresponding computing node N according to the allocation scheme i for execution. During the execution process, continuously monitor the task progress and system status, including: task execution time, node load; if an anomaly is detected or the system status changes, re-invoke the GNN model for inference, adjust the allocation scheme to ensure the stability and efficiency of the system; iterate until the allocation scheme no longer changes, and complete the optimized deployment of the GNN.

7. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S7 are as follows: S71. Use the graph neural network model GNN to jointly model the global information of computing nodes and data blocks; Based on the graph structure G1=(V1, E1) constructed in step S6, initialize the feature vector x for each node in G1 v ; Among them, each node in G1 includes: a computing node and a data block node; the computing node attributes include: processing performance, storage capacity, network bandwidth, resource occupancy, and the data block node features include: data volume, processing complexity, and expected computing overhead; the computing node attributes and the data block node features are concatenated to obtain the node initialization feature vector x v ; S72. Use the graph neural network model GNN to perform multiple iterative trainings on the graph structure G1; Set the hidden representation vector of node v′ It is calculated comprehensively based on its own features and the features of neighboring nodes; the calculation method is as follows: where k′ represents the iteration layer number in the neural network, represents the set of neighbor nodes of node v′, W′1 and W′2 represent trainable parameter matrices, b′ represents the bias term, and σ represents the non-linear activation function; Among them, the number of iterations is set according to the actual situation; after the iteration ends, all nodes obtain the embedded representation integrating the context information S73. Generate an allocation plan for data block nodes to computing nodes; After completing the node representation learning, for each data block node D y calculate its fitness score S' x assigned to the computing node M xy ; this fitness score is mapped through a parameterized function, and the expression is as follows: where K′ represents the total number of iterations, and respectively represent the final node representations of computing node M x and data block D y Θ′ represents the trainable parameters of function f′; According to the adaptation degree score S' xy Select the optimal computing node allocation scheme for each data block; by analyzing each data block D j Find the computing node M' that maximizes S' xy to achieve the optimal allocation, and the expression is as follows: o p timal ​ Among them, M' o p timal represents the optimal computing node; allocate the data block D y to M' o p timal corresponding computing node, so as to realize the intelligent allocation and efficient parallel processing of data blocks in the distributed system; S74. Improve the efficiency of data processing based on step S73; During the actual execution process, continuously monitor the task progress and system status; when the data block processing process, node load, or network condition changes, re-invoke the GNN model for inference and dynamically adjust the data allocation plan; iterate until the allocation plan no longer changes, and complete the optimized deployment of GNN.

8. A distributed multi-node algorithm and data automatic parallel method according to claim 1, characterized in that The specific steps of step S8 are as follows: After each computing node completes the assigned subtasks and data processing, it reports the intermediate results to the coordination node or the master node in a unified format. It is set to collect the intermediate result set {R1, R2,..., R e} from the data block or subtask distributed computing nodes; where each R i includes the intermediate data, statistical metrics, or model parameters after the corresponding subtask processing; e takes p or q according to the data block or subtask corresponding to the distributed computing node, respectively; Then the coordination node performs integration and fusion operations on the collected intermediate results; according to the actual task requirements and algorithm strategies, an appropriate aggregation function or integration rule is used to merge {R i}; for other types of data integration, the corresponding integration strategy is selected according to the task characteristics and application scenarios; After obtaining the final integration result, the coordination node further processes and formats the result according to the output requirements of the task, and performs quality inspection and verification steps to ensure that the final result meets the expected accuracy and performance standards.

Citation Information

Patent Citations

  • Language model parallel reasoning method and system

    CN119089901A

  • Wireless policy optimization method and apparatus

    WO2024065476A1