Storage cluster and intelligent computing cluster interconnection communication method based on VRB communication framework
By combining communication interaction requests for memory read handling tasks of intelligent computing clusters and storage clusters in the VRB communication framework, the problem of low communication efficiency under high-frequency tasks is solved, and more efficient communication processing is achieved.
Patent Information
- Application Number
- CN202510161540.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-27
AI Technical Summary
In the communication scenario where smart computing clusters and storage clusters are interconnected, when facing high-frequency startup of memory read and handling tasks, the existing communication methods are under great pressure, which affects the efficiency of communication.
Using the interconnected communication method of the storage cluster and intelligent computing cluster based on the VRB communication framework, by acquiring the memory read and handling tasks of multiple computing nodes, the target tasks for the same storage node and in the same communication interaction stage are determined, and the communication interaction requests of these tasks are merged to reduce duplicate communication and improve efficiency.
By combining communication interaction requests, the frequency of QP resources is reduced, the system load is reduced, the communication efficiency is improved, and the interconnection communication capabilities of the storage cluster and the intelligent computing cluster are improved.
Smart Images

Figure CN120223759A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and particularly to an interconnection communication method between a storage cluster and an intelligent computing cluster based on a VRB communication framework. Background Art
[0002] In the communication scenario of interconnection between an intelligent computing cluster and a storage cluster, the computing nodes in the intelligent computing cluster perform memory reading and transfer tasks on the storage nodes in the storage cluster through memory reading. However, when facing the high-frequency startup of memory reading and transfer tasks, the existing communication methods are under great pressure, affecting the communication efficiency. Summary of the Invention
[0003] In view of the above problems, an interconnection communication method between a storage cluster and an intelligent computing cluster based on a VRB communication framework is proposed to overcome or at least partially solve the above problems, including:
[0004] An interconnection communication method between a storage cluster and an intelligent computing cluster based on a VRB communication framework, the method includes:
[0005] Obtain the memory reading and transfer tasks of multiple computing nodes in the intelligent computing cluster;
[0006] Determine, from the memory reading and transfer tasks, the target memory reading and transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction stage;
[0007] Combine and communicate at least some of the communication interaction requests corresponding to the target memory reading and transfer tasks.
[0008] Optionally, the combining and communicating at least some of the communication interaction requests corresponding to the target memory reading and transfer tasks includes:
[0009] Combine at least some of the communication interaction requests corresponding to the target memory reading and transfer tasks;
[0010] Transmit the combined communication interaction requests to the same storage node in the same batch, and perform batch processing on the combined communication interaction requests at the storage node.
[0011] Optionally, before transmitting the combined communication interaction requests to the same storage node in the same batch and performing batch processing on the combined communication interaction requests at the storage node, it further includes:
[0012] Determine the stage identifier of the communication interaction stage where each communication interaction request is currently located, and set the stage identifier in the communication interaction request;
[0013] Transmitting the merged communication interaction requests to the same storage node in the same batch and performing batch processing on the merged communication interaction requests at the storage node, including:
[0014] Transmitting the merged communication interaction requests to the same storage node in the same batch, and performing batch processing on the merged communication interaction requests at the storage node according to the phase identifier.
[0015] Optionally, the performing batch processing on the merged communication interaction requests at the storage node according to the phase identifier includes:
[0016] Determining the target operation task corresponding to the merged communication interaction requests according to the communication interaction phase indicated by the phase identifier, and performing batch processing on the merged communication interaction requests at the storage node according to the target operation task.
[0017] Optionally, the transmitting the merged communication interaction requests to the same storage node in the same batch includes:
[0018] Obtaining batch transmission data for the merged communication interaction requests;
[0019] Transmitting the batch transmission data to the same storage node through Remote Direct Memory Access.
[0020] Optionally, before merging the communication interaction requests corresponding to at least part of the target memory read and transfer tasks, it further includes:
[0021] Determining the network bandwidth situation, the storage node load situation, and the task attribute information of the target memory read and transfer tasks; wherein, the task attribute information includes task priority;
[0022] Determining at least part of the target memory read and transfer tasks according to the network bandwidth situation, the storage node load situation, and the task attribute information of the target memory read and transfer tasks.
[0023] Optionally, it further includes:
[0024] Obtaining multiple completion confirmation messages; wherein, the completion confirmation message is a message sent by the storage node after the data transmission to the computing node is completed, and the completion confirmation message carries a task identifier and a phase identifier, and the task identifier is the identifier of the memory read and transfer task;
[0025] Batch - sending the multiple completion confirmation messages to the corresponding computing nodes according to the task identifier and the phase identifier.
[0026] A storage cluster and intelligent computing cluster inter - connection communication device based on the VRB communication framework, the device includes:
[0027] A memory read and transfer task acquisition module, configured to acquire memory read and transfer tasks of multiple computing nodes in an intelligent computing cluster;
[0028] A target memory read and transfer task determination module, configured to determine, from the memory read and transfer tasks, target memory read and transfer tasks for the same storage node in a storage cluster and currently in the same communication interaction phase;
[0029] A combined communication module, configured to perform combined communication on communication interaction requests corresponding to at least some of the target memory read and transfer tasks.
[0030] An electronic device, including a processor, a memory, and a computer program stored on the memory and capable of running on the processor, where when the computer program is executed by the processor, the method described above is implemented.
[0031] A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method described above is implemented.
[0032] The embodiments of the present invention have the following advantages:
[0033] In the embodiments of the present invention, by acquiring memory read and transfer tasks of multiple computing nodes in an intelligent computing cluster, determining, from the memory read and transfer tasks, target memory read and transfer tasks for the same storage node in a storage cluster and currently in the same communication interaction phase, and performing combined communication on communication interaction requests corresponding to at least some of the target memory read and transfer tasks, combined communication is achieved for target memory read and transfer tasks of multiple computing nodes for the same storage node and currently in the same communication interaction phase, improving communication efficiency. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the description of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0035] Figure 1 is a schematic diagram of a system architecture provided by some embodiments of the present invention;
[0036] Figure 2 is a flowchart of steps of a method for interconnected communication between a storage cluster and an intelligent computing cluster based on a VRB communication framework provided by some embodiments of the present invention;
[0037] Figure 3It is a flowchart of the steps of another method for interconnected communication between a storage cluster and an intelligent computing cluster based on the VRB communication framework provided by some embodiments of the present invention;
[0038] Figure 4 It is a structural block diagram of a device for interconnected communication between a storage cluster and an intelligent computing cluster based on the VRB communication framework provided by some embodiments of the present invention. Detailed implementation manners
[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] In the communication scenario of interconnected intelligent computing clusters and storage clusters, the Remote Direct Memory Access (RDMA) technology can be adopted. RDMA enables the computing nodes in the intelligent computing cluster to directly access the memory data of the storage nodes in the storage cluster, realizing low-latency and high-bandwidth data transmission, while reducing the participation of the host CPU and significantly improving the system communication efficiency. With its efficient memory data transfer ability, the RDMA technology has become one of the key technologies for interconnected communication between storage clusters and intelligent computing clusters. Of course, other communication methods can also be used for communication.
[0041] In the storage cluster, efficient data sharing and distribution are required among storage nodes, while in the intelligent computing cluster, a large amount of parallel computing and model training are needed. The interconnected communication between these two types of clusters poses extremely high requirements on network bandwidth, latency, throughput, etc. However, when applying RDMA in the communication scenario of interconnected intelligent computing clusters and storage clusters, some challenges still exist, which are as follows:
[0042] Repeated interactions of high-frequency tasks: In RDMA communication, a memory read and transfer task needs to go through multiple communication interactions, such as connection establishment, data transfer request, completion confirmation, etc. Each task needs to independently complete these interaction processes. When the system needs to process multiple high-frequency tasks simultaneously, this repeated interaction will occupy a large amount of QP (Queue Pair, a virtual interface between hardware and software, a resource used in RDMA communication to establish a connection between a source node and a target node, including a send queue and a receive queue) resources, resulting in a decline in communication efficiency.
[0043] Low resource utilization efficiency: Although RDMA can provide low-latency and high-bandwidth transmission capabilities, in high-concurrency scenarios (i.e., the frequent initiation of memory read and transfer tasks), the independent interactions of multiple tasks will cause resource waste. For example, during the communication process between the storage cluster and the intelligent computing cluster, there are many similar interaction stages, and related technologies cannot fully utilize these common characteristics.
[0044] Limited system scalability: As the scales of the storage cluster and the intelligent computing cluster continue to expand, the communication task volume increases exponentially. When dealing with high-frequency tasks, the resource requirements and system overheads of related technologies increase exponentially, which limits the system scalability.
[0045] Specifically, in RDMA communication, a complete memory read and transfer task (the data transfer operation from the source node's memory to the target node's memory in RDMA communication) includes the following key interaction stages (multiple communication steps involved in a data transfer task, such as connection establishment, memory registration, data transfer request, etc.):
[0046] 1. Connection establishment stage
[0047] Function: Establish a QP connection between the source node and the target node to complete communication initialization.
[0048] Technical bottleneck: When high-frequency tasks simultaneously initiate connection establishment requests, it will significantly increase the occupancy of QP resources, resulting in a decrease in resource allocation efficiency. Moreover, frequent connection establishment operations increase network latency, which is more obvious especially in scenarios that require cross-cluster communication.
[0049] Actual impact: In the interconnection between the storage cluster and the intelligent computing cluster, the resource overhead in the connection establishment stage will lead to too long communication initialization time, affecting the task startup efficiency.
[0050] 2. Memory area registration stage:
[0051] Function: Register the memory areas that the source node and the target node need to access into the RDMA network to generate a Memory Key.
[0052] Technical bottleneck: The process of registering memory areas requires the participation of the operating system. When the task concurrency is high, it will cause an increase in the operating system load. The problem of repeated registration of memory areas may waste precious memory resources.
[0053] Actual impact: In a high-concurrency task environment, frequent memory registration operations may lead to memory resource tension, thereby affecting the overall performance of the system.
[0054] 3. Data transfer request stage
[0055] Function: The source node sends a data transfer request to initiate data transfer.
[0056] Technical bottleneck: In high-concurrency scenarios, multiple tasks simultaneously initiating data transfer requests may cause network congestion, increasing the transfer waiting time, and the management complexity of data transfer requests grows linearly with the increase in the number of tasks.
[0057] Actual impact: For the collaboration between the storage cluster and the intelligent computing cluster, the bottleneck in this stage will directly lead to an increase in the latency of memory read and transfer tasks.
[0058] 4. Data Transfer Stage
[0059] Function: Complete the actual transfer of memory data.
[0060] Technical bottleneck: Insufficient bandwidth of the network link may become a bottleneck, especially during the transfer of large amounts of data, and the lack of dynamic routing optimization during data transfer will lead to a decrease in link utilization.
[0061] Actual impact: When multiple tasks share the same link, competition for network bandwidth may lead to a significant reduction in transfer efficiency, affecting task throughput.
[0062] 5. Completion Confirmation Stage
[0063] Function: The source node and the destination node confirm the completion of the memory read and transfer task.
[0064] Technical bottleneck: The Completion Queue (CQ) may become congested under a high task volume, affecting the timeliness of confirmation messages, and overly frequent confirmation operations increase the overall communication overhead.
[0065] Actual impact: For tasks with high real-time requirements, the latency in the completion confirmation stage will have a negative impact on the final execution time of the tasks.
[0066] When completing the above multiple interaction stages, the relevant technologies require each task to independently execute a complete process. When the task volume and concurrency increase geometrically, this design will lead to the following bottlenecks:
[0067] High-frequency interactions cause resource occupation: The independent interaction process of each task consumes additional QP resources, especially in scenarios where tasks are frequently started and completed, and the resource load increases rapidly.
[0068] High repetition in the same stage: There are many identical interaction stages during the execution of different tasks. For example, multiple tasks need to establish connections or send data transfer requests simultaneously, and these stages cannot be effectively merged. The repetition of the interaction process leads to waste of system resources and an increase in task latency.
[0069] In the interconnection scenario between the storage cluster and the intelligent computing cluster, the high-frequency startup of communication tasks makes the system's requirements for low latency and high bandwidth more stringent. Users hope to reduce the repeated overhead of communication during multi-task concurrency, improve the overall efficiency of data transfer tasks, and thereby enhance the communication efficiency of high-frequency tasks.
[0070] Moreover, users need a more efficient mechanism to reduce resource waste. For example, in current RDMA tasks, a large number of tasks have similar communication interaction processes, but lack a merging mechanism, resulting in repeated resource occupation. Solving this problem can significantly improve the utilization rate of hardware resources.
[0071] In the embodiments of the present invention, by applying merged communication in the interaction phase of application tasks in the VRB (V2V RDMA Bandwidth, virtual machine-to-virtual machine remote direct memory access bandwidth) framework, the communication efficiency is improved. Specifically, the embodiments of the present invention analyze the interaction characteristics of multiple tasks, integrate the interaction phases with the same function into one communication operation, and the design of merged communication significantly reduces the usage frequency of QP resources and at the same time reduces the system load caused by repeated communication.
[0072] Among them, the VRB system is an architecture for network data transmission and organization. It is a system related to virtual resource allocation and data grouping, and is used to efficiently process data aggregation and transmission in a network environment including a central node and terminal nodes. In the VRB system network environment, there is a central node and multiple terminal nodes, and some terminal nodes have the need for data aggregation, that is, to transmit the data they own to a specific location for integration and processing.
[0073] Among them, V2V (Video to Video) is a protocol in the visual networking.
[0074] Moreover, to support the efficient management of large-scale tasks, the embodiments of the present invention propose a sub-numbering mechanism. In the merged communication in the task interaction phase, a unique sub-number (i.e., task identifier) is dynamically assigned to each task as an independent identifier of the task, and by adding a sub-number (i.e., phase identifier) to the protocol field, the embodiments of the present invention achieve precise differentiation of multiple task data streams under shared resources, ensuring the independence of tasks and the correct transmission of data. Compared with the related technology, the sub-numbering mechanism reduces the dependence on QP connection resources and significantly improves the scalability of the system.
[0075] In addition, combined with the dynamic characteristics of tasks, the sub-numbering mechanism and the merged communication strategy work together, and the system can support the communication requirements of ultra-large-scale nodes with low management complexity in high-frequency task scenarios, successfully solving the efficiency bottleneck and insufficient resource utilization problems in RDMA high-frequency task communication.
[0076] In the embodiments of the present invention, for the memory read and transfer of the same storage node in the storage cluster by multiple different computing nodes. To improve efficiency, a merge communication strategy is adopted, and different tasks and task phases are managed through a sub-coding mechanism. As Figure 1 , the design framework is as follows:
[0077] 1. Task Scheduler&Resource Allocator: The task scheduling module is responsible for scheduling the tasks of each computing cluster and reasonably allocating resources (computing resources, network bandwidth, memory resources). When multiple clusters access the same storage node, the task scheduling will dynamically manage the resource conflicts and collaborations between the clusters.
[0078] 2. Merge Communication Strategy: In the merge communication strategy module, the system will perform communication merging according to the access requirements of multiple computing nodes to the same storage node simultaneously. The request tasks of multiple computing nodes will be merged into the same memory read and transfer phase process, reducing the repeated transmission overhead and improving the bandwidth utilization rate.
[0079] 3. Sub-Numbering Mechanism: The key role of the sub-coding mechanism in this design is to assign independent sub-codes (including task identifiers and phase identifiers) to the memory read and transfer tasks between each computing node and the target storage node. The sub-codes can identify different phases of the tasks (such as connection establishment, memory area registration, data transmission, etc.).
[0080] 4. Transport&Network Layers: The transport layer is responsible for data transmission through the efficient V2V RDMA protocol, supporting low-latency memory transfer. The network layer is responsible for routing selection, bandwidth allocation, and path optimization to ensure that there are no bottlenecks when data is transmitted from the computing node to the storage node.
[0081] The present invention will be further described below with reference to the accompanying drawings:
[0082] Refer to Figure 2 , which shows a flowchart of the steps of a method for interconnection and communication between a storage cluster and an intelligent computing cluster based on a VRB communication framework provided by some embodiments of the present invention. Specifically, it may include the following steps:
[0083] Step 201, obtain the memory read and transfer tasks of multiple computing nodes in the intelligent computing cluster;
[0084] Step 202: From the memory read transfer tasks, determine the target memory read transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction phase.
[0085] Step 203: Combine and communicate the communication interaction requests corresponding to at least some of the target memory read transfer tasks.
[0086] In the embodiments of the present invention, by obtaining the memory read transfer tasks of multiple computing nodes in the intelligent computing cluster, determining the target memory read transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction phase from the memory read transfer tasks, and combining and communicating the communication interaction requests corresponding to at least some of the target memory read transfer tasks, it realizes the combined communication of the target memory read transfer tasks of multiple computing nodes for the same storage node and currently in the same communication interaction phase, improving the communication efficiency.
[0087] In some embodiments of the present invention, before combining and communicating the communication interaction requests corresponding to at least some of the target memory read transfer tasks, it further includes: determining the network bandwidth situation, the storage node load situation, and the task attribute information of the target memory read transfer tasks; determining at least some of the target memory read transfer tasks according to the network bandwidth situation, the storage node load situation, and the task attribute information of the target memory read transfer tasks.
[0088] Wherein, the task attribute information includes task priority.
[0089] In the Task Scheduling&Resource Allocation phase, the main task of the task scheduling and resource allocation module is to reasonably plan and identify the memory read transfer requests of multiple computing nodes for the same storage node, ensuring efficient and parallel communication in one communication. The core goal is to cover different task stages of multiple computing nodes in one communication through reasonable task division and scheduling, ensuring efficient transmission and reducing management complexity.
[0090] The task scheduling module analyzes the memory read tasks of each computing node to identify which computing nodes need to access the same storage node at the same time. The tasks of these computing nodes usually perform memory read operations in the same stage, such as the "memory area registration stage" or the "data transmission stage". If multiple computing nodes are in the same stage (for example, both are performing memory area registration or data transmission requests), the task scheduling module combines these tasks into the same batch. In this case, the system tries to combine the memory read transfer requests of computing nodes for the same storage node and complete the communication interaction of these tasks within one stage, reducing the repeated management overhead.
[0091] Phase Identification and Task Division: Based on the task lifecycle (connection establishment, memory registration, data transfer, completion confirmation, etc.), the task scheduling module divides different computing node requests into phases. The communication requirements for each phase are reasonably planned and optimized according to the task priority and network bandwidth.
[0092] Computing Node Division Strategy: For each phase, the task scheduling module divides which computing nodes will access the storage node at the same time according to the resource allocation rules. The scheduling module plans and ensures that the requests of multiple computing nodes will not cause bandwidth bottlenecks or conflicts within the same phase based on factors such as network bandwidth and storage node load.
[0093] Parallelism and Bulk Communication: During the communication process between the computing nodes and the storage node, the task scheduling module tries to increase the parallelism and process the requests of multiple computing nodes in batches to reduce communication latency and improve bandwidth utilization.
[0094] In practical applications, the memory read and transfer tasks submitted by users include the task identifiers, task priorities, resource requirements (bandwidth, memory area) of each computing node, and the memory area identifiers of the storage nodes to be accessed. The task scheduling module starts to schedule the tasks according to the task priority, computing resources, and network bandwidth requirements. The scheduling module will analyze which computing nodes need to access the same storage node at the same time and classify these tasks into phases (such as connection establishment, memory area registration, data transfer), and the specific operations are as follows:
[0095] 1. Identify which computing nodes are in the same phase and merge the requests of these nodes.
[0096] 2. According to the task lifecycle, determine in which phases each computing node participates in communication. The task is divided into multiple phases, and the requests for each phase are scheduled to the appropriate timing.
[0097] 3. The requests of the computing nodes are reasonably planned based on factors such as network bandwidth and the load of the storage node to ensure the optimal allocation of bandwidth and resources.
[0098] Through the above operations, a task scheduling plan is generated, which plans which computing nodes will access the storage node at a specific phase, and outputs the network resources and task phase information allocated to each computing node.
[0099] In some embodiments of the present invention, the merging communication of the communication interaction requests corresponding to at least part of the target memory read and transfer tasks includes: merging the communication interaction requests corresponding to at least part of the target memory read and transfer tasks; transmitting the merged communication interaction requests in the same batch to the same storage node, and performing batch processing on the merged communication interaction requests at the storage node.
[0100] In some embodiments of the present invention, before transmitting the merged communication interaction requests in the same batch to the same storage node and performing batch processing on the merged communication interaction requests at the storage node, it further includes: determining the stage identifier of the communication interaction stage where each communication interaction request is currently located, and setting the stage identifier on the communication interaction request;
[0101] In the Merge Communication Strategy phase, since multiple computing nodes initiate similar memory read and transfer requests in the same phase, the merge communication strategy combines these requests for processing in one batch. This can not only reduce the independent communication between each computing node, but also improve the data transmission efficiency and bandwidth utilization rate. For example, if multiple computing nodes initiate data requests to the storage node in the "Data Transmission Request Phase", the merge communication strategy combines these requests and performs data transfer simultaneously in batches, avoiding the bandwidth waste caused by each computing node's independent requests.
[0102] In merge communication, the sub-encoding mechanism is used to identify different phases, rather than assigning independent sub-encodings to the requests of each computing node. The requests of each computing node will follow the original task identifier (such as the original RDMA task ID), and different tasks are distinguished by sub-encoding (i.e., stage identifier) for different communication phases (such as "Memory Region Registration", "Data Transmission Request", "Completion Confirmation").
[0103] Each phase (such as connection establishment, memory region registration, data transmission, etc.) will be assigned a unique sub-encoding. This encoding is used to distinguish the requests of different computing nodes in the same phase. Through the sub-encoding mechanism, even if the requests of multiple computing nodes are carried out simultaneously, the system can accurately distinguish the operations of each computing node in different phases. The tasks of each computing node still retain the original identifier (such as the RDMA task ID) to distinguish the priorities between tasks and specific communication contexts. The sub-encoding mechanism is only used for phase differentiation during merge communication to ensure the orderly execution of each phase.
[0104] In practical applications, for the task scheduling plan received from the task scheduling and resource allocation module, tasks are divided into phases, and it is determined which computing nodes will access the storage node in the same phase. The merged communication policy module merges the requests of multiple computing nodes according to the communication phase. When multiple computing nodes initiate similar memory read requests within the same phase, the merged communication policy merges these requests into one batch for processing, specifically including the following operations:
[0105] 1. Phase merging: Merge the requests of multiple computing nodes belonging to the same phase (such as multiple "data transfer request phase" tasks) into one batch, reducing the repeated overhead during independent communication of each computing node and improving bandwidth utilization.
[0106] 2. Merged requests: The merged requests are sent to the sub - encoding mechanism module for identification to ensure that the requests of each computing node can be correctly executed during merged communication.
[0107] Through the above operations, merged communication requests are generated, including the merged task information and sub - encoding. These requests will be passed as input to the sub - encoding mechanism module for processing.
[0108] In practical applications, for the merged requests received from the merged communication policy module. Each request contains the task information of multiple computing nodes, and the merged requests need to identify the phases of each task (such as connection establishment, memory area registration, data transfer, etc.). The sub - encoding mechanism module assigns a unique sub - encoding to each phase. The tasks of each computing node continue to use the original task identifier (such as RDMA task ID). The sub - encoding mechanism distinguishes the phases of tasks by assigning sub - encodings to each phase, specifically including the following:
[0109] 1. Sub - encoding assignment: Each phase is assigned an independent sub - encoding. The sub - encoding is used to distinguish requests in different phases and ensure the isolation between tasks.
[0110] 2. Task identifier: The requests of each computing node will continue to use the original task identifier (for example, RDMA task ID) to distinguish different tasks. The sub - encoding mechanism ensures the correct execution of tasks in different phases by assigning different sub - encodings to each task.
[0111] Through the above operations, the requests of each computing node will be attached with sub - encoding identification, indicating which phase the request is in (such as "data transfer request phase"). These requests with sub - encodings will be passed to the transport layer and network layer modules for processing.
[0112] In some embodiments of the present invention, the step of transmitting the merged communication interaction requests in the same batch to the same storage node includes: obtaining batch transmission data for the merged communication interaction requests; and transmitting the batch transmission data to the same storage node through Remote Direct Memory Access (RDMA).
[0113] In the data transfer request stage, the computing node initiates a data transfer request through a sub-encoding mechanism and maps it to the memory area of the storage node. In a high-frequency task scenario where multiple computing nodes need to access the memory of the storage node, the sub-encoding mechanism ensures that the requests for each task are differentiated according to the stage. In the same stage, the merged communication strategy merges all the requests from the computing nodes that need to make data transfer requests into one batch. Through the sub-encoding identifier, the requests within these batches can be processed separately to ensure that the independent task characteristics of each computing node are not affected.
[0114] In practical applications, for the requests with sub-encoding received from the sub-encoding mechanism module, each request represents a task in a stage, including the data to be transmitted, a task identifier, and a stage identifier (sub-encoding). The transport layer transmits the data through an efficient V2V RDMA protocol and optimizes the data transfer path through the network layer. The network layer is responsible for selecting the best transfer path and performing bandwidth allocation according to the stage of the task (identified by the sub-encoding) and the network load conditions. The specific operations are as follows:
[0115] 1. Transport layer: The transport layer moves the data of the merged requests through the V2V RDMA protocol. When processing the requests, the transport layer uses the sub-encoding mechanism to ensure that the requests from multiple computing nodes in the same stage can be correctly transmitted in parallel.
[0116] 2. Network layer: The network layer ensures that when the data is transmitted from the computing node to the storage node, the network bottleneck is minimized and the bandwidth is reasonably allocated through dynamic route selection. It also optimizes the priority and path selection of the data stream according to the sub-encoding (stage information) of different tasks to ensure that the tasks in different stages are processed in parallel.
[0117] Through the above operations, the data is transmitted from the computing node to the storage node through the RDMA protocol to complete the memory read and transfer task, and the system ensures the correctness of the task stage and the orderliness of the data according to the sub-encoding identifier.
[0118] In some embodiments of the present invention, the step of transmitting the merged communication interaction requests in the same batch to the same storage node and batch-processing the merged communication interaction requests at the storage node includes: transmitting the merged communication interaction requests in the same batch to the same storage node, and batch-processing the merged communication interaction requests at the storage node according to the stage identifier.
[0119] In some embodiments of the present invention, batch processing of the merged communication interaction requests at the storage node according to the phase identifier includes: determining a target operation task corresponding to the merged communication interaction requests according to the communication interaction phase indicated by the phase identifier, and batch processing the merged communication interaction requests at the storage node according to the target operation task.
[0120] In the Storage Node Access Stage, requests from multiple computing nodes are batch processed according to the merged communication policy. Requests from each computing node are phase - differentiated according to its sub - encoding to ensure the orderly execution of the requests. The storage node transmits data according to the merged batch, and the system processes different tasks according to different sub - encoding identifiers. For example, if a computing node is in the data transmission phase, the system can distinguish, according to the sub - encoding, which requests are for "data transmission" and which are for "memory area registration" to ensure that the phase requests are correctly processed.
[0121] In practical applications, requests with sub - encodings received from the transport layer and the network layer have been merged and divided according to task phases. The storage node processes according to the received request batches. Each request contains a sub - encoding used to identify the specific phase of the task. The storage node decides whether to perform "memory area registration", execute "data transmission", or perform "completion confirmation" according to the sub - encoding of the request. The specific steps are as follows:
[0122] 1. Request processing: The storage node divides the requests from multiple computing nodes according to the sub - encoding and performs corresponding operations in the memory area registration phase, data transmission phase, etc. Depending on the phase, the storage node may need to perform tasks such as data transfer and memory mapping.
[0123] 2. Batch transmission: The storage node batch - processes the requests from multiple computing nodes to reduce the independent access overhead of a single computing node. Through the sub - encoding mechanism, the storage node can identify and process requests in different task phases.
[0124] After the storage node finishes processing according to the sub - encoding of the request through the above operations, it returns the processing result, including the confirmation information of the completion of data transmission.
[0125] In some embodiments of the present invention, it further includes:
[0126] Obtain multiple completion confirmation messages; wherein, the completion confirmation message is a message sent by the storage node after the data transmission to the computing node is completed, and the completion confirmation message carries a task identifier and a phase identifier, and the task identifier is the identifier of the memory read and transfer task; according to the task identifier and the phase identifier, batch send the multiple completion confirmation messages to the corresponding computing nodes.
[0127] In the Completion Confirmation Stage, when all data transmission tasks are completed, the storage node will send a completion confirmation message. The task completion confirmations of multiple computing nodes will be merged and processed. The sub-encoding mechanism ensures that each computing node's task can correctly receive its own completion confirmation without being interfered by the tasks of other nodes. The system merges the confirmation information of multiple computing nodes together for batch processing, reducing redundant communication overhead. Through the sub-encoding mechanism, the system can accurately identify which tasks have been completed and which tasks still need to wait for the completion of data transmission.
[0128] In practical applications, after the data transmitted from the storage node to the computing node is completed, the storage node sends a completion confirmation message. In the completion confirmation stage, the storage node sends completion confirmation messages to multiple computing nodes. The system will ensure that each computing node's task is correctly confirmed through the sub-encoding mechanism, specifically including the following steps:
[0129] 1. Confirmation and merging: The completion confirmation messages of multiple computing nodes are batch processed, reducing redundant communication overhead. Each confirmation message carries sub-encoding for identifying specific tasks and phases.
[0130] 2. Task status update: The system updates the status of each task according to the confirmation information to ensure that the completion status of the task is accurately reflected in the scheduling system.
[0131] Through the above operations, the completion confirmation information is batch sent to each computing node, the execution status of the task is updated, and the computing node is notified that the task has been completed.
[0132] Refer to Figure 3 , which shows a flowchart of steps of another method for interconnection and communication between a storage cluster and an intelligent computing cluster based on a VRB communication framework provided by some embodiments of the present invention, specifically including the following steps:
[0133] Step 301, obtain the memory read and transfer tasks of multiple computing nodes in the intelligent computing cluster.
[0134] Step 302, determine the target memory read and transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction phase from the memory read and transfer tasks.
[0135] Step 303: Merge at least some of the communication interaction requests corresponding to the target memory read and transfer tasks.
[0136] Step 304: Determine the stage identifier of the communication interaction stage in which each communication interaction request is currently located, and set the stage identifier in the communication interaction request.
[0137] Step 305: Transmit the merged communication interaction requests to the same storage node in the same batch, and batch process the merged communication interaction requests at the storage node according to the stage identifier.
[0138] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0139] Referring to Figure 4 , a schematic structural diagram of an interconnection communication device between a storage cluster and an intelligent computing cluster based on a VRB communication framework provided by some embodiments of the present invention is shown, which may specifically include the following modules:
[0140] A memory read and transfer task acquisition module 401, configured to acquire memory read and transfer tasks of multiple computing nodes in the intelligent computing cluster;
[0141] A target memory read and transfer task determination module 402, configured to determine, from the memory read and transfer tasks, target memory read and transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction stage;
[0142] A merged communication module 403, configured to perform merged communication on at least some of the communication interaction requests corresponding to the target memory read and transfer tasks.
[0143] Optionally, the performing merged communication on at least some of the communication interaction requests corresponding to the target memory read and transfer tasks includes:
[0144] Merging at least some of the communication interaction requests corresponding to the target memory read and transfer tasks;
[0145] Transmitting the merged communication interaction requests to the same storage node in the same batch, and batch processing the merged communication interaction requests at the storage node.
[0146] Optionally, before transmitting the communication interaction requests to be merged in the same batch to the same storage node and batch-processing the merged communication interaction requests at the storage node, the method further includes:
[0147] Determining a stage identifier of the communication interaction stage in which each communication interaction request currently resides, and setting the stage identifier in the communication interaction request;
[0148] The transmitting the merged communication interaction requests to the same storage node in the same batch and batch-processing the merged communication interaction requests at the storage node includes:
[0149] Transmitting the merged communication interaction requests to the same storage node in the same batch, and batch-processing the merged communication interaction requests at the storage node according to the stage identifier.
[0150] Optionally, the batch-processing the merged communication interaction requests at the storage node according to the stage identifier includes:
[0151] Determining a target operation task corresponding to the merged communication interaction requests according to the communication interaction stage indicated by the stage identifier, and batch-processing the merged communication interaction requests at the storage node according to the target operation task.
[0152] Optionally, the transmitting the merged communication interaction requests to the same storage node in the same batch includes:
[0153] Obtaining batch transmission data for the merged communication interaction requests;
[0154] Transmitting the batch transmission data to the same storage node through remote direct memory access.
[0155] Optionally, before merging the communication interaction requests corresponding to at least part of the target memory read and transfer tasks, the method further includes:
[0156] Determining the network bandwidth condition, the storage node load condition, and the task attribute information of the target memory read and transfer tasks; wherein the task attribute information includes task priority;
[0157] Determining at least part of the target memory read and transfer tasks according to the network bandwidth condition, the storage node load condition, and the task attribute information of the target memory read and transfer tasks.
[0158] Optionally, the method further includes:
[0159] Obtain multiple completion confirmation messages; wherein, the completion confirmation message is a message sent by the storage node after the data transmission to the computing node is completed, and the completion confirmation message carries a task identifier and a phase identifier, and the task identifier is the identifier of the memory read and transfer task;
[0160] According to the task identifier and the phase identifier, batch send the multiple completion confirmation messages to the corresponding computing nodes.
[0161] In the embodiments of the present invention, by obtaining the memory read and transfer tasks of multiple computing nodes in the intelligent computing cluster, from the memory read and transfer tasks, determine the target memory read and transfer tasks for the same storage node in the storage cluster and currently in the same communication interaction phase, and merge the communication interaction requests corresponding to at least some of the target memory read and transfer tasks, it realizes the merged communication of the target memory read and transfer tasks of multiple computing nodes for the same storage node and currently in the same communication interaction phase, improving the communication efficiency.
[0162] Some embodiments of the present invention also provide an electronic device, including a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the above method is implemented.
[0163] Some embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by the processor, the above method is implemented.
[0164] Some embodiments of the present invention also provide a computer program product, including a computer program. When the computer program is executed by the processor, the above method is implemented.
[0165] For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.
[0166] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0167] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, refer to each other.
[0168] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, an apparatus, or a computer program product. Therefore, the embodiments of the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.
[0169] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0170] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0171] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0172] Although the preferred embodiments of the embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
[0173] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the above elements.
[0174] The above provides a detailed introduction to a method, apparatus, electronic device and medium for trunking communication processing. Specific examples are used in this text to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for interconnecting and communicating a storage cluster and an intelligent computing cluster based on a VRB communication framework, characterized in that: The method comprises: Obtain memory read and transfer tasks for multiple computing nodes in the intelligent computing cluster; Determine, from the memory read and carry tasks, a target memory read and carry task for the same storage node in the storage cluster and currently in the same communication interaction phase; The communication interaction requests corresponding to at least part of the target memory read and transfer tasks are merged for communication.
2. The method according to claim 1, characterized in that The step of merging and communicating the communication interaction requests corresponding to at least part of the target memory read and transfer tasks includes: Merging communication interaction requests corresponding to at least some of the target memory read and carry tasks; The merged communication interaction requests are transmitted in a batch to a same storage node, and the merged communication interaction requests are processed in batches at the storage node.
3. The method according to claim 2, characterized in that Before the combined communication interaction requests are transmitted in a batch to a same storage node, and before the storage node processes the combined communication interaction requests in batches, the method further includes: Determine a phase identifier of a communication interaction phase currently in which each communication interaction request is located, and set the phase identifier to the communication interaction request; The step of transmitting the combined communication interaction requests in the same batch to the same storage node, and batch processing the combined communication interaction requests at the storage node, includes: The merged communication interaction requests are transmitted in the same batch to the same storage node, and the merged communication interaction requests are batch processed at the storage node according to the stage identifier.
4. The method according to claim 3, characterized in that The batch processing of the combined communication interaction requests at the storage node according to the phase identifier includes: According to the communication interaction stage indicated by the stage identifier, a target operation task corresponding to the merged communication interaction request is determined, and the merged communication interaction request is batch processed at the storage node according to the target operation task.
5. The method according to claim 2, characterized in that: The step of transmitting the merged communication interaction requests in the same batch to the same storage node includes: obtaining batch transmission data for the merged communication interaction request; The batch transfer data is transmitted to the same storage node via remote direct memory access.
6. The method according to any one of claims 1 to 5, characterized in that: Before merging and communicating the communication interaction requests corresponding to at least part of the target memory read and transfer tasks, the method further includes: Determine the network bandwidth, storage node load, and task attribute information of the target memory read and transfer task; wherein the task attribute information includes task priority; At least part of the target memory read and transfer tasks are determined according to the network bandwidth condition, the storage node load condition, and the task attribute information of the target memory read and transfer tasks.
7. The method according to any one of claims 1 to 5, characterized in that: Also includes: Acquire multiple completion confirmation messages; wherein the completion confirmation message is a message sent after the storage node completes data transmission to the computing node, and the completion confirmation message carries a task identifier and a stage identifier, and the task identifier is an identifier of the memory read and transfer task; The multiple completion confirmation messages are sent in batches to corresponding computing nodes according to the task identifier and the stage identifier.
8. A storage cluster and intelligent computing cluster interconnection communication device based on VRB communication framework, characterized in that: The device comprises: The memory read and transfer task acquisition module is used to obtain the memory read and transfer tasks of multiple computing nodes in the intelligent computing cluster; A target memory read and carry task determination module, used to determine, from the memory read and carry tasks, a target memory read and carry task for the same storage node in the storage cluster and currently in the same communication interaction phase; The merging communication module is used to merge the communication interaction requests corresponding to at least part of the target memory reading and carrying tasks.
9. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the method according to any one of claims 1 to 7 when executed by the processor.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.