Communication method, device, system, equipment and medium
By dynamically matching communication controllers and priority scheduling in a multi-node environment, transmission channel conflicts and bandwidth contention problems are resolved, communication efficiency and stability are improved, and smooth system operation and rapid fault recovery are ensured.
Patent Information
- Application Number
- CN202510983586.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-07-16
AI Technical Summary
In a multi-node environment, there are transmission channel conflicts and bandwidth contention issues among the devices in the server. Traditional management component transmission protocols in a dual-node environment face challenges such as incomplete device discovery, high cross-node communication latency, severe bandwidth contention, and low fault recovery efficiency.
By determining the target time period as the target communication controller of the communication period among multiple communication controllers, dual isolation of the communication period and the transmission path is achieved. A dynamic matching and priority scheduling mechanism is adopted to ensure the timely transmission of high-priority messages, optimize the transmission path selection, and reduce conflicts and resource contention.
It has achieved improved communication transmission efficiency and enhanced stability, solved transmission channel conflicts and bandwidth contention problems, and ensured smooth system operation and rapid fault recovery.
Smart Images

Figure CN120492396B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a communication method, device, system, electronic equipment, medium and program product. Background Art
[0002] With the popularization of multi-core processors and heterogeneous computing architectures, the demand for highly reliable communication in multi-node server systems is becoming increasingly urgent. However, traditional single-node services usually have single-channel exclusivity, and the communication mechanism between devices in the server has significant limitations in a multi-node environment.
[0003] In the process of realizing the concept of the present invention, the inventors discovered that there are at least the following problems in the related art: in a multi-node environment, there are transmission channel conflicts and bandwidth contention problems among the devices in the server. Summary of the Invention
[0004] In view of the above problems, the present invention provides a communication method, a communication device, a communication system, an electronic device, a medium and a program product.
[0005] According to one aspect of the present invention, a communication method is provided, comprising: in response to a target time period having been reached, determining a target communication controller from a plurality of communication controllers that uses the target time period as a communication time period, wherein the communication time period is used for the communication controller to communicate with a computing node, which is different from the computing nodes with which the plurality of communication controllers each communicate; in response to the target communication controller receiving a target message from a target computing node, processing the target message based on a message type included in the target message to obtain a message processing result for the message type; and returning the message processing result to the target computing node via the target communication controller.
[0006] Another aspect of the present invention provides a communication system, including: a baseboard management controller, used to execute the above-mentioned communication method; multiple computing nodes, respectively connected to the baseboard management controller for communication, the computing nodes are used to manage multiple computing components, and are used to process message processing results sent by the baseboard management controller.
[0007] Another aspect of the present invention provides a communication device, including: a determination module, for determining, in response to the target time period having been reached, a target communication controller from multiple communication controllers that uses the target time period as a communication time period, wherein the communication time period is used for the communication controller to communicate with a computing node, which is different from the computing nodes with which the multiple communication controllers communicate respectively; a processing module, for processing the target message based on the message type included in the target message in response to the target communication controller receiving a target message from the target computing node, and obtaining a message processing result for the message type; and a sending module, for returning the message processing result to the target computing node through the target communication controller.
[0008] Another aspect of the present invention provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0009] Another aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0010] Another aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0011] According to the communication method of the present invention, upon reaching the target time period, a target communication controller is determined to use the target time period as the communication time period. After receiving the target message sent by the target computing node through the target communication controller, the target communication controller processes the message according to its type and returns the message processing result. Because the multiple communication controllers each communicate with different computing nodes and each communication controller has a different communication time period, dual isolation of the communication time period and the transmission path is achieved, allowing the communication needs of different computing nodes to be distinguished in time and space, at least partially resolving the technical problems of transmission channel conflicts and bandwidth contention, and achieving the technical effects of improved communication transmission efficiency and enhanced stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings.
[0013] Figure 1 An application scenario diagram of a communication method according to an embodiment of the present invention is shown.
[0014] Figure 2 A flow chart of a communication method according to an embodiment of the present invention is shown.
[0015] Figure 3 A schematic diagram of inter-device communication according to an embodiment of the present invention is shown.
[0016] Figure 4 A block diagram of a communication system according to an embodiment of the present invention is shown.
[0017] Figure 5 FIG. 1 is a block diagram of a communication system according to another embodiment of the present invention.
[0018] Figure 6 A structural block diagram of a communication device according to an embodiment of the present invention is shown.
[0019] Figure 7 A block diagram of an electronic device suitable for implementing a communication method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0021] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0023] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0024] During the research process, we discovered that with the prevalence of multi-core processors and heterogeneous computing architectures, the need for highly reliable communication in dual-node server systems is becoming increasingly urgent. While the traditional Management Component Transport Protocol (MCTP) protocol has proven successful in single-node management scenarios, it faces core challenges in dual-node environments, including incomplete device discovery, high cross-node communication latency, severe bandwidth contention, and inefficient fault recovery. Among related technologies, single-channel high-speed Peripheral Component Interconnect Express (PCIe) virtualization solutions suffer from performance losses in multi-node environments, static topology protocols cannot adapt to dynamic network changes, and fixed bandwidth allocation strategies lead to low resource utilization.
[0025] Related single-channel virtualization solutions use software virtualization to simulate multi-channel communication in a multi-node environment. However, software virtualization requires frequent context switching, resulting in reduced throughput and significant performance loss. Furthermore, dual nodes are incompatible, leading to channel address conflicts. Furthermore, the default topology protocol is used, resulting in poor dynamic adaptability. This results in low device discovery rates and a high cross-node device loss rate.
[0026] Therefore, related technologies have problems such as dual-channel conflicts and resource contention, incomplete device discovery, high latency in cross-node communication, uncontrollable delays in critical services, and low fault recovery efficiency.
[0027] In view of this, an embodiment of the present invention provides a communication method, comprising: in response to reaching a target time period, determining a target communication controller from multiple communication controllers that uses the target time period as a communication time period, wherein the communication time period is used for the communication controller to communicate with a computing node, which is different from the computing nodes with which the multiple communication controllers communicate respectively; in response to the target communication controller receiving a target message from a target computing node, processing the target message based on a message type included in the target message to obtain a message processing result for the message type; and returning the message processing result to the target computing node through the target communication controller.
[0028] Figure 1 An application scenario diagram of a communication method according to an embodiment of the present invention is shown.
[0029] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a baseboard management controller 110 and multiple computing nodes 120 .
[0030] The baseboard management controller 110 may be a baseboard management controller (BMC) in a multi-node server. The multi-node server may be a dual-node server. The baseboard management controller may be used to execute the communication method of the embodiment of the present application.
[0031] The baseboard management controller 110 may include a communication controller corresponding to each of the multiple compute nodes. The communication controller is configured to receive target messages sent by the corresponding compute node and return message processing results to the corresponding compute node. For example, communication controller 111 corresponds to compute node 121, and communication controller N112 corresponds to compute node N122. The communication controller may be a PCIe controller integrated into the BMC.
[0032] A computing node may be a central processing unit (CPU). Multiple computing nodes may be isolated from each other by network environment.
[0033] It should be noted that the communication method provided in the embodiment of the present invention can generally be executed by the baseboard management controller 110. Accordingly, the communication device provided in the embodiment of the present invention can generally be set in the baseboard management controller 110. The communication method provided in the embodiment of the present invention can also be executed by a baseboard management controller different from the baseboard management controller 110 and capable of communicating with multiple computer nodes 120 and / or the baseboard management controller 110. Accordingly, the communication device provided in the embodiment of the present invention can also be set in a baseboard management controller different from the baseboard management controller 110 and capable of communicating with multiple computer nodes 120 and / or the baseboard management controller 110.
[0034] It should be understood that Figure 1 The number of baseboard management controllers 110 and computing nodes 120 in the embodiment is merely illustrative. Any number of baseboard management controllers 110 and computing nodes 120 may be provided according to implementation requirements.
[0035] The following will be based on Figure 1 The scene described by Figure 2~Figure 3 The communication method of the application embodiment is described in detail.
[0036] Figure 2 A flow chart of a communication method according to an embodiment of the present invention is shown.
[0037] like Figure 2 As shown, the method includes operations S210 to S230.
[0038] In operation S210 , in response to reaching a target period, a target communication controller is determined from among multiple communication controllers to use the target period as a communication period, wherein the communication period is used for the communication controller to communicate with a computing node different from the computing nodes with which the multiple communication controllers communicate.
[0039] In operation S220 , in response to the target communication controller receiving the target message from the target computing node, the target message is processed based on the message type included in the target message to obtain a message processing result for the message type.
[0040] In operation S230 , the message processing result is returned to the target computing node through the target communication controller.
[0041] The PCIe controller is responsible for managing PCIe bus communications, that is, it is used to implement communication between the BMC and multiple computing nodes.
[0042] In the case where there are two computing nodes, the BMC and the computing nodes may be integrated into a dual-node server.
[0043] The target communication controller that uses the target period as the communication period can be detected regularly and activated, so that the target communication controller can normally send and receive messages during the communication period.
[0044] The configuration space of the target communication controller can be switched according to the clock cycle, for example, the base address register (BAR) is switched, such as BAR0-BAR5, so that the target communication controller is activated by mapping the address of the target communication controller with its own base address in the configuration space and activating its communication permission, that is, it can communicate with the corresponding computing node, thereby realizing the allocation of independent communication windows for computing nodes such as CPU0 and CPU1.
[0045] When the target communication controller receives a target message from a target computing node, it may adopt different processing strategies for the message based on the message type and obtain a message processing result for the message type.
[0046] After obtaining the message processing result, the target communication controller may return the message processing result to the computing node in the current communication period or the next communication period.
[0047] According to an embodiment of the present invention, upon reaching the target time period, a target communication controller is determined that uses the target time period as the communication time period. After receiving the target message sent by the target computing node through the target communication controller, the target communication controller processes the message according to type and returns the message processing result. Because multiple communication controllers each communicate with different computing nodes and each communication controller has a different communication time period, dual isolation of the communication time period and the transmission path is achieved, allowing the communication needs of different computing nodes to be distinguished in time and space, at least partially resolving the technical problems of transmission channel conflicts and bandwidth contention, and achieving the technical effects of improved communication transmission efficiency and enhanced stability.
[0048] According to an embodiment of the present invention, determining a target communication controller using a target period as a communication period from among a plurality of communication controllers may include the following operations.
[0049] The time period position of the target time period in the current clock cycle is matched with the preset time period positions corresponding to the multiple communication controllers, and the target communication controller is determined from the multiple communication controllers.
[0050] Clock cycles can be pre-divided into time slots. Each time slot can be considered a time period, corresponding to a communication controller. For example, in the case of dual computing nodes, each clock cycle is divided into two time slots, Phase 0 and Phase 1, corresponding to the communication windows of the PCIe0 controller and PCIe1 controller, respectively. During Phase 0, the configuration space of the PCIe0 controller is in effect, processing communication requests from CPU0. During Phase 1, the PCIe1 controller switches to processing requests from CPU1.
[0051] According to an embodiment of the present invention, by pre-establishing the preset time slot positions for each communication controller within a clock cycle, the target time slot position in the current clock cycle is dynamically matched with the preset time slot position of each communication controller. This allows for rapid targeting of the target communication controller within the clock cycle, reducing invalid communication attempts. Synchronization conflicts between multiple communication controllers are also avoided. Because the preset time slot position of each communication controller is unique, the matching process ensures that it has exclusive access to communication resources within its dedicated time slot.
[0052] According to an embodiment of the present invention, the base addresses of the plurality of communication controllers are different.
[0053] Different base addresses can be assigned to multiple communication controllers. For example, independent physical base addresses such as 0x12C06000 and 0x12C07000 can be assigned to the PCIe0 controller and PCIe1 controller, respectively. This allows for independent, non-overlapping base addresses to be configured for multiple communication controllers, avoiding address conflicts. Furthermore, since one communication controller corresponds to one compute node, the BMC can determine the identity of the compute node sending the message by identifying the controller base address, thus avoiding the issue of being unable to determine the identity of the sender's CPU.
[0054] According to an embodiment of the present invention, returning the message processing result to the target computing node through the target communication controller may include the following operations.
[0055] Based on the content of the message processing result, the priority of the message processing result is determined; when the priority is determined to be the first priority, the bandwidth to be occupied by the message processing result and the available bandwidth of the transmission channel are determined; when it is determined that the bandwidth to be occupied is greater than the available bandwidth, a target sending task with a second priority is determined from at least one current sending task, and the second priority is lower than the first priority; the target sending task is suspended based on the random waiting time, and a new available bandwidth is determined; when it is determined that the bandwidth to be occupied is less than or equal to the new available bandwidth, the message processing result is returned to the target computing node through the target communication controller.
[0056] The first priority may be a higher priority, and the second priority may be a lower priority than the first priority.
[0057] Data or tasks of the first priority may require higher real-time performance, for example, data or tasks used for key operations such as Platform Level Data Model (PLDM) control instructions and firmware updates.
[0058] Data or tasks of the second priority may require lower real-time performance, such as sensor data collection, log transmission and other non-real-time tasks.
[0059] Fixed bandwidth can be pre-allocated for different priority levels. The higher the priority, the larger the fixed bandwidth. For example, 70% of the bandwidth is allocated to the first priority level, and 30% to the second priority level. If the bandwidth corresponding to the first priority level is insufficient to accommodate the remaining bandwidth for the message processing result, a target sending task with the second priority level is determined from at least one current sending task.
[0060] Fields or instructions included in the content of the message processing result can be matched with preset fields or instructions to determine the message type of the message processing result. The message type is then used to determine the priority of the message processing result. Message types can include real-time type and batch type. Real-time type corresponds to the first priority, and batch type corresponds to the second priority.
[0061] The bandwidth to be occupied by the message processing result can be calculated based on the total data volume of the message processing result and the maximum allowed transmission time, and the available bandwidth can be determined based on the currently occupied bandwidth and the total bandwidth of the transmission channel.
[0062] When it is determined that the bandwidth to be occupied is greater than the available bandwidth, a target sending task with a priority lower than the first priority can be determined from the currently transmitting tasks and the target sending task can be suspended, thereby achieving priority and timely processing of urgent tasks.
[0063] When sending tasks with a tentative target, a random waiting time can be introduced for the low-priority message to avoid channel congestion.
[0064] When it is determined that the priority of the message processing result is the second priority, the message processing result may be added to a priority queue so that tasks are processed in order.
[0065] When it is determined that the bandwidth to be occupied is less than or equal to the available bandwidth, the message processing result may be returned to the target computing node through the target communication controller.
[0066] If it is determined that the bandwidth to be occupied is still greater than the new available bandwidth, a target sending task with a second priority level can be determined again from the remaining at least one current sending task, and subsequent steps can be performed until the bandwidth to be occupied is less than or equal to the new available bandwidth.
[0067] According to an embodiment of the present invention, by determining the priority of a message before transmitting the message, and when it is the first priority but the available bandwidth of the transmission channel is insufficient to transmit the message, the low-priority task is interrupted until the high-priority message is sent, thereby ensuring that high-priority and urgent messages can be sent in a timely manner, thereby ensuring the smooth operation of the system.
[0068] According to an embodiment of the present invention, pausing a target sending task based on a random waiting time may include the following operations.
[0069] Based on the number of conflicts between the message processing result and at least one current sending task, a target random number is determined; based on the target random number and a preset waiting time, a random waiting time is determined as a pause time of the target sending task.
[0070] The number of conflicts may be calculated by a conflict counter. When it is determined that the bandwidth to be occupied by the message processing result is greater than the available bandwidth and there is a target sending task, the number of conflicts is increased by 1 based on the original value.
[0071] The pause duration of the target sending task can be determined according to the following formula (1).
[0072] t=rand(1,N)×T_slot; Formula (1)
[0073] Where N is the number of conflicts, rand(1, N) is the target random number, T_slot is the preset waiting time, and t represents the pause time of the target sending task.
[0074] According to an embodiment of the present invention, the pause duration of the target sending task is dynamically determined by combining the message processing result with the number of conflicts of the current sending task, and the number of conflicts is combined with the preset waiting time, so that the waiting time can be longer when the number of conflicts increases, thereby effectively reducing the probability of continuous conflicts between tasks, and avoiding synchronous congestion of multiple tasks due to fixed waiting time through randomized pause time, so that channel resources can be more evenly distributed in the time dimension, thereby improving overall utilization.
[0075] According to an embodiment of the present invention, based on the message type included in the target message, processing the target message to obtain a message processing result for the message type may include the following operations.
[0076] When the message type is a path determination type, based on the device identifier included in the target message, at least one candidate transmission path matching the device identifier is determined from the preset path topology; based on the multiple transmission nodes included in each of the at least one candidate transmission paths, and the transmission time consumption between the multiple transmission nodes, the target transmission path is determined from the at least one candidate transmission path; and the message processing result is generated based on the target transmission path.
[0077] The target message may be an MCTP message. The preset path topology may be tree-like, including multiple transmission nodes and edges between the transmission nodes. The transmission nodes may represent computing nodes or computing components used for computing nodes. The edges between the transmission nodes may represent the communication relationship between the computing nodes. The edges may have weights, where the weights represent the transmission time between two transmission nodes.
[0078] When the message type is determined to be a path determination type, the routing engine can extract the device identification field of the target message and determine the target transmission path that matches the device identification. The device identification field can be, for example, a Node ID field.
[0079] Based on the multiple transmission nodes included in each candidate transmission path and the transmission time consumption between the multiple transmission nodes, the path characteristics of each candidate transmission path can be analyzed to determine the optimal target transmission path.
[0080] According to an embodiment of the present invention, by adopting device identification matching and matching with preset path topology, the interference of irrelevant paths can be reduced, the accuracy of path matching can be improved, and candidate paths can be evaluated based on transmission nodes and time consumption between nodes, so as to select paths with fewer nodes and shorter time consumption according to needs, thereby shortening message transmission delay and improving communication speed.
[0081] According to an embodiment of the present invention, the transmission node includes a computing node and a computing component for the computing node; the device identification includes a starting device identification and a destination device identification; the starting device identification includes a first component identification of the starting computing component and a first node identification of the computing node for which the starting computing component is used; the destination device identification includes a second component identification of the destination computing component and a second node identification of the computing node for which the destination component is used.
[0082] The computing components can be PCIe devices under the computing nodes. The CPU can manage and communicate with the computing components, such as graphics processing unit (GPU) cards, network cards, and redundant array of independent disks (RAID) cards.
[0083] Both a computing node and a computing component may be a type of device; therefore, the originating device and the destination device involved in the originating device identifier and the destination device identifier may be a computing node or a computing component.
[0084] When the initiating device is a computing node, the first component identifier in the initiating device identifier may be empty; when the destination device is a computing node, the second component identifier in the destination device identifier may be empty.
[0085] According to embodiments of the present invention, by splicing computing components and the computing nodes they are used for, on the one hand, computing nodes and computing components can be bound together, quickly identifying the computing nodes they are used for and clarifying hierarchical relationships. On the other hand, devices can be automatically associated with corresponding PCIe channels, i.e., PCIe controllers.
[0086] For example: CPU0 → PCIe0 controller, CPU1 → PCIe1 controller.
[0087] According to an embodiment of the present invention, based on the device identification included in the target message, determining at least one transmission path matching the device identification from a preset path topology may include the following operations.
[0088] Based on the first node identifier and the second node identifier, multiple initial transmission paths that match both the first node identifier and the second node identifier are determined from a preset path topology; based on the first component identifier and the second component identifier, at least one candidate transmission path that matches the device identifier is determined from the multiple initial transmission paths.
[0089] Since the preset path topology can be tree-like, and there is usually a communication relationship between the computing node and the computing component used for the computing node, the tree branches where the starting device and the destination device are located can be filtered out through the first node identifier and the second node identifier, that is, multiple initial transmission paths can be determined.
[0090] Therefore, by matching the first component identifier and the second component identifier with the corresponding tree structures respectively, the transmission nodes corresponding to the first component and the second component can be quickly determined, and at least one candidate transmission path can be determined.
[0091] According to an embodiment of the present invention, through hierarchical screening from nodes to components, at least one candidate transmission path can be quickly screened out from a preset path topology with a large number of transmission nodes, while reducing a certain amount of calculation.
[0092] According to an embodiment of the present invention, determining a target transmission path from at least one candidate transmission path based on multiple transmission nodes included in each of the at least one candidate transmission path and transmission time consumption between the multiple transmission nodes may include the following operations.
[0093] Based on multiple transmission nodes included in each of the at least one candidate transmission paths, the number of transmission hops of the at least one candidate transmission path is determined; the transmission time consumptions between multiple transmission nodes of the same candidate transmission path are accumulated to obtain the transmission duration of the at least one candidate transmission path; based on the number of transmission hops and the transmission duration of each of the at least one candidate transmission paths, the transmission overhead of the at least one transmission path is determined; based on the transmission overhead of the transmission path, a target transmission path is determined from the at least one transmission path.
[0094] The transmission hop count may be the intermediate transmission nodes among the multiple transmission nodes excluding the starting transmission node and the destination transmission node.
[0095] By accumulating the transmission time between every two transmission nodes in each candidate transmission path, the transmission duration of at least one candidate transmission path can be accurately obtained.
[0096] Weight values corresponding to the number of transmission hops and the transmission time can be set respectively, and the transmission overhead of each candidate transmission path can be obtained by performing weighted summation on the number of transmission hops and the transmission time of each candidate transmission path.
[0097] For example, if the weight of hops is 0.6 and the weight of latency is 0.4, the path cost is 0.6 × hops + 0.4 × latency.
[0098] The weights for transmission hop count and transmission time can be adjusted as needed, enabling dynamic adjustments to transmission cost calculations. For example, if low latency is prioritized over fewer hops in path selection, a higher weight can be assigned to transmission hop count and a lower weight to transmission time. This allows the candidate transmission path with the lowest transmission cost to be selected as the target path during path selection.
[0099] According to an embodiment of the present invention, transmission hop count and transmission duration are used as core evaluation parameters for candidate transmission paths, enabling multi-dimensional quantitative evaluation of transmission paths. The number of hops reflects the node complexity of the path, while the transmission duration reflects the actual transmission efficiency. This avoids path selection bias caused by a single metric and improves the comprehensiveness of path evaluation. Furthermore, transmission overhead is used to screen out paths with fewer hops and shorter transmission durations, reducing data forwarding delays and losses between nodes and improving overall transmission efficiency.
[0100] According to an embodiment of the present invention, the preset path topology is obtained in the following manner.
[0101] Broadcast a component discovery message to multiple computing nodes, wherein the component discovery message includes a node identifier and a time determination instruction, wherein the time determination instruction is used to instruct the computing node to calculate the transmission time between any transmission nodes with communication connections, and the transmission node includes a root node and a leaf node, the root node represents the computing node, and the leaf node represents the computing component or other computing node used for the computing node; in response to receiving a component discovery response returned by multiple computing nodes, obtain the communication relationship and transmission time between the computing node and the multiple computing components used for the computing node, between the multiple computing nodes, and between the multiple computing components from the component discovery response, and construct a preset path topology.
[0102] Component discovery messages can be broadcast to multiple compute nodes simultaneously. Components can send messages that follow the Management Component Transport Protocol (MCTP), which can be MCTP discovery packets. For example, if there are two compute nodes, MCTP discovery packets can be sent to both the PCIe0 controller and the PCIe1 controller simultaneously to trigger device responses.
[0103] Extended fields such as Node ID, Latency, or PathCost can be embedded in the original MCTP packet. The Node ID can be two bits long and identifies the compute node and component, e.g., 0x01 for CPU0 and 0x02 for CPU1. Latency is a timing determination instruction, and Path Cost is a cost determination instruction. It can be four bits long and instructs the CPU to dynamically calculate transmission costs.
[0104] Before transmitting the component sending message packet to the communication controller, multiple communication controllers may be activated simultaneously so that the multiple communication controllers broadcast the component discovery message to the corresponding computing nodes.
[0105] After receiving a component discovery message, a computing node can send a component request to a device that has a communication connection with the computing node, such as a component used by the computing node or other computing nodes other than the computing node, and receive a response packet from the device. The transmission time of the transmission link can be determined based on the time the response packet is sent and the time the computing node receives it. At the same time, the component discovery request can include a forwarding instruction to instruct the device to continue forwarding the component discovery message to other devices that have a communication relationship with the device and to return the response packet to the computing node along the original path.
[0106] Through the above component discovery messages and forwarding of component discovery messages, each computing node can obtain the relationship and transmission time between each transmission node that has direct or indirect communication connections with the computing node, and then summarize and send it to the BMC.
[0107] When the BMC receives component discovery responses from each compute node, it extracts the Node ID field, Path Cost field, or Latency field. It then summarizes and aggregates the communication relationships and transmission times between the compute node and its multiple computing components, the communication relationships and transmission times between multiple compute nodes, and the communication relationships and transmission times between multiple computing components, to obtain a preset path topology.
[0108] In addition, when labeling a transmission node representing a computing component, the target computing component identifier, obtained by concatenating the computing component identifier of the computing component with the computing node identifier used by the computing component, can be used as the identifier of the transmission node. The computing component identifier can be the universally unique identifier (UUID) of the computing component or the component type.
[0109] In some embodiments, the component discovery message may carry a BMC signature.
[0110] In some embodiments, the BMC may pre-calculate the transmission cost of each path in a preset path topology, and after obtaining at least one candidate transmission path each time, quickly determine the target transmission path based on the transmission cost of each candidate transmission path.
[0111] The preset path topology and transmission cost may be stored in a topology database.
[0112] According to an embodiment of the present invention, by broadcasting component discovery messages to multiple computing nodes, the broadcast method can reach all computing nodes at the same time, so that they can synchronously process component discovery messages. Not only can the local computing components of each computing node be discovered, but also the communication connections between computing nodes and the interactive relationships between computing components across nodes can be used to comprehensively capture cross-node communication links and transmission time consumption, ensure that the preset path topology covers computing nodes, computing components and all possible connection relationships between them, and solve the problem of omission of cross-node components caused by one-by-one communication in the unicast mechanism, resulting in incomplete path topology information.
[0113] According to an embodiment of the present invention, obtaining the communication relationship and transmission time between a computing node and multiple computing components used for the computing node, between multiple computing nodes, and between multiple computing components from a component discovery response, and constructing a preset path topology may include the following operations.
[0114] Through multiple sub-threads respectively used for multiple computing nodes, the communication relationship and transmission time are obtained from the component discovery response and stored in the respective buffers; through the main thread, the communication relationship and transmission time are respectively obtained from the respective buffers of at least two sub-threads, and a preset path topology is constructed.
[0115] The main thread can be used for global coordination, such as managing the pre-defined path topology, resource scheduling, and error recovery. It maintains a topology database in shared memory, recording and updating the node ID, transmission time, and link status of each device. In some embodiments, the main thread can maintain the pre-defined path topology and record the optimal path.
[0116] Subthreads: These can correspond to multiple compute nodes and are used for transactions on the corresponding compute nodes. For example, the PCIe0 subthread is bound to the PCIe0 channel and handles MCTP data transmission and reception for CPU0. The PCIe1 subthread is bound to the PCIe1 channel and handles MCTP data transmission and reception for CPU1.
[0117] An independent ring buffer can be allocated to each sub-thread, such as Buffer0 of PCIe0 sub-thread and Buffer1 of PCIe1 sub-thread, for temporarily storing received and to-be-sent data.
[0118] Child threads can exchange data with the main thread via a lock-free ring buffer, avoiding thread contention. For example, a child thread will obtain the communication relationship and transmission time from the component discovery response and store it in its own ring buffer. During processing, the main thread can obtain information about the communication relationship and transmission time from each child thread's ring buffer and build a pre-set path topology.
[0119] The read and write pointers of the buffer can be synchronized through atomic operations to ensure data consistency in high-concurrency scenarios.
[0120] In some embodiments, before constructing the preset path topology, sub-threads corresponding to respective computing nodes may be employed to simultaneously perform link training, thereby suggesting a physical layer connection.
[0121] Link training can be implemented using the PCIe physical (PHY) layer interface in the BMC. For example, the PCIe PHY interface detects voltage changes on differential signal lines to identify the device at the other end of the link, such as the PCIe device corresponding to CPU0 or CPU1, and confirm whether a physical connection exists. Alternatively, the PCIe controller sends a training sequence to the CPU through the PCIe PHY interface, allowing both parties to negotiate transmission rates, channel widths, and other parameters.
[0122] In some embodiments, when communicating across nodes, messages must first be sent to the BMC, which then forwards them to the other node. For example, when device CPU0 needs to access device CPU1, the PCIe0 child thread forwards the received message to the PCIe1 child thread, which then forwards it to CPU1 via the PCIe1 controller.
[0123] According to an embodiment of the present invention, a master-slave thread architecture is employed. Independent sub-threads are assigned to multiple computing nodes. Communication relationships and transmission times are concurrently retrieved from component discovery responses and stored in their respective buffers. The master thread then aggregates the data from the multiple sub-thread buffers to construct a pre-defined path topology. This enables parallel processing of cross-node device discovery, avoiding the inefficiencies caused by single-thread bottlenecks. Furthermore, by storing data in independent sub-thread buffers, thread contention is reduced and data consistency is guaranteed in high-concurrency scenarios, improving device discovery integrity and topology construction efficiency.
[0124] According to an embodiment of the present invention, the communication method may further include the following operations.
[0125] In response to triggering the first periodic task, link status reading messages are sent to multiple computing nodes respectively, wherein the link status reading messages are used to instruct the computing nodes to read the communication link status of the computing components from the status storage space of the computing components; in response to triggering the second periodic task, heartbeat detection messages are sent to multiple computing nodes respectively, wherein the heartbeat detection messages are used to instruct the computing nodes to perform heartbeat detection on the computing components to obtain the component status; in response to receiving the communication link status and component status of the computing components sent by multiple computing nodes respectively, path status annotation is performed in the preset path topology.
[0126] There is no limitation on the specific execution period of the first periodic task and the second periodic task, which can be determined according to actual needs, and the two can be the same or different.
[0127] The PCIe Link Training and Status State Machine (LTSSM) state machine can be used to implement link training, negotiate link parameters, and detect link status.
[0128] By sending a link status read message to a computing node, the computing node can obtain the communication link status between the computing component and the device with which it is communicating in real time from the state storage space of the PCIe LTSSM state machine of the computing component. The communication link status can include bit error rate, signal loss, etc.
[0129] By sending a heartbeat detection message to a computing node, the computing node can send a heartbeat detection packet to the computing component of the computing node, and the component status of the computing component can be determined by the response and response duration of the computing component.
[0130] By receiving the communication link status and component status of the computing components of each computing node, the path status of the transmission path in the preset path topology and the component status of each computing component can be incrementally marked in the topology database, so that the available links and available components can be quickly determined.
[0131] For example, you can mark computing components as online or offline, and mark path status as faulty or available.
[0132] According to an embodiment of the present invention, a dual-periodic task triggering mechanism is employed. A first-periodic task sends link status read messages to multiple computing nodes to obtain the communication link status of computing components. A second-periodic task sends heartbeat detection messages to obtain component status. Both types of status information are then integrated into a pre-set path topology for annotation. This achieves multi-layered monitoring of both the communication link status and the operational status of computing components, avoiding the limitations of a single detection mechanism. Status information is annotated in real time within the pre-set path topology, dynamically updating the link health status of the topology database. This provides a foundation for subsequent intelligent routing and fast failover, enhancing the reliability and real-time nature of system communications.
[0133] According to an embodiment of the present invention, the communication method may further include the following operations.
[0134] When it is determined that a first computing node among the multiple computing nodes has a faulty component based on the communication link status and component status of the computing components of each of the multiple computing nodes, a task used for processing by the faulty component is bound to a target component managed by a second computing node, wherein the target component has the same function as the faulty component, and the second computing node is determined from the multiple computing nodes other than the first computing node; and a repair message is sent to the first computing node to instruct the first computing node to update and repair the faulty component.
[0135] When it is determined that the communication link status or component status of the target computing component is abnormal, an automatic alarm can be triggered and a backup component or backup link can be switched within a preset time period.
[0136] For example, if there is a link interruption in the PCIe0 channel, such as a faulty component in the computing unit of CPU0, the main thread can migrate the device traffic bound to the PCIe0 channel to the PCIe1 channel.
[0137] The first computing node may be notified via a system management bus (SMBus) to reset or initialize the faulty component and record a fault log.
[0138] After the faulty component recovers, the main thread can resume binding the tasks processed by the faulty component and update the preset path topology.
[0139] According to an embodiment of the present invention, based on the monitoring of the communication link status and component status of computing components across multiple computing nodes, upon determining that a faulty component exists on a first computing node, the faulty component's tasks are bound to the target component with the same function on a second computing node, and a repair message is synchronously sent to the first computing node to trigger an update and repair of the faulty component. This enables rapid task migration and parallel repair of faulty components. Binding tasks to components with the same function ensures service continuity, avoiding system interruptions caused by single-node component failures. The synchronously triggered repair mechanism shortens the recovery cycle of faulty components, thereby reducing the impact of the failure on overall communication efficiency.
[0140] According to an embodiment of the present invention, the target message is processed based on the message type included in the target message to obtain a message processing result for the message type, including: when the target message type is a data transmission type, based on the data source address included in the target message, determining the data to be transmitted; based on the memory address included in the preset descriptor, transferring the data to be transmitted to the user space corresponding to the memory address to obtain a message processing result.
[0141] When the target computing node receives a data transfer instruction related to the computing component under the target computing node, the target computing node can send a target message to the BMC so that the BMC can use the data source address included in the target message through the Direct Memory Access (DMA) engine to obtain the data to be transmitted without passing through the target computing node, and transfer the data to be transmitted to the user space based on the memory address of the user space included in the preset descriptor, thereby reducing memory copy overhead and improving bandwidth utilization.
[0142] The data transmission instruction may be sent by an application program, a BMC, etc. The data source address is the address of the data storage space of the target computing node or the computing component of the target computing node.
[0143] Independent DMA descriptor tables may be allocated for the PCIe0 controller and the PCIe1 controller.
[0144] According to an embodiment of the present invention, based on the data source address in the target message and the user space address of the preset descriptor, direct data transmission to the user space without passing through the target computing node is achieved, which can bypass the intervention of the target computing node, eliminate the memory copy and context switching overhead involved in the computing node in traditional data transmission, and improve data transmission efficiency; at the same time, it releases the computing power resources of the target computing node, allowing it to focus on business processing, further reducing communication delays.
[0145] According to an embodiment of the present invention, determining the data to be transmitted based on the data source address included in the target message may further include the following operations.
[0146] Initial transmission data is obtained from a storage space corresponding to a data source address; and when it is determined that the priority of the initial transmission data is the third priority and the data volume is greater than or equal to a preset threshold, the initial transmission data is used as data to be transmitted.
[0147] The third priority may be the same as or different from the second priority, and may be a lower priority, such as logs, sensor data, and other data with lower real-time requirements.
[0148] According to an embodiment of the present invention, when performing data transmission, if it is determined that the priority of the initial transmission data is low, the transmission can be performed when the data volume of the initial transmission data is greater than or equal to a preset threshold, thereby saving protocol header overhead and reducing the waste of computing resources and communication resources.
[0149] Figure 3 A schematic diagram of inter-device communication according to an embodiment of the present invention is shown.
[0150] like Figure 3 As shown, in the case where the plurality of computing nodes include CPU0 and CPU1, in operation S1, the BMC broadcasts a component discovery message to CPU0. In operation S2, the BMC broadcasts a component discovery message to CPU1.
[0151] After receiving the component discovery message, CPU0 sends a component discovery request to all computing components under CPU0 and receives response packets from the computing components. For example, in operation S3, CPU0 sends a component discovery request to computing component a. In operation S4, computing component a returns a response packet to CPU0.
[0152] In addition, CPU0 aggregates all received response packets to generate a component discovery response, and in operation S5 , CPU0 returns the component discovery response to the BMC.
[0153] After receiving the component discovery message, CPU1 sends a component discovery request to all computing components under CPU1 and receives response packets from the computing components. For example, in operation S6, CPU1 sends a component discovery request to computing component b. In operation S7, computing component b returns a response packet to CPU1.
[0154] In addition, the CPU 1 aggregates all received response packets to generate a component discovery response, and in operation S8 , the CPU 1 returns the component discovery response to the BMC.
[0155] When CPU0 needs to communicate with CPU1, in some embodiments, communication can be forwarded through the BMC. For example, when CPU0 wants to obtain data about CPU1's computing components, in operation S9, CPU0 sends a second message to the BMC instructing it to forward the first message. The second message includes the first message and instructions to forward the first message. After receiving the second message, the BMC forwards the first message to CPU1 in operation S10. The first message may be the data about CPU1's computing components that CPU0 wants to obtain.
[0156] Figure 4 A block diagram of a communication system according to an embodiment of the present invention is shown.
[0157] like Figure 4 As shown, the communication system 400 includes: a baseboard management controller 410 and a plurality of computing nodes 420. The plurality of computing nodes 420 may include computing node 1 421 ... computing node N 422.
[0158] The baseboard management controller 410 is configured to execute the above communication method.
[0159] The plurality of computing nodes 420 are respectively connected to the baseboard management controller 410 for communication. The computing nodes are used to manage the plurality of computing components and to process the message processing results sent by the baseboard management controller 410 .
[0160] Figure 5 FIG. 1 is a block diagram of a communication system according to another embodiment of the present invention.
[0161] like Figure 5 As shown, the communication system includes: a baseboard management controller, multiple computing nodes, and a computing component for each computing node.
[0162] The baseboard management controller may include multiple communication controllers, where the communication controllers are used to communicate with corresponding computing nodes, wherein the multiple communication controllers have different communication controller addresses respectively.
[0163] The baseboard management controller may further include multiple PCIe PHY interfaces, multiple communication controllers such as communication controller 0 and communication controller 1, and multiple computing nodes such as CPU0 and CPU1.
[0164] Each computing node may have multiple computing components, such as computing components 1, ..., computing components n of CPU0 and computing components 1, ..., computing components n of CPU1.
[0165] There is no limitation on the model of the baseboard management controller, for example, it can be any baseboard management controller that can execute the communication method.
[0166] The PCIe PHY interface manages PCIe physical layer signal transmission, such as link training and rate negotiation. The PCIe PHY interface also supports PCIe 4.0 or PCIe 5.0 protocols, providing high-bandwidth communication capabilities.
[0167] The communication controller and the computing nodes can use the Compute ExpressLink (CXL) for transmission, which can achieve protocol compatibility design and multi-node expansion interface.
[0168] Based on the above communication method, the present invention also provides a communication device. Figure 6 The device is described in detail.
[0169] Figure 6 A structural block diagram of a communication device according to an embodiment of the present invention is shown.
[0170] like Figure 6 As shown, the communication device 600 of this embodiment includes a determination module 610 , a processing module 620 and a sending module 630 .
[0171] The determination module 610 is used to determine, in response to reaching the target time period, a target communication controller from multiple communication controllers to use the target time period as a communication time period, wherein the communication time period is used for the communication controller to communicate with the computing node, which is different from the computing node with which the multiple communication controllers communicate respectively.
[0172] The processing module 620 is configured to process the target message based on the message type included in the target message in response to the target communication controller receiving the target message from the target computing node, and obtain a message processing result for the message type.
[0173] The sending module 630 is configured to return the message processing result to the target computing node through the target communication controller.
[0174] According to an embodiment of the present invention, the determination module includes: a matching submodule.
[0175] The matching submodule is used to match the time period position of the target time period in the current clock cycle with the preset time period positions corresponding to the multiple communication controllers, and determine the target communication controller from the multiple communication controllers.
[0176] According to an embodiment of the present invention, the sending module includes: a priority determination submodule, a bandwidth determination submodule, a task determination submodule, a waiting submodule and a return submodule.
[0177] The priority determination submodule is used to determine the priority of the message processing result based on the content of the message processing result.
[0178] The bandwidth determination submodule is configured to determine the bandwidth to be occupied by the message processing result and the available bandwidth of the transmission channel when the priority is determined to be the first priority.
[0179] The task determination submodule is configured to determine, when it is determined that the bandwidth to be occupied is greater than the available bandwidth, a target sending task with a second priority from at least one current sending task, where the second priority is lower than the first priority.
[0180] The waiting submodule is used to suspend the target sending task based on the random waiting time and determine the new available bandwidth.
[0181] The return submodule is used to return the message processing result to the target computing node through the target communication controller when it is determined that the bandwidth to be occupied is less than or equal to the new available bandwidth.
[0182] According to an embodiment of the present invention, the waiting submodule includes: a random number determination unit and a duration determination unit.
[0183] The random number determining unit is used to determine a target random number based on the number of conflicts between the message processing result and at least one currently sent task.
[0184] The duration determining unit is used to determine the random waiting duration based on the target random number and the preset waiting duration as the pause duration of the target sending task.
[0185] According to an embodiment of the present invention, the processing module includes a first processing submodule. The first processing submodule includes: a first path determination unit, a second path determination unit, and a result determination unit.
[0186] The first path determination unit is configured to determine, when the message type is a path determination type, at least one candidate transmission path matching the device identifier from a preset path topology based on the device identifier included in the target message.
[0187] The second path determining unit is configured to determine a target transmission path from at least one candidate transmission path based on a plurality of transmission nodes included in each of the at least one candidate transmission paths and a transmission time consumed between the plurality of transmission nodes.
[0188] The result determination unit is used to generate a message processing result based on the target transmission path.
[0189] According to an embodiment of the present invention, the second path determination unit includes: a hop count determination subunit, an accumulation subunit, a cost determination subunit, and a path determination subunit.
[0190] The hop count determining subunit is configured to determine the transmission hop count of at least one candidate transmission path based on a plurality of transmission nodes included in each of the at least one candidate transmission paths.
[0191] The accumulation subunit is configured to accumulate the transmission times between multiple transmission nodes of the same candidate transmission path to obtain the transmission duration of at least one candidate transmission path.
[0192] The overhead determination subunit is configured to determine the transmission overhead of at least one transmission path based on the number of transmission hops and the transmission duration of each of the at least one candidate transmission paths.
[0193] The path determination subunit is configured to determine a target transmission path from at least one transmission path based on the transmission overhead of the transmission path.
[0194] According to an embodiment of the present invention, the first path determination unit includes: a first screening subunit and a second screening subunit.
[0195] The first screening subunit is configured to determine, based on the first node identifier and the second node identifier, a plurality of initial transmission paths that match both the first node identifier and the second node identifier from a preset path topology.
[0196] The second screening subunit is configured to determine, based on the first component identifier and the second component identifier, at least one candidate transmission path that matches the device identifier from the plurality of initial transmission paths.
[0197] According to an embodiment of the present invention, the communication device 600 includes: a broadcast module and a topology generation module.
[0198] A broadcast module is used to broadcast component discovery messages to multiple computing nodes, wherein the component discovery messages include node identifiers and time determination instructions, wherein the time determination instructions are used to instruct the computing nodes to calculate the transmission time between any transmission nodes with communication connections, and the transmission nodes include root nodes and leaf nodes, the root nodes represent computing nodes, and the leaf nodes represent computing components or other computing nodes used for the computing nodes.
[0199] A topology generation module is used to, in response to receiving component discovery responses returned by multiple computing nodes, obtain from the component discovery responses the communication relationships and transmission times between the computing nodes and the multiple computing components used for the computing nodes, between the multiple computing nodes, and between the multiple computing components, and to construct a preset path topology.
[0200] According to an embodiment of the present invention, the topology generation module includes: a storage submodule and a construction submodule.
[0201] The storage submodule is used to obtain the communication relationship and transmission time from the component discovery response through multiple sub-threads respectively used for multiple computing nodes, and store them in respective buffers.
[0202] The sub-module is constructed to obtain the communication relationship and transmission time from the buffers of at least two sub-threads through the main thread, and to construct a preset path topology.
[0203] According to an embodiment of the present invention, a communication device includes: a first detection module, a second detection module and a labeling module.
[0204] The first detection module is used to send link status reading messages to multiple computing nodes in response to triggering the first periodic task, wherein the link status reading message is used to instruct the computing node to read the communication link status of the computing component from the status storage space of the computing component.
[0205] The second detection module is used to send heartbeat detection messages to multiple computing nodes in response to triggering the second periodic task, wherein the heartbeat detection message is used to instruct the computing node to perform heartbeat detection on the computing component to obtain the component status.
[0206] The marking module is used to mark the path status in the preset path topology in response to receiving the communication link status and component status of the computing component respectively sent by multiple computing nodes.
[0207] According to an embodiment of the present invention, a communication device includes: a binding module and a repair module.
[0208] A binding module is used to bind tasks used for processing by a faulty component to a target component managed by a second computing node when it is determined that a first computing node among the multiple computing nodes has a faulty component based on the communication link status and component status of the computing components of the multiple computing nodes, wherein the target component has the same function as the faulty component and the second computing node is determined from computing nodes other than the first computing node among the multiple computing nodes.
[0209] The repair module is used to send a repair message to the first computing node to instruct the first computing node to update and repair the faulty component.
[0210] According to an embodiment of the present invention, the processing module: the second processing submodule comprises: a transmission data acquisition unit and a transmission unit.
[0211] The transmission data acquisition unit is used to determine the data to be transmitted based on the data source address included in the target message when the target message type is a data transmission type.
[0212] The transmission unit is used to transmit the data to be transmitted to the user space corresponding to the memory address based on the memory address included in the preset descriptor, and obtain a message processing result.
[0213] According to an embodiment of the present invention, the transmission data acquisition unit includes: an initial data acquisition subunit and a transmission data determination unit.
[0214] The initial data acquisition subunit is used to acquire the initial transmission data from the storage space corresponding to the data source address.
[0215] The transmission data determining unit is configured to use the initial transmission data as data to be transmitted when it is determined that the priority of the initial transmission data is the third priority and the data volume is greater than or equal to a preset threshold.
[0216] According to embodiments of the present invention, any multiple modules among the determination module 610, the processing module 620, and the sending module 630 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the determination module 610, the processing module 620, and the sending module 630 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of these. Alternatively, at least one of the determination module 610, the processing module 620, and the sending module 630 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0217] Figure 7 A block diagram of an electronic device suitable for implementing a communication method according to an embodiment of the present invention is shown.
[0218] like Figure 7 As shown, an electronic device 700 according to an embodiment of the present invention includes a processor 701, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 702 or programs loaded from a storage unit 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0219] The RAM 703 stores various programs and data required for the operation of the electronic device 700. The processor 701, ROM 702, and RAM 703 are connected to each other via a bus 704. The processor 701 executes the programs in the ROM 702 and / or RAM 703 to perform the various operations of the method flow according to the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 702 and RAM 703. The processor 701 may also execute the programs stored in one or more memories to perform the various operations of the method flow according to the embodiment of the present invention.
[0220] According to an embodiment of the present invention, electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to bus 704. Electronic device 700 may also include one or more of the following components connected to I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a LAN card or modem. Communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to I / O interface 705 as needed. Removable media 711, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 710 as needed, so that computer programs read from the removable media can be installed into storage section 708 as needed.
[0221] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0222] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above, and / or one or more memories other than ROM 702 and RAM 703.
[0223] The embodiments of the present invention further include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the communication method provided by the embodiments of the present invention.
[0224] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when the computer program is executed by the processor 701. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0225] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 709, and / or installed from a removable medium 711. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0226] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the processor 701, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0227] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0228] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0229] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.
[0230] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A communication method, characterized in that: The method comprises: In response to a target time period being reached, determining a target communication controller from a plurality of communication controllers to use the target time period as a communication time period, wherein the communication time period is for the communication controller to communicate with a computing node different from the computing nodes with which the plurality of communication controllers are each communicating; In response to the target communication controller receiving a target message from a target computing node, processing the target message based on a message type included in the target message to obtain a message processing result for the message type; Returning the message processing result to the target computing node through the target communication controller; The processing of the target message based on the message type included in the target message to obtain a message processing result for the message type includes: In a case where the message type is a path determination type, based on the device identifier included in the target message, at least one candidate transmission path matching the device identifier is determined from a preset path topology; based on a plurality of transmission nodes included in each of the at least one candidate transmission paths and transmission times between the plurality of transmission nodes, a target transmission path is determined from the at least one candidate transmission path; and the message processing result is generated based on the target transmission path. When the message type is a data transmission type, the data to be transmitted is determined based on the data source address included in the target message; based on the memory address included in the preset descriptor, the data to be transmitted is transferred to the user space corresponding to the memory address to obtain the message processing result.
2. The method according to claim 1, characterized in that The step of determining, from a plurality of communication controllers, a target communication controller that uses the target time period as a communication time period comprises: The time period position of the target time period in the current clock cycle is matched with the preset time period positions corresponding to the multiple communication controllers, and the target communication controller is determined from the multiple communication controllers.
3. The method according to claim 1 or 2, characterized in that The returning the message processing result to the target computing node through the target communication controller includes: Determining the priority of the message processing result based on the content of the message processing result; When it is determined that the priority is the first priority, determining a bandwidth to be occupied by the message processing result and an available bandwidth of a transmission channel; In a case where it is determined that the bandwidth to be occupied is greater than the available bandwidth, determining a target sending task with a second priority from at least one current sending task, where the second priority is lower than the first priority; Pausing the target sending task based on a random waiting time and determining a new available bandwidth; In a case where it is determined that the bandwidth to be occupied is less than or equal to the new available bandwidth, the message processing result is returned to the target computing node through the target communication controller.
4. The method according to claim 3, characterized in that The pausing the target sending task based on the random waiting time includes: determining a target random number based on the number of conflicts between the message processing result and at least one of the currently sent tasks; Based on the target random number and the preset waiting time, the random waiting time is determined as the pause time of the target sending task.
5. The method according to claim 1 or 2, characterized in that The base addresses of the plurality of communication controllers are different.
6. The method according to claim 1, characterized in that The determining of the target transmission path from at least one of the candidate transmission paths based on a plurality of transmission nodes included in each of the at least one candidate transmission paths and transmission time consumptions between the plurality of transmission nodes includes: determining a transmission hop count of at least one of the candidate transmission paths based on a plurality of transmission nodes included in each of the at least one candidate transmission paths; Accumulating transmission times between multiple transmission nodes of the same candidate transmission path to obtain a transmission duration of at least one candidate transmission path; determining a transmission overhead of at least one transmission path based on a number of transmission hops and a transmission duration of each of the at least one candidate transmission paths; A target transmission path is determined from at least one of the transmission paths based on the transmission costs of the transmission paths.
7. The method according to claim 1, characterized in that The transmission node includes a computing node and a computing component used for the computing node; the device identifier includes a starting device identifier and a destination device identifier; the starting device identifier includes a first component identifier of the starting computing component and a first node identifier of the computing node used by the starting computing component; the destination device identifier includes a second component identifier of the destination computing component and a second node identifier of the computing node used by the destination computing component.
8. The method according to claim 7, characterized in that The determining, based on the device identifier included in the target message, at least one transmission path matching the device identifier from a preset path topology includes: Based on the first node identifier and the second node identifier, determining, from the preset path topology, a plurality of initial transmission paths that match both the first node identifier and the second node identifier; At least one candidate transmission path matching the device identification is determined from a plurality of the initial transmission paths based on the first component identification and the second component identification.
9. The method according to claim 1, characterized in that The preset path topology is obtained in the following manner: Broadcasting a component discovery message to the plurality of computing nodes, wherein the component discovery message includes a node identifier and a time-consuming determination instruction, wherein the time-consuming determination instruction is used to instruct the computing node to calculate the transmission time between any transmission nodes with communication connections, wherein the transmission nodes include a root node and a leaf node, the root node represents the computing node, and the leaf node represents a computing component or other computing node for the computing node; In response to receiving component discovery responses returned by multiple computing nodes, the communication relationships and transmission times between the computing node and multiple computing components used for the computing node, between multiple computing nodes, and between multiple computing components are obtained from the component discovery responses to construct the preset path topology.
10. The method according to claim 9, characterized in that The acquiring, from the component discovery response, the communication relationships and transmission times between the computing node and the plurality of computing components used for the computing node, between the plurality of computing nodes, and between the plurality of computing components, and constructing the preset path topology includes: Obtaining the communication relationship and transmission time from the component discovery response through a plurality of sub-threads respectively used for a plurality of the computing nodes, and storing the communication relationship and transmission time in respective buffers; The communication relationship and the transmission time consumption are respectively obtained from the buffers of at least two sub-threads through the main thread, and the preset path topology is constructed.
11. The method according to claim 1, wherein The method further comprises: In response to triggering the first periodic task, sending a link status read message to each of the plurality of computing nodes, wherein the link status read message is used to instruct the computing node to read the communication link status of the computing component from the status storage space of the computing component; In response to triggering the second periodic task, sending a heartbeat detection message to each of the plurality of computing nodes, wherein the heartbeat detection message is used to instruct the computing node to perform a heartbeat detection on the computing component to obtain a component status; In response to receiving the communication link status and component status of the computing component respectively sent by the plurality of computing nodes, path status annotation is performed in the preset path topology.
12. The method according to claim 11, characterized in that The method further comprises: When it is determined that a first computing node among the plurality of computing nodes has a faulty component based on communication link states and component states of respective computing components of the plurality of computing nodes, binding a task processed by the faulty component to a target component managed by a second computing node, wherein the target component has the same function as the faulty component, and the second computing node is determined from computing nodes other than the first computing node among the plurality of computing nodes; Sending a repair message to the first computing node to instruct the first computing node to update and repair the faulty component.
13. The method according to claim 1, wherein The determining of the data to be transmitted based on the data source address included in the target message includes: Obtaining initial transmission data from a storage space corresponding to the data source address; When it is determined that the priority of the initial transmission data is the third priority and the data volume is greater than or equal to a preset threshold, the initial transmission data is used as the data to be transmitted.
14. A communication system, characterized in that: The communication system comprises: A baseboard management controller, configured to execute the method according to any one of claims 1 to 13; A plurality of computing nodes are respectively connected to the baseboard management controller for communication. The computing nodes are used to manage a plurality of computing components and to process message processing results sent by the baseboard management controller.
15. The communication system according to claim 14, wherein: The baseboard management controller includes: Multiple communication controllers are used to communicate with corresponding computing nodes, wherein the multiple communication controllers have different communication controller addresses respectively.
16. A communication device, characterized in that: The device comprises: a determining module configured to, in response to reaching a target time period, determine, from a plurality of communication controllers, a target communication controller that uses the target time period as a communication time period, wherein the communication time period is used for the communication controller to communicate with a computing node that is different from the computing node with which each of the plurality of communication controllers communicates; a processing module, configured to, in response to the target communication controller receiving a target message from a target computing node, process the target message based on a message type included in the target message to obtain a message processing result for the message type; A sending module, configured to return the message processing result to the target computing node through the target communication controller; The processing module includes: a first processing submodule and a second processing submodule; The first processing submodule includes: a first path determination unit for determining, when the message type is a path determination type, at least one candidate transmission path matching the device identifier from a preset path topology based on the device identifier included in the target message; a second path determination unit for determining, based on a plurality of transmission nodes included in each of the at least one candidate transmission paths and transmission times between the plurality of transmission nodes, a target transmission path from the at least one candidate transmission path; and a result determination unit for generating the message processing result based on the target transmission path. The second processing sub-module includes: a transmission data acquisition unit, which is used to determine the data to be transmitted based on the data source address included in the target message when the message type is a data transmission type; a transmission unit, which is used to transfer the data to be transmitted to the user space corresponding to the memory address included in the preset descriptor to obtain the message processing result.
17. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 13.
18. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Method and system for multi-domain heterogeneous interconnected network management based on PCE (path computation element)
CN105187225A
Distributed computing environment using real-time scheduling logic and time deterministic architecture
CN1303497A