Computing device, data communication method, computing device cluster, and chip
Patent Information
- Application Number
- CN202510387508.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-09-29
AI Technical Summary
目前管控面网络架构的硬件成本和维护成本高
[0073]关于第二方面至第六方面中任一种可选的实现方式所带来的技术效果可以参见第一方面或第一方面任一种可选的实现方式所带来的技术效果。此处不再赘述。本申请在上述各方面提供的实现方式的基础上,还可以进一步组合以提供更多实现方式。
Smart Images

Figure CN122845320A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of server technology, and in particular to a computing device, a data communication method, a computing device cluster, and a chip. Background Technology
[0002] A data center, as a data integration system, typically consists of components such as computing systems, storage systems, communication systems, network systems (including multiple network devices), environmental control systems, and security systems. It is commonly used for computing and storing core business data at the company level, or for computing and storing operational data within a company organization. With the continuous development of technology and the emergence of technologies such as cloud computing / distributed computing, computing power and storage have also become information technology services offered to customers. These services can be provided to customers by leveraging the network, information technology, and security capabilities of a data center.
[0003] Network devices, as a critical component of data centers, enable data plane and control plane communication, providing the necessary connectivity and security for application operation and data processing. Control plane communication refers to the administrator's commands, control, and management of various network protocols on network devices, providing the network protocols and routing tables required for data plane communication. Data plane communication involves network devices using network protocols and routing tables to forward data. Currently, control plane network architectures have high hardware and maintenance costs. Summary of the Invention
[0004] This application provides a computing device, a data communication method, a computing device cluster, and a chip to reduce the hardware and maintenance costs of the control plane network architecture.
[0005] Firstly, this application provides a computing device comprising multiple computing units. The different computing units are interconnected via a bus. Each computing unit includes a root node and multiple leaf nodes. When the computing device implements management plane communication, the root node sends management-type messages to the leaf nodes via the bus, these messages carrying a first identifier. When the computing device implements data plane communication, the root node sends service-type messages to the leaf nodes via the bus, these messages carrying a second identifier. The leaf nodes receive either the management-type messages or the service-type messages.
[0006] Based on the first aspect, both management and service messages are forwarded via the bus, allowing management plane communication and data plane communication to reuse a single physical network without the need for additional management plane network deployment, thereby reducing the hardware cost of the management plane network architecture. Furthermore, by using a first or second identifier carried in the first message to distinguish between management and service messages, secure isolation between management plane communication and data plane communication is achieved, ensuring the reliability and rationality of reusing a single physical network for management plane and data plane communication. In addition, setting up a root node for each computing device enables unified management of multiple leaf nodes, reducing the maintenance cost of the management plane network architecture.
[0007] In one alternative implementation, the root node in multiple computing units is set with an identifier; the identifier is a first value; wherein the first value is used to indicate the permission to issue management-type messages.
[0008] Based on this optional implementation, the root node and leaf node are distinguished by whether or not an identifier is set, and the permissions of the root node and leaf node in the management plane communication are configured from the hardware level.
[0009] In one optional implementation, each of the multiple computing units is assigned an identifier. The root node is the computing unit with the identifier having a first value, and the leaf nodes are the computing units with the identifier having a fourth value. The first value indicates that the unit has permission to send management messages; the fourth value indicates that the unit has permission to receive management messages but not permission to send them.
[0010] Based on this optional implementation, the root node and leaf node are distinguished by the value of the identifier, and the permissions of the root node and leaf node in the management plane communication are configured from the hardware.
[0011] In one alternative implementation, the management message carries a first flag bit that indicates the execution of management plane communication. When the first message is a management message, the value of the first flag bit is related to the identifier of the root node.
[0012] Optionally, the value of the first flag bit is the value of the root node's identifier. Optionally, the value of the first flag bit is the encoded value of the root node's identifier.
[0013] Based on this optional implementation, management messages and service messages are distinguished by whether or not they carry a first flag bit. In this way, by differentiating between management messages and service messages, secure isolation is achieved between management plane communication and data plane communication, ensuring the reliability and rationality of reusing a single physical network for both. Furthermore, the value of the first flag bit is related to the value of the root node's identifier, ensuring the security of management plane communication from a hardware perspective.
[0014] In one alternative implementation, business messages and management messages are implemented based on the same data communication protocol, with the first flag bit carried by the extended header field of the data communication protocol.
[0015] Optionally, the header of a management message includes a first identifier and an extension header field. The extension header field carries a first flag bit.
[0016] Optionally, the header of a business message may include a second identifier.
[0017] Based on this optional implementation, management messages and service messages are implemented using the same data communication protocol to ensure that management plane communication can reuse the data plane network. Furthermore, management messages and service messages are distinguished by a first flag bit and an identifier (either a first identifier or a second identifier). Thus, by distinguishing between management messages and service messages, secure isolation between management plane communication and data plane communication is achieved, ensuring the reliability and rationality of reusing a single physical network for both.
[0018] In one alternative implementation, management messages and business messages carry a second flag bit.
[0019] Optionally, the second flag bit of management messages is a second value, and the second flag bit of business messages is a third value.
[0020] The second value is used to indicate the execution of management plane communication, and the third value is used to indicate the execution of data plane communication.
[0021] Optionally, the second value is related to the identifier of the root node.
[0022] In this way, it is possible to determine whether the first message is a management message based on the value of the second flag bit carried in the first message, and thus use the value of the second flag bit to achieve logical isolation between data plane communication and management plane communication.
[0023] In one alternative implementation, business messages and management messages are implemented based on the same data communication protocol, and the second flag bit is carried by the extension header field of the data communication protocol.
[0024] The header of a management message includes: a first extended header field; the first extended header field contains a second value.
[0025] The header of a business message includes: a second extended header field; the second extended header field contains a third value.
[0026] Based on this optional implementation, management messages and service messages are implemented using the same data communication protocol to ensure that management plane communication can reuse the data plane network. The distinction between management messages and service messages is achieved through the value of the second flag bit carried in the first message. In this way, logical isolation between data plane communication and management plane communication can be achieved within the same physical network.
[0027] In an alternative approach, the root node is further specifically used to: receive a first instruction and send a first message to the leaf nodes via the bus, the first message carrying the first instruction.
[0028] Optionally, if the first instruction instructs management plane communication, the first message is a management message.
[0029] Optionally, if the first instruction indicates data plane communication, the first message is a service-type message.
[0030] Based on this optional approach, the root node can implement management plane communication and data plane communication. In the management plane communication process between the root node and the leaf nodes, the data plane communication network of the computing device is reused, without the need to create an additional management plane network, thereby achieving in-band control and reducing the hardware cost of the computing device.
[0031] In one alternative implementation, the root node is further specifically used to: receive the second message and send the first message to the leaf nodes via the bus. The second message and the first message have the same message body, and the first message is a business-type message.
[0032] Optionally, the second message is a business message.
[0033] In this way, the root node can achieve management plane communication and data plane communication through the bus without the need to create an additional management plane network, thus achieving in-band control and reducing the hardware cost of computing devices.
[0034] In one optional implementation, the root node is configured with a first port and a second port. The first port is used to send management messages, and the second port is used to send business messages.
[0035] Based on this optional implementation method, management messages and service messages are distinguished by the port through which the first message is transmitted, thereby achieving isolation between management plane communication and data plane communication and reducing software overhead.
[0036] In one alternative implementation, the computing device further includes a first switching unit. The first switching unit connects the root node and the leaf nodes via a bus. The root node is used to send a first message to the leaf nodes via the bus and the first switching unit. The first switching unit is used to forward the first message to the leaf nodes if the first message is a management message issued by the root node, or if the first message is a service message.
[0037] Based on this optional implementation method, the first switching unit verifies the legitimacy of the first message by identifying whether it is a management message sent by the root node during management plane communication, thereby ensuring the security of management plane communication.
[0038] In one optional implementation, the first message carries a first source address; the first switching unit stores the second source address of the root node. Specifically, the first switching unit is further configured to: determine that the first message is a management message sent by the root node when the first message carries a first identifier and the first source address matches the second source address of the root node.
[0039] Based on this optional implementation method, by judging the source address and message type in two rounds, it can be determined whether the first message is a management message sent by the root node, which can ensure the security of management plane communication.
[0040] In one alternative implementation, the computing device includes a first switching unit and a second switching unit. The first switching unit stores first routing data, and the second switching unit stores second routing data. The first routing data includes the second source address of the root node and the destination address of the first leaf node in the leaf nodes. The second routing data includes the second source address of the root node and the destination address of the second leaf node in the leaf nodes.
[0041] The root node is used to: send a first message to the first leaf node via the bus and the first switching unit, and to send a first message to the second leaf node via the bus and the second switching unit. The first switching unit is used to: forward the first message to the first leaf node using first routing data. The second switching unit is used to: forward the first message to the second leaf node using second routing data.
[0042] Based on this optional implementation, when the root node sends a first message to multiple leaf nodes, the first message can be sent to different leaf nodes through the first switching unit and the second switching unit respectively. This allows the first and second switching units to forward the received message to the corresponding leaf nodes. In this way, the first and second switching units can perform message forwarding in parallel, reducing communication overhead.
[0043] In one alternative implementation, the root node is further configured to: send second routing data to the first switching unit, or send first routing data to the second switching unit. The first switching unit is further configured to: forward a first message to the second switching unit, so that the second switching unit forwards the first message to the first leaf node via the first routing data. The second switching unit is further configured to: forward a first message to the first switching unit, so that the first switching unit forwards the first message to the second leaf node via the second routing data.
[0044] Based on this optional approach, multiple leaf nodes are dual-homed to two switching units. Data is routed synchronously between the first and second switching units, enabling multiple leaf nodes and the root node to form a ring topology through the first and second switching units. Any node can access other leaf nodes through a switching unit (either the first or second switching unit). In the event of a communication link failure between a switching unit (either the first or second switching unit) and the root node, or a communication link failure between a leaf node and a switching unit (either the first or second switching unit), a backup channel can be used for data communication, thereby improving the reliability of data communication within the computing device.
[0045] Secondly, this application provides a data communication method. The computing device includes multiple computing units interconnected via a bus. The multiple computing units include a root node and multiple leaf nodes. In the data communication, the root node sends a first message to the leaf nodes via the bus. This first message includes either a management message or a service message. The leaf nodes receive the first message. Optionally, the first message carries a first identifier or a second identifier. The first identifier indicates that the first message is a management message. The second identifier indicates that the first message is a service message.
[0046] Based on the second aspect, since both management and service messages are forwarded via the bus, management plane communication and data plane communication can reuse a single physical network without the need for additional management plane network deployment, thereby reducing the hardware cost of the management plane network architecture. Furthermore, by distinguishing between management and service messages, secure isolation between management plane communication and data plane communication is achieved, ensuring the reliability and rationality of reusing a single physical network for both. In addition, unified management of multiple leaf nodes by a single root node reduces the maintenance cost of the management plane network architecture.
[0047] In one alternative implementation, the root node in multiple computing units is set with an identifier; the identifier is a first value; wherein the first value is used to indicate the permission to issue management-type messages.
[0048] In one optional implementation, each of the multiple computing units is assigned an identifier. The root node is the computing unit with the identifier having a first value, and the leaf nodes are the computing units with the identifier having a fourth value. The first value indicates that the unit has permission to send management messages; the fourth value indicates that the unit has permission to receive management messages but not permission to send them.
[0049] In one alternative implementation, the management message carries a first flag bit that indicates the execution of management plane communication. When the first message is a management message, the value of the first flag bit is related to the identifier of the root node.
[0050] Optionally, the value of the first flag bit is the value of the root node's identifier. Optionally, the value of the first flag bit is the encoded value of the root node's identifier.
[0051] In one alternative implementation, business messages and management messages are implemented based on the same data communication protocol, with the first flag bit carried by the extended header field of the data communication protocol.
[0052] Optionally, the header of a management message includes a first identifier and an extension header field. The extension header field carries a first flag bit.
[0053] Optionally, the header of a business message may include a second identifier.
[0054] In one optional implementation, both management and business messages carry a second flag bit. Optionally, the second flag bit of the management message has a second value, and the second flag bit of the business message has a third value. The second value indicates the execution of management plane communication, and the third value indicates the execution of data plane communication.
[0055] In one alternative implementation, business messages and management messages are implemented based on the same data communication protocol, and the second flag bit is carried by the extension header field of the data communication protocol.
[0056] The header of a management message includes: a first extended header field; the first extended header field contains a second value.
[0057] The header of a business message includes: a second extended header field; the second extended header field contains a third value.
[0058] In one alternative approach, the root node sends a first message to the leaf node via the bus, including: the root node receiving a first instruction and sending the first message, which carries the first instruction, to the leaf node via the bus.
[0059] Optionally, if the first instruction instructs management plane communication, the first message is a management message.
[0060] Optionally, if the first instruction indicates data plane communication, the first message is a service-type message.
[0061] In one optional implementation, the root node sends a first message to the leaf nodes via the bus, including: the root node receiving a second message and sending the first message to the leaf nodes via the bus. The second message and the first message have the same message body, and the first message is a business-type message.
[0062] Optionally, the second message is a business message.
[0063] In one alternative implementation, the computing device further includes a switching unit. The switching unit connects the root node and leaf nodes via a bus. The root node sends a first message to the leaf nodes via the bus, including: the root node sending the first message to the leaf nodes via the bus and the first switching unit. The first switching unit forwards the first message to the leaf nodes.
[0064] In one optional implementation, the first switching unit forwards the first message to the leaf node, including: the switching unit receiving the first message. If the first message is a management message sent by the root node, the switching unit forwards the first message to the leaf node. If the first message is a service message, the switching unit forwards the first message to the leaf node.
[0065] In one optional implementation, the first message carries a first source address; the switching unit stores a second source address of the root node. The switching unit determines that the first message is a management message sent by the root node if the first message carries a first identifier, the first source address of the first message matches the source address of the root node, the first message carries a first flag bit, and the first flag bit matches a first value.
[0066] In one optional implementation, the switching unit includes a first switching unit and a second switching unit. The first switching unit stores first routing data, and the second switching unit stores second routing data. The first routing data includes the second source address of the root node and the destination address of the first leaf node. The second routing data includes the second source address of the root node and the destination address of the second leaf node. The root node sends a first message to the leaf nodes via the bus, including: the root node sending the first message to the first leaf node via the bus and the first switching unit, and sending the first message to the second leaf node via the bus and the second switching unit. The first switching unit forwards the first message to the first leaf node using the first routing data. The second switching unit forwards the first message to the second leaf node using the second routing data.
[0067] In one optional implementation, the root node sends second routing data to the first switching unit, or sends first routing data to the second switching unit. If the communication link between the first switching unit and the first leaf node fails, the first switching unit forwards a first message to the second switching unit, enabling the second switching unit to forward the first message to the first leaf node using the first routing data. If the communication link between the second switching unit and the second leaf node fails, the second switching unit forwards a first message to the first switching unit, enabling the first switching unit to forward the first message to the second leaf node using the second routing data.
[0068] Thirdly, this application provides a computing device cluster, including: a first computing device and at least one second computing device connected to the first computing device. The second computing device is a computing device provided in the first aspect or any optional implementation thereof. The first computing device is used to set a root node in the at least one second computing device and to send a first instruction to the root node in the at least one second computing device. The root node in the at least one second computing device is used to receive the first instruction and coordinate with the leaf nodes within the second computing device to implement the method provided in the second aspect or any optional implementation thereof.
[0069] In one alternative implementation, the first computing device is further configured to: select a new root node from a plurality of leaf nodes of the target computing device in the event of a root node failure. The target computing device includes at least one second computing device whose root node has failed.
[0070] Fourthly, this application provides a processing chip, including a control circuit and an interface circuit. The interface circuit is used to receive data from other devices outside the processing chip and transmit it to the control circuit, or to send data from the control circuit to other devices outside the processing chip. The control circuit implements the function of the root node in the first aspect or any optional implementation of the first aspect through logic circuits or executed code instructions and the interface circuit. Alternatively, the control circuit implements the function of the leaf node in the first aspect or any optional implementation of the first aspect through logic circuits or executed code instructions and the interface circuit.
[0071] Fifthly, this application provides a computer-readable storage medium, comprising: computer software instructions. When invoked by a computing device, the computer software instructions implement the method described in the second aspect or any of the optional implementations of the second aspect.
[0072] Sixthly, this application provides a computer program product, which, when run on a computing device, executes the method in the second aspect or any optional implementation of the second aspect.
[0073] The technical effects of any of the optional implementations of aspects two through six can be found in the first aspect or any of the optional implementations of the first aspect. Further details are omitted here. Based on the implementations provided in the above aspects, this application can be further combined to provide more implementations. Attached Figure Description
[0074] Figure 1A A schematic diagram of a data center network architecture;
[0075] Figure 1B A schematic diagram of a data center network architecture Figure 2 ;
[0076] Figure 2 A schematic diagram of a data center structure is provided for this application;
[0077] Figure 3 A software schematic diagram of a data center provided for this application;
[0078] Figure 4 A schematic diagram of the structure of a computing device provided in this application;
[0079] Figure 5 A schematic diagram of the structure of a computing device provided in this application Figure 2 ;
[0080] Figure 6 A schematic diagram of the structure of a computing device provided in this application Figure 3 ;
[0081] Figure 7 A flowchart illustrating a data communication method provided in this application;
[0082] Figure 8 A schematic diagram of a message format provided for this application;
[0083] Figure 9 A schematic diagram of a data communication process provided for this application;
[0084] Figure 10 A message format illustration provided for this application Figure 2 ;
[0085] Figure 11 A schematic diagram of a data communication process provided in this application Figure 2 . Detailed Implementation
[0086] Network devices, as a critical component of data centers, typically include switches, routers, and other hardware components. The computing devices within these network devices enable management plane communication, control plane communication, and data plane communication. Management plane communication and control plane communication can be collectively referred to as management plane communication, administrative plane communication, or control plane communication; this application's embodiments do not limit the specific terminology used. The following description uses management plane communication and data plane communication as examples.
[0087] Because control plane communication and data plane communication transmit different data, data centers typically need to set up two independent network architectures, such as... Figure 1A As shown, in the control plane network architecture, the central processing unit (CPU) sends routing table entries and other configurations to the switches via the high-speed peripheral component interconnect express (PCIe) bus. In the data plane network architecture, the switches carry Operation Administration and Maintenance (OMA) messages via Ethernet. In the management plane network architecture, the CPU obtains fault information of the switches and hardware devices (such as computing unit 1, computing unit 2, or computing unit 3) via PCIe, and reports the fault information of the CPU, switches, and hardware devices (such as computing unit 1, computing unit 2, or computing unit 3) to the Baseboard Management Controller (BMC) via the Inter-Integrated Circuit (I2C) bus.
[0088] In data center computing devices where components are connected via a peer-to-peer network, a dedicated unified bus (UB) device management plane network architecture is constructed. Each UB device has a control CPU directly connected via the bus. This bus can be a UB bus, PCIe bus, I2C bus, or other type of bus. Figure 1BAs shown, UB device 1 is connected to control CPU 1 via a bus, UB device 2 is connected to control CPU 2 via a bus, and UB device 3 is connected to control CPU 3 via a bus. The switch is connected to control CPU 4 via a bus. Control CPU 1, control CPU 2, control CPU 3, and control CPU 4 are each connected to the data center's management and control equipment via Ethernet (ETH). In the management plane network architecture, the management and control equipment manages the UB devices (UB device 1, UB device 2, or UB device 3) individually through the control CPUs (control CPU 1, control CPU 2, or control CPU 3). In the data plane network architecture, the switch is connected to UB device 1, UB device 2, and UB device 3 via the UB bus. The switch transmits messages with the UB devices (UB device 1, UB device 2, or UB device 3) via the UB bus, and transmits messages with the data center's management and control equipment via ETH.
[0089] Peer-to-peer (P2P) networks: Communication networks based on peer-to-peer bus connections. In P2P networks, endpoint devices support hot-swapping and can automatically recover and retry network connections, ensuring high availability and fault tolerance. Furthermore, P2P networks can employ various network topologies, such as mesh, star, or other types, to meet the needs of more complex scenarios.
[0090] From the above Figure 1A and Figure 1B It can be seen that, to facilitate the management of computing devices and command issuance in the data center, an additional control plane network architecture needs to be created. Separating the control plane network from the service plane network increases the hardware cost of the control plane network architecture. Furthermore, from... Figure 1B As we know, each hardware device in a computing device has a management node to manage each hardware device individually, which makes it impossible to achieve centralized management of UB devices and results in high maintenance costs.
[0091] Based on this, to reduce the hardware and maintenance costs of the control plane network architecture, this application provides a computing device comprising multiple computing units interconnected via a bus, and each computing unit includes a root node and multiple leaf nodes. Both management and service messages are forwarded via the bus, thus enabling control plane and data plane communication to reuse a single physical network without requiring additional deployment of a control plane network, thereby reducing the hardware cost of the control plane network architecture. Furthermore, by using a first and second identifier to distinguish between management and service messages, secure isolation between control plane and data plane communication is achieved, ensuring the reliability and rationality of reusing a single physical network for both. In addition, each computing device is configured with a root node, enabling unified management of multiple leaf nodes, which further reduces the maintenance cost of the control plane network architecture.
[0092] Specifically, the root node sends management messages to the leaf nodes via the bus to achieve control plane communication of the computing device, or the root node sends service messages to the leaf nodes via the bus to achieve data plane communication of the computing device. The management messages carry a first identifier, and the service messages carry a second identifier.
[0093] The technical solutions involved in this application may be applied not only to current storage technologies, cloud computing technologies, data centers, big data technologies, or computing devices, but also to future storage technologies, cloud computing technologies, big data technologies, or computing devices, or data centers or computing device clusters that include computing devices. The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in this application is provided below.
[0094] A supernode refers to a computing node in a large cluster environment. A supernode typically includes multiple storage boards and multiple computing resource boards (e.g., CPU boards, neural processing unit (NPU) boards, or graphics processing unit (GPU) boards). Storage boards comprise pooled storage media. Computing resource boards (CPU boards, NPU boards, or GPU boards) comprise pooled chips (CPU, NPU, or GPU).
[0095] Resource pooling refers to centralizing the resources of all computing devices in a computing cluster into a single resource pool, which is then shared by different services. These resources include computing resources, storage resources, and network resources.
[0096] Data plane: Used for data transmission and forwarding. For example, it processes and forwards service data, or forwards data packets (or messages) according to instructions generated by the control plane. In some alternative approaches, the data plane may also be called the service plane or other names.
[0097] Control plane: Used to control and manage the network. For example, it enables route configuration or network topology management.
[0098] Management plane: Used for configuration management, performance monitoring, fault management, and security management of network devices. In some optional implementations, the control plane and management plane can be collectively referred to as the control and management plane, used to control and manage the network and its devices. This article uses the control and management plane as an example.
[0099] In-band: Also known as in-band management, this refers to the transmission of management and control information and network-bearing service information through a single logical channel. In this application, in-band can refer to the control plane and data plane sharing a single network architecture.
[0100] Out-of-band: Also known as out-of-band management, this refers to the transmission of management and control information and network-bearing service information through different logical channels. In this application, out-of-band may refer to the use of different network architectures for the control plane and the data plane.
[0101] To make the objectives, technical solutions, and advantages of this application clearer, the application will now be described in further detail with reference to the accompanying drawings.
[0102] In the following description, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0103] Furthermore, in this application, directional terms such as "upper" and "lower" are defined relative to the orientation of the components shown in the accompanying drawings. It should be understood that these directional terms are relative concepts, used for relative description and clarification, and can change accordingly depending on the orientation of the components in the accompanying drawings.
[0104] Figure 2This application provides a schematic diagram of a data center structure. The data center 100 includes a management and control device 10 and multiple computing devices. The management and control device 10 communicates with the multiple computing devices via a network 20. The communication function of the network 20 can be implemented by a switch or a router. In an optional example, the management and control device 10 can also communicate with the multiple computing devices via a wired connection, such as a PCIe bus, compute express link (CXL), universal serial bus (USB) protocol, or other protocol buses.
[0105] Please continue reading. Figure 2 , Figure 2 The control device 10 includes a host 101 and a computing device, represented by a CPU 102, but this should not be construed as limiting the scope of this application. The computing device may include one or more processing units, which can be not only a CPU, but also a data processing unit (DPU), other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor. The computing device may also be a dedicated processor for artificial intelligence (AI), such as a neural processing unit (NPU) or a graphics processing unit (GPU). Physically, the one or more processing units included in the computing device can be packaged into a card, such as... Figure 2 The CPU 102 can be connected to the host 101 via a PCIe interface, CXL interface, UB interface, NVlink interface or other communication interface.
[0106] For example, host 101 is a computer running an application. For instance, if the computer running the application is a physical computing device, the physical computing device could be a server. The embodiments of this application do not limit the specific technology or device form used by host 101.
[0107] In one alternative implementation, the control device 10 can send control commands to the computing device to realize control plane communication between the control device 10 and the computing device, and receive data sent by the computing device, or send data to the computing device to realize data plane communication between the control device 10 and the computing device.
[0108] In one optional scenario, the control device 10 and multiple computing devices are interconnected. Only during control plane communication does the control device 10, acting as the central node (or master node) of the data center 100, send control commands to the computing devices, thus enabling control plane communication between the control device 10 and the computing devices. During data plane communication, the endpoints of the data center 100 are interconnected, allowing different devices (control device 10 or computing devices) to communicate with each other.
[0109] In an alternative scenario, the control device may also be referred to as the first computing device, the main device, or other names, which are not limited in this application.
[0110] Please continue reading. Figure 2 Multiple computing devices in Data Center 100, such as Figure 2 The computing devices 1 to 3 in the middle.
[0111] In the first alternative scenario, the computing device (computing device 1, computing device 2, or computing device 3) can be a physical server, which includes a network interface card (NIC), a memory, a RAM, and multiple computing units. The computing units, memory, NIC, and RAM are connected via a bus. This bus can be PCIe, CXL, an inter-integrated circuit (I2C) interface, a controller area network (MAN) bus, a serial peripheral interface, a queued serial peripheral interface, a full-duplex asynchronous serial interface, or a half-duplex differential serial interface, or other types of buses. This application does not limit the specific type of bus.
[0112] Among them, the Inter-Integrated Circuit (I2C) bus is a source-synchronous serial bus used for short-distance communication between different integrated circuits. I2C uses two lines for data transmission: a serial data line (SDL) and a serial clock line (SCL). The Controller Area Network (CAN) bus is a serial communication protocol bus used for real-time applications; it can use twisted-pair cables to transmit signals. The Serial Peripheral Interface (SPI) bus is a 3-wire synchronous serial full-duplex communication interface, offering advantages such as simple circuitry, high speed, and reliable communication. The Queued Serial Peripheral Interface (QSPI) bus adds a queued transmission mechanism to SPI; QSPI uses a dedicated communication interface to connect single, dual, or four data lines. A full-duplex asynchronous serial interface, also known as a universal asynchronous receiver / transmitter (UART) interface, is a universal serial data bus used for asynchronous communication. This UART bus can be a bidirectional communication bus, converting the data to be transmitted between serial and parallel communication modes. For example, a UART interface refers to an RS-232 interface. A half-duplex differential serial interface is a serial communication bus interface that uses a two-wire, differential transmission, and half-duplex mode, such as an RS-485 interface.
[0113] A computing device includes multiple computing units and memory to provide computing resources. In some alternative embodiments, the computing unit can be a chip, such as a CPU, GPU, DPU, or NPU. In other alternative embodiments, the computing unit can be a system-on-a-chip comprising one or more chips. This application does not limit the specific form of the computing unit.
[0114] A network interface card (NIC) is used to provide network resources. A memory device (RAM) is used to provide storage resources. In some alternatives, the computing unit may also be called a computing node, computing card, service node, chip, or other names.
[0115] In the second alternative scenario, the computing device (computing device 1, computing device 2, or computing device 3) can be a server rack composed of multiple physical servers. After pooling the resources of multiple computing units of the multiple physical servers within the server rack, a computing resource pool is formed; after pooling the storage resources of the multiple physical servers, a storage resource pool is formed; and after pooling the network interface card (NIC) resources of the multiple physical servers, a network resource pool is formed. During business processing, each business can share the computing resources provided by the computing resource pool within the server rack. For example, multiple computing units within the computing resource pool can distribute the processing of each business, or different computing units within the computing resource pool can process different businesses.
[0116] In the third optional scenario, the computing device (computing device 1, computing device 2, or computing device 3) can be a resource pool composed of multiple resource boards. Resource boards include computing resource boards, storage resource boards, or network resource boards. Each computing resource board has multiple computing units, which are networked in a full mesh mode, meaning that each computing unit is directly connected to the others. In some optional methods, the communication link between different computing units can also be called an inter-computing unit communication link. For example, different computing units may be connected using one or more of the following methods: High-Speed Custom Communication System (HCCS) interface, High-Speed GPU Interconnect Bandwidth interface, Integrated Circuit interface, Controller Area Network bus, Serial Peripheral Device interface, Queued Serial Peripheral Device interface, Full-Duplex Asynchronous Serial Interface, or Half-Duplex Differential Serial Interface, etc.
[0117] The HCCS interface is a high-speed connection channel between computing units, used to accelerate data and computation to produce executable results. For example, in a computing resource board, different computing units are connected in pairs using HCCS technology.
[0118] High-speed GPU interconnect bandwidth interface is a high-speed interconnect technology between GPUs, which is usually implemented by multiple pairs of wires printed on the computer board, with each pair of wires connecting to different GPUs at their respective ends.
[0119] For example, the computing device includes a switching board with multiple slots that are fully interconnected. Each resource board is inserted into a corresponding slot to connect it to the computing device. When a resource board needs to be replaced, it is removed from its corresponding slot.
[0120] In a fourth alternative scenario, computing device 1 can be a supernode. This supernode includes multiple computing resource boards, each containing multiple computing units.
[0121] In the fifth alternative scenario, the computing devices (computing device 1, computing device 2, or computing device 3) can be a computing cluster. This computing cluster can include multiple computing nodes within an availability zone. The multiple computing nodes are interconnected, and each computing node includes multiple computing units, which are interconnected.
[0122] The above five options are merely alternatives to different hardware structures of the computing device. In other embodiments, the computing device may have other hardware structures, which are not limited in this application.
[0123] For example, taking a computing device (computing device 1, computing device 2, or computing device 3) as a resource pool composed of multiple resource boards, the computing device includes multiple computing resource boards, network interface cards (NICs), and storage resource boards. The following description uses computing device 1 as an example to illustrate the computing resource board 11, computing resource board 12, computing resource board 13, NIC 14, and storage unit 15 included in computing device 1.
[0124] Each computing resource board (computing resource board 11, computing resource board 12, or computing resource board 13) is used to process data access requests from outside the computing device 1 (control device 10, server, user terminal, or other storage system), and also to process requests generated internally by the computing device 1. For example, when a computing resource board (computing resource board 11, computing resource board 12, or computing resource board 13) receives a data access request from a user terminal through a communication interface, it temporarily stores the data in these data access requests in the memory of the computing device 1. When the total amount of data in memory reaches a certain threshold, the computing resource board (computing resource board 11, computing resource board 12, or computing resource board 13) sends the data stored in memory to the storage unit 15 for persistent storage through the communication interface.
[0125] In an alternative scenario, each of the multiple computing resource boards can be of the same type, for example, computing resource boards 11, 12, and 13 can all be equipped with a CPU, GPU, or NPU.
[0126] In another alternative scenario, at least two of the multiple computing resource boards have different types. For example, computing resource board 11 is a CPU board, computing resource board 12 is a GPU board, and computing resource board 13 is an NPU board. Figure 2As shown, computing resource board 11 includes N CPUs: CPU1101, CPU1102, ..., CPU110N. Computing resource board 12 includes N NPUs: NPU1201, NPU1202, ..., NPU120N. Computing resource board 13 includes N GPUs: GPU1301, GPU1302, ..., GPU130N. N can be a positive integer greater than 2. In other embodiments, the number of computing units on computing resource boards 11, 12, and 13 may be different.
[0127] The above two are merely different optional methods for the computing resource board. In some embodiments, the computing resource board may also have other forms, which are not limited in this application.
[0128] Network interface card 14 can be a standard network interface card (NIC) or a smart NIC. A smart NIC, also known as a smart network adapter, not only performs the network transmission functions of a standard NIC but also provides a built-in programmable and configurable hardware acceleration engine. This improves application performance and significantly reduces the communication consumption of the connected computing resource boards (computing resource board 11, 12, or 13), providing more computing resources for the application. For example, in a highly virtualized environment, the computing resource boards (computing resource board 11, 12, or 13) in computing device 1 need to run open virtual switch (OVS) related tasks. Simultaneously, these boards also handle storage, online or offline encryption / decryption of data packets, deep packet inspection, firewalls, complex routing, and other operations. These operations not only consume significant computing resources but also, due to resource contention between different services, prevent optimal service performance. Network interface card 14 (NIC 14) serves as a hub connecting various services, and these services are accelerated on NIC 14. For example, when NIC 14 is a smart NIC, it includes a processor for performing acceleration functions. This processor may include, but is not limited to, processing chips such as DPU, GPU, and NPU, or it may refer to FPGA, tensor processing unit (TPU), microprocessor chip DSP, application-specific integrated circuit (ASIC), or one or more integrated circuit chips.
[0129] Optionally, the computing device 1 may also include memory, which refers to internal storage that directly exchanges data with the computing resource boards (computing resource board 11, computing resource board 12, or computing resource board 13). It can read and write data at any time and is very fast, serving as temporary data storage for the operating system or other running programs. Memory includes at least two types of storage; for example, it can be random access memory (RAM) or read-only memory (ROM). For example, random access memory is DRAM or SCM. DRAM is a semiconductor memory and, like most random access memory (RAM), is a type of volatile memory. However, DRAM and SCM are merely illustrative examples in this embodiment; memory may also include other random access memories, such as static random access memory (SRAM). For read-only memory, for example, it can be programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), etc. In addition, the memory can also be a dual in-line memory module (DIMM), i.e., a module composed of dynamic random access memory (DRAM), or an SSD. In practical applications, computing device 1 can be configured with multiple memory modules and different types of memory. This embodiment does not limit the number and type of memory. Furthermore, the memory can be configured to have a power-saving function. The power-saving function means that when the operating system in computing device 1 loses power and then is powered on again, the data stored in the memory will not be lost. Memory with a power-saving function is called non-volatile memory. The memory stores software programs, and the computing resource board (computing resource board 11, computing resource board 12, or computing resource board 13) can manage the hard disk by running the software programs in the memory. For example, the hard disk (such as storage unit 15) can be abstracted as a storage resource pool, and the storage resource pool can be provided to the server or users in the form of logical unit number (LUN). Here, LUN is actually the hard disk seen on the server. Of course, some storage systems are also file servers themselves, and can provide shared file services to the server.
[0130] In some optional scenarios, when the computing device 1 communicates in the management plane, it may include a root node and multiple leaf nodes. The root node, acting as the central node for management plane communication, sends management messages to one or more leaf nodes, thereby managing the multiple leaf nodes. In the case of data plane communication, the root node and leaf nodes are interconnected, and each node (root node or leaf node) can send service-type messages to other nodes (leaf nodes or root nodes) and also receive service-type messages sent by other nodes.
[0131] The following is an exemplary description of the root node and leaf node in conjunction with different physical forms of computing device 1.
[0132] In the first alternative approach, computing device 1 can be a server. Accordingly, the root node can be a computing unit within the server, and the leaf nodes can be the remaining computing units within the server excluding the root node. For example, root node 1 can be the CPU within the server, and the leaf nodes can be the GPU and NPU within the server.
[0133] In the second alternative approach, computing device 1 can be a computing resource board. The root node can be a chip, or a system-on-a-chip consisting of multiple chips. Similarly, leaf nodes can also be chips, or a system-on-a-chip consisting of multiple chips.
[0134] The root node and leaf nodes can be deployed on the same physical device; for example, they can reside on the same server. Alternatively, the root node and leaf nodes can be deployed on different physical devices; for example, they can reside on different servers. Or, if computing device 1 includes multiple leaf nodes, the root node and the multiple leaf nodes can be deployed on multiple physical devices. For example, the root node and a portion of the multiple leaf nodes are deployed on a first server. Another portion of the multiple leaf nodes are deployed on a second server. The second server can include one or more servers, with the other portion of the multiple leaf nodes deployed on one second server, or the other portion of the multiple leaf nodes deployed on multiple second servers.
[0135] In a third alternative approach, computing device 1 can be a server rack. This server rack can include multiple servers. In one alternative approach, the root node can be one of the multiple servers, and the leaf nodes can be the remaining servers excluding the root node. In other alternative approaches, the root node can be a computing unit on one of the multiple servers. The leaf nodes can be the remaining computing units excluding the root node.
[0136] In a fourth alternative approach, computing device 1 can be a supernode. In one alternative approach, the root node can be a computing resource board within the supernode, and the leaf nodes can be the remaining computing resource boards in the supernode excluding the root node. For example, the root node can be computing resource board 11, and the leaf nodes can be computing resource boards 12 and 13. In another alternative approach, the root node can be a computing unit within a computing resource board of the supernode, and the leaf nodes can be the remaining computing units in the supernode excluding the root node. For example, the root node can be CPU 1101 on computing resource board 11, and the leaf nodes can be the computing units on computing resource boards 11, 12, and 13 excluding CPU 1101.
[0137] In the fifth alternative approach, computing device 1 can be a computing cluster. The root node can be a computing device in the computing device cluster. Leaf nodes can be the remaining computing devices in the computing device cluster excluding the root node. Each computing device can be a server, a server rack, or a supernode.
[0138] The above five optional scenarios are merely different options for the root node and leaf node under different physical forms of the computing device 1. In other embodiments, there may be other optional methods for the root node and leaf node, which this application does not limit.
[0139] The following explanation uses the example where the root node is a computing unit on any one of multiple computing resource boards.
[0140] In some alternative implementations, the control device 10 can set the root node in the computing device.
[0141] For example, when a new computing device is connected to the data center, the management device 10 can select a root node from multiple computing units of the computing device and set the remaining computing units other than the root node from the multiple computing units provided by multiple computing resource boards as leaf nodes.
[0142] In the first alternative example, the control device 10 can select any one of the multiple computing units provided by the multiple computing resource boards of the computing device as the root node.
[0143] In the second alternative example, the control device 10 can determine the computing unit whose port number or sequence number is first among the multiple computing units provided by the multiple computing resource boards of the computing device as the root node.
[0144] In a third optional example, the control device 10 can also designate a CPU-type computing unit among the multiple computing units provided by the multiple computing resource boards of the computing device as the root node. Optionally, when the computing device includes multiple CPU-type computing units, any one CPU-type computing unit can be selected as the root node. Alternatively, the computing unit with the earliest port number or sequence number among the multiple CPU-type computing units can be designated as the root node.
[0145] The above three optional examples are merely different implementation methods for setting the root node in the control device 10. In other embodiments, the control device 10 may also use other implementation methods to set the root node. This application does not limit this.
[0146] In one alternative implementation, the root node and leaf nodes can be changed. For example, if the root node (CPU 1101) in computing device 1 fails, the control device 10 can select a new root node from multiple leaf nodes of computing device 1. After the failed root node restarts, the failed root node is used as the new leaf node. The method for selecting the new root node is similar to the method for selecting the root node described above, and will not be elaborated upon here.
[0147] In some alternative implementations, computing device 1 may select a root node and a leaf node from multiple computing units provided by multiple computing resource boards when computing device 1 starts up.
[0148] For example, when computing device 1 is first powered on, the computing unit whose port number or sequence number is first among the multiple computing units provided by multiple computing resource boards is determined as the root node. For example, the CPU 1101 mentioned above is used as the root node in computing device 1.
[0149] The two optional implementation methods mentioned above are only different ways of selecting the root node. In other embodiments, the root node can also be selected in other ways, and this application does not limit this.
[0150] Please continue reading. Figure 2 , Figure 2 The root node (CPU1101) is connected to multiple leaf nodes via bus 40. This root node is used to implement data communication between the root node and at least one leaf node via bus 40, as well as data communication between computing device 1 and management device 10. Bus 40 may also be referred to as the first bus, communication channel, or other names.
[0151] In one alternative implementation, the bus 40 can be, but is not limited to: PCIe, CXL, I2C, CAN, SPI, QSPI, UART, or a unified bus (Ubus or UB).
[0152] In some optional scenarios, where the endpoint devices of bus 40 are interconnected peer-to-peer, bus 40 may include, but is not limited to: unified bus (Ubus or UB), NVLink, etc. TM Computer Express Link (CXL) or other possible peer-to-peer interconnect buses are not limited to this application.
[0153] Peer-to-peer (PC) is a high-bandwidth, high-quality cloud resource interconnection service that enables routing communication between devices. By configuring routing policies at both ends of the device, interconnection can be achieved between devices in the same or different regions, or between the same or different users. Peer-to-peer (PC) does not rely on any single hardware component, eliminating single points of failure or bandwidth bottlenecks. When multiple endpoint devices are connected via the peer-to-peer bus, they can communicate directly without data packets needing to be relayed through the public internet or CPU.
[0154] A communication network based on a peer-to-peer bus connection is called a peer-to-peer network.
[0155] In the first optional scenario, the Unified Bus (UB) is also known as Lingqu. TM Bus, in Lingqu TM In a network composed of buses, various devices such as GPUs, DPUs, and CPUs can directly communicate with each other without needing to pass through the CPU in computing device 1 to relay data, which greatly improves the communication efficiency between different devices.
[0156] In the second optional scenario, NVLink TM It is a vertically scalable interconnect bus that enables faster input of large datasets into models and rapid data exchange between CPUs and GPUs. Specifically, NVLink... TM It adopts a point-to-point structure and serial transmission, and is used for connection between CPU and GPU, and can also be used for interconnection between different GPUs.
[0157] In the third alternative scenario, CXL is a high-speed interconnect technology that provides higher data throughput and lower latency, helping to address the problem of high latency between CPUs and devices, and between devices.
[0158] This document uses bus 40 with UB as an example for illustrative purposes, but it should not be construed as meaning that the computing device provided in this application embodiment can only achieve peer-to-peer interconnection via a UB-connected network (UB network). In other words, the computing device provided in this application embodiment supports applications not only on UB-connected networks, but also on networks using NVLink. TM The connected network can also be used in CXL-connected networks, or other buses or networks that support peer-to-peer interconnection of endpoint devices.
[0159] In this article, an endpoint device refers to the receiving or sending end of data communication, which includes control command communication and service data communication. Control command communication is also called control plane communication or management plane communication, and service data communication is also called data plane communication.
[0160] Please see Figure 2 Bus 40 supports control plane communication between multiple computing units. For example, the root node (CPU111) and leaf nodes communicate via bus 40. The control plane, also known as the management plane or control interface, refers to the transmission of control signaling or commands, rather than actual business data (such as voice data, image data, or others). Control commands are used to configure routing and forwarding, control the establishment of a call process, maintain processes, or release process resources. For example, control commands can be used to create routing tables, start computing units, obtain the running status of computing units, and initiate data processing procedures, including but not limited to data reading, data writing, or other data processing procedures.
[0161] Data reading refers to the process by which a computing unit reads data from a storage medium.
[0162] Data writing refers to the process by which the computing unit writes data into the storage medium.
[0163] The above data processing procedures are merely optional examples provided in the embodiments of this application and should not be construed as limiting this application.
[0164] In this article, the data plane is also called the user plane or the service plane. Data plane communication refers to the transmission of actual service data (also called real service data) by different computing units through the bus 40, such as voice data, image data, operating parameters or other types of service data.
[0165] Please see Figure 2The bus 40 also supports data plane communication between computing units. Specifically, any two computing resource boards among computing resource boards 11, 12, and 13 can communicate via the bus 40. For example, different computing resource boards can transmit actual service data (also called real service data), such as voice data, image data, or other types of service data, through this UB network. Furthermore, different computing units within the same computing resource board can communicate via the bus 40.
[0166] Please continue reading. Figure 2 The computing device 1 also includes a storage unit 15. The storage unit 15, also called a first storage unit, storage resource board, or storage resource pool, includes: a media controller and a storage medium (…). Figure 2 (Not shown in the image).
[0167] Storage media are used to store data, and may include, but are not limited to, one or more of the following: magnetic tape, optical disc, HDD media, SSD media, DRAM media, SRAM media, DIMM media, SCM media, non-volatile magnetic random access memory (MRAM) media, resistive random access memory (RRAM) media, ferroelectric memory (FeRAM) media, high bandwidth memory (HBM) media, phase change memory (PCM) media, or other types of media, such as cache chips, flash memory chips (flash memory particles), or others. SCM is a composite storage technology that combines the characteristics of traditional storage devices and memory. Storage-class memory can provide faster read and write speeds than hard drives, but its access speed is slower than DRAM, and its cost is also lower than DRAM. For example, flash memory chips may include: XL-LAND, single-level cell (SLC), multi-level cell (MLC), trinary-level cell (TLC), quad-level cell (QLC), enterprise multi-level cell (eMLC), or others.
[0168] A media controller is used to manage the storage medium. For example, the media controller includes one or more processors and a cache. The processor is a CPU or other type of processing chip, used to process data access requests from outside the storage unit 15, and also to process requests generated internally within the storage unit 15. For example, when the processor receives a write data request from the computing resource board 11, it temporarily stores the data in these write data requests in the cache. When the total amount of data in the cache reaches a certain threshold, the processor writes the data stored in the memory cache to the storage medium.
[0169] In this embodiment of the application, the computing resource board 11, computing resource board 12, computing resource board 13, network card 14, and storage unit 15 in the computing device 1 communicate with each other via bus 40.
[0170] In one optional scenario, data plane communication is implemented between the computing resource board and other storage units, or between multiple storage units. Specifically, any two storage units among storage units 15, 25, and 35 communicate via bus 40.
[0171] It is worth noting that the structure of the computing device 1 described above is merely an example provided in the embodiments of this application and should not be construed as limiting the scope of this application. In some optional cases, a computing device may include multiple storage units, or a computing device may also include other memory, such as including but not limited to: HDD, SSD, DRAM, SRAM, DIMM, SCM, MRAM, RRAM, FeRAM, HBM, PCM, cache chips, flash memory chips (flash memory particles), or others, such as optical discs or magneto-electric disks (MED).
[0172] Please continue reading Figure 2 Computing device 2 includes: computing resource board 21, computing resource board 22, computing resource board 23, network card 24, and storage unit 25. Computing device 3 includes: computing resource board 31, computing resource board 32, computing resource board 33, network card 34, and storage unit 35.
[0173] The specific implementation of computing device 2 and computing device 3 can be referred to the description of computing device 1 above, and will not be repeated here.
[0174] The above Figure 2 This is only one optional structure for data center 100. In other embodiments, data center 100 may have other structures. For example, data center 100 may also include multiple computing devices. During control plane communication, one or more first computing devices are present among the multiple computing devices. These first computing devices are used to send control commands to second computing devices among the multiple computing devices to realize control plane communication of data center 100. The first computing device can be any one of the multiple computing devices in data center 100, or a computing device designated by operations and maintenance personnel. The second computing device is any computing device among the multiple computing devices other than the first computing device. The first or second computing device is related to the above... Figure 2 The structure of computing device 1 is similar to that of computing device 1.
[0175] The above combination Figure 2The structure of the data center and the bus 40 used by each device are illustrated below. Figure 3 The management and control software for the data center and the peer-to-peer interconnection network corresponding to bus 40 provided in the embodiments of this application will be described. Figure 3 A software schematic diagram of a data center provided for this application. (About...) Figure 3 The hardware structure of each computing node shown can be referenced. Figure 2 The description of that will not be repeated here.
[0176] Figure 2 The data communication network provided by the bus 40 shown includes: Figure 3 The data communication network shown is an abstract network implemented based on bus 40. In some optional configurations, this data communication network is a peer-to-peer interconnection network. For information on the characteristics and supported functions of peer-to-peer interconnection networks, please refer to [link to relevant documentation]. Figure 2 The description of [the specific network type] will not be repeated here. In some alternative approaches, the data communication network is a communication network, such as an Ethernet network, or other types of networks.
[0177] In the first alternative approach, the data communication network enables control plane communication between different endpoint devices. For example, taking computing device 1 as an example, the root node (CPU1101) and leaf nodes (CPU1102, NPU1201, NPU1202, GPU1301 or GPU1302) transmit control commands or management commands via the data communication network. For instance, the control command could be a routing configuration command, and the management command could be a performance monitoring command.
[0178] In the second alternative approach, the data communication network can also enable data plane communication between different endpoint devices. For example, taking computing device 1 as an example, the leaf node (CPU1102, NPU1201, NPU1202, GPU1301 or GPU1302) and the root node (CPU1101) transmit business data through the data network.
[0179] In addition, each computing device can also access a data communication network. For example, the computing devices can communicate with each other via the data plane. For instance, when computing device 1 receives a read / write request from another computing device (e.g., computing device 2 or computing device 3), computing device 1 can communicate with one or more storage units via the computing resource board (computing resource board 11, computing resource board 12 or computing resource board 13) or network interface card 14 to realize the access service corresponding to the read / write request.
[0180] Please continue reading. Figure 3Regarding the software structure in each computing device, the following example of computing device 1 is used for illustration: CPU1101 (root node) is equipped with an operating system (OS) and management software.
[0181] The operating system (OS) is a program built into the computing device 1. The OS coordinates with various hardware components within the computing device 1 to interact with the user (or the aforementioned control device 10). Depending on the operating environment, the operating system (OS) can be categorized as: desktop operating system, mobile operating system, server operating system, embedded operating system, or others. The interface type provided by the operating system (OS) for user interaction can include, but is not limited to: command-line interface, graphical user interface, touch interface, natural user interface (NUI), or other interfaces.
[0182] Based on their operating environment, operating systems (OS) can be categorized as follows: Desktop Operating Systems (OS), Mobile Operating Systems (OS), Server Operating Systems (OS), Embedded Operating Systems (OS), and others. A desktop OS provides a "black box" for developers, allowing them to access its functionality through a series of standard system call functions. Server operating systems generally refer to those installed on mainframe computers, such as web servers, application servers, and database servers. Within a specific network, a server OS undertakes additional management, configuration, stability, and security functions. An embedded operating system is used in embedded systems. It is a versatile type of system software, typically including hardware-related low-level driver software, a system kernel, device driver interfaces, communication protocols, a graphical interface, and a standardized browser. An embedded operating system is responsible for allocating all software and hardware resources, scheduling tasks, and controlling and coordinating concurrent activities within the embedded system. It reflects the characteristics of the system in which it resides and can achieve the required functions by installing or removing certain modules. Mobile operating systems are mainly used in devices such as smartphones or smart tablets. They are developed from embedded operating systems and are specifically designed for mobile phones. In addition to the functions of embedded operating systems (such as process management, file system, network protocol stack, etc.), mobile operating systems also need to have power management for battery-powered systems, input / output for user interaction, embedded graphical user interface services that provide calling interfaces for upper-layer applications, low-level encoding and decoding services for multimedia applications, Java runtime environment, core wireless communication functions for mobile communication services, and upper-layer applications for smartphones.
[0183] Please see Figure 3 The CPU1101 contains management software that is used to manage various hardware components in the computing device 1 (such as...). Figure 3 Storage unit 15, network card 14, CPU 1102, NPU 1201, NPU 1202, GPU 1301 or GPU 1302) or other storage devices ( Figure 3 (Not shown in the image), it can also be used to realize data interaction between the control device 10 and the computing device 1.
[0184] For example, the management software is an application, such as a software module or unit that provides management functions to the outside world. This software module can provide logical functional modules with management granularity that meet the management needs of the control device 10. After obtaining a management request from the control device 10 (or the client), the management software can also send a device management response corresponding to the management request to the control device 10 (or the client).
[0185] Optionally, the management software code is installed on the CPU1101, or the CPU1101 can use the relevant management functions of the management software by providing one or more calling interfaces. The management functions supported by the management software include, but are not limited to, one or more of the following: node operating status monitoring, node startup, node shutdown, network port allocation, hard disk drive letter allocation, network configuration, routing table configuration, data reading, data writing, data migration, data redundancy backup, or other management functions. Therefore, this management software is also called in-band management software.
[0186] Please see Figure 3 Computing device 2 includes computing resource boards 21, 22, and 23, a network interface card (NIC) 24, and a storage unit 25. Computing device 3 includes computing resource boards 31, 32, and 33, a NIC 34, and a storage unit 35. Similar to computing device 1, computing resource boards 21, 22, and 23, NIC 24, and storage unit 25 communicate with each other via a data communication network, as do computing resource boards 31, 32, 33, NIC 34, and storage unit 35. In some optional examples, computing resource board 21 includes CPUs 2101 and 2102, computing resource board 22 includes NPUs 2201 and 2202, computing resource board 23 includes GPUs 2301 and 2302, and computing resource board 31 includes CPUs 3101 and 3102. The computing resource board 32 includes NPU3201 and NPU3202. The computing resource board 33 includes GPU3301 and GPU3302. Among them, CPU2101 is the root node of computing device 2, and CPU3101 is the root node of computing device 3. Similar to CPU1101 mentioned above, management software runs in the OS of CPU2101 and CPU3101.
[0187] above Figure 2 and Figure 3The content shown is merely an example of a data center provided in the embodiments of this application and should not be construed as limiting this application. Depending on changes in user needs or adjustments to the computing requirements of the data center, the data center provided in the embodiments of this application may include more or fewer computing devices, and other storage devices or memories may also be provided in the data center. This application does not limit this.
[0188] In one alternative implementation, for each computing device in the data center, a single network can be shared to enable data plane communication and control plane communication within the computing device, eliminating the need to build a separate network dedicated to control plane communication and reducing the hardware cost of the computing device. This application provides a method of setting a root node in the computing device, using this management node to uniformly manage the leaf nodes within the computing device, and enabling the root node to reuse the data plane communication network between computing units within the computing device to achieve control plane communication with the leaf nodes. The following is based on the above... Figure 2 and Figure 3 The provided data center provides an exemplary description of the structure of computing equipment.
[0189] For example, taking the computing unit in the computing device 1 above as an example, the hardware implementation of the computing device will be described in an exemplary manner. Similarly, the implementation of computing device 2 and computing device 3 can refer to the implementation of computing device 1.
[0190] Please see Figure 4 , Figure 4 The present application provides a schematic diagram of the structure of a computing device 1, which includes a root node 41, a switching board 42, leaf nodes 1 and 2. The root node 41, leaf nodes 1 and 2 are different computing units within the computing device 1.
[0191] In some alternative configurations, the root node 41, leaf node 1, and leaf node 2 can be located on the same computing resource board. For example, the root node 41, leaf node 1, and leaf node 2 can all be located on the aforementioned computing resource board 11.
[0192] In other alternative configurations, the root node 41, leaf node 1, and leaf node 2 can be located on different computing resource boards. For example, the root node 41 can be located on computing resource board 11, leaf node 1 on computing resource board 12, and leaf node 2 on computing resource board 13. Alternatively, the root node 41 can be located on computing resource board 11, and both leaf node 1 and leaf node 2 can be located on computing resource board 12 or computing resource board 13.
[0193] The switching board 42 communicates with the root node, leaf node 1, and leaf node 2 via bus 40. For example... Figure 4As shown, the switch board is connected to the root node, leaf node 1, and leaf node 2 via the UB network.
[0194] In one optional implementation, the root node 41 is used to manage all components (leaf nodes or switching boards) within the computing device 1. "Management" may include, but is not limited to, issuing control commands, configuring the network, switching operating states, acquiring operating parameters, or detecting abnormal states. Issuing control commands may involve the root node 41 sending control commands to the switching board 42, or sending control commands from the root node 41 to the leaf nodes. This application does not limit the specific type of control commands. For example, the control command may be a start command, an interrupt command, an operating parameter acquisition command, or a network configuration command, etc.
[0195] In some alternative approaches, the network configuration may be routing data established by the switching board 42 for nodes (leaf nodes or root node 41) within the computing device 1. In some alternative approaches, the routing data may also be referred to as a routing table, data forwarding table, network forwarding table, or other names, which are not limited in this application.
[0196] The routing data includes multiple routing paths. Each routing path includes the device identifier of the node sending management messages and its corresponding network address, as well as the device identifier of the leaf node receiving management messages and its corresponding network address. In some optional configurations, the node sending management messages in multiple routing paths can be the root node 41 within computing device 1. That is, the node sending management messages is the same in multiple routing paths. Since each routing path runs from the root node to the leaf node, in the case of control plane communication, management messages can only be sent by the root node to ensure that the root node is the same as the management leaf node during management plane communication, thus guaranteeing the reliability of management plane communication.
[0197] In an alternative approach, the routing path may also be referred to as a routing record, routing subdata, or other names, which are not limited in this application.
[0198] In some optional configurations, the routing data may also include routing data for control plane communication and routing data for data plane communication. The routing data for control plane communication includes the multiple routing paths mentioned above. The routing data for data plane communication includes the device identifiers of each node (leaf node or root node 41) within computing device 1 and the network address corresponding to each device identifier. Compared to the routing data for control plane communication, the routing data for data plane communication does not limit the source and destination nodes of the routing path; instead, it provides the device identifier and network address of each node within computing device 1, enabling interconnection between nodes. Therefore, during data plane communication, the switching board 42 can achieve data transmission between different nodes through the routing data for data plane communication.
[0199] In some optional examples, to ensure the reliability of data communication, the switching board 42 may also store identification codes for management messages and service messages. The routing data for data plane communication is associated with the identification codes for service messages, and the routing data for control plane communication is associated with the identification codes for management messages. After receiving a message, if the identification code indicating the message type carried in the message matches the identification code for a management message, the switching board 42 forwards the message using the routing data for control plane communication. If the identification code indicating the message type carried in the message matches the identification code for a service message, the switching board 42 forwards the message using the routing data for data plane communication. In this way, the routing data for control plane communication and the routing data for data plane communication are distinguished by the identification code indicating the message type, thereby achieving isolation between management plane communication and data plane communication and ensuring the reliability of data communication.
[0200] In some alternative implementations, the network address can be assigned by the root node 41. For example, after startup, the root node 41 obtains the device identifiers of the nodes (leaf nodes or root node 41) within the computing device 1 and assigns a network address to each device identifier. The root node 41 sends the device identifiers of the nodes (leaf nodes or root node 41) within the computing device 1 and the network address corresponding to each device identifier to the switching board, so that the switching board 42 can establish routing data based on the device identifiers and the network addresses corresponding to the device identifiers. Alternatively, after startup, the root node 41 obtains the device identifiers of the nodes (leaf nodes or root node 41) within the computing device 1 and assigns a network address to each device identifier, establishing routing data based on the device identifiers of the nodes (leaf nodes or root node 41) within the computing device 1 and the network address corresponding to each device identifier. The root node 41 then sends the routing data to the switching board 42.
[0201] In this context, the device identifier refers to the unique identifier of a node (leaf node or root node) within the computing device 1. The network address, also known as the network identifier, indicates the network access address of the node (leaf node or root node) indicated by the device identifier. Since the root node assigns a network address to each device identifier, the association between the node and the network address is weakened through the device identifier. In the event of a leaf node failure or replacement within the computing device, the network configuration of the leaf node can be achieved by replacing the node corresponding to the device identifier, without the need to reassign the leaf node's network address.
[0202] In other alternative methods, network configuration can also involve configuring the port on which leaf node 41 receives control messages. For example, after startup, root node 41 assigns a network address to each device identifier. The root node then sends an access request to a leaf node (leaf node 1 or leaf node 2), which includes the leaf node's device identifier and the corresponding network address. During the Basic Input Output System (BIOS) startup process, leaf node 1 or leaf node 2 associates its first port (among multiple ports) with the network address corresponding to its device identifier, enabling the first port to receive control messages, while the other ports do not. The access request can indicate a request for the leaf node to access the control plane, allowing root node 41 to communicate with the leaf node via the control plane. The port can also be referred to as an entity, entity, or other name.
[0203] In one alternative approach, a node (root node or leaf node) includes multiple ports, each used to receive (or process) different types of messages (or data), such as ports for receiving management messages or ports for processing data reads or writes. Based on whether a node's (root node or leaf node) ports accept management messages, the multiple ports can be divided into first-class ports and second-class ports. The first-class ports are used to receive (or send) management messages. The second-class ports are used to receive (or send) service messages.
[0204] In the first optional approach, to ensure the security of message transmission, the first type of port is used only for processing management messages and not for processing business messages. Correspondingly, the second type of port is used only for processing business messages and not for processing management messages.
[0205] In the second optional method, the first type of port is used to process both management and business messages. Correspondingly, the second type of port is used only to process business messages and not management messages. The first type of port can also be called the first port, the management message receiving port, or other names, and the second type of port can also be called the second port, the business message interface port, or other names. The following explanation uses the first type of port as the first port and the second type of port as the second port as an example.
[0206] This application does not limit the number of first ports. For example, a leaf node (leaf node 1 or leaf node 2) may include one or more first ports.
[0207] The two options above are merely alternatives for different network configurations. In other embodiments, the network configuration may have other implementations, which are not limited in this application.
[0208] Operating state switching can involve controlling a component (leaf node or switching unit) within a computing device to switch from a stopped state to a running state (e.g., controlling a component to start up), or controlling a component within a computing device to switch from a running state to a stopped state (e.g., controlling a component within a computing device to stop running). Operating parameters can include the component's power consumption, resource utilization, number of tasks, or number of processes during operation. Abnormal state detection can involve detecting whether a computing unit is operating abnormally, and, if an abnormality is found, switching the operating state of the computing unit within the computing device accordingly.
[0209] This application does not limit the specific type of anomaly in the computing unit. For example, an anomaly could be that the response rate of the computing unit is less than a response speed threshold. Another example is that the temperature of the computing unit exceeds a temperature threshold. Yet another example is that the computing unit experiences a power outage, disconnection, or startup failure.
[0210] In one alternative scenario, the root node 41 may be a single CPU, or the root node 41 may include multiple CPUs; this application does not limit this.
[0211] In some alternative implementations, such as Figure 4 As shown, the root node 41 in the software structure also includes a software layer 411 and a hardware layer 412.
[0212] In some alternative embodiments, hardware layer 412 includes an I / O core 4121, which contains an integrated management platform, an integrated device (Idev), and a security configuration unit.
[0213] The IO chip 4121 is used to implement input / output (I / O) operations of the root node, such as forwarding data (or messages, or instructions) and receiving data.
[0214] The integrated management platform is used to implement one or more of the following: resource management and allocation, protocol conversion and standardization, security management, performance monitoring, firmware updates and upgrades, or user interface configuration. Resource management and allocation may involve allocating I / O resources. Protocol conversion and standardization may involve converting between different protocols. Security management ensures secure data transmission between the root node and leaf nodes. Performance monitoring obtains the performance data of the root node. Firmware updates and upgrades update the hardware or software configuration of the root node. User interface configuration provides the user interface.
[0215] Internally integrated devices include multiple interfaces, interface controllers, serial communication controllers, parallel communication controllers, interrupt controllers, buffers and registers, or power management units, etc.
[0216] The interface, also known as a UB interface, communication interface, port, function entity (FE), or other names, is used to transmit data to the outside of the root node (such as leaf nodes or the aforementioned management device 10) and to receive data sent from the outside of the root node. In some optional embodiments, multiple interfaces include a first interface and a second interface. The first interface may include one or more interfaces for implementing management plane communication. The second interface includes one or more interfaces for implementing data plane communication. In some optional embodiments, the first interface may also be called a first-type interface, a first-type port, or a first port; this application does not limit this. The second interface may also be called a second-type interface, a second-type port, a second port, or other names; this application does not limit this.
[0217] The interface controller manages data transmission through the interface. The serial communication controller can be, but is not limited to, UB, SPI, I2C, etc., to implement serial communication. The parallel communication controller implements parallel communication. The interrupt controller manages interrupt requests from external devices of the computing device and sends interrupt commands to components within the computing device. Buffers and registers, also known as buffers, are used to temporarily store data received by the root node. The power management unit provides power management for the computing device or for the root node.
[0218] The security configuration unit is used to control the management plane communication of the root node and to distinguish between the root node and leaf nodes. In some optional embodiments, the security configuration unit of the root node is equipped with a hardware protocol flag. The presence or absence of the hardware protocol flag distinguishes the root node from the leaf nodes. For example, if a node's security configuration unit has the hardware protocol flag set, that node can initiate management plane communication. If a node's security configuration unit does not have the hardware protocol flag set, that node cannot initiate management plane communication and can only receive management messages. In this way, the computing device can ensure, through the security configuration unit, that only the root node can initiate management plane communication within the entire computing device, thereby improving the security of management plane communication. In some optional embodiments, the hardware protocol flag may also be called an identifier or other names, which are not limited in this application.
[0219] In some alternative configurations, both leaf and root node security configuration units contain hardware protocol flags. The root node's hardware protocol flag is set to a first value, while the leaf node's is set to a fourth value. For example, if the hardware protocol flag is set to the first value, the node containing the security configuration unit is the root node, and this node can initiate control plane communication. If the hardware protocol flag is set to the fourth value, the node containing the security configuration unit is a leaf node, and this node cannot initiate control plane communication. Thus, the computing device, through its security configuration units, ensures that only the root node can initiate control plane communication within the entire computing device, improving the security of control plane communication.
[0220] This application does not limit the specific values of the first and fourth values. For example, the first value can be "01" and the fourth value can be "00". Another example is that the first value can be "11" and the fourth value can be "00". Yet another example is that the first value can be "EE" and the fourth value can be "EF".
[0221] In an alternative approach, hardware layer 412 may also include computing chips. Figure 4 (Not shown in the image), this computational core is used to perform data calculations.
[0222] Please continue reading. Figure 4 The software layer 411 consists of program code running on the hardware layer 412. The software layer 411 can be divided into several layers, which communicate with each other through software interfaces. For example... Figure 4 As shown, the software layer 411 includes, from top to bottom, a proxy layer 4111, an interface management layer 4112, an OS layer 4113, and a bus switching unit layer 4114.
[0223] The proxy layer 4111 is used to transmit data with external devices (such as the control device 10 or the client).
[0224] The interface management layer 4112 is used to provide multiple application programming interfaces (APPs) so that the proxy layer 4111 can call the APP to implement data transmission.
[0225] The bus switching unit layer 4114 is a low-level running program below the OS layer 4113.
[0226] OS layer 4113, also known as OS, includes the OS program code and the management software program code. OS layer 4113 runs above the root node, while the management software runs within the OS.
[0227] In this application, the OS can be stored on the hard disk of the computing device, or it can be stored in the storage medium of the root node 41, or it can be stored on a remote device. When the root node 41 starts, the OS will be loaded into the memory of the root node 41 and run.
[0228] The above is only one illustration of the software structure in the root node 41. In other embodiments, the root node 41 may have other software structures, which are not limited in this application.
[0229] Please continue reading. Figure 4 In computing device 1, each leaf node (leaf node 1 or leaf node 2) includes one or more dies. For example, leaf node 1 includes die 1 and die 2, and leaf node 2 includes die 3 and die 4. Leaf node 1 can be, but is not limited to, a CPU chip, a GPU chip, an NPU chip, or a TPU chip. Each leaf node's chip is a GPU within a GPU chip, an NPU within an NPU chip, or a TPU within a TPU chip. For example, if a die is a GPU within a GPU chip, then that die can be used to perform mathematical and geometric calculations to achieve tasks such as image rendering.
[0230] In one alternative implementation, similar to the root node 41 described above, the software structure of the leaf nodes also includes a software layer and a hardware layer. For example, leaf node 1 is used as an example to illustrate the software structure of the leaf node; similarly, the software structure of leaf node 2 can refer to the software structure of leaf node 1.
[0231] Please see Figure 4 Leaf node 1 includes a software layer 42 and a hardware layer 44. Unlike the software structure of the root node 41, in leaf node 1, the software layer 43 includes, from top to bottom, an OS layer 431 and a driver layer 421. The OS layer 431 functions similarly to the OS layer 4113 in the root node, and will not be described in detail here. The driver layer 432 is used to implement hardware initialization, data transmission, or resource management in leaf node 1. "Resource management" can be, but is not limited to, memory allocation, bandwidth allocation, or computing resource allocation.
[0232] Similar to the hardware layer of the root node 41, the hardware layer 44 of leaf node 1 also includes an I / O core 442 and a computing core 441. The I / O core 441 contains an integrated management platform, an integrated device (Idev), and a security configuration unit. Unlike the hardware layer of the root node, in the first optional method, the security configuration unit in leaf node 1 does not have a hardware protocol flag set. In the second optional method, the hardware protocol flag set in the security configuration unit of leaf node 1 has a fourth value.
[0233] Additionally, in some alternative methods, the OS of OS layer 431 in leaf node 1 can be the same as the OS of OS layer 4113 in root node 41. Alternatively, in other alternative methods, the OS of OS layer 431 in leaf node 1 can be different from the OS of OS layer 4113 in root node 41. In practical applications, the operating system running in the leaf node (leaf node 1 or leaf node 2) can be determined according to the specific task performed by the computing device, and this application does not impose any limitations on this.
[0234] In addition, a first interface and a second interface can also be set in the internal integrated device of leaf node 1. The first interface is used to receive management messages, and the second interface is used to receive or send service messages.
[0235] Please continue reading. Figure 4 The switching board 42 is used to implement data forwarding between the root node 41 and the leaf nodes (leaf node 1 or leaf node 2), and to implement data conversion between leaf node 1 and leaf node 2. For example, the switching board 42 forwards control commands (or data) sent by the root node 41 to the leaf nodes (leaf node 1 or leaf node 2). Alternatively, the switching board 42 forwards data sent by the leaf nodes (leaf node 1 or leaf node 2) to the root node 41. Or, the switching board 42 sends data sent by leaf node 2 (or leaf node 1) to leaf node 1 (or leaf node 2).
[0236] In an alternative approach, the switching board may also be referred to as a switching unit, switching device, UB switching device, UB bus interaction device, or other names, and the embodiments of this application do not limit this.
[0237] In one alternative implementation, the switching board 42 may store routing data. After receiving a message, the switching board 42 forwards the message according to the routing data and the destination address indicated by the message.
[0238] In an alternative implementation, to reduce the load on the root node, the computing device 1 may also include a control CPU ( Figure 4 (Not shown in the diagram), the control CPU is used to control the switch board 42, enabling control plane communication of the switch board 42. In some alternative embodiments, the control CPU can be any one of multiple leaf nodes. Alternatively, the control CPU can be a computing unit within a computing device other than a leaf node. This computing unit can be a CPU or other chip with control functions.
[0239] In some optional cases, the control CPU can be integrated on the switching board 42.
[0240] In some alternative scenarios, the switch board 42 is provided with an eSPI (external interface), through which the control CPU accesses the switch board 42. For example, the control CPU can be located on the same computing resource board as the root node 41, and the control CPU on this computing resource board accesses the peripheral interface of the switch board 42 via PCIe. Another example is that the control CPU accesses the peripheral interface of the switch board 42 via a programmable field-programmable device (FPGA). This application embodiment does not limit the implementation method of the control CPU accessing the peripheral interface of the switch board 42.
[0241] The two optional scenarios mentioned above are merely different connection methods between the control CPU and the switching board 42. In other embodiments, the control CPU and the switching board 42 may also adopt different connection methods, which are not limited in this application.
[0242] The above is only one possible structure for a computing device. In other embodiments, the computing device 1 may have other structures. For example, if the computing device 1 is a memory-computing integrated computing node, the computing device 1 may also include the aforementioned storage unit. As another example, if the computing device 1 is a memory-computing separated computing node, the computing device 1 may also include the aforementioned network interface card (NIC). This application does not limit the scope of the application in this regard.
[0243] based on Figure 4 The following is an exemplary description of the data communication process between the root node 41 and the leaf node in the provided computing device.
[0244] First, taking the data plane communication between root node 41 and leaf node 1 as an example, the data communication process between root node 41 and leaf node 1 in computing device 1 will be illustrated. This data plane communication process includes stages ① to ④, wherein:
[0245] Phase ①: Root node 41 receives data access requests sent by external devices.
[0246] The external device can be the aforementioned control device 10, or other computing devices.
[0247] In one alternative approach, the data access request may carry the data to be written, or the storage address of the data to be read in the storage unit.
[0248] Phase ②: In response to the data access request, node 41 calls the second interface to send a service-type message to switch board 42. Phase ③: switch board 42 obtains the service-type message sent by root node 41 from the UB network, and forwards the service-type message to leaf node 1 indicated by the destination address carried in the service-type message.
[0249] Phase 4: Leaf node 1 obtains service-type packets from the UB network.
[0250] Secondly, taking the control plane communication between root node 41 and leaf node 1 as an example, the data communication process between root node 41 and leaf node 1 in computing device 1 is illustrated. This control plane communication includes stages ① to ④, wherein:
[0251] Phase 1: The management software sends control commands to the root node 41.
[0252] In phase ②, root node 41 responds to the control command and calls the second interface to send a management message to switch board 42.
[0253] Phase 3: Switchboard 42 obtains the management message sent by root node 41 from the UB network, and forwards the management message to leaf node 1 indicated by the destination address through the destination address carried in the management message.
[0254] Phase 4: Leaf node 1 obtains management messages from the UB network and executes the control commands carried in the management messages.
[0255] from Figure 4 It can be seen that during the data plane communication and control plane communication processes between the root node 41 and the leaf nodes, the same network is used for both data plane and control plane communication. The computing device does not need to establish an additional network dedicated to control plane communication. This reduces the hardware cost of the computing device. Furthermore, in conjunction with the above... Figure 1A Compared to the existing architecture, the computing device provided in this application allows the root node to centrally manage all leaf nodes, eliminating the need for a separate management node (or root node) for each leaf node and reducing hardware costs. Furthermore, since only one root node is used, maintenance personnel can manage the entire device through this single node, reducing operational difficulty and complexity. For the data communication process between the root node and leaf nodes, please refer to the diagram. Figures 7 to 11 The provided embodiments are not described in detail here.
[0256] In one alternative implementation, when the number of leaf nodes in computing device 1 is large, to reduce the communication overhead between the root node and leaf nodes, computing device 1 can be equipped with multiple switching boards. Different switching boards are used to implement data communication between different leaf nodes and the root node. Alternatively, multiple switching units can be set on the switching boards, with different switching units used to implement data communication between different leaf nodes and the root node. In this way, in the event of a switching board (or switching unit) failure, or a communication link failure between a node (root node or leaf node) and a switching board (or switching unit), data communication can be achieved through other switching boards (or switching units), thereby improving the reliability of data communication within the computing device.
[0257] For example, taking computing device 1 as an example, two switching boards can be set up, such as Figure 5 As shown, Figure 5 A schematic diagram of the structure of a computing device provided in this application Figure 2 The computing device 1 includes a root node 41, leaf node 1, leaf node 2, leaf node 3, leaf node 4, a switching board 1, and a switching board 2.
[0258] Root node 41 is communicatively connected to switching board 1 and switching board 2. Switching board 1 connects to leaf node 1 and leaf node 2, and switching board 2 connects to leaf node 3 and leaf node 4. In some alternative configurations, switching board 1 may also be referred to as the main switching unit, the first switching unit, or other names, and switching board 2 may also be referred to as the standby switching unit, the second switching unit, or other names; this application does not limit this. Correspondingly, leaf node 1 and leaf node 2 may also be referred to as the first leaf node, the first type of leaf node, or other names, and leaf node 3 and leaf node 4 may also be referred to as the second leaf node, the second type of leaf node, or other names; this application does not limit this.
[0259] In the case of data communication between root node 41 and leaf nodes 1, 2, 3, and 4, root node 41 can achieve data communication with the first leaf node through the channel of root node → switch board 1 → first leaf node (leaf node 1 and leaf node 2). Root node 41 can also achieve data communication with the second leaf node through the channel of root node → switch board 2 → second leaf node (leaf node 3 and leaf node 4).
[0260] based on Figure 5 As can be seen from the provided embodiments, when the root node 41 sends control commands or data to leaf nodes 1, 2, 3, and 4, different messages can be sent to switching board 1 and switching board 2 respectively, so that switching board 1 and switching board 2 forward the messages to the corresponding leaf nodes according to the received messages. In this way, switching board 1 and switching board 2 can perform message forwarding in parallel, reducing communication overhead.
[0261] In one alternative implementation, to improve the reliability of data communication within the computing device, a primary communication link and a backup communication link are established between the leaf node and the root node in the computing device. In the event of a switch failure or a failure of the communication link between the node (root node or leaf node) and the switch, data communication is achieved by switching the communication link, thereby improving the reliability of data communication within the computing device.
[0262] like Figure 6As shown, the root node 41 establishes a primary / backup communication channel with leaf nodes 1, 2, 3, and 4 through switchboards 1 and 2.
[0263] For example, taking switchboard 1 as an example, the main communication channel of switchboard 1 is a communication link consisting of root node → switchboard 1 → leaf node (leaf node 1 and leaf node 2). The backup communication channels of switchboard 1 include backup channel 1, backup channel 2 and backup channel 3.
[0264] Among them, backup channel 1 is a communication link consisting of root node → switching board 2 → leaf node (leaf node 1 and leaf node 2).
[0265] Backup Channel 2: The communication link consisting of root node → switch board 1 → switch board 2 → leaf node (leaf node 1 and leaf node 2).
[0266] Backup channel 3: The communication link consisting of root node → switch board 2 → switch board 1 → leaf node (leaf node 1 and leaf node 2).
[0267] The following examples illustrate the implementation of primary and backup communication link selection based on the aforementioned backup channel 1, backup channel 2, and backup channel 3.
[0268] In the first alternative approach, if the root node 41 detects that the communication link between the root node and the switching board 1 is disconnected or abnormal, the root node 41 can use backup channel 1 or backup channel 3 to communicate with leaf node 1 and leaf node 2.
[0269] In the second alternative approach, if switchboard 1 detects a disconnection or abnormality in the communication link between switchboard 1 and leaf node 1 or leaf node 2, switchboard 1 can use the switchboard 1→switchboard 2→leaf node (leaf node 1 and leaf node 2) communication link for data communication. Alternatively, switchboard 1 reports to root node 41, and root node 41 uses backup channel 2 to communicate with leaf node 1 and leaf node 2.
[0270] In the third alternative, if the root node 41 detects a failure of the switch board 1, the root node 41 can use the backup channel 1 to communicate with the leaf nodes 1 and 2.
[0271] The above three optional implementation methods are only different ways of selecting the primary and backup communication links. In other embodiments, the root node 41 or the switching board 1 can also use other implementation methods to select the primary and backup communication links.
[0272] The above backup channels 1, 2, and 3 are illustrated with the root node as the starting node. In actual applications, when a leaf node sends a service message to the root node 41, the leaf node can also detect whether the communication link between the leaf node and the switching board it belongs to is abnormal. If the communication link is abnormal, the leaf node will communicate data through the backup channel.
[0273] For example, leaf node 1 is used as an example. The primary communication channel between leaf node 1 and the root node is: leaf node 1 → switch board 1 → root node. The backup channels for leaf node 1 include: leaf node 1 → switch board 2 → root node or leaf node 1 → switch board 2 → switch board 1 → root node. If leaf node 1 detects that the communication link between leaf node 1 and switch board 2 is broken, leaf node 1 sends a service message to switch board 2, so that switch board 2 forwards the service message to root node 41.
[0274] The above explanation uses switchboard 1 as an example. The following explanation uses switchboard 2 as an example to illustrate the selection method of primary and backup communication links.
[0275] like Figure 6 As shown, the main communication channel of the switch board 2 is a communication link consisting of root node → switch board 2 → leaf nodes (leaf node 3 and leaf node 4). The backup communication channels of the switch board 2 include backup channel 1, backup channel 2 and backup channel 3.
[0276] Among them, backup channel 1 is a communication link consisting of root node → switching board 1 → leaf nodes (leaf node 3 and leaf node 4).
[0277] Backup Channel 2: A communication link consisting of root node → switch board 2 → switch board 1 → leaf nodes (leaf node 3 and leaf node 4).
[0278] Backup channel 3: The communication link consisting of root node → switch board 1 → switch board 2 → leaf node (leaf node 3 and leaf node 4).
[0279] The selection method for the primary and backup communication links of switch board 2 can be referred to the selection method for the primary and backup communication links of switch board 1 mentioned above, and will not be elaborated upon in this application.
[0280] In one optional implementation, to facilitate the switching between primary and backup channels, the management software running in the root node 41 can, during network configuration, send multiple routing data to each of the multiple switching boards, enabling each switching board to establish communication connections with all computing units within computing device 1. For example, the root node 41 sends first routing data and second routing data to switching board 1, and sends first routing data and second routing data to switching board 2. The first routing data includes the device identifier S1 of leaf node 1, the device identifier S2 of leaf node 2, and the network address W1 of device identifier S1 and the network address W2 of device identifier S2. The second routing data includes the device identifier S3 of leaf node 3, the device identifier S4 of leaf node 4, and the network address W3 of device identifier S3 and the network address W4 of device identifier S4.
[0281] In another optional implementation, to facilitate the switching between primary and backup channels, the management software running in the root node 41 can, during network configuration, send the same routing data to each of the multiple switching boards. This routing data includes the device identifiers of all computing units within computing device 1 and the network addresses corresponding to those device identifiers. For example, the root node sends third routing data to switching board 1 and to switching board 2. This third routing data includes: the device identifier S1 of leaf node 1, the device identifier S2 of leaf node 2, the device identifier S3 of leaf node 3, the device identifier S4 of leaf node 4, and the network addresses W1, W2, W3, and W4 of device identifier S1, S2, W3, and S4.
[0282] The above Figure 6 This is only one possible structure for a computing device; in practical applications, computing devices can be configured with structures more complex than those described above. Figure 6 More switching boards, and each switching board can connect to more than the above. Figure 6 The number of leaf nodes is not limited in this embodiment.
[0283] based on Figure 6 The provided computing device allows for the configuration of two switching boards, with multiple leaf nodes dual-homed to both boards. In the event of a communication link failure between the switching board and the root node, or between a leaf node and the switching board, a backup channel can be used for data communication, thereby improving the reliability of data communication within the computing device.
[0284] above Figures 2 to 6The structure of the data center and computing device provided in this application is described. During data communication by the computing device, control plane communication and data plane communication reuse the same network, and the data communication method provided in this application can be implemented during data communication. The implementation method of the data communication method can be found below. Figures 7 to 11 The provided examples.
[0285] The following is based on the above Figures 2 to 6 The provided computing device serves as an example to illustrate the implementation of the data communication method provided in this application.
[0286] Taking the application of data communication methods to computing device 1 as an example, the structure of computing device 1 can be referred to the above. Figure 4 , Figure 5 or Figure 6 The computing device shown is not described in detail herein. Figure 7 As shown, Figure 7 A flowchart illustrating the data communication method provided in this application. Figure 7 The computing device 1 includes a root node 1, leaf nodes 1, and a switchboard 1. Management software runs on the OS of the root node 1. The data communication method includes steps S810 to S840.
[0287] S810, root node 1 sends the first message to leaf node 1.
[0288] In the first alternative approach, after receiving the first instruction, root node 1 sends the first message to leaf node 1, such as... Figure 7 As shown.
[0289] The first instruction can be a control instruction, which instructs root node 1 to perform management plane communication, and the corresponding first message is a management message.
[0290] The following provides several implementation methods to illustrate the triggering method of the first instruction.
[0291] In a first alternative implementation, the first instruction may be sent by an external device. This external device may be the aforementioned control device 10, a client, or an out-of-band BMC. This application embodiment does not limit the specific type of external device.
[0292] In the second alternative implementation, the first instruction can be sent by the application.
[0293] For example, the application can be deployed on a single device, which is either the aforementioned control device 10 or the aforementioned root node 1.
[0294] For example, the application can be deployed on a distributed system, which includes multiple devices, each with a complete application deployed on it. Alternatively, each device may deploy a portion of the application's code. For instance, the application can include, but is not limited to, management applications, operational applications, or distributed applications. A distributed application, for example, refers to an application distributed across different computers (or computing units) to collaboratively complete a task via a network.
[0295] In the third optional implementation, the first instruction can be task-triggered. This task can be an operation and maintenance task, a status detection task, or other management and control task. Specifically, this task can be a management and control task issued by a single application, or it can be a management and control task of multiple applications managed by a single interface; this application does not limit this.
[0296] In the fourth alternative implementation, the first instruction can be triggered by root node 1. For example, as described above. Figure 4 As shown, after root node 1 starts, it runs management software, which sends a first instruction to the hardware layer of root node 1. For example, when the management software receives a detection request or running status request from an external device, it sends a first instruction to the hardware layer of root node 1. For another example, after root node 1 runs the management software, it periodically sends a first instruction to its hardware layer. For yet another example, when the management software detects an abnormal running status of leaf node 1 or needs to switch the running status of leaf node 1, the management software sends a first instruction to the hardware layer of root node 1.
[0297] The four optional implementation methods described above are merely different triggering methods for the first instruction when the first instruction is a control instruction. In other embodiments, the first instruction may also be triggered by other methods, which are not limited in this application.
[0298] The above explanation uses a control command as the first instruction as an example. In another optional implementation, the first instruction can also be an I / O instruction, such as a data write instruction or a data read instruction, or a data calculation instruction or a task scheduling instruction. Accordingly, when the first instruction is an I / O instruction, the first instruction indicates data plane communication, and the first message is a service-type message.
[0299] Similar to the triggering method of the first instruction under the control instruction, when the first instruction is an I / O instruction, the first instruction can also have multiple triggering methods. Regarding the triggering method of the first instruction under the I / O instruction, please refer to the triggering method of the first instruction under the control instruction. This application will not elaborate on this.
[0300] In some alternative implementations, when the first instruction is an I / O instruction, it can carry data to be written, data to be computed, or a task to be scheduled. Root node 1 parses the first instruction to obtain the corresponding data or task.
[0301] In some alternative implementations, where the first instruction is an I / O instruction, it may carry a data identifier (or task identifier) instead of actual data. The data identifier indicates the data to be written or computed. The task identifier indicates the task to be scheduled. Root node 1 retrieves the corresponding data (or task) according to the data identifier (or task identifier).
[0302] The following explains how to distinguish between management messages and business messages.
[0303] In the first optional implementation, management messages and business messages are distinguished through the message transmission interface. For example, if the first instruction is a control instruction, root node 1 calls the aforementioned first interface to send the first message. If the first instruction is an I / O instruction, root node 1 calls the aforementioned second interface to send the first message.
[0304] In the second optional implementation, management messages and business messages are distinguished by their message formats. For details on how to distinguish management messages from business messages using message formats, please refer to the following. Figure 9 or Figure 11 The provided embodiments are not described in detail here.
[0305] In the third optional implementation, management messages and business messages are distinguished by message format and the interface for transmitting messages.
[0306] The above three optional implementation methods are only different ways for the root node 1 to send the first message. In other embodiments, the root node may also use other implementation methods to send the first message.
[0307] In the second alternative approach, upon receiving a second message forwarded by an external device, root node 1 sends a first message to leaf node 1 via bus 40. For example, root node 1 receives the second message and sends the first message to leaf node 1 via bus 40.
[0308] The second message can be a message instructing data reading or writing. Alternatively, the second message can also be a response message. This application does not limit the specific form of the second message. It should be understood that the message type of the second message is a business message.
[0309] In some alternative approaches, the second message may be sent by an external device (such as the computing unit in computing device 2, or the computing unit in computing device 3, or the control device 10). This application does not limit the external device that sends the second message.
[0310] The two optional methods mentioned above are only different triggering methods for root node 1 to send the first message. In other embodiments, root node 1 may have other triggering methods for sending the first message, and this application does not limit them.
[0311] The following is an example of how the root node 1 sends the first message to the leaf node 1.
[0312] In a first alternative approach, root node 1 sends the first message to leaf node 1. For example, root node 1 sends the first message to leaf node 1 via the aforementioned bus 40. Another example is that root node 1 sends the first message to leaf node 1 via the aforementioned data network. Yet another example is that root node 1 sends the first message to leaf node 1 via Ethernet. This application does not limit this approach.
[0313] In a second alternative approach, root node 1 sends the first message to leaf node 1 via switch board 1. For example, root node 1 sends the first message to leaf node 1 via the aforementioned bus 40 and switch board 1. Another example is that root node 1 sends the first message to leaf node 1 via the aforementioned data network and switch board 1. Yet another example is that root node 1 sends the first message to leaf node 1 via Ethernet and switch board 1. This application does not limit the scope of this approach.
[0314] The two optional methods mentioned above are merely different implementations of the root node 1 sending the first message to the leaf node 1. In other embodiments, the root node 1 may also use other implementations to send the first message to the leaf node 1, and this application does not limit this.
[0315] In one optional implementation, to facilitate the distinction between business-type messages and management-type messages, when root node 1 implements data plane communication, the first message carries a second identifier. When root node 1 implements control plane communication, the first message carries a first identifier. The first identifier indicates that the message type of the first message is a management-type message, and the second identifier indicates that the first message is a business-type message. This application does not limit the specific form of the first and second identifiers; for example, the first and second identifiers can be characters, strings, numbers, or other forms of content. The following provides guidance on how the first message carries the first or second identifier. Figure 8 or Figure 10 The provided embodiments are not described in detail here.
[0316] S820, Switchboard 1 forwards the first message to Leaf Node 1 (optionally).
[0317] Compared to the processing procedure in S810, in S820, when root node 1 sends the first message to switch board 1 via bus 40, switch board 1 receives the first message sent by root node 1 via bus 40, such as... Figure 7 As shown.
[0318] In one optional implementation, switchboard 1 can determine the network address of the leaf node receiving the first packet based on the routing data and the destination address indicated by the first packet. The first packet is then forwarded based on the network address of the leaf node receiving the first packet. For example... Figure 7 As shown, when the leaf node receiving the first message is leaf node 1, the switching board 1 forwards the first message to leaf node 1 through bus 40.
[0319] In one optional implementation, to ensure network security within computing device 1, after receiving the first message, switch board 1 can perform a validity verification on the first message. If the validity verification of the first message passes, switch board 1 forwards the first message to leaf node 1 via bus 40. If the validity verification of the first message fails, switch board 1 discards the first message. Thus, switch board 1 intercepts the first message when its validity verification fails, blocking risky messages within the in-band network and ensuring the security of the in-band network. The implementation method for the validity verification of the first message can be referred to below. Figure 9 or Figure 11 The provided embodiments are not described in detail here.
[0320] In some alternative methods, switchboard 1 performs validity verification on the first received message during both data plane and management plane communications. Alternatively, switchboard 1 performs validity verification on the first received message during management plane communications.
[0321] For example, after receiving the first message, switch board 1 identifies the message type of the first message. If the message type of the first message is a management message, it performs a validity verification on the first message. If the message type of the first message is a service message, it forwards the first message to leaf node 1 via bus 40. For example, if the first message carries a first identifier code, switch board 1 performs a validity verification on the first message. If the first message carries a second identifier code, switch board 1 forwards the first message to leaf node 1 via bus 40.
[0322] S830, leaf node 1 receives the first message.
[0323] In the first optional approach, if the first message is a management message, leaf node 1 receives the first message and performs the corresponding control operation. This "corresponding control operation" may include, but is not limited to: switching the running state, starting leaf node 1, sending running parameters, or configuring interfaces, etc.
[0324] In the second alternative approach, if the first message is a business-type message, leaf node 1 can simply receive the first message. Alternatively, leaf node 1 can perform the corresponding business-type operation after receiving the first message.
[0325] The "business-related operations" can include, but are not limited to, data writing, data reading, data sending, or task scheduling. For example, if the business-related operation is data writing, "performing the corresponding business-related operation" can mean writing the data carried by the first message into the memory. If the business-related operation is data reading, "performing the corresponding business-related operation" can mean reading the data indicated by the first message from the memory. If the business-related operation is data sending, "performing the corresponding business-related operation" can mean transmitting data. If the business-related operation is task scheduling, "performing the corresponding business-related operation" can mean executing the task indicated by the first message. Furthermore, "performing the corresponding business-related operation" can also involve processing the data carried by the first message, such as performing data analysis on the data carried by the first message, or using the data carried by the first message as input data to perform AI calculations (or AI training). This application embodiment does not limit the specific operations performed by leaf node 1.
[0326] In the third alternative method, leaf node 1 sends a response message after receiving the first message.
[0327] This response message is used to indicate that leaf node 1 has received the first message. In some alternative methods, this response message may also be called an IO response message, a third message, or other names, which are not limited in this application.
[0328] Alternatively, in an alternative implementation, after performing the corresponding operation of the first message, leaf node 1 sends a response message to root node 1, which indicates that leaf node 1 has completed the corresponding operation indicated by the first message.
[0329] Alternatively, in an optional implementation, after leaf node 1 executes the corresponding operation indicated by the first message and obtains the processing result, it sends a response message to root node 1. This response message carries the processing result. The processing result can be, but is not limited to, the read data, the calculation result of the data, or the data write address, etc.
[0330] The response message mentioned above can be a business-type message. In some optional methods, leaf node 1 can send a response message to switch board 1 through its second interface. Switch board 1 then forwards the response message to root node 1. For example... Figure 7 As shown, leaf node 1 sends a response message (S840) to switch board 1 via bus 40.
[0331] based on Figure 7 In the provided embodiments, since both management and service messages are forwarded through the same network (such as bus 40), control plane communication and data plane communication can reuse a single physical network without the need for additional deployment of a control plane network, thereby reducing the hardware cost of the control plane network architecture. Furthermore, by distinguishing between management and service messages, secure isolation is achieved between control plane communication and data plane communication, ensuring the reliability and rationality of reusing a single physical network for both. In addition, setting up a root node for each computing device enables unified management of multiple leaf nodes, reducing the maintenance cost of the control plane network architecture.
[0332] above Figure 7 The data communication method provided in this application is illustrated using the interaction between the root node 1, leaf node 1, and switch board 1 in computing device 1 as an example. In other optional implementations, the data communication method provided in this application can also be executed by other computing devices, such as, but not limited to, the aforementioned computing device 2, management device 10, a server capable of management plane communication, a server cluster, a system-on-a-chip, or a cluster. Furthermore, the data communication method provided in this application can also be used to implement data communication between different devices in a cluster. It is understood that when the data communication method is used to implement data communication between different devices in a cluster, the root node can be a computing device, and the corresponding leaf nodes can also be computing devices.
[0333] In one alternative implementation, since the control plane communication and data plane communication in the computing device reuse a single network, in order to reduce the impact of control plane communication on data plane communication, control plane communication and data plane communication are isolated in computing device 1. In this way, control plane communication does not affect data plane communication, thus ensuring the service performance within computing device 1.
[0334] In the first alternative implementation, computing device 1 can achieve isolation between control plane communication and data plane communication through an interface.
[0335] For example, root node 1 implements control plane communication through a first interface and data plane communication through a second interface. For instance, when the first instruction is a control instruction, the first message generated by root node 1 based on the first instruction is a management message, and root node 1 sends the first message through the first interface. Switch board 1 determines that the first message is a management message based on the transmission interface of the first message and forwards the first message to the first interface of leaf node 1. Leaf node 1 receives the first message through the first interface. Computing device 1 implements control plane communication. As another example, when the first instruction is an I / O instruction, the first message generated by root node 1 based on the first instruction is a service message, and root node 1 sends the first message through the second interface. Switch board 1 determines that the first message is a service message based on the transmission interface of the first message and forwards the first message to the second interface of leaf node 1. Leaf node 1 receives the first message through the second interface. Computing device 1 implements data plane communication. Yet another example, when root node 1 forwards a service message, root node 1 sends the first message through the first interface. Leaf node 1 receives the first message through the second interface. Computing device 1 implements data plane communication.
[0336] In the second optional implementation, computing device 1 isolates control plane communication from data plane communication through fields carried in the message header. For example, root node 1 encapsulates the first message using different message header formats for control plane communication and data plane communication. For instance, if the first instruction is a control instruction, root node 1 encapsulates the first instruction according to the message format of a management message to obtain the first message. Root node 1 sends the first message. Switch board 1 determines the first message is a management message based on the message header and forwards the first message to leaf node 1. Leaf node 1 receives the first message, determines it is a management message based on the fields carried in the message header, and processes the first message, thus enabling control plane communication. As another example, if the first instruction is an I / O instruction, root node 1 encapsulates the first instruction according to the message format of a service message to obtain the first message. Root node 1 sends the first message. Switch board 1 determines the first message is a service message based on the fields carried in the message header and forwards the first message to leaf node 1. Leaf node 1 receives the first message and determines it to be a service-type message based on the fields carried in the message header, thus enabling data plane communication for computing device 1. Alternatively, for example, if root node 1 forwards the second message, leaf node 1 receives the first message and determines it to be a service-type message based on the fields carried in the message header.
[0337] In a third optional implementation, computing device 1 isolates control plane communication from data plane communication through interfaces and fields carried in the message header. For example, taking the root node 1 triggering the transmission of a first message based on a first instruction as an example, root node 1 encapsulates the first instruction using different message formats to obtain the first message in the case of control plane communication and data plane communication, and then implements control plane communication through the first interface, or data plane communication through the second interface. For example, when the first instruction is a control instruction, root node 1 encapsulates the first instruction according to the message format of a management message to obtain the first message. Root node 1 sends the first message through the first interface. Switch board 1 determines that the first message is a management message based on the fields carried in the message header and the transmission interface of the first message, and forwards the first message to the first interface of leaf node 1. Leaf node 1 receives the first message, determines that the first message is a management message based on the fields carried in the message header and the transmission interface of the first message, and processes the first message, thus computing device 1 implements control plane communication. For example, when the first instruction is an I / O instruction, root node 1 encapsulates the first instruction according to the message format of a business-type message to obtain the first message. Root node 1 sends the first message through the second interface. Switch board 1 determines that the first message is a business-type message based on the fields carried in the message header and the transmission interface of the first message, and forwards the first message to the second interface of leaf node 1. Leaf node 1 receives the first message, determines that the first message is a business-type message based on the fields carried in the message header and the transmission interface of the first message, and computing device 1 implements data plane communication. The above three optional implementation methods are only different ways to achieve isolation between control plane communication and data plane communication. In other embodiments, computing device 1 can also use other implementation methods to achieve isolation between control plane communication and data plane communication, and this application does not limit this.
[0338] The following example illustrates how to implement isolation between control plane communication and data plane communication by using fields carried in the message header.
[0339] In one alternative implementation, the protocol followed by control plane communication is the same as that followed by data plane communication.
[0340] For example, taking the data communication between the components in computing device 1 via a data network as an example, computing device 1 implements data plane communication and control plane communication through a data communication protocol.
[0341] In some alternative approaches, the data communication protocol may be an Ethernet communication protocol, a UB communication protocol, or other types of communication protocols. This application does not limit the specific type of communication protocol.
[0342] In some optional configurations, the message format of the data communication protocol includes a source address, a destination address, and an identifier (first identifier or delegate identifier). The source address indicates the node sending the management message; for example, the source address can be, but is not limited to, the IP address, MAC address, or device identifier of the node sending the management message. The destination address indicates the node receiving the management message; for example, the destination address can be, but is not limited to, the IP address, MAC address, or device identifier of the node receiving the management message.
[0343] In other alternative approaches, to further distinguish between management plane communication and data plane communication, an extension header field can be added to the header of management plane messages compared to service plane messages. This extension header field carries a first flag bit. The first flag bit is used to indicate the execution of management plane communication.
[0344] The first flag bit may also be called a high security bit, security flag bit, control plane communication flag bit, or other names, and this application does not limit it to any particular name.
[0345] For example, taking the UB protocol as the data communication protocol, the message format of management messages will be illustrated.
[0346] like Figure 8 As shown, the message format of this management message includes an Ethernet header, an IP header, a UDP header, a basic transport header, an extended header, and the actual data.
[0347] like Figure 8 As shown, the extended header includes an extended header field, which contains a first flag bit. In some alternative implementations, this extended header field may also be referred to as an extended field, a reserved field, or other names, which are not limited in this application.
[0348] The Ethernet header contains the destination device identifier and the source device identifier. For example, if the root node 1 sends the first message to the leaf node 1, the destination device identifier is the device identifier of the leaf node 1, and the source device identifier is the device identifier of the root node 1.
[0349] The IP header contains the source network address and the destination network address. For example, taking the first packet sent from root node 1 to leaf node 1 as an example, the source network address is the network address corresponding to the source device identifier (the device identifier of root node 1), and the destination network address is the network address corresponding to the destination device identifier (the device identifier of leaf node 1).
[0350] In some alternative methods, the destination device identifier and destination network address are used to provide the destination address, indicating the node receiving the management message. The source device identifier and source network address are used to provide the source address, indicating the node sending the management message. The UDP header contains the source port, destination port, and length. The length is the word length of the UDP header and data. Taking the root node 1 sending the first message to leaf node 1 as an example, the source port is the port of the root node 1, such as FE. The destination port is the port of the leaf node 1, such as FE.
[0351] The basic transport header includes at least a message type (MTYPE). The message type (MTYPE) indicates the message type. In the message format of the control plane hardware protocol, the message type (MTYPE) carries a first identifier.
[0352] The actual data contains at least an opcode, which indicates the type of management operation.
[0353] In data plane communication, the message format of service-type messages includes the Ethernet header, IP header, UDP header, and basic transport header mentioned above. Furthermore, the message type (MTYPE) in the basic transport header carries a second identifier.
[0354] In one alternative implementation, when the first instruction is a control instruction, root node 1 follows... Figure 8 The management message format shown encapsulates the first instruction to form the first message. When the first instruction is an I / O instruction, root node 1 encapsulates the first instruction according to the UB protocol to form the first message.
[0355] from Figure 8 As can be seen from the provided embodiments, by adding an extended header to carry a first flag bit and carrying a first identifier code through the message type (MTYPE), the switch board or leaf node can determine whether the first message is a management message based on the message type (MTYPE) field in the first message and whether the first message carries the first flag bit. In this way, the logical isolation between data plane communication and control plane communication is achieved by using the first flag bit and the message type (MTYPE).
[0356] In one optional implementation, to ensure the security of control plane communication, root node 1 uses the value of its hardware protocol flag as the value of the first flag bit in the first message during the control plane communication processing. Thus, switch board 1 or leaf node 1 can determine whether the first message is a management message sent by the root node based on the value of the first flag bit, thereby determining the legitimacy of the first message. This enables switch board 1 or leaf node 1 to block invalid first messages, ensuring the security of control plane communication.
[0357] In some alternative embodiments, the first flag bit can be a first value or a fourth value. This application does not limit the specific form of the first and fourth values. For example, the first value may be "00" or "11", and the fourth value may be "11" or "00". Alternatively, the fourth value may be "00" or "11", and the first value may be "11" or "00". Furthermore, the first and fourth values can also be characters, such as the first value being "F" and the fourth value being "T", or the first value being "T" and the fourth value being "F".
[0358] Please see Figure 9 , Figure 9 A schematic diagram of the data communication process provided in this application is shown. The data communication process shown includes stages ① to ⑥, wherein:
[0359] Phase ①: Root node 1 encapsulates the first instruction to form the first message.
[0360] In one optional implementation, the first instruction is a control instruction, where root node 1 uses the value of the hardware protocol flag bit as the value of the first flag bit in the first message, and follows the above... Figure 8 The provided message format is used to encapsulate the first instruction, resulting in the first message. For example... Figure 9 As shown, the value of the first flag bit carried in the first message is the first value.
[0361] In one optional implementation, the first instruction is an I / O instruction. Root node 1 encapsulates the first instruction according to the message format provided by the UB protocol to obtain the first message.
[0362] Phase ②: Root node 1 sends the first message.
[0363] In one alternative implementation, root node 1 sends a first message to switching board 1 with reference to the above S810, which will not be elaborated upon in this application.
[0364] In stage ③, the switching board 1 determines that the first message is a management message based on the message format of the first message, and performs a legality verification on the first message based on the value of the first tag of the first message.
[0365] Corresponding to the processing in stage ②, in stage ③, switching board 1 receives the first message sent by root node 1.
[0366] like Figure 9 As shown, switchboard 1 determines whether the first message includes the first flag bit (S101). If the first message includes the first flag bit, it determines whether the validity verification passes (S102). If the first message does not include the first flag bit, switchboard 1 forwards the first message to leaf node 1 (S102). If the validity verification of the first message fails, switchboard 1 discards the first message (S103).
[0367] In the first optional implementation, if the message type (MTYPE) of the first message carries a first identifier, the switch board 1 identifies whether the first message carries a first flag bit. If the first message contains the first flag bit, the first message is determined to be a management message. The value of the first flag bit is compared with a reference value. If the value of the first flag bit matches the reference value, the first message is determined to be a management message sent by root node 1, i.e., the validity verification of the first message passes. If the value of the first flag bit does not match the reference value, the first message is determined not to be a management message sent by root node 1, i.e., the validity verification of the first message fails. The reference value is the value of the hardware protocol flag bit of root node 1; for example, the reference value can be the aforementioned first value, or the reference value can be the aforementioned value of the hardware protocol flag bit of root node 1.
[0368] The phrase "the first message is not a management message sent by root node 1" can be, but is not limited to: the first message being a management message sent by a leaf node, or the first message being a management message sent by another root node, or the first message being a forged management message.
[0369] Here, "the value of the first tag bit matches the reference value" can mean that the value of the first tag bit is the same as the reference value. Alternatively, "the value of the first tag bit matches the reference value" can also mean that the encoded value of the first tag bit is the same as the encoded value of the reference value.
[0370] In the second optional implementation, if the message type (MTYPE) of the first message carries a first identifier, includes a first flag bit, and the first flag bit matches the base value, the switch board 1 queries the routing data according to the destination network address in the first message. If the routing data contains a first network address that matches the destination network address in the first message, and the device identifier corresponding to the first network address matches the destination device identifier in the first message, the validity verification of the first message is determined to be successful. If the routing data does not contain a first network address that matches the destination network address in the first message, or if the routing data contains a first network address that matches the destination network address in the first message, but the device identifier corresponding to the first network address does not match the destination device identifier in the first message, the validity verification of the first message is determined to be unsuccessful.
[0371] In the third optional implementation, if the message type (MTYPE) of the first message carries a first identifier, includes a first flag bit, and the first flag bit matches the base value, the switching board 1 queries the routing data according to the first source address of the first message. If the second source address of the root node recorded in the routing data matches the first source address, the first message is determined to be a management message sent by root node 1, i.e., the validity verification of the first message passes. If the second source address of the root node recorded in the routing data does not match the first source address, the first message is determined not to be a management message sent by root node 1, i.e., the validity verification of the first message fails.
[0372] The phrase "the second source address is the same as the first source address" can include both the source network address in the second source address being the same as the source network address in the first source address, and the source device node in the second source address being the same as the source device node in the first source address. Conversely, the phrase "the second source address is not the same as the first source address" can include both the source network address in the second source address being different from the source network address in the first source address, and / or the source device node in the second source address being different from the source device node in the first source address.
[0373] The three optional implementation methods described above are merely different ways for the switching board 1 to perform legality verification on the first message. In other embodiments, the switching board 1 may also employ other implementation methods to perform legality verification on the first message. For example, if the first message contains a first flag bit, the first flag bit matches the reference value, and the transmission interface of the first message is the first interface, the legality verification of the first message passes. Conversely, if the first message does not contain a first flag bit, or the transmission interface of the first message is the second interface, the legality verification of the first message fails. This application does not impose any limitations on this.
[0374] Phase 4: If the validity of the first message passes the verification, the switching board 1 forwards the first message to the leaf node 1.
[0375] In one alternative implementation, switchboard 1 can refer to the above S820 to forward the first message to leaf node 1.
[0376] Phase 5: Leaf node 1 receives the first message, determines that the first message is a management message based on its message format, and performs a validity verification on the first message based on the value of the first tag.
[0377] In one alternative implementation, leaf node 1 can refer to the validity verification method of the exchange board 1 to verify the validity of a message, which will not be elaborated here.
[0378] In one alternative implementation, if the validity verification of the first message fails, leaf node 1 discards the first message (S104).
[0379] In one alternative implementation, if the first message does not contain the first flag bit, leaf node 1 processes the first message (S105).
[0380] In stage 6, if the validity of the first message passes the verification, leaf node 1 processes the first message.
[0381] In one alternative implementation, leaf node 1 can process the first message with reference to S830 described above. Further details are omitted here.
[0382] based on Figure 9 In the provided embodiment, during control plane communication, the root node sets the value of the first flag bit of the first packet based on the value of the root node's hardware protocol flag bit, blocking other nodes from impersonating and sending management packets. Furthermore, the switch board can determine the legitimacy of the first packet based on the value of the first flag bit, thereby blocking in-band network impersonation of management packets. Additionally, leaf nodes can also determine the legitimacy of the first packet based on the value of the first flag bit, preventing injection attacks of management packets from the end side. Thus, through these three security mechanisms, the security of control plane communication can be guaranteed.
[0383] The above Figure 8 and Figure 9 This section uses the presence or absence of a first flag bit as an example to illustrate the isolation methods between control plane communication and data plane communication. In another optional implementation, isolation between control plane communication and data plane communication is also achieved by distinguishing the values of the flag bits carried in the message.
[0384] The following example illustrates the isolation method between control plane communication and data plane communication, using the example that the protocol followed by control plane communication is the same as that followed by data plane communication.
[0385] For example, taking data communication between components in computing device 1 via a UB network as an example, computing device 1 implements data plane communication and control plane communication through a control plane hardware protocol. This control plane hardware protocol is as follows: Figure 10 As shown above. Figure 8 The message formats shown are similar, in Figure 10 The message format of the control plane hardware protocol also includes an Ethernet header, IP header, UDP header, basic transport header, extended header, and the actual data. However, it differs from the above. Figure 8 The message format of the control plane hardware protocol shown is as follows: In control plane communication, it is as follows: Figure 10As shown in Figure (a), the header of a management message includes an Ethernet header, an IP header, a UDP header, a basic transport header, a first extended header, and the actual data. The first extended header includes a second flag bit. This second flag bit is used to distinguish between control plane communication and data plane communication.
[0386] In some alternative methods, such as Figure 10 As shown in Figure (a), the second flag bit is set with a second value, which is used to indicate the execution of control plane communication. In some optional examples, the second value is the identifier of the node that sends the management class message. In other optional examples, the second value is the encoded value of the identifier of the node that sends the management class message. Furthermore, in some optional examples, the second value may also be other values. This application does not limit the specific content of the second value.
[0387] During data plane communication, such as Figure 10 As shown in Figure (b), the message header of a service-type message includes an Ethernet header, an IP header, a UDP header, a basic transport header, a second extended header, and the actual data. The first extended header includes a second flag bit. This second flag bit is used to distinguish between control plane communication and data plane communication. For example, a value of the second flag bit (the third value) indicates that data plane communication is being performed.
[0388] In some alternative approaches, during management plane communication, the first extended header may also include the aforementioned first flag bit, such as... Figure 10 As shown in Figure (a), the value of the first flag bit is the first value mentioned above. Thus, the first message is verified multiple times using the identifier (first identifier or second identifier), the first flag bit, and the second flag bit to ensure the reliability of management plane communication.
[0389] In some alternative methods, the first extension header may also be referred to as the first extension header field, and the second extension header may also be referred to as the second extension header field.
[0390] In some alternative approaches, during data plane communication, the second extended header of the business class message does not include the first flag bit.
[0391] In other alternative methods, during data plane communication, the second extended header of the service-type message header may also include a first flag bit, the field of which can be empty, such as... Figure 10 As shown in Figure (b) of the document.
[0392] In this way, the switching board or leaf node can determine whether the first message is a management message based on the value of the second flag bit carried by the first message, thus using the value of the second flag bit to achieve logical isolation between data plane communication and control plane communication.
[0393] This application does not specify the exact numerical values of the second and third values.
[0394] With the above Figure 8 and Figure 9 Similar to the provided embodiments, when the protocol followed by control plane communication is the same as that followed by data plane communication, to ensure the security of control plane communication, the legitimacy of the first message can also be verified by the value of the first flag bit carried in the first message. For example... Figure 11 As shown, Figure 11 Schematic diagram of data communication provided for this application Figure 2 The data communication process shown includes stages ① to ⑥, wherein:
[0395] Phase ①: Root node 1 encapsulates the first instruction to form the first message.
[0396] In one optional implementation, the first instruction is a control instruction. Root node 1 uses the value of the hardware protocol flag bit as the value of the first flag bit in the first message and uses the second value as the value of the second flag bit, and follows the above... Figure 10 The provided message format is used to encapsulate the first instruction to obtain the first message.
[0397] In one optional implementation, the first instruction is an I / O instruction. Root node 1 sets the value of the first flag bit in the first message to null, uses the fourth value as the value of the second flag bit in the first message, and follows the above... Figure 10 The provided message format is used to encapsulate the first instruction to obtain the first message.
[0398] Phase ②: Root node 1 sends the first message.
[0399] In one alternative implementation, such as Figure 11 As shown, root node 1 sends the first message to switching board 1 with reference to the above S810, which will not be described in detail in this application.
[0400] In stage ③, the switching board 1 determines that the first message is a management message based on the message format of the first message, and performs a legality verification on the first message based on the value of the first tag of the first message.
[0401] Corresponding to the processing in stage 2, in stage 3, switching board 1 receives the first message.
[0402] In the first optional implementation, if the first message carries a first identifier and the second flag bit is set to a second value, the switch board 1 determines that the first message is a management message. If the value of the first flag bit in the first message matches the reference value, the switch board 1 determines that the validity verification of the first message has passed. If the value of the first flag bit does not match the reference value, the validity verification of the first message has failed. The reference value is the value of the hardware protocol flag bit of the root node 1; for example, the reference value can be the aforementioned first value.
[0403] like Figure 11 As shown, the switching board 1 determines whether the second flag bit of the first message is the second value (S121). If the second flag bit of the first message is the second value, the switching board 1 performs a validity verification based on the value of the first flag bit of the first message (S122). If the validity verification of the first message fails, the switching board 1 discards the first message (S123).
[0404] In some optional examples, if the first message carries a first identifier but the second flag bit is not a second value, the switch 1 determines that the validity verification of the first message fails. Alternatively, if the first message does not carry a first identifier but the second flag bit is a second value, the switch 1 determines that the validity verification of the first message fails. Here, "the second flag bit is not a second value" can be a third value or any other value inconsistent with the second value.
[0405] In the second optional implementation, if the first message carries a first identifier, the second flag bit is a second value, and the value of the first flag bit in the first message matches the base value, the switch board 1 can also query routing data according to the destination network address in the first message. If a first network address matching the destination network address in the first message exists in the routing data, and the device identifier corresponding to the first network address matches the destination device identifier in the first message, the validity verification of the first message is determined to be successful. If no first network address matching the destination network address in the first message exists in the routing data, or if a first network address matching the destination network address exists in the routing data, but the device identifier corresponding to the first network address does not match the destination device identifier in the first message, the validity verification of the first message is determined to be unsuccessful.
[0406] In the third optional implementation, if the first message carries a first identifier, the second flag bit is set to a second value, and the value of the first flag bit in the first message matches the base value, the switching board 1 queries the routing data according to the first source address of the first message. If the second source address of the root node recorded in the routing data matches the first source address, the first message is determined to be a management message sent by root node 1, i.e., the validity verification of the first message passes. If the second source address of the root node recorded in the routing data does not match the first source address, the first message is determined not to be a management message sent by root node 1, i.e., the validity verification of the first message fails.
[0407] The above three optional implementation methods are merely different ways for the switching board 1 to perform legality verification on the first message. In other embodiments, the switching board 1 may also use other implementation methods to perform legality verification on the first message. This application does not limit this.
[0408] In one alternative implementation, such as Figure 10 As shown, in one optional implementation, when the switching board 1 receives the first message, and the first message carries a second identifier and the value of the second flag bit is a third value, it determines that the first message is a service message, and the switching board 1 forwards the first message to the leaf node 1 (S124).
[0409] In some alternative methods, if the first message does not carry a second identifier or the value of the second flag bit is not a third value, the switching board 1 discards the first message. Here, "the first message does not carry a second identifier" can mean that the first message does not carry any identifier, or that the identifier carried by the first message is inconsistent with the second identifier. "The value of the second flag bit is not a third value" can mean that the value of the second flag bit is the aforementioned second value, or that the value of the second flag bit is inconsistent with the aforementioned third value.
[0410] Phase 4: If the validity of the first message passes the verification, the switching board 1 forwards the first message to the leaf node 1.
[0411] In one alternative implementation, switchboard 1 can refer to the above S820 to forward the first message to leaf node 1.
[0412] Phase 5: Leaf node 1 receives the first message, determines that the first message is a management message based on its message format, and performs a validity verification on the first message based on the value of the first tag.
[0413] In one optional implementation, if the first message carries a first identifier and the value of the second flag bit is the second value, leaf node 1 determines that the first message is a management message. Leaf node 1 compares the value of the first flag bit in the first message with the value of the hardware protocol flag bit of leaf node 1. If the value of the first flag bit matches the value of the hardware protocol flag bit of leaf node 1, the validity verification of the first message is determined to have failed. If the value of the first flag bit does not match the value of the hardware protocol flag bit of leaf node 1, but the value of the first flag bit matches the baseline value, the validity verification of the first message is determined to have passed.
[0414] In one alternative implementation, if the validity verification of the first message fails, leaf node 1 discards the first message (S125).
[0415] In one optional implementation, if the first message carries a second identifier and the value of the second flag bit is a third value, the first message is determined to be a service message (S126).
[0416] In stage 6, if the validity of the first message passes the verification, leaf node 1 processes the first message.
[0417] In one alternative implementation, leaf node 1 can process the first message with reference to S830 described above. Further details are omitted here.
[0418] based on Figure 11 The provided embodiment isolates control plane communication from data plane communication through an identifier (either a first identifier or a second identifier) and a second flag bit. In the case of control plane communication, the root node sets the value of the first flag bit of the first message based on the value of the root node's hardware protocol flag bit, blocking other nodes from impersonating and sending management messages. Furthermore, the switch board can determine the legitimacy of the first message based on the value of the first flag bit, thereby blocking in-band network impersonation of management messages. Additionally, leaf nodes can also determine the legitimacy of the first message based on the value of the first flag bit, preventing injection attacks of management messages from the end-side. Thus, through a three-round security mechanism using the identifier (either a first identifier or a second identifier), the second flag bit, and the first identifier, the security of control plane communication can be guaranteed.
[0419] above Figures 7 to 11 This paper describes the data communication method provided in this application using a single switchboard in computing device 1 as an example. In practical applications, multiple switchboards can be configured in computing device 1, with different switchboards managing different leaf nodes. As described above... Figure 5 and Figure 6As shown. In a scenario with multiple switching boards, root node 1 can send a first message to each switching board, so that each switching board forwards the received first message to the leaf node it manages. As described above. Figure 5 As shown.
[0420] Furthermore, unlike the data communication method in a single-switchboard scenario, in a multi-switchboard scenario, if the communication link between root node 1 and the first switchboard among the multiple switches fails or is disconnected, root node 1 can send the first message through a backup channel. Alternatively, if the communication link between the second switchboard among the multiple switches and the leaf nodes it manages is disconnected or abnormal, the second switchboard can forward the first message to the leaf nodes it manages through a backup channel, as described above. Figure 6 As shown. In this way, in the event of a communication link failure between the switch board and the root node, or a communication link failure between the leaf node and the switch board, a backup channel can be used to achieve data communication, thereby improving the reliability of data communication within the computing device.
[0421] To achieve the functions of the above embodiments, the computing device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the units and method steps described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0422] The above Figures 2 to 11 This paper details the computing device cluster, computing device, and data communication method described in this application. A root node is set up in the computing device to manage and control the leaf nodes within the computing device. Furthermore, during the management plane communication between the root node and leaf nodes, the data plane communication network of the computing device is reused, eliminating the need to create an additional management plane network, achieving in-band management and reducing the hardware cost of the computing device. Moreover, the permissions of the root node and leaf nodes in management plane communication are configured through hardware protocol flags, and the isolation between management plane communication and data plane communication is achieved using message flags, reducing software overhead. In addition, during management plane communication, message validity is verified through hardware protocol flags to ensure the security of management plane communication. Furthermore, a primary and backup channel is established for management plane communication. In the event of a communication link failure between the switch board and the root node, or a communication link failure between a leaf node and the switch board, the backup channel can be used for data communication, thereby improving the reliability of data communication within the computing device.
[0423] This application also provides a communication device capable of implementing the above-described data communication method. This communication device can be implemented using software units. For example, the communication device can be applied to the above-described computing device, or to the above-described root node, or to the above-described leaf node. In some optional scenarios, the communication device may include: a communication module, a processing module, and a storage module. The storage module is used to store program code and data generated by the communication device during data communication, such as protocol, routing data, or other data. In a first optional scenario, the communication module is used to send management-type messages or service-type messages, or to receive service-type messages. The processing module is used to control the communication module to send management-type messages during management plane communication, or the processing module is also used to control the communication module to send service-type messages, or to control the communication module to receive service-type messages during data plane communication.
[0424] In the second optional scenario, the communication module is also used to receive management messages or service messages, and to send service messages. The processing module is also used to control the communication module to receive management messages during management plane communication, or to control the communication module to send service messages or receive service messages during data plane communication.
[0425] In some optional cases, the communication device of this application embodiment can be implemented by a processing chip. The communication device according to the embodiment of this application can correspond to the execution of the methods described in the embodiment of this application, and the above and other operations and / or functions of each unit and module in the communication device are respectively to implement the corresponding processes of the various methods in the foregoing drawings. For the sake of brevity, they will not be described again here.
[0426] For example, the processing chip may include control circuitry and interface circuitry. The interface circuitry is used to receive data from other devices outside the processing chip and transmit it to the control circuitry, or to send data from the control circuitry to other devices outside the processing chip. The control circuitry implements the functions of the aforementioned root node through logic circuitry or executed code instructions and the interface circuitry, or the control circuitry implements the functions of the aforementioned leaf node through logic circuitry or executed code instructions and the interface circuitry.
[0427] The method steps in this embodiment can be implemented in hardware or by a processing chip executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processing chip, enabling the processing chip to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processing chip. The processing chip and the storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a computing device. Of course, the processing chip and the storage medium can also exist as discrete components in a network device, computing device, or terminal device.
[0428] This application also provides a chip system including a processing chip for implementing the functions of the root node or leaf node in the above-described method. In one possible design, the chip system further includes a memory for storing program instructions and / or data. This chip system can be composed of chips or may include chips and other discrete devices.
[0429] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product runs on at least one computing device (processor chip), it causes the at least one computing device to perform the aforementioned data communication method.
[0430] For example, when a computer program product is run on at least one computing device, it causes the at least one computing device to perform... Figure 7 The data communication method shown.
[0431] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be an internal storage unit of the computing device in any of the foregoing embodiments, such as a hard disk or memory of the computing device. The computer-readable storage medium can also be an external storage device of the computing device, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computing device. Further, the computer-readable storage medium can include both internal storage units and external storage devices of the computing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computing device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0432] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions of the embodiments of this application are performed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0433] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A computing device, characterized in that, The computing device includes multiple computing units, which are interconnected via a bus; each computing unit includes a root node and multiple leaf nodes. The root node is used to: send a first message to the leaf node via the bus; the first message includes a service message or a management message; the first message carries a first identifier or a second identifier, the first identifier being used to indicate that the first message is a management message, and the second identifier being used to indicate that the first message is a service message; The leaf node is used to receive the first message.
2. The computing device according to claim 1, characterized in that, The root node of the plurality of computing units is provided with an identifier; the identifier is a first value; wherein the first value is used to indicate the permission to send management messages.
3. The computing device according to claim 2, characterized in that, The management message carries a first flag bit, which is used to indicate the execution of management plane communication; when the first message is the management message, the value of the first flag bit is related to the identifier of the root node.
4. The computing device according to claim 3, characterized in that, The service-type messages and the management-type messages are implemented based on the same data communication protocol, and the first flag bit is carried by the extension header field of the data communication protocol; The header of the management message includes: the first identifier and an extended header field; the extended header field is used to carry the first flag bit; The header of the business message includes: the second identifier.
5. The computing device according to claim 1 or 2, characterized in that, The management message and the service message carry a second flag bit; the second flag bit of the management message has a second value, and the second flag bit of the service message has a third value; the second value is used to indicate the execution of management plane communication, and the third value is used to indicate the execution of data plane communication.
6. The computing device according to claim 5, characterized in that, The service-type messages and the management-type messages are implemented based on the same data communication protocol, and the second flag bit is carried by the extension header field of the data communication protocol; The header of the management message includes: the first identifier and a first extended header field; the first extended header field contains the second value; The header of the service message includes: the second identifier and the second extended header field; the second extended header field contains the third value.
7. The computing device according to any one of claims 1 to 6, characterized in that, The root node is further specifically used for: receiving a first instruction and sending a first message to the leaf node via the bus; when the first instruction indicates management plane communication, the first message is a management type message; when the first instruction indicates data plane communication, the first message is a service type message; the first message carries the first instruction. Alternatively, the root node may be specifically used to: receive a second message and send a first message to the leaf node via the bus; the first message is the service type message; the second message and the first message have the same message body.
8. The computing device according to any one of claims 1 to 7, characterized in that, The root node is configured with a first port and a second port. The first port is used to send management messages, and the second port is used to send service messages.
9. The computing device according to any one of claims 1 to 8, characterized in that, The computing device further includes a first switching unit, which is connected to the root node and the plurality of leaf nodes via the bus; The root node is used to: send a first message to the leaf node through the bus and the first switching unit; The first switching unit is configured to: forward the first message to the leaf node when the first message is a management message sent by the root node, or forward the first message to the leaf node when the first message is a service message.
10. The computing device according to claim 9, characterized in that, The first message carries a first source address; the first switching unit stores the second source address of the root node; The first switching unit is further configured to: determine that the first message is a management message sent by the root node when the first message carries the first identifier and the first source address is consistent with the second source address of the root node.
11. The computing device according to any one of claims 1 to 10, characterized in that, The computing device includes a first switching unit and a second switching unit. The first switching unit stores first routing data, and the second switching unit stores second routing data. The first routing data includes the second source address of the root node and the destination address of the first leaf node among the leaf nodes. The second routing data includes the second source address of the root node and the destination address of the second leaf node among the leaf nodes. The root node is used to: send a first message to the first leaf node through the bus and the first switching unit, and send a first message to the second leaf node through the bus and the second switching unit; The first switching unit is configured to: forward the first packet to the first leaf node using the first routing data; The second switching unit is used to forward the first packet to the second leaf node using the second routing data.
12. The computing device according to claim 11, characterized in that, The root node is also used to: send the second routing data to the first switching unit, or send the first routing data to the second switching unit; The first switching unit is further configured to: forward the first message to the second switching unit, so that the second switching unit forwards the first message to the first leaf node through the first routing data; The second switching unit is further configured to: forward the first message to the first switching unit, so that the first switching unit forwards the first message to the second leaf node through the second routing data.
13. A data communication method, characterized in that, The method is applied to a computing device, the computing device comprising multiple computing units interconnected via a bus; each computing unit includes a root node and multiple leaf nodes; the method includes: The root node sends a first message to the leaf node via the bus; the first message includes a service message or a management message; the first message carries a first identifier or a second identifier, the first identifier being used to indicate that the first message is a management message, and the second identifier being used to indicate that the first message is a service message. The leaf node receives the first message.
14. The method according to claim 13, characterized in that, The root node is configured with an identifier; the identifier is a first value; wherein the first value is used to indicate permission to send management-type messages; the step of sending the first message to the leaf node via the bus includes: In the case of management plane communication, the root node sends a first message to the leaf node through the bus based on the root node's identifier; the first message carries a first flag bit, the value of which is related to the root node's identifier. In the case of data plane communication, the root node sends a first message to the leaf node through the bus.
15. The method according to claim 13 or 14, characterized in that, The computing device further includes a switching unit, which connects the root node and the plurality of leaf nodes via the bus; the first message is sent to the leaf nodes via the bus. The first message is sent to the leaf node via the bus and the switching unit; The switching unit forwards the first message to the leaf node.
16. The method according to claim 15, characterized in that, Before forwarding the first packet to the leaf node, the method further includes: The switching unit receives the first message; If the first message is a management message sent by the root node, the switching unit performs the operation of forwarding the first message to the leaf node; If the first message is a service-type message, the switching unit performs the operation of forwarding the first message to the leaf node.
17. The method according to claim 16, characterized in that, After the switching unit receives the first message, the method further includes: The switching unit determines that the first message is the management message sent by the root node when the first message carries the first identifier code, the first source address of the first message is consistent with the source address of the root node, the first message carries a first flag bit, and the first flag bit matches the first value.
18. A computing device cluster, characterized in that, It includes a first computing device and one or more second computing devices connected to the first computing device, wherein the second computing device is the computing device according to any one of claims 1 to 12; The first computing device is configured to: set a root node in at least one second computing device, and send a first instruction to the root node in the at least one second computing device, the first instruction being used to instruct the at least one second computing device to implement data plane communication or management plane communication.
19. The cluster according to claim 18, characterized in that, The first computing device is further configured to: select a new root node from a plurality of leaf nodes of the target computing device in the event of a root node failure of the target computing device; the target computing device includes a second computing device in which the root node of the at least one second computing device has failed.
20. A processing chip, characterized in that, The processing chip includes a control circuit and an interface circuit, wherein the control circuit, in conjunction with the interface circuit, implements the function of the root node in the computing device according to any one of claims 1 to 12; or, the control circuit, in conjunction with the interface circuit, implements the function of the leaf node in the computing device according to any one of claims 1 to 12.
21. A computer-readable storage medium, characterized in that, include: Computer software instructions, when invoked by a computing device, wherein the computing device implements the method of any one of claims 13 to 17.
22. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device performs the method of any one of claims 13 to 17.