A computing power network system suitable for ultra-large-scale intelligent computing center industrial parks
By designing a computing network system suitable for ultra-large-scale intelligent computing center industrial parks, the problem that traditional data center network architecture cannot adapt to ultra-large-scale intelligent computing centers has been solved, a flexible network architecture and cost-effectiveness have been achieved, and the stability and security of the network have been ensured.
Patent Information
- Application Number
- CN202411380580.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-09-30
AI Technical Summary
Traditional data center network architecture lacks computing power network design and is not suitable for ultra-large-scale intelligent computing center industrial parks. It is difficult to meet the requirements of flexible deployment and has high costs.
A computing network system suitable for ultra-large-scale intelligent computing center industrial parks has been designed, including a computing network for training, a computing network for inference, front-end and back-end storage networks, an in-band management network, and an out-of-band management network. Through the combination of components such as the Internet cloud platform, core switches, and dedicated lines, a unified computing network resource pool and intelligent scheduling are realized.
It provides a flexible network architecture, supports elastic expansion and compatibility, reduces costs, ensures network stability and security, and improves computing efficiency and resource utilization.
Smart Images

Figure CN119276729B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data center network architecture, and specifically relates to a computing power network system suitable for ultra-large-scale intelligent computing center industrial parks. Background Art
[0002] In recent years, with the increasing development of intelligent computing centers (which create new growth models for industrial parks through a variety of services such as cloud computing, artificial intelligence application development, high-performance computing outsourcing, and edge computing), traditional IDCs (Internet Data Centers) are facing the need to transform into intelligent computing centers. For developers, it is necessary to propose a top-level computing network design architecture for the entire ultra-large-scale intelligent computing center industrial park to meet the construction needs and the changing needs of the final business implementation.
[0003] At present, the traditional data center network architecture has the following disadvantages: (1) Lack of overall planning of computing power network and its integration with IDC computing network; (2) Inability to adapt to the project characteristics of ultra-large-scale intelligent computing center industrial park; (3) The common computing power network on the market is based on the products of different manufacturers and is not based on the perspective of the builder. The functions are not targeted, repeated and bloated, making it difficult to meet the needs of flexible deployment and the cost is high; (4) There is a lack of functional integration and a unified architecture to simplify the different computing power requirements for training and inference, different computing power cards, different storage configurations, computing power management and / or in-band and out-of-band management functions. Summary of the Invention
[0004] The purpose of the present invention is to provide a computing power network system suitable for ultra-large-scale intelligent computing center industrial parks, so as to solve the problems of the existing data center network architecture, such as lack of computing power network, inability to adapt to the project characteristics of ultra-large-scale intelligent computing center industrial parks, difficulty in meeting flexible deployment and high cost.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] In the first aspect, a computing power network system suitable for an ultra-large-scale intelligent computing center industrial park is provided, comprising a computing power network, a wide area network, and a platform area, wherein the computing power network comprises a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network, and an out-of-band management network; the in-band management network comprises a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch, and a second spine switch; the wide area network comprises an Internet cloud platform, a core switch, and a dedicated line; the core switch is respectively connected to the Internet cloud platform and the first switch; the dedicated line is connected to the first switch through the fourth leaf switch; Machine, the first switch is communicatively connected to the platform area, the first switch is communicatively connected to the front-end and back-end storage networks through the second switch, the first switch is communicatively connected to the out-of-band management network through the third switch, the front-end and back-end storage networks are communicatively connected to the training computing network and the inference computing network respectively through the third leaf switch, any two training servers in the training computing network are communicatively interconnected through the first leaf switch and the first spine switch, any two inference servers in the inference computing network are communicatively interconnected through the second leaf switch and the second spine switch, and the training computing network and the inference computing network are both compatible with the RoCE protocol and the IB protocol;
[0007] The Internet cloud platform is used to aggregate and form a unified computing power network resource pool;
[0008] The core switch is responsible for establishing connections with the Internet cloud platform and the in-band management network;
[0009] The first switch is used to forward source data from the Internet cloud platform to the front-end and back-end storage networks together with the second switch, and to forward management data from the platform area to the out-of-band management network together with the third switch;
[0010] The front-end and back-end storage networks are used to provide the source data required for training to the training computing network, and to provide the source data required for reasoning to the reasoning computing network.
[0011] Based on the above invention content, a comprehensive network architecture solution designed for the characteristics of ultra-large-scale intelligent computing center industrial parks is provided, namely, it includes a computing power network, a wide area network and a platform area, wherein the computing power network includes a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, the in-band management network includes a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch and a second spine switch, and the wide area network includes an Internet cloud platform, a core switch and a dedicated line, and through their mutual connection relationship and functional limitation, it can solve the problems of the existing data center network architecture, such as lack of computing power network, inability to be applicable to the project characteristics of ultra-large-scale intelligent computing center industrial parks, difficulty in meeting flexible deployment and high cost, and facilitate practical application and promotion.
[0012] In one possible design, the training computing network includes a plurality of the training servers, a plurality of the first leaf switches, and a plurality of the first spine switches, wherein each of the first leaf switches is communicatively connected to each of the training servers, and each of the first spine switches is communicatively connected to each of the first leaf switches.
[0013] And / or, the computing network for inference includes multiple servers for inference, multiple second leaf switches and multiple second spine switches, wherein each second leaf switch is respectively communicated with each server for inference, and each second spine switch is respectively communicated with each second leaf switch.
[0014] In one possible design, the front-end and back-end storage networks include multiple data storage servers and multiple third leaf switches;
[0015] Each of the third leaf switches is respectively communicatively connected to the second switch, each of the plurality of data storage servers, and each server in the computing network, wherein the computing network refers to the training computing network and the inference computing network.
[0016] In one possible design, the front-end and back-end storage networks use storage virtualization technology to abstract a portion of the multiple data storage servers into a logical storage resource pool to store hot data, and abstract another portion of the multiple data storage servers into another logical storage resource pool to store cold data.
[0017] And / or, each data storage server in the front-end and back-end storage networks supports single-port network cards and dual-port network cards.
[0018] In one possible design, the out-of-band management network includes multiple OOB channels, server management interfaces, and switch management interfaces, wherein the server management interfaces refer to management interfaces of servers in the training computing network, the inference computing network, and the front-end and back-end storage networks, and the switch management interfaces refer to management interfaces of switches in the training computing network, the inference computing network, and the front-end and back-end storage networks.
[0019] Each OOB channel in the plurality of OOB channels is communicatively connected to the third switch, and each OOB channel is communicatively connected to the server management interface or the switch management interface.
[0020] In one possible design, the first switches are leaf switches and there are at least two of them, the second switches are spine switches and there are multiple of them, and the third switches are leaf switches and there are multiple of them.
[0021] In one possible design, an IDC network is included that is communicatively connected to the core switch, wherein the IDC network includes a computer room core switch, a computer room aggregation switch, a computer room access switch, and a computer room server that are communicatively connected in sequence, the computer room core switch is communicatively connected to the core switch, and the computer room server is used to store various types of files, undertake general computing tasks, and perform data encryption and decryption operations.
[0022] In one possible design, the wide area network includes a traffic monitoring, analysis, and cleaning module communicatively connected to the core switch;
[0023] The traffic monitoring, analysis and cleaning module is used to monitor the network traffic flowing through the core switch in real time to identify whether there is abnormal traffic. If so, the network traffic is pulled to a traffic cleaning device or area to separate the abnormal traffic from the normal traffic in the network traffic, and the normal traffic is re-injected into the original path of the network traffic.
[0024] In one possible design, the wide area network includes a firewall disposed between the core switch and the first switch;
[0025] And / or, the dedicated line includes a telecom network dedicated line, a Unicom network dedicated line and / or a mobile / third-party network dedicated line.
[0026] In one possible design, the platform area includes an SDN controller, a heterogeneous computing power debugging platform, a monitoring platform, a storage management platform, a logging platform, and / or a bastion host;
[0027] The SDN controller is used to configure, manage, and monitor the network through a unified platform, and is responsible for managing and scheduling network traffic. It dynamically adjusts the path and distribution of traffic according to different application requirements and network conditions, and uniformly manages network devices through the southbound interface protocol OpenFlow, and uniformly distributes various configuration scripts required for the operation of network devices to the first switch.
[0028] The heterogeneous computing power debugging platform is used to monitor computing resources in real time and display monitoring results through dashboards and graphical interfaces. It also uses intelligent scheduling algorithms to dynamically allocate computing resources based on the needs and priorities of different tasks. After receiving various computing tasks, it places the tasks into a task queue and queues them according to their priorities and resource requirements. It also displays the status and progress of tasks to users. It also provides intelligent performance optimization suggestions through real-time monitoring and analysis of computing resources and task execution. It is also used to authenticate and authorize users to ensure that only legitimate users can access computing resources.
[0029] The monitoring platform is used to monitor the temperature and humidity of network equipment, hardware utilization, port traffic and / or network quality in real time. Through real-time data collection and analysis, it determines whether potential faults and / or abnormalities have occurred. If so, it triggers alarm actions before the problem occurs through a pre-set early warning mechanism. In addition, through monitoring and analysis of system performance indicators, it provides administrators with system performance bottlenecks and / or performance optimization directions.
[0030] The storage management platform is a software system or tool used to centrally manage and control storage resources. By building an intelligent operation and maintenance management platform, it can realize intelligent operation of the storage platform in terms of automated deployment, status monitoring, capacity prediction, performance optimization, remote inspection, fault diagnosis and / or hard disk failure prediction, as well as real-time monitoring of storage device performance indicators to promptly identify performance bottlenecks and problems, and improve storage system performance by adjusting storage configuration and optimizing storage layout;
[0031] The log platform is used to monitor log data in real time and promptly trigger alarm notifications to administrators when specific events or abnormal situations are discovered. In addition, when a system failure occurs, it responds to administrator inquiries about relevant log records to quickly locate the root cause of the problem. It also records various system operations and access logs to provide a basis for security audits, and through log analysis, it discovers potential security vulnerabilities and violations to enhance system security.
[0032] The bastion host is used to protect the network and data from intrusion and damage from internal and external users.
[0033] Beneficial effects of the above scheme:
[0034] (1) The present invention provides a comprehensive network architecture solution designed for the characteristics of ultra-large-scale intelligent computing center industrial parks, namely, it includes a computing power network, a wide area network and a platform area, wherein the computing power network includes a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, the in-band management network includes a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch and a second spine switch, and the wide area network includes an Internet cloud platform, a core switch and a dedicated line, and through their interconnection relationship and functional limitations, it can solve the problems of the existing data center network architecture, such as lack of computing power network, inability to be applicable to the project characteristics of ultra-large-scale intelligent computing center industrial parks, difficulty in meeting the requirements of flexible deployment and high cost, and facilitates practical application and promotion;
[0035] (2) It is elastically scalable, that is, the architecture design allows for "elastic" changes between private and public protocols, avoiding the problem of its intelligent computing network architecture being unable to "elastically expand" due to the binding of hardware and software. In contrast, existing technologies may be limited to the products or protocols of specific manufacturers and are difficult to flexibly adapt to changes in different scales and needs;
[0036] (3) Compatibility means it is compatible with computing power cards from different manufacturers at home and abroad, such as those from NVIDIA, Huawei, and other domestic card manufacturers, meeting the needs of different cluster sizes and easily expanding from small-scale clusters to large-scale computing clusters; this means that customers can freely choose the hardware that best suits their needs, rather than being restricted to a single supplier;
[0037] (4) Storage network customization, that is, customizing the storage network architecture according to user needs, supporting "single-port" and "dual-port" network application scenarios, and realizing partitioned storage of hot data and cold data. This provides higher flexibility and efficiency, while existing technologies may only support standard configurations and are not flexible enough;
[0038] (5) Out-of-band management network, which is physically isolated from the business network. This ensures that the out-of-band management channel can still operate normally when the business network is paralyzed, ensuring remote management of ultra-large cluster networks. This is an important security feature that can prevent management communications from being affected by business traffic;
[0039] (6) Intelligent scheduling and monitoring: the heterogeneous computing power debugging platform monitors and intelligently schedules computing resources in real time. The monitoring platform provides real-time data collection and analysis to promptly identify problems and optimize performance. This intelligent resource management can improve the overall efficiency and responsiveness of the system.
[0040] (7) Centralized management and security auditing: The SDN controller uniformly configures, manages, and monitors the network. The bastion host provides monitoring and recording of operational behaviors, enhancing system security. Centralized management and security auditing functions can help simplify operations, reduce human errors, and improve system transparency and security.
[0041] (8) Cost-effectiveness: Due to its elastic scalability and compatibility, this technology can help users save costs and avoid additional expenses caused by hardware binding. By optimizing resource usage, it can reduce waste and further reduce operating costs.
[0042] (9) Adaptability: This architecture takes into account the characteristics of ultra-large-scale intelligent computing center industrial parks and can adapt to the special needs of this environment, while existing technologies may not be specifically designed for this scale and complexity;
[0043] (10) It has the characteristics of system integration. That is, the solution provides an integrated platform area, including SDN controller, heterogeneous computing power debugging platform and monitoring platform. The tight integration of these components provides a unified management interface, simplifying the management and maintenance of the network;
[0044] (11) It has the characteristics of performance optimization, that is, through real-time monitoring and intelligent scheduling, the architecture can optimize network performance, ensure that key tasks obtain resources first, and improve overall computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 A schematic diagram of a computing power network system suitable for an ultra-large-scale intelligent computing center industrial park provided in an embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structures of the drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0048] It should be understood that although the terms first, second, etc. may be used herein to describe various objects, these objects should not be limited by these terms. These terms are merely used to distinguish one object from another. For example, a first object can be referred to as a second object, and similarly, a second object can be referred to as a first object without departing from the scope of the exemplary embodiments of the present invention.
[0049] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that there may be three relationships. For example, A and / or B can indicate three situations: A exists alone, B exists alone, or A and B exist at the same time. For another example, A, B and / or C can indicate the existence of any one of A, B and C or any combination of them. The term " / and" that may appear in this document describes another type of association object relationship, indicating that there may be two relationships. For example, A / and B can indicate two situations: A exists alone or A and B exist at the same time. In addition, the character " / " that may appear in this document generally indicates that the previous and next associated objects are in an "or" relationship.
[0050] Example 1
[0051] like Figure 1As shown, the computing power network system provided in this embodiment and applicable to the ultra-large-scale intelligent computing center industrial park includes but is not limited to a computing power network, a wide area network and a platform area, etc., wherein the computing power network includes but is not limited to a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, etc., the in-band management network includes but is not limited to a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch and a second spine switch, etc., the wide area network includes but is not limited to an Internet cloud platform, a core switch and a dedicated line, etc., the core switch is respectively connected to the Internet cloud platform and the first switch for communication, and the dedicated line is connected through the fourth leaf switch The switch is communicatively connected to the first switch, the first switch is communicatively connected to the platform area, the first switch is communicatively connected to the front-end and back-end storage networks through the second switch, the first switch is communicatively connected to the out-of-band management network through the third switch, the front-end and back-end storage networks are communicatively connected to the training computing network and the inference computing network through the third leaf switch, any two training servers in the training computing network are interconnected through the first leaf switch and the first spine switch, any two inference servers in the inference computing network are interconnected through the second leaf switch and the second spine switch, the training computing network and the inference computing network are both compatible with RoCE (RDMA over Converged Ethernet, which is a remote direct memory access technology based on Ethernet) protocol and IB (InfiniBand, literally translated as "infinite bandwidth" technology, abbreviated as IB, is a computer network communication standard for high-performance computing, it has extremely high throughput and extremely low latency, and is used for data interconnection between computers) protocol. Figure 1 As shown, the computing power network can be arranged in the intelligent computing building of the intelligent computing center industrial park, and the wide area network and platform network can be arranged in the equipment information building of the intelligent computing center industrial park.
[0052] The internet cloud platform is used to aggregate and form a unified computing network resource pool, that is, to connect originally dispersed and heterogeneous computing resources to form a large and unified computing network resource pool. This can facilitate the following uses: regardless of the physical location and specific form of computing resources, they can be uniformly managed and allocated on the platform; and it can intelligently analyze indicators such as the status (such as whether they are idle or in use), performance (such as computing speed or processing power), and utilization rate of various computing resources. By comprehensively considering these indicators, it can accurately and uniformly schedule computing power; and it can also intuitively display resource usage through dashboards and graphical interfaces, allowing administrators to understand the overall resource status at a glance (this visual presentation method greatly improves management efficiency and facilitates administrators to quickly identify potential problem areas); and it can also analyze resource usage patterns and historical trends to provide administrators with in-depth insights. By mining this data, it can understand the resource demand characteristics of different tasks and the usage patterns of resources in different time periods; and it can also provide optimization suggestions based on the analysis results, such as adjusting server configuration and optimizing application code. These suggestions help improve computing performance and resource utilization efficiency and reduce system operating costs.
[0053] The core switch is responsible for establishing connections with the internet cloud platform and the in-band management network. Specifically, it is used to select the optimal transmission path for data entering and exiting the intelligent computing center network based on the network topology and routing protocol, thereby improving network transmission efficiency and reliability. Through routing functions, it guides and distributes traffic of different types and / or priorities, ensuring that critical business traffic is processed and transmitted first and foremost, and rationally allocating network resources. Furthermore, to improve the stability of the entire wide area network, the number of core switches is preferably at least two, and each core switch is separately connected to the internet cloud platform and the first switch.
[0054] Dedicated lines are designed to provide high bandwidth, meeting the needs of businesses or institutions for large-scale data transmission, such as high-definition video conferencing and large-scale file transfers. They ensure fast and stable data transmission, effectively avoiding delays and lags caused by network congestion. They also feature low latency, which is crucial for applications with high real-time requirements, such as financial transactions, online gaming, and telemedicine. They ensure timely data transmission and processing, reducing response times. Compared to ordinary network connections, dedicated lines offer greater stability and minimal network fluctuations, effectively reducing packet loss during data transmission and ensuring continuous and reliable business operations. Dedicated lines can also achieve physical isolation, isolating the enterprise or institution's network from the public internet and mitigating security threats from external networks, such as hacker attacks and malware distribution. Specifically, dedicated lines include, but are not limited to, dedicated lines for China Telecom, China Unicom, and / or China Mobile / third-party networks, allowing users to select one based on resource availability. In addition, since the fourth leaf switch is the network interface between the in-band management network and the wide area network, the fourth leaf switch can also be regarded as belonging to the wide area network. Figure 1 shown.
[0055] The first switch is used, together with the second switch, to forward source data from the internet cloud platform to the front-end and back-end storage networks, and is used, together with the third switch, to forward management data from the platform area to the out-of-band management network. Through the aforementioned design, the in-band management network can import large amounts of source data to meet the needs of ultra-large-scale computing cluster customers, and manage the import of customer source data in-band, as well as enable data communication between the storage network and the platform area. This enables the computing nodes in the training computing network to begin training models based on algorithms (such as existing deep learning algorithms and optimization algorithms) so that they can be applied to the inference computing network for inference. Specifically, the first switch adopts a leaf switch (i.e., a Leaf switch, which is a next-level switch connected to a Spine switch, serving as an entry point for connecting to servers and other network devices. It is responsible for providing an interface connected to the data center network and providing network connections for servers and other devices; Leaf switches usually have more ports, support connections to multiple servers and devices, and provide high-bandwidth and low-latency data forwarding capabilities) and there are at least two of them, the second switch adopts a spine switch (i.e., a Spine switch, which is a core switch in a data center network, usually used to build a high-performance, scalable network architecture. It has high bandwidth and low latency characteristics, is used to connect multiple Leaf switches, and provides horizontal high-speed data transmission and forwarding functions; the Spine switch is responsible for achieving high availability and fault tolerance of the data center network, and achieves load balancing and redundancy through a multi-path network architecture) and there are multiple of them, and the third switch adopts a leaf switch and there are multiple of them (at this time, relative to the third switch, the second switch can be regarded as a spine switch). In addition, each of the first switches needs to communicate and connect with each of the second switches and each of the third switches separately to improve the stability of the entire in-band management network.
[0056] The front-end and back-end storage networks are used to provide the source data required for training for the training computing network, and to provide the source data required for reasoning for the inference computing network. The training computing network is used to implement specific model training functions, while the inference computing network is used to implement specific reasoning functions based on the trained model. Since they are both compatible with the RoCE protocol and the IB protocol, they are not limited to any one domestic or foreign hardware manufacturer's solution (currently domestic Ethernet manufacturers can only support RoCE protocol networking, not IB protocol networking). They can thus have extremely strong flexibility and adaptability for ultra-large-scale intelligent computing power network architectures, enabling "elastic" changes between private and public protocols, avoiding the problem of its intelligent computing network architecture being unable to "elastically expand" due to the binding of hardware and software. That is, they can achieve "elasticity" capabilities for large-scale intelligent computing power network architectures, greatly reducing the cost of intelligent computing networks.
[0057] Based on the detailed description of the above computing power network system, it can be seen that this embodiment provides a comprehensive network architecture solution designed for the characteristics of ultra-large-scale intelligent computing center industrial parks, namely, it includes a computing power network, a wide area network and a platform area, wherein the computing power network includes a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, the in-band management network includes a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch and a second spine switch, the wide area network includes an Internet cloud platform, a core switch and a dedicated line, and through their mutual connection relationship and functional limitation, it can solve the problems of the existing data center network architecture, such as lack of computing power network, inability to be applicable to the project characteristics of ultra-large-scale intelligent computing center industrial parks, difficulty in meeting flexible deployment and high cost, and facilitate practical application and promotion.
[0058] Preferably, the training computing network includes, but is not limited to, multiple training servers, multiple first leaf switches, and multiple first spine switches, wherein each first leaf switch is separately connected to each training server, and each first spine switch is separately connected to each first leaf switch. Specifically, the training servers are primarily used for, but not limited to, the following functions: simultaneously processing large amounts of data, greatly improving computational efficiency and enabling rapid feature extraction and analysis of large amounts of image data; since deep learning models typically require a significant amount of computing resources for training, the server significantly shortens training time by performing operations such as matrix multiplication and convolution in parallel. The first leaf switch serves as an edge device of the training computing network (i.e., the first leaf switch also belongs to the training computing network, a dual-state scenario), responsible for connecting various terminal devices, such as the training servers, and providing network access ports for these devices, enabling device-to-network connectivity. This means aggregating traffic from terminal devices and distributing it to appropriate destinations based on network topology and routing policies. It acts as both a traffic ingress and egress, ensuring efficient data transmission within the intelligent computing network. The first spine switch is located at the core of the training computing network (that is, the first spine switch also belongs to the training computing network, and there are two cases), which is responsible for forwarding a large amount of data traffic at high speed. It is usually equipped with a high-bandwidth port (such as 400Gbps or higher) and can quickly process data from each leaf switch. Through the fully connected topology of the aforementioned first spine switch and each first leaf switch, data can be transmitted in the network through multiple paths, which improves the reliability and bandwidth utilization of the network. That is, no matter which leaf switch the data enters the network from, it can reach the destination through the shortest path, reducing the delay in data transmission. It is also considered that a large amount of data needs to be transmitted quickly between different devices. For example, in large-scale deep learning training, massive training data needs to be transmitted from the storage device to the computing node (such as the training server). Therefore, through the aforementioned fully connected topology design, high bandwidth support can be provided for data transmission, meeting the high requirements of intelligent computing tasks for data throughput, ensuring that data can flow quickly and stably between nodes, avoiding data transmission bottlenecks, and improving overall computing efficiency.In addition, since the number of the first leaf switches and the first spine switches is adjustable, it can also adapt to different cluster sizes to achieve scalable construction and more flexible and changeable purposes, that is, adapt to the deployment of small clusters and large clusters (such as 32POD, 64POD, 128POD, 256POD, 512POD and 1024POD clusters; the aforementioned POD is an existing concept, usually containing high-performance computing chips, such as graphics processors and tensor processing units, etc., which are specially optimized for artificial intelligence and high-performance computing), fully meeting the positioning of 10,000-card-level clusters, and having the advantage of being easily expanded from small-scale clusters to large-scale computing clusters.
[0059] Preferably, the computing network for inference includes but is not limited to a plurality of the inference servers, a plurality of the second leaf switches and a plurality of the second spine switches, etc., wherein each of the second leaf switches is respectively communicated with each of the inference servers, and each of the second spine switches is respectively communicated with each of the second leaf switches. The inference server is mainly used for, but not limited to, realizing the following functions: large-scale data parallel processing when performing inference tasks (for example, in the inference scenario of image recognition, faced with the input of thousands of pictures, the graphics processor can simultaneously extract features and classify these pictures and quickly give results); complex model parallel computing (for example, for complex models such as deep neural networks, the graphics processor can efficiently perform matrix multiplication, convolution and other operations in parallel; in the inference process, these operations occupy most of the computing time, and the parallel processing capabilities of the graphics processor can significantly shorten the inference time, making real-time applications possible); in the inference process, input data needs to be continuously transmitted to the graphics processor for processing, and high-bandwidth memory can ensure fast data transmission to avoid data transmission becoming a bottleneck (for example, in the inference scenario of video stream analysis, the graphics processor can process high-resolution video streams in real time without delays due to slow data transmission speeds). ; When model inference begins, the pre-trained model parameters need to be loaded into the memory (the high memory bandwidth of the graphics processor can complete this process quickly, reducing startup time). The second leaf switch is used as an edge device of the computing network for inference (that is, the second leaf switch also belongs to the computing network for inference, and there are two cases), responsible for connecting various terminal devices, such as the server for inference, etc., and its specific details can be obtained by referring to the conventional derivation of the aforementioned first leaf switch, which will not be repeated here. The second spine switch is located at the core position of the computing network for inference (that is, the second spine switch also belongs to the computing network for inference, and there are two cases), and its specific details can be obtained by referring to the conventional derivation of the aforementioned first spine switch, which will not be repeated here. In addition, since the number of the second leaf switch and the second spine switch can be adjusted, it can also adapt to different cluster sizes to achieve scalable construction and more flexible and changeable purposes.
[0060] Preferably, the front-end and back-end storage networks include, but are not limited to, multiple data storage servers and multiple third leaf switches; each third leaf switch is communicatively connected to the second switch, each of the multiple data storage servers, and each server in the computing network, wherein the computing network refers to the training computing network and the inference computing network. The data storage servers are used to provide large-capacity storage space for the computing network, which typically needs to process large amounts of data (specifically, training data, model parameters, and intermediate results), to meet the storage needs of this data. To ensure data security and reliability, data backup is typically performed. The functions of the third leaf switches are similar to those of the first and second leaf switches and are not further described here. Specifically, each data storage server in the front-end and back-end storage networks supports both single-port and dual-port network cards. This allows for customization of the storage network architecture and compatibility with storage server selection requirements based on end-user needs. This customized network architecture can flexibly accommodate both single-port network service needs and dual-port network application scenarios, avoiding the manufacturer's rigid requirement for dual-port network architectures, thereby providing "elasticity" in the storage area network architecture. Specifically, the front-end and back-end storage networks use storage virtualization technology (which is an existing technology that can abstract multiple physical storage devices into a logical storage resource pool and manage and allocate them through a unified management interface. It can integrate storage devices of different types and brands to achieve unified management and allocation of storage resources; for example, storage area network virtualization and network attached storage virtualization are storage virtualization technologies) to abstract a portion of the multiple data storage servers into a logical storage resource pool to store hot data (i.e., data that is frequently accessed and requires real-time or near real-time processing), and to abstract another portion of the multiple data storage servers into another logical storage resource pool to store cold data (i.e., data that is infrequently accessed and has a long storage time). In this way, partitioning can be performed through resource pooling to achieve storage space for "hot data" and "cold data", thereby solving the different storage requirements of hot data and cold data, and avoiding the traditional vendors' strong binding and bloated configuration of hot storage servers and cold storage servers, thereby reducing unnecessary storage server expenses; at the same time, it is compatible with more storage server manufacturers (such as IBM, DDN and Netapp, etc.), and is more suitable for ultra-large-scale intelligent computing power cluster application scenarios.
[0061] Preferably, the out-of-band management network includes but is not limited to multiple OOB channels, server management interfaces and switch management interfaces, etc., wherein the server management interface refers to the management interface of the server in the training computing network, the inference computing network and the front-end and back-end storage network, and the switch management interface refers to the management interface of the switch in the training computing network, the inference computing network and the front-end and back-end storage network; each OOB channel in the multiple OOB channels is respectively communicatively connected to the third switch, and each OOB channel is respectively communicatively connected to the server management interface or the switch management interface. The aforementioned OOB (Out of Band) channel is used to provide a communication channel independent of the data network, so as to allow administrators to remotely access and control devices in the intelligent computing network at any time and any place; even if the data network fails or the device cannot start normally, the administrator can still connect to the device through the OOB channel to perform fault diagnosis and repair. In this way, the out-of-band management network can be physically isolated from the business networks such as the training computing network, the inference computing network, and the front-end and back-end storage networks within the computing power network. Even if the business network is completely paralyzed, the out-of-band management channel can still work normally, ensuring that the administrator can perform emergency processing and fault recovery on the equipment (this is necessary for ultra-large-scale intelligent computing power clusters, otherwise there will be the problem of "disconnection" of remote management of ultra-large cluster networks).
[0062] Preferably, it includes an IDC network that is communicatively connected to the core switch, wherein the IDC network includes but is not limited to a computer room core switch, a computer room aggregation switch, a computer room access switch and a computer room server that are communicatively connected in sequence, the computer room core switch is communicatively connected to the core switch, and the computer room server is used to store various types of files, undertake general computing tasks and perform data encryption and decryption operations. The computer room core switch is responsible for quickly forwarding data from the aggregation layer and other network areas; it has the characteristics of high bandwidth and low latency, and can ensure the efficient transmission of large amounts of data; for example, in a large enterprise network, this core switch can quickly process data interactions between departments and with external networks; in the backbone of the network, this core switch can connect various aggregation switches and other important network devices (such as servers and routers, etc.) together to form a unified network entity, so that networks in different areas can communicate with each other. The computer room aggregation switch is used to connect multiple computer room access switches, collect a large amount of data from the access layer, and perform preliminary processing and integration, and then forward this data to the computer room core switch. The computer room access switch is used to directly connect end-user devices, such as CPU (Central Processing Unit) servers and storage servers, etc., to provide network access ports for these devices so that they can access the network and realize communication and resource sharing with other devices. Specific examples of the aforementioned various types of files include but are not limited to documents, pictures and videos, so that users can access files on the computer room server through the network to realize file sharing and collaboration. Specific examples of the aforementioned general computing tasks include but are not limited to running operating systems and executing applications, etc. It can process various types of data, including text, numbers, images and audio, etc. And because data security is of vital importance in network communications, the confidentiality and integrity of data can be protected by performing data encryption and decryption operations. Therefore, through the configuration of the aforementioned IDC network, the computing power network and the IDC general computing network can also be coordinated to meet the business needs of the two major customers of computing power and IDC. In addition, the IDC network can be arranged in the IDC building of the intelligent computing center industrial park.
[0063] Preferably, the wide area network includes a traffic monitoring, analysis and cleaning module that is communicatively connected to the core switch; the traffic monitoring, analysis and cleaning module is used to monitor the network traffic flowing through the core switch in real time to identify whether there is abnormal traffic. If so, the network traffic is pulled to a traffic cleaning device or area to separate the abnormal traffic from the normal traffic in the network traffic, and the normal traffic is re-injected into the original path of the network traffic. The aforementioned abnormal traffic situation is specific but not limited to abnormal traffic with signs of DDoS (Distributed Denial of Service) attacks, and the type of attack can be further analyzed and determined: whether it is an attack at the network layer (such as SYN Flood), an attack at the application layer (such as HTTP Flood) or an attack at other specific protocols, etc., so as to provide a basis for subsequent targeted cleaning measures. In this way, through the configuration of the aforementioned traffic monitoring, analysis and cleaning module, normal business can be ensured to be unaffected, and normal traffic can be accurately delivered to the target server or application, ensuring the continuity and stability of the business. In addition, the traffic monitoring, analysis and cleaning modules need to correspond one-to-one with the core switches. That is, when there are at least two core switches, there also need to be at least two traffic monitoring, analysis and cleaning modules that correspond one-to-one with each other.
[0064] Preferably, the wide area network includes a firewall arranged between the core switch and the first switch. The firewall is used to strictly control access rights and data flows between different areas when dividing network areas (for example, dividing the network of the IDC computer room into different security areas, such as the internal server area, isolation area and external network area, etc.) according to preset rules (such as based on IP address, port number and protocol type, etc., to allow qualified network data packets to enter and exit the IDC computer room), and also prevent the spread of internal threats: even when a security incident occurs or malicious behavior occurs inside the IDC computer room, the firewall can limit the communication between the infected or problematic host and other devices, avoiding the rapid spread of security threats throughout the computer room. In addition, the firewall also needs to correspond one-to-one with the core switch, that is, when there are at least two core switches, the firewall also needs to be at least two and correspond one-to-one with one.
[0065] Preferably, the platform area includes but is not limited to an SDN controller, a heterogeneous computing power debugging platform, a monitoring platform, a storage management platform, a log platform and / or a bastion host, etc.
[0066] The SDN (Software Defined Network) controller is used to configure, manage and monitor the network through a unified platform (this is very necessary for ultra-large-scale campus equipment, and there is no need to operate each network device separately), and is responsible for the management and scheduling of network traffic, and dynamically adjusts the path and distribution of traffic according to different application requirements and network conditions, and uniformly manages network devices through the southbound interface protocol Openflow, and uniformly sends various configuration scripts required for the operation of network devices to the first switch.
[0067] The heterogeneous computing power debugging platform is used to monitor computing resources in real time and display monitoring results (such as GPU usage, CPU usage, memory usage, disk space and / or network bandwidth, etc.) through dashboards and graphical interfaces, and adopt intelligent scheduling algorithms (such as existing greedy algorithms, etc.) to dynamically allocate computing resources according to the needs and priorities of different tasks (for example, for high-priority tasks, such as critical artificial intelligence training tasks, more GPU resources and memory resources can be allocated first to ensure that the tasks can be completed quickly; and for low-priority tasks, they can be scheduled when resources are idle to avoid resource waste), and after receiving the submitted After completing various computing tasks (such as data analysis tasks, model training tasks, and scientific computing tasks submitted through a simple interface), the tasks are placed in a task queue and queued according to their priority and resource requirements. The status and progress of the tasks are displayed to the user. Through real-time monitoring and analysis of computing resources and task execution, intelligent performance optimization suggestions are provided (for example, based on the task execution time, resource usage, and performance indicators, the bottleneck of the task can be analyzed, and corresponding optimization suggestions can be made, such as adjusting task parameters, optimizing algorithms, and increasing resource allocation, etc.). It is also used to authenticate and authorize users to ensure that only legitimate users can access computing resources.
[0068] The monitoring platform is used to monitor the temperature and humidity of network equipment, hardware (such as graphics processors, central processing units and memory, etc.) usage, port traffic and / or network quality (such as packet loss rate and delay time) in real time, and through real-time data collection and analysis, determine whether potential fault problems and / or abnormal problems have occurred. If so, a pre-set early warning mechanism is used to trigger an alarm action before the problem occurs (this can help administrators have enough time to take preventive measures), and through monitoring and analysis of system performance indicators, provide administrators with system performance bottlenecks and / or performance optimization directions. For example, by analyzing network traffic data, the time period and cause of network congestion can be discovered, so that corresponding optimization measures can be taken, such as increasing bandwidth and optimizing network configuration.
[0069] The storage management platform is a software system or tool for centrally managing and controlling storage resources (which can help users utilize storage devices more efficiently, improve the utilization, reliability and manageability of storage resources, and meet the growing data storage and management needs of enterprises or individuals), and through the construction of an intelligent operation and maintenance management platform, it can realize intelligent operation of the storage platform in terms of automated deployment, status monitoring, capacity prediction, performance optimization, remote inspection, fault diagnosis and / or hard disk failure prediction, as well as real-time monitoring of storage device performance indicators (such as read and write speeds, the number of read and write operations per second and / or latency, etc.) to promptly discover performance bottlenecks and problems, and improve the performance of the storage system by adjusting storage configuration and optimizing storage layout.
[0070] The log platform is used to monitor log data in real time and trigger alarm notifications to administrators in a timely manner when specific events or abnormal situations are discovered. When a system failure occurs, it responds to administrator inquiries about relevant log records to quickly locate the root cause of the problem. It also records various system operations and access logs to provide a basis for security audits, and through log analysis, discovers potential security vulnerabilities and violations to enhance system security.
[0071] The bastion host is used to protect the network and data from intrusion and damage from internal and external users. Specifically, various technical means can be used to monitor and record operations performed by maintenance personnel on servers, network equipment, security devices, databases, and other devices within the network, enabling centralized alerting, timely processing, and auditing to determine accountability.
[0072] In summary, the computing power network system provided by this embodiment has the following technical effects:
[0073] (1) This embodiment provides a comprehensive network architecture solution designed for the characteristics of a super-large-scale intelligent computing center industrial park, namely, it includes a computing power network, a wide area network and a platform area, wherein the computing power network includes a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, the in-band management network includes a first switch, a second switch, a first leaf switch, a second leaf switch, a third leaf switch, a fourth leaf switch, a first spine switch and a second spine switch, and the wide area network includes an Internet cloud platform, a core switch and a dedicated line, and through their interconnection relationship and functional limitations, it can solve the problems of the existing data center network architecture, such as lack of computing power network, inability to be applicable to the project characteristics of a super-large-scale intelligent computing center industrial park, difficulty in meeting the requirements of flexible deployment and high cost, and facilitate practical application and promotion;
[0074] (2) It is elastically scalable, that is, the architecture design allows for "elastic" changes between private and public protocols, avoiding the problem of its intelligent computing network architecture being unable to "elastically expand" due to the binding of hardware and software. In contrast, existing technologies may be limited to the products or protocols of specific manufacturers and are difficult to flexibly adapt to changes in different scales and needs;
[0075] (3) Compatibility means it is compatible with computing power cards from different manufacturers at home and abroad, such as those from NVIDIA, Huawei, and other domestic card manufacturers, meeting the needs of different cluster sizes and easily expanding from small-scale clusters to large-scale computing clusters; this means that customers can freely choose the hardware that best suits their needs, rather than being restricted to a single supplier;
[0076] (4) Storage network customization, that is, customizing the storage network architecture according to user needs, supporting "single-port" and "dual-port" network application scenarios, and realizing partitioned storage of hot data and cold data. This provides higher flexibility and efficiency, while existing technologies may only support standard configurations and are not flexible enough;
[0077] (5) Out-of-band management network, which is physically isolated from the business network. This ensures that the out-of-band management channel can still operate normally when the business network is paralyzed, ensuring remote management of ultra-large cluster networks. This is an important security feature that can prevent management communications from being affected by business traffic;
[0078] (6) Intelligent scheduling and monitoring: the heterogeneous computing power debugging platform monitors and intelligently schedules computing resources in real time. The monitoring platform provides real-time data collection and analysis to promptly identify problems and optimize performance. This intelligent resource management can improve the overall efficiency and responsiveness of the system.
[0079] (7) Centralized management and security auditing: The SDN controller uniformly configures, manages, and monitors the network. The bastion host provides monitoring and recording of operational behaviors, enhancing system security. Centralized management and security auditing functions can help simplify operations, reduce human errors, and improve system transparency and security.
[0080] (8) Cost-effectiveness: Due to its elastic scalability and compatibility, this technology can help users save costs and avoid additional expenses caused by hardware binding. By optimizing resource usage, it can reduce waste and further reduce operating costs.
[0081] (9) Adaptability: This architecture takes into account the characteristics of ultra-large-scale intelligent computing center industrial parks and can adapt to the special needs of this environment, while existing technologies may not be specifically designed for this scale and complexity;
[0082] (10) It has the characteristics of system integration. That is, the solution provides an integrated platform area, including SDN controller, heterogeneous computing power debugging platform and monitoring platform. The tight integration of these components provides a unified management interface, simplifying the management and maintenance of the network;
[0083] (11) It has the characteristics of performance optimization, that is, through real-time monitoring and intelligent scheduling, the architecture can optimize network performance, ensure that key tasks obtain resources first, and improve overall computing efficiency.
[0084] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A computing network system suitable for ultra-large-scale intelligent computing center industrial park, characterized in that: It includes a computing power network, a wide area network and a platform area, wherein the computing power network includes a computing network for training, a computing network for inference, a front-end and back-end storage network, an in-band management network and an out-of-band management network, the computing network for training includes multiple training servers, multiple first leaf switches and multiple first spine switches, the front-end and back-end storage networks include multiple third leaf switches and multiple data storage servers, the computing network for inference includes multiple inference servers, multiple second leaf switches and multiple second spine switches, the out-of-band management network includes multiple OOB channels, server management interfaces and switch management interfaces, the in-band management network includes a first switch, a second switch and a third switch, the wide area network includes an Internet cloud platform, a core switch and a dedicated line, and the core switches are respectively connected to the Internet for communication. The cloud platform and the first switch, the dedicated line is connected to the first switch through the fourth leaf switch, the first switch is connected to the platform area, the first switch is connected to the front-end and back-end storage networks through the second switch, the first switch is connected to the out-of-band management network through the third switch, the front-end and back-end storage networks are respectively connected to the training computing network and the inference computing network through the third leaf switch, any two training servers in the training computing network are interconnected through the first leaf switch and the first spine switch, any two inference servers in the inference computing network are interconnected through the second leaf switch and the second spine switch, and the training computing network and the inference computing network are both compatible with the RoCE protocol and the IB protocol; The Internet cloud platform is used to aggregate and form a unified computing power network resource pool; The core switch is responsible for establishing connections with the Internet cloud platform and the in-band management network; The first switch is used to forward source data from the Internet cloud platform to the front-end and back-end storage networks together with the second switch, and to forward management data from the platform area to the out-of-band management network together with the third switch; The front-end and back-end storage networks are used to provide the source data required for training to the training computing network, and to provide the source data required for reasoning to the reasoning computing network.
2. The computing power network system according to claim 1, wherein: Each of the first leaf switches is communicatively connected to each of the training servers, and each of the first spine switches is communicatively connected to each of the first leaf switches. And / or, each of the second leaf switches is respectively communicatively connected to each of the inference servers, and each of the second spine switches is respectively communicatively connected to each of the second leaf switches.
3. The computing power network system according to claim 1, wherein: Each of the third leaf switches is respectively communicatively connected to the second switch, each of the plurality of data storage servers, and each server in the computing network, wherein the computing network refers to the training computing network and the inference computing network.
4. The computing power network system according to claim 3, wherein: The front-end and back-end storage networks use storage virtualization technology to abstract a portion of the multiple data storage servers into a logical storage resource pool for storing hot data, and abstract another portion of the multiple data storage servers into another logical storage resource pool for storing cold data; And / or, each data storage server in the front-end and back-end storage networks supports single-port network cards and dual-port network cards.
5. The computing power network system according to claim 1, wherein: The server management interface refers to the management interface of the server in the training computing network, the inference computing network and the front-end and back-end storage network; the switch management interface refers to the management interface of the switch in the training computing network, the inference computing network and the front-end and back-end storage network; Each OOB channel in the plurality of OOB channels is communicatively connected to the third switch, and each OOB channel is communicatively connected to the server management interface or the switch management interface.
6. The computing power network system according to claim 1, wherein: The first switches are leaf switches and there are at least two of them, the second switches are spine switches and there are multiple of them, and the third switches are leaf switches and there are multiple of them.
7. The computing power network system according to claim 1, wherein: It includes an IDC network that is communicatively connected to the core switch, wherein the IDC network includes a computer room core switch, a computer room aggregation switch, a computer room access switch, and a computer room server that are communicatively connected in sequence, the computer room core switch is communicatively connected to the core switch, and the computer room server is used to store various types of files, undertake general computing tasks, and perform data encryption and decryption operations.
8. The computing power network system according to claim 1, wherein: The wide area network includes a traffic monitoring, analysis and cleaning module that is communicatively connected to the core switch; The traffic monitoring, analysis and cleaning module is used to monitor the network traffic flowing through the core switch in real time to identify whether there is abnormal traffic. If so, the network traffic is pulled to a traffic cleaning device or area to separate the abnormal traffic from the normal traffic in the network traffic, and the normal traffic is re-injected into the original path of the network traffic.
9. The computing power network system according to claim 1, wherein: The wide area network includes a firewall arranged between the core switch and the first switch; And / or, the dedicated line includes a telecom network dedicated line, a Unicom network dedicated line and / or a mobile / third-party network dedicated line.
10. The computing power network system according to claim 1, wherein: The platform area includes an SDN controller, a heterogeneous computing power debugging platform, a monitoring platform, a storage management platform, a logging platform and / or a bastion host; The SDN controller is used to configure, manage, and monitor the network through a unified platform, and is responsible for managing and scheduling network traffic. It dynamically adjusts the path and distribution of traffic according to different application requirements and network conditions, and uniformly manages network devices through the southbound interface protocol OpenFlow, and uniformly distributes various configuration scripts required for the operation of network devices to the first switch. The heterogeneous computing power debugging platform is used to monitor computing resources in real time and display monitoring results through dashboards and graphical interfaces. It also uses intelligent scheduling algorithms to dynamically allocate computing resources based on the needs and priorities of different tasks. After receiving various computing tasks, it places the tasks into a task queue and queues them according to their priorities and resource requirements. It also displays the status and progress of tasks to users. It also provides intelligent performance optimization suggestions through real-time monitoring and analysis of computing resources and task execution. It is also used to authenticate and authorize users to ensure that only legitimate users can access computing resources. The monitoring platform is used to monitor the temperature and humidity of network equipment, hardware utilization, port traffic and / or network quality in real time. Through real-time data collection and analysis, it determines whether potential faults and / or abnormalities have occurred. If so, it triggers alarm actions before the problem occurs through a pre-set early warning mechanism. In addition, through monitoring and analysis of system performance indicators, it provides administrators with system performance bottlenecks and / or performance optimization directions. The storage management platform is a software system or tool used to centrally manage and control storage resources. By building an intelligent operation and maintenance management platform, it can realize intelligent operation of the storage platform in terms of automated deployment, status monitoring, capacity prediction, performance optimization, remote inspection, fault diagnosis and / or hard disk failure prediction, as well as real-time monitoring of storage device performance indicators to promptly identify performance bottlenecks and problems, and improve storage system performance by adjusting storage configuration and optimizing storage layout; The log platform is used to monitor log data in real time and promptly trigger alarm notifications to administrators when specific events or abnormal situations are discovered. In addition, when a system failure occurs, it responds to administrator inquiries about relevant log records to quickly locate the root cause of the problem. It also records various system operations and access logs to provide a basis for security audits, and through log analysis, it discovers potential security vulnerabilities and violations to enhance system security. The bastion host is used to protect the network and data from intrusion and damage from internal and external users.
Citation Information
Patent Citations
Industrial internet network architecture based on intelligent fusion of general inductance calculation and virtual controller
CN116962409A
Computing power routing method for white-box switch
CN118449897A