Intelligent computing server for GPU communication acceleration and I / O storage acceleration
Patent Information
- Application Number
- CN202510892309.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
Smart Images

Figure CN120803989A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of high-performance computing and artificial intelligence computing devices, and relates to a smart computing server. BACKGROUND
[0002] At present, the performance of domestic smart computing servers still has a significant gap compared with the international leading level, especially in large-scale data processing and high-concurrency application scenarios. The main problems are low GPU computing power efficiency and CPU resource bottlenecks.
[0003] Regarding the problem of low GPU computing power efficiency: in traditional GPU servers, graphics cards are interconnected with the main CPU through PCIE channels. This communication mode needs to pass through the CPU and even cross the CPU NUMA node, thereby causing communication delay and increasing the load of the main CPU. When domestic CPUs are used to build GPU servers, if the traditional server architecture is followed, it will be limited by the number and generation of PCIE channels, resulting in limited interconnection bandwidth between GPU cards and limited number of GPU cards that a single system can support. Moreover, some domestic CPUs only support PCIE 3.0 channels, and some support PCIE 4.0 channels but with limited number, which cannot meet the demand of high-speed interconnection between GPU cards. This problem not only slows down the execution speed of computing tasks, but also increases energy consumption and cost, thereby reducing the performance-price ratio of the overall product.
[0004] As for the CPU resource bottleneck problem: the X86 and ARM architectures commonly used in domestic smart computing servers consume too much CPU computing resources when processing hard disk data access, memory data interaction, network card packet forwarding and other peripheral operations. These peripheral operations occupy a large time slice of the CPU, which limits the CPU performance and increases the system resource occupancy. This not only weakens the business carrying capacity, but also reduces the robustness and compatibility of the system, which seriously restricts the performance of smart computing servers in complex application scenarios. SUMMARY
[0005] In view of the above performance bottleneck, the application proposes an innovative smart computing server whole machine structure design, which combines GPU communication acceleration and I / O storage acceleration technology, and can effectively solve the performance deficiency problem commonly existing in domestic smart computing servers.
[0006] The technical scheme of the application is a GPU communication acceleration and I / O storage acceleration intelligent calculation server, which comprises a calculation module, a network module and a power management module, is arranged on a server mainboard, the calculation module comprises a processor CPU and an intelligent calculation acceleration card, the CPU is used for managing and using data, the intelligent calculation acceleration card comprises a processing unit, a memory unit, a storage unit and an IO unit, the processing unit is a core control chip of the intelligent calculation acceleration card, the memory unit is a memory bank and is used as a cache, and is interconnected with the processing unit through a cache channel, the storage unit is at least one SAS / SSD hard disk and is interconnected with the processing unit through a SAS / SSD high-speed channel, and the IO unit comprises a network IO, a disk IO and a management IO, the network IO is interconnected with a network chip through a network high-speed channel, the network chip is provided with a plurality of 25GE network ports, the disk IO is interconnected with a hard disk control system through a chip hard disk channel, the hard disk control system is connected with a server data disk through a hard disk interface line, and the management IO is interconnected with a management control system through a chip management channel, and the management control system is provided with a peripheral interface; The intelligent calculation acceleration card is provided with a GPU-PCIE expansion board, at least one GPU card is connected through a PCIE channel, high-speed interconnection between the GPU cards is realized through the PCIE channel, the intelligent calculation acceleration card is interconnected with the server mainboard through a PCIE Gen3, the intelligent calculation acceleration card is combined with a plurality of GPU cards to construct a standardized independent interconnection channel.
[0007] Further, the intelligent calculation acceleration card is provided with an EXP-1 expansion board and an X8-1 adapter card, which are used for connection with various types of signal creation GPU servers.
[0008] Further, the management control system peripheral interface has a VGA, an LED, a TYPE C port and a gigabit network port.
[0009] The network module supports four 100Gbps network interfaces, integrates an RDMA unit, and has a dedicated network interface NIC card without occupying physical host CPU resources.
[0010] Further, the intelligent calculation acceleration card is provided with an NVME storage interface.
[0011] Preferably, the intelligent calculation server is provided with two intelligent calculation acceleration cards, each of which provides two 100G external connection ports, one intelligent calculation server forms a 400G cross-cabinet interconnection bandwidth, and a cross-cabinet multi-node training cluster interconnection relationship is constructed.
[0012] The server mainboard is provided with 3.5-inch or 2.5-inch mechanical hard disks, and the hard disks are interconnected through independent distributed storage interfaces to realize multipoint collaborative work.
[0013] The power management module supports a redundant power supply configuration, ensures continuous power supply in the event of a single power supply failure, and can dynamically adjust the power output according to the load condition, thereby maintaining stable operation of the server.
[0014] The beneficial effects of the present application are: 1. Improve GPU computing efficiency: by adopting advanced GPU communication acceleration technology, improve the interconnection bandwidth and communication speed between GPU cards, reduce the delay and bottleneck in parallel computing, thereby fully exerting the hardware performance of GPU and improving the overall computing efficiency.
[0015] 2. Relieve CPU resource bottleneck: by introducing I / O storage acceleration technology, optimize the process of hard disk data access, memory data interaction and network card packet forwarding, etc. peripheral operations, reduce the time consumption of CPU in these operations, release the computing resources of CPU, improve the business carrying capacity and robustness of the system. BRIEF DESCRIPTION OF DRAWINGS
[0016] Figure 1 is a block diagram of the overall structure of the intelligent computing server; Figure 2 is the rack structure of the intelligent computing server; Figure 3 is a screenshot of the compatibility test scheduling of the intelligent computing server; Figure 4 is a screenshot of the inference test after the completion of the compatibility test scheduling of the intelligent computing server; Figure 5 is a screenshot of the inference process monitoring of the compatibility test of the intelligent computing server; Figure 6 is a screenshot of the compatibility test results of the intelligent computing server; Figure 7 is a screenshot of the ChatGLM3-6B model inference performance test of the intelligent computing server; Figure 8 is a screenshot of the LLaMA2-7B model inference performance test of the intelligent computing server; Figure 9 is a screenshot of the Baichuan2-7B model inference performance test of the intelligent computing server; Figure 10 is a screenshot of the Stable Diffusion2.1 model inference performance test of the intelligent computing server; Figure 11 is a high-speed interconnection channel architecture diagram of the present application; Figure 12 is an internal system diagram of the intelligent computing acceleration card; Figure 13 is a simplified architecture diagram of the server with an acceleration card; Figure 14 is a processing flow diagram of the intelligent computing server; Figure 15 is an appearance diagram of an intelligent computing server. DETAILED DESCRIPTION
[0017] Noun explanation: GPU acceleration card (intelligent acceleration card): a software-defined, hardware-accelerated, self-controllable card designed to improve the efficiency of domestic GPU computing power. Through the EXP-1 expansion board and the X8-1 adapter card, it is conveniently deployed in various trust computing GPU servers to improve the overall computing power efficiency of trust computing GPU servers. Unlike industry DPU, RAID, HBA, and IPU cards, it not only integrates their strengths but also has the characteristic function of IO acceleration and meets the requirements of self-control. It innovatively solves the problem of changing from the Von Neumann architecture to the numerical control separation Harvard architecture of domestic servers, solves the fundamental problem of the performance and function of the whole machine being affected by PCIE, data channel link, RAID card, etc., and comprehensively improves the CPU, memory business resource utilization and performance of trust computing servers.
[0018] RDMA (Remote Direct Memory Access): a remote direct memory access technology designed to reduce processing delay and resource consumption during data transmission, providing high-bandwidth and low-latency network communication.
[0019] The overall structure design of the intelligent computing server is based on the core modules: the computing module, the network module, and the power management module, which integrate GPU communication acceleration and I / O storage acceleration technology to maximize overall performance, thereby providing high-performance, low-latency, and highly reliable computing services for various application scenarios and meeting the needs of complex computing tasks.
[0020] The overall design builds an independent distributed data acceleration platform through the acceleration unit + NVME, which does not occupy CPU resources, meets the needs of fast loading / interaction of data / models, optimizes the GPU server architecture, and weakens the performance requirements of domestic CPUs. The inter-card rate reaches 500 GB / s, the inter-node rate reaches 400 Gbps, and it has better performance advantages. Through the intelligent acceleration card GPU expansion board, it realizes the mixed insertion of different brands and models of domestic GPU cards from Cambrian, Tianzhi Chip, Denglin, and Xim brands. The built-in driver realizes the bottom adaptation and unified resource scheduling of domestic GPU cards, realizes the connection of the domestic GPU ecosystem, and realizes the full open ecosystem architecture.
[0021] The computing module comprises a processor CPU and an intelligent calculation acceleration card, the CPU is used for managing and using data, and the intelligent calculation acceleration card comprises a processing unit, a memory unit, a storage unit and an IO unit, the processing unit is a core control chip of the intelligent calculation acceleration card, the memory unit is a memory bank and is used as a cache, and the memory unit is interconnected with the processing unit through a cache channel, the storage unit is at least one SAS / SSD hard disk, and the storage unit is interconnected with the processing unit through a SAS / SSD high-speed channel, the IO unit comprises a network IO, a disk IO and a management IO, the network IO is interconnected with a network chip through a network high-speed channel, the network chip is provided with a plurality of 25GE network ports, the disk IO is interconnected with a hard disk control system through a chip hard disk channel, the hard disk control system is connected with a server data disk through a hard disk interface line, and the management IO is interconnected with a management control system through a chip management channel, and the management control system is provided with a peripheral interface; The intelligent calculation acceleration card is provided with a GPU-PCIE expansion board, at least one GPU card is connected through a PCIE channel, the GPU cards are interconnected through a PCIE channel, the intelligent calculation acceleration card is interconnected with a server mainboard through a PCIE Gen3, one intelligent calculation module is composed of one GPU acceleration card and four GPU cards, and a standardized independent interconnection channel is constructed. Higher data transmission bandwidth and lower communication delay are provided, so that the server is more suitable for AI training and inference.
[0022] The GPU card is flexibly configured with 4-8 cards, and supports domestic GPU cards such as Cambrian, Tianchi, Denglin and Xim.
[0023] The peripheral interface of the management control system comprises a VGA, an LED, a TYPE C port and a gigabit network port.
[0024] The network module supports four 100Gbps network interfaces without occupying physical host CPU resources, and ensures high-speed data transmission capability. The network module integrates an RDMA unit, reduces data transmission delay, and further optimizes network performance. In addition, a dedicated network interface NIC card is provided without occupying physical host CPU resources, and the network throughput is optimized.
[0025] The intelligent calculation acceleration card is provided with an NVME storage interface.
[0026] The server mainboard is provided with 3.5" or 2.5" mechanical hard disks, and the hard disks are interconnected through independent distributed storage interfaces. Through the cooperation of multiple points, the reliability and availability of data are enhanced, and the safety of data is ensured.
[0027] To ensure the stable operation of the server, the power management module uses a high-performance power supply that supports redundant power supply configuration, ensuring continuous power supply in the event of a single power failure and dynamically adjusting power output based on load conditions to maintain stable server operation.
[0028] The intelligent computing server uses a standard domestic server motherboard, which can use server motherboards with different CPUs such as Loongson, Feiteng, and Kunpeng. The intelligent computing acceleration card realizes high-speed interconnection between the intelligent computing acceleration card and the domestic GPU card through the GPU-PCIE expansion board. Two intelligent computing acceleration cards can be configured, and each intelligent computing acceleration card can provide two 100G external ports for building an independent external computing power network. Each 100G external port connected to a 100G computing power network switch can realize cross-node cluster data interaction. A single intelligent computing server can form a 400G cross-cabinet interconnection bandwidth and build a cross-cabinet multi-node training cluster interconnection relationship. The business network uses a gigabit interface for interconnection to build a business cluster interconnection relationship. The high-speed cache system provided on the intelligent computing acceleration card provides a high-speed data processing mechanism for data interaction during inference and training of the intelligent computing server, meeting the high-speed data processing needs of intelligent computing.
[0029] Adaptation test of intelligent computing server: (1) Test items: functional test of heterogeneous GPU card compatibility, and inference performance test of ChatGLM3-6B, Llama-7B, Baichuan2-7B, and Stable Diffusion2.1. Tools: MaaS platform, inference performance dashboard.
[0030] (2) Test cases and test records ① Heterogeneous GPU card compatibility test Test number: 1 Test purpose: To verify that in the environment with GPU acceleration cards, domestic CPUs and GPUs from multiple different manufacturers can be compatible with each other.
[0031] Evaluation criteria: MaaS multi-card scheduling is available, GPU cards need to come from different manufacturers, and inference is achieved.
[0032] Test machine: 1 intelligent computing server, 2 GPU acceleration cards, 4 Cambrian GPU cards, and 4 Tianshizhi GPU cards.
[0033] Environment setup: operating system openEuler 22.04, intelligent computing MaaS platform.
[0034] Test steps: Build Cambrian intelligent computing cluster and Tianshizhi intelligent computing cluster; Deploy the MaaS platform; Create an inference task; Inference test; Expected result: The inference process schedules the intelligent computing resources of two clusters and provides a unified inference interface.
[0035] Test record: Create an inference task and select a scheduling strategy, such as Figure 3 ; The scheduling is in progress, such as Figure 4 ; The scheduling is complete, and the inference test is performed, such as Figure 5 ; Inference process monitoring, such as Figure 6 ; Test result: The inference task creation according to the scheduling strategy is completed, the inference test is successful, and the monitoring data is normal.
[0036] ② ChatGLM3-6B model inference performance test Test purpose: Test the inference performance of ChatGLM3-6B model in two environments; Test conditions: Intelligent computing server, 4 XPU cards Control group: Intel CPU, 4 XPU cards.
[0037] Due to the insufficient GPU card resources of Tianzhi Chips, the control experiment cannot be formed, so this test only completes the inference test of XPU.
[0038] The manufacturer provides code and scripts for inference, including: input and output token length combinations are divided into (256, 64) / (256, 256) / (512, 512) / (1024, 1024) / (4096, 4096) 5 different combinations, all combinations BatchSize is not limited; the inference result encounters an end symbol and cannot be directly ended, and needs to be inferred to the specified length of token quantity, if the input / output length exceeds the specified length, it will be truncated; the number of performance test data is not limited, and each manufacturer can process according to the optimal data number, and the performance test result is calculated according to the specified method.
[0039] Test process: 1. Input test file, set batchsize, calculate: 1) End-to-end token generation per second (unit: tokens / second); 2) Record the average return time of the first token (unit: milliseconds); 2. Record test results.
[0040] Test record: 1. Create ChatGLM3-6B inference task through intelligent computing MaaS platform.
[0041] 2. Record the inference time and token number by simulating user direct inference test on the platform.
[0042] Note: This way of testing data is lower than directly testing data on the intelligent computing server, but closer to user experience.
[0043] Test output as Table 1 Table 1 ChatGLM3-6B model inference test data Test results: The inference speed of the two environments is shown in Table 2. After using the acceleration card, the average return time of the first token is reduced by 7.3 ms (37.5%), and the number of tokens generated per second is increased by 8.9 (36.7%).
[0044] Table 2 ChatGLM3-6B model inference test results ③ LLaMA2-7B model inference performance test Test purpose: Test the inference performance of LLaMA2-7B model in two environments Test conditions: Intelligent computing server, 4 XPU cards Control group: Intel CPU, 4 XPU cards.
[0045] Due to insufficient GPU card resources of Tianzhi chips, the control experiment cannot be formed, so this test only completes the inference test of XPU cards.
[0046] The manufacturer provides code and scripts for inference, including the following functions: input and output token length combinations are divided into (256, 64) / (256, 256) / (512, 512) / (1024, 1024) / (4096, 4096) 5 different combinations, all combinations BatchSize is not limited; the inference result encounters the end symbol, which cannot be directly ended, and needs to be inferred to the specified length of token number, if the input / output length exceeds the specified length, it is truncated; the number of performance test data is not limited, each manufacturer can process according to the optimal data number, and the performance test results are counted according to the specified method.
[0047] Test process: 1. Input test file, set batchsize, calculate: 1) End-to-end tokens generated per second (unit: tokens / second); 2) Record the average return time of the first token (unit: ms); 2. Record test results.
[0048] Test Record: 1. Create LLaMA-7B inference tasks through the intelligent MaaS platform.
[0049] 2. Test and record inference time and token count by simulating user direct inference on the platform.
[0050] Note: This method tests data lower than directly testing data on the intelligent server, but closer to user experience.
[0051] Test output as Table 3: Table 3 LLaMA2-7B Model Inference Performance Test Data Test Results: Inference speed in two environments as shown in Table 4, after using the acceleration card, the average return time of the first token decreased by 8.66ms (35.5%), and the number of tokens generated per second increased by 8.4 (50.4%).
[0052] Table 4 LLaMA2-7B Model Inference Performance Test Results IV. Baichuan2-7B Model Inference Performance Test Test Purpose: Test the inference performance of Baichuan2-7B model in two environments Test Conditions: Intelligent Server, 4 XPU cards Control Group: Intel CPU, 4 XPU cards
[0053] Due to insufficient GPU card resources of Tianzhi Chips, the control experiment cannot be formed, so this test only completes the inference test of XPU cards.
[0054] The manufacturer provides code and scripts for inference, including the following functions: input and output token length combinations are divided into (256, 64) / (256, 256) / (512, 512) / (1024, 1024) / (4096, 4096) five different combinations, all combinations BatchSize is not limited; Inference results encounter end symbols and cannot be directly ended, and need to be inferred to the specified length of token number, if the input / output length exceeds the specified length, it will be truncated; The number of performance test data is not limited, and each manufacturer can process the optimal number of data, and the performance test results are counted according to the specified method.
[0055] Test Process: 1. Input test file, set batchsize, calculate: 1) The number of tokens generated per second end-to-end (unit: tokens / second); 2) Record the average return time of the first token (unit: milliseconds); 2. Record the test results.
[0056] Testing Log: 1. Create the Baichuan2-7B inference task through the Zhisuan MaaS platform.
[0057] 2. Simulate user direct reasoning tests on the platform and record the reasoning time and token count.
[0058] Note: The test data in this way is lower than the test data directly on the intelligent computing server, but it is closer to the user experience.
[0059] The test output is shown in Table 5 Table 5 Baichuan2-7B model inference performance test data Test results: The inference speeds of the two environments are shown in Table 6. After using the accelerator card, the average return time of the first token was reduced by 7.9ms (30.1%), and the number of tokens generated per second increased by 5.55 (44.5%).
[0060] Table 6 Baichuan2-7B model inference performance test results ⑤ Stable Diffusion2.1 model inference performance test Test purpose: Test the inference performance of the Stable Diffusion 2.1 model in two environments Test conditions: Intelligent computing server, 4 Cambrian GPU cards Control group: Intel CPU, 4 Cambrian GPU cards.
[0061] Due to insufficient GPU card resources of Tianshu Zhixin, a control experiment could not be formed, so this test only completed the Cambrian inference test.
[0062] The manufacturer provides code and scripts for inference. Functions include: generating one, two, three, four, and five 512*512 images according to user needs, and recording the single generation time; there is no limit on the number of performance test data items, and each manufacturer can process the data according to the optimal number of data items and calculate the performance test results according to the specified method.
[0063] Testing process: 1. Input the test file, set the batch size, and calculate the total generation time; 2. Record test results.
[0064] Test records: 1. Create a Stable Diffusion2.1 inference task through the intelligent computing MaaS platform.
[0065] 2. Test and record the inference time and token number by simulating user direct inference on the platform.
[0066] Note: This method tests data that is lower than direct testing on the intelligent computing server, but closer to user experience.
[0067] The test output is as follows: Table 7 Stable Diffusion2.1 model inference performance test data Test results: The inference speed of the two environments is shown in Table 8. After using the acceleration card, the generation time of 1-5 512*512 images is reduced by 3.45 seconds (41.2%).
[0068] Table 8 Stable Diffusion2.1 model inference performance test results (3) Test conclusion According to the above test results, the intelligent computing server can support heterogeneous GPU inference tasks in terms of function. Meanwhile, in terms of performance, using the same Cambrian GPU, the configuration of the intelligent computing server is superior to the Intel configuration without an acceleration card in terms of latency, bandwidth, and underlying hardware performance. The actual inference performance is greatly improved. Testing ChatGLM3-6B, LLaMA2-7B, and Baichuan2-7B three text generation large models, the average number of generated tokens per second increases by 7.62, with an average improvement rate of 44.87%; testing Stable Diffusion2.1, after using the intelligent computing server, the generation time of 1-5 512*512 images is reduced by 3.45 seconds (41.2%). In addition, the inference deployment efficiency is improved from 2 minutes to about 30 seconds.
[0069] Therefore, the intelligent computing server has the following advantages: (1) Heterogeneous computing power scenarios. Through testing, it is found that the intelligent computing server supports Loongson and other domestic CPUs, is compatible with different manufacturers' domestic GPUs, and can perform unified scheduling between heterogeneous cards through the MaaS platform.
[0070] (2) High-speed reasoning demand scenarios. The intelligent computing server can accelerate the interconnection between domestic GPU cards, thereby accelerating the speed of multi-card GPU reasoning and improving by 44.87%; image generation is improved by 41.2%.
[0071] (3) One-stop service scenario. The intelligent computing server uses a large cache technology to improve deployment efficiency to about 30 seconds, greatly improving the deployment efficiency of large models.
[0072] Intelligent computing acceleration card: Based on the parallel numerical control separation framework, a special acceleration card for the national security server is designed and developed for disk I / O data acceleration and X86 compatible space, including data input, data processing, data output, X86 compatible space, and high-speed cache. The acceleration card uses a FPGA special chip, internally connected through a SAS interface with the server hard disk backplane to take over the server hard disk; through PCIE3.0 X8 to realize interconnection with the server mainboard and realize interconnection with the CPU; through IO exchange disk addressing and IO parallel bus technology, forming a parallel processing effect, improving data transmission speed and efficiency; through the data control engine to take charge of the control task of data transmission, including data reading, writing, cache management, etc. And integrate advanced technologies such as cache optimization, virtualization, I / O acceleration, multi-core fusion, etc. into the national X86 compatible space.
[0073] By connecting to the PCIE interface of the low-end national security server, the internal structure of the server is changed to a "numerical control separation architecture", and the out-of-band center point of data processing and exchange is constructed through the acceleration card, which is responsible for data interaction and data management between storage media. CPU only participates in management and use of data, while having IO performance acceleration and RAID card functions, offloading the occupation of computing resources by external devices such as hard disk data IO and data network card of the national security server, reducing the load of the original CPU and bus, and reducing the pressure on the original CPU and bus., so that it can focus on computing tasks, all peripherals and application services directly access data through the corresponding interface and driver, while the special chip internally constructs multiple IO channels, forming a parallel processing effect, improving data transmission speed and efficiency, and effectively improving data transmission efficiency and processing speed., the CPU resource utilization is about 2 times growth, fundamentally guarantees that the performance of the domestic CPU server can carry more business system load, reduces the delay, improves the stability and efficiency in high-concurrency business scenarios, can support IO-intensive access, CPU-intensive, and other high-concurrency production business scenarios Efficient operation, effectively support various key applications in the national security environment.
[0074] At the same time, the acceleration card adopts heterogeneous compatible technology, improves the standardization of hardware interface and the compatibility of software driver, shields the differences between different hardware platforms by developing a unified abstraction layer, allows different hardware platforms to work together in the data center, realizes the wide compatibility of multi-brand domestic CPU, and promotes the high-performance construction and integration of domestic data center ecological industry.
[0075] To realize the above scene application, the acceleration card in the application adopts the following technical scheme: Parallel numerical control separation framework: the acceleration card is responsible for processing data interaction and data management between storage media, and internally implements a parallel numerical control separation framework, which divides the data processing flow into a control path and a data path. The control path is responsible for processing instructions and data transmission requests from the CPU, while the data path focuses on efficiently executing I / O operations such as disk reading and writing, memory access, and network data transmission, achieving fast data throughput.
[0076] Hardware design: this design is a special acceleration card for signal creation servers, including data input, data processing, data output, X86 compatible space, and cache. The acceleration card uses FPGA special chips, built-in domestic X86 compatible space, integrates parallel processing mechanism, cache optimization, virtualization, I / O acceleration, multi-core fusion, and other advanced technologies of modern computer systems, and is connected to the signal creation server motherboard through a standard PCIE interface, ensuring wide compatibility. The data control engine is responsible for data transmission control tasks, including data reading, writing, cache management, etc.
[0077] Interface interconnection: the internal interconnection of the acceleration card is connected to the server hard disk backplane through the SAS interface, realizes the takeover of the server hard disk, and is connected to the server motherboard through PCIE3.0 X8, and is connected to the CPU. The acceleration card is externally connected to the server nodes through a 25GE port to realize high-speed interconnection between server nodes. The high-load processing of data interaction and data balancing between server nodes does not pass through the main CPU, which can greatly reduce the load of the server main CPU. The acceleration card also provides independent backup ports and independent backup channels, which do not pass through the CPU for processing, and can ensure that the server data backup does not consume main CPU resources. It can fundamentally solve the problem of idle-time backup due to the consumption of a large amount of CPU resources in traditional backup, and realize full-time data backup through the acceleration card without affecting the operation of server business.
[0078] Resource offloading and acceleration: Utilize the high-performance computing capabilities of the acceleration card to offload I / O operations that would otherwise be handled by server hard drives, memory, network cards, and other peripherals. Integrate a dedicated IO processing engine and optimized data path to reduce CPU resource usage and accelerate these operations through hardware-level optimization, improving data transfer efficiency and processing speed, resulting in a 3-10 times improvement in IO performance and a 2 times increase in CPU resource usage for the ChinaSoft server.
[0079] Software adaptation and optimization: Develop supporting software drivers to ensure seamless integration of the acceleration card with the ChinaSoft server operating system, supporting high-concurrency production business scenarios such as IO-intensive access (e.g., database operations) and CPU-intensive tasks (e.g., large-scale computing tasks), significantly improving overall system performance.
[0080] Heterogeneous compatibility: This technology achieves widespread compatibility with multiple brands of domestic CPU chips by developing a unified abstraction layer to mask differences between different hardware platforms. This implementation approach standardizes hardware interfaces and improves software driver compatibility, allowing different hardware platforms to work together in data centers, promoting high-performance construction of domestic data centers and integration of the industrial ecosystem.
[0081] The technical points of the structure design of the intelligent algorithm server include the following two aspects: 1. GPU communication acceleration (1) GPU card-to-GPU card high-speed interconnection Provide GPU-to-GPU card high-speed interconnection channels to achieve standard interconnection between 8 GPU cards, ensuring high bandwidth and low latency for data transmission to meet the high-speed data exchange requirements within the GPU server, ensuring optimal performance in training and inference tasks.
[0082] (2) Inter-node high-speed interconnection The intelligent algorithm acceleration card can provide powerful data exchange capabilities for GPU high-performance computing environments.
[0083] Cross-server node interconnection: Build a 400G high-speed network with the intelligent algorithm acceleration card to support efficient data exchange between cross-server nodes. The integrated RDMA technology allows network adapters to directly read and write GPU memory, reducing memory consumption and further optimizing system performance.
[0084] Parallel processing capability: In a multi-GPU server environment, the cross-server node interconnection implemented by the intelligent algorithm acceleration card ensures efficient data exchange between different GPU servers GPU Host-to-GPU Host. Support for distributing tasks to multiple GPU server nodes for parallel processing improves the overall system's processing power efficiency.
[0085] High-speed data cache capability: the intelligent calculation acceleration card supports TB-level high-speed NVME cache, which is used to accelerate model acceleration tasks in training and inference environments, and supports multi-card sharing of cache data in single training and inference tasks, and cache data exists between multiple cards in a strongly consistent manner.
[0086] 2. I / O storage acceleration In a GPU server, GPUs are connected to CPUs through PCIE channels, and data exchange between GPUs traditionally relies on CPU mediation, which is particularly highlighted as a bottleneck problem when domestic CPU performance and PCIE channel resources are limited.
[0087] The intelligent calculation server of the application realizes a breakthrough in structural design, realizes the separation of CPU and GPU architecture and the independence of the computing power network and the business network. The intelligent calculation module and the acceleration card jointly construct the core platform for data interaction between GPU cards and between server nodes in the server, bypassing the main CPU, thereby avoiding the increase of CPU load caused by a large amount of data interaction, and reducing the impact on domestic CPU. At the same time, by combining the acceleration card with the NVME technology, an independent distributed data acceleration platform is constructed, which can realize the rapid loading and interaction of data / models without occupying CPU resources. In addition, the 25G data network and the 4*100G computing power network equipped with the acceleration card jointly create an independent distributed data network and node computing power interaction network, realizing the hierarchical and network interaction of data, further reducing the load of the main CPU, and effectively alleviating the data processing bottleneck caused by CPU performance limitation.
[0088] The intelligent calculation server designed by the application can be developed based on the Loongson 3C5000 chip, and the core components are made of domestic hardware. It is an artificial intelligence AI server suitable for model training, inference, HPC, rendering and other applications, which can meet the needs of computing power centers, industry markets and Internet and other types of large-scale AI business.
Claims
1. The intelligent computing server with GPU communication acceleration and I / O storage acceleration includes a computing module, a network module, and a power management module, and is located on the server motherboard. Its features include: The computing module contains a processor CPU and an intelligent computing acceleration card. The CPU is used to manage and use data. The intelligent computing acceleration card contains a processing unit, a memory unit, a storage unit, and an IO unit. The processing unit is the core control chip of the intelligent computing acceleration card. The memory unit is a memory bar, which is used as a high-speed cache and is interconnected with the processing unit through a high-speed cache channel. The storage unit is at least one SAS / SSD hard disk, which is interconnected with the processing unit through a SAS / SSD high-speed channel. The IO unit includes network IO, disk IO, and management IO. The network IO is interconnected with the network chip through a network high-speed channel. The network chip is provided with multiple 25GE network ports. The disk IO is interconnected with the hard disk control system through the chip hard disk channel. The hard disk control system is connected to the server data disk through the hard disk interface line. The management IO is interconnected with the management control system through the chip management channel. The management control system is provided with a peripheral interface; The intelligent computing accelerator card is equipped with a GPU-PCIE expansion board, which is connected to at least one GPU card through the PCIE channel. The GPU cards are interconnected at high speed through the PCIE channel. The intelligent computing accelerator card and the server motherboard are interconnected through PCIE Gen3. The intelligent computing accelerator card is combined with multiple GPU cards to build a standardized and independent interconnection channel.
2. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The intelligent computing acceleration card is equipped with an EXP-1 expansion board and an X8-1 adapter card for connecting to various types of trusted computing GPU servers.
3. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The management and control system peripheral interfaces include VGA, LED, TYPE C port, and Gigabit network port.
4. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The network module supports four 100Gbps network interfaces, integrates an RDMA unit, and has a dedicated network interface NIC card without occupying the CPU resources of the physical host.
5. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The intelligent computing accelerator card is equipped with an NVME storage interface.
6. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The intelligent computing server is equipped with two intelligent computing acceleration cards. Each intelligent computing acceleration card provides two 100G external connection ports. One intelligent computing server constitutes 400G cross-cabinet interconnection bandwidth, building a training cluster interconnection relationship between multiple nodes across cabinets.
7. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The server motherboard is equipped with a 3.5” or 2.5” mechanical hard drive, which is interconnected through independent distributed storage interfaces to achieve multi-point collaborative work.
8. The intelligent computing server with GPU communication acceleration and I / O storage acceleration according to claim 1, characterized in that: The power management module supports redundant power supply configuration to ensure continuous power supply when a single power supply fails, and can dynamically adjust power output according to load conditions to keep the server running stably.