Modularized reasoning and training integrated server
The modularly designed inference and training integrated server solves the problem of mismatch between existing server hardware and requirements by flexibly switching computing modules and dynamically allocating resources, thus achieving efficient and flexible AI task processing and device stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2026-04-03
AI Technical Summary
The existing server hardware is mismatched with the requirements, resulting in serious resource redundancy and an inability to expand as needed. Furthermore, when training tasks change, the entire hardware needs to be replaced, which cannot meet the requirements of high computing power and dynamically switching AI tasks.
The modularly designed inference and training integrated server flexibly switches computing modules through standardized interfaces. Combined with storage module groups and intelligent control modules, it achieves dynamic resource allocation and status monitoring, has active heat dissipation and protection functions, and supports closed-loop adaptation of training and inference tasks.
It improves the efficiency of AI task processing, reduces resource waste, lowers hardware upgrade costs, ensures stable equipment operation, and adapts to the needs of multiple application scenarios.
Smart Images

Figure CN121785436A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server technology, specifically to a modular inference and training integrated server. Background Technology
[0002] Servers, through centralized resource management, efficient request response, and guaranteed service stability, have become the "core infrastructure" of internet operation and enterprise IT architecture. Our daily online activities and enterprise digital operations are inseparable from their support. Servers are divided into two key layers: hardware and software. Only through the collaboration of the two can complete service functions be realized. On the hardware layer, dedicated computer equipment optimized for "long-term high-load operation" uses enterprise-grade CPUs such as Intel Xeon and AMD EPYC, paired with highly stable memory and hard drives, supporting 24-hour uninterrupted operation. At the same time, it has strong scalability and can flexibly add hard drives, memory, and network cards to cope with data growth. With the deep penetration of artificial intelligence technology in fields such as computer vision, natural language processing, and autonomous driving, model training and inference tasks exhibit the core characteristics of "high computing power requirements, frequent dynamic switching, and complex resource adaptation".
[0003] Although inference has low computational power per request, it needs to support thousands to tens of thousands of concurrent requests per second, with a response latency of less than 100ms. Edge scenarios are also sensitive to power consumption. The current industry adopts a "separate" hardware architecture, which leads to a mismatch between hardware and requirements. Training servers are often equipped with high-performance components, resulting in insufficient GPU utilization and serious resource redundancy during inference. Inference servers use low-power GPUs or ASIC chips, which cannot support training of models with ultra-large parameters. At the same time, the architecture has poor scalability. Conventional servers are integrated designs, and when training tasks change, the entire hardware needs to be replaced, and it cannot be expanded as needed. To address this, the invention provides a modular integrated inference and training server. Summary of the Invention
[0004] The purpose of this invention is to provide a modular inference and training integrated server to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a modular inference and training integrated server, comprising a protective cover, an inner wall of which is connected to a mounting base, an inner wall of which is connected to a server body, a storage module group connected to the upper surface of the server body, the storage module group being electrically connected to a cache module via wires, the storage module group being electrically connected to a large-capacity storage module via wires, an inference computing module connected to the front of the server body, and a training computing module connected to the front of the server body, both the inference computing module and the training computing module being connected to the server body via wires. The server main body is electrically connected. A server expansion module is connected to the right side of the protective cover. A set of expansion interfaces is connected to the left side of the server expansion module. A floating quick connector is connected to the right side of the server expansion module. The floating quick connector includes two limiting slides connected to the server expansion module. A mounting bracket is slidably connected to the inner wall of each limiting slide. A cable management tube is embedded in the inner wall of the mounting bracket. Two floating components are embedded in the inner wall of the mounting bracket. A connecting plate is connected to the right side of each floating component. Two electric push rods are connected to the side of the server expansion module that is close to the mounting bracket.
[0006] Preferably, the bottom surface of the protective cover is connected to a mounting base, the bottom surface of the mounting base is provided with a groove, and the inner wall of the groove is connected to a rubber pad.
[0007] Preferably, the protective cover has a protective door hinged to its front, a potential plate is connected to the front of the protective door, and two sets of status indicator lights are connected to the front of the potential plate.
[0008] Preferably, the protective door has a lock attached to its front side, and an observation window is provided on the front side of the protective door.
[0009] Preferably, a heat-conducting plate is connected to the back of the protective cover, and a set of heat dissipation fins are connected to the back of the heat-conducting plate. Heat dissipation grooves are provided on both the left and right sides of the protective cover, and protective filter plates are bolted to the sides of the two heat dissipation grooves that are far apart from each other.
[0010] Preferably, the inner walls of both heat dissipation slots are connected to a set of mounting plates, and the sides of the two sets of mounting plates that are close to each other are connected to heat dissipation drivers. The output end of each heat dissipation driver is connected to a set of heat dissipation blades.
[0011] Preferably, a network adapter integrated panel is fixedly connected to the left side of the protective cover.
[0012] Preferably, the server body is electrically connected to an intelligent control module via wires, and the intelligent control module is electrically connected to a cache module, a large-capacity storage module, an inference computing module, and a training computing module via wires respectively.
[0013] Preferably, a server expansion docking module is connected to the right side of the connecting plate, and the server expansion docking module is electrically connected to an expansion outdoor unit via wires.
[0014] Preferably, a power interface is embedded on the back of the protective cover.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0016] The server's main body connects to an inference computing module and a training computing module, which can be flexibly switched via a standardized interface. This allows for handling the high computational demands of large-scale model training using the training computing module, while simultaneously enabling the inference computing module to handle high-concurrency real-time inference tasks. Combined with a storage module group that integrates high-speed cache modules and large-capacity storage modules for differentiated storage support, it achieves closed-loop adaptation for training, inference, and storage tasks. This avoids the cumbersome multi-device collaboration required by traditional single-function servers, significantly improving AI task processing efficiency. The intelligent control module, acting as a central regulator, can identify task types in real time and dynamically allocate resources, preventing resource waste or insufficiency. The status is also displayed via a power supply panel. The indicator lights provide intuitive feedback on the operating status, allowing operators to understand the equipment's condition without complex testing. The protective cover and mounting base create a closed protective space, and the rubber pads on the mounting base effectively reduce vibration and minimize external environmental interference to the core module. The active cooling driver and heat dissipation fins, combined with the passive cooling heat conduction plate and heat dissipation fins, quickly dissipate heat. The protective filter plate prevents dust from entering, extending the module's lifespan. The server expansion module's expansion interface allows for flexible module addition as the task scales up without replacing the entire machine, effectively controlling hardware upgrade costs. The network adapter's integrated panel and expansion interface feature a dual-access design, adapting to different data and request sources and further expanding application scenarios. Attached Figure Description
[0017] Figure 1 The schematic diagram shows a three-dimensional structure of a modular inference and training integrated server according to an embodiment of the present invention.
[0018] Figure 2 The diagram illustrates a rear-view stereoscopic structure of a modular inference and training integrated server according to an embodiment of the present invention.
[0019] Figure 3 The schematic diagram shows a three-dimensional structure of a heat dissipation driver in a modular inference and training integrated server according to an embodiment of the present invention.
[0020] Figure 4 The diagram schematically shows a front cross-section of a protective shield in a modular inference and training integrated server according to an embodiment of the present invention.
[0021] Figure 5 The illustration shows a three-dimensional structural diagram of the main body of a modular inference and training integrated server according to an embodiment of the present invention.
[0022] Figure 6 The diagram illustrates the system structure of the server body in a modular inference and training integrated server according to an embodiment of the present invention.
[0023] Figure 7 The illustration shows a three-dimensional structural diagram of a floating quick-connect in a modular inference and training integrated server according to an embodiment of the present invention.
[0024] In the diagram: 1. Protective cover; 2. Mounting base; 3. Groove; 4. Rubber pad; 5. Protective door; 6. Potential plate; 7. Status indicator light; 8. Observation window; 9. Lock; 10. Network adapter integrated panel; 11. Heat sink; 12. Protective filter plate; 13. Server expansion module; 14. Expansion interface; 15. Heat conduction plate; 16. Heat sink fins; 17. Power interface; 18. Mounting plate; 19. Heat dissipation driver; 20. Heat dissipation blades; 21. Server body; 22. Inference computing module; 23. Training computing module; 24. Mounting base; 25. Storage module group; 26. Cache module; 27. Large capacity storage module; 28. Intelligent control module; 29. Floating quick connector; 30. Expansion outdoor unit; 31. Limiting slide; 32. Floating component; 33. Connecting plate; 34. Mounting bracket; 35. Server expansion docking module; 36. Cable tray; 37. Electric push rod. Detailed Implementation
[0025] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0026] Please see Figure 1-7 As shown.
[0027] Example 1
[0028] A modular inference and training integrated server includes a protective enclosure 1. A mounting base 24 is connected to the inner wall of the protective enclosure 1. A server body 21 is connected to the inner wall of the mounting base 24. A storage module group 25 is connected to the upper surface of the server body 21. The storage module group 25 is electrically connected to a cache module 26 via wires. A large-capacity storage module 27 is also electrically connected to the storage module group 25 via wires. An inference computing module 22 and a training computing module 23 are connected to the front of the server body 21. Both the inference computing module 22 and the training computing module 23 are electrically connected to the server body 21 via wires. A server expansion module 13 is connected to the right side of the protective enclosure 1. A set of expansion interfaces 14 is connected to the left side of the server expansion module 13. A floating quick-connect component 29 is connected to the right side of the server expansion module 13. The floating quick-connect component 29 includes two limiting slides 31 connected to the server expansion module 13. Each limiting slide... The inner wall of the frame 31 is slidably connected to the mounting frame 34. The inner wall of the mounting frame 34 is inlaid with a cable tray 36. The inner wall of the mounting frame 34 is inlaid with two floating parts 32. The right side of each floating part 32 is connected to a connecting plate 33. The side of the server expansion module 13 and the mounting frame 34 that are close to each other are connected to two electric push rods 37. When the task scale expands and the expansion outdoor unit 30 needs to be connected, the intelligent control module 28 sends an "expansion docking command". The electric push rods 37 start and drive the mounting frame 34 to move along the limit slide 31 in the docking direction, which drives the server expansion docking module 35 on the connecting plate 33 to approach the server expansion module 13. During the docking process, the floating quick connector 29 automatically compensates for the installation deviation. The intelligent control module 28 detects the docking pressure in real time. After the threshold is reached, the electric push rods 37 stop running. The "expansion docking light" on the potential plate 6 lights up green, indicating that the docking is successful. The expansion outdoor unit 30 starts to work together, and the cable tray 36 straightens the wires in sync.
[0029] Example 2
[0030] The bottom of the protective cover 1 is connected to the mounting base 2. The bottom of the mounting base 2 has a groove 3, and the inner wall of the groove 3 is connected to a rubber pad 4. The front of the protective cover 1 is hinged to a protective door 5. The front of the protective door 5 is connected to a potential board 6. The front of the potential board 6 is connected to two sets of status indicator lights 7. The front of the protective door 5 is latched with a lock 9. The front of the protective door 5 has an observation window 8. The protective door 5 is made of wear-resistant steel plate, which can resist workshop dust and minor collisions. It is locked by the lock 9 to prevent non-maintenance personnel from accidentally touching the internal modules. The staff can directly view the status of the server indicator lights through the observation window 8 without opening the protective door 5, reducing the frequency of dust entering due to opening the door. The two sets of status indicator lights 7 on the potential board 6 have clear functions. One set displays the operating status of the inference computing module 22. A solid green light indicates normal operation, and a flashing yellow light indicates overload. The other set reflects the data interaction status of the storage module group 25. A solid red light indicates a storage fault. Quality inspectors can quickly identify equipment problems.
[0031] Example 3
[0032] A heat-conducting plate 15 is connected to the back of the protective cover 1, and a set of heat dissipation fins 16 are connected to the back of the heat-conducting plate 15. Heat dissipation slots 11 are provided on both the left and right sides of the protective cover 1. Protective filter plates 12 are bolted to the sides of the two heat dissipation slots 11 that are far apart from each other. A set of mounting plates 18 is connected to the inner walls of both heat dissipation slots 11, and heat dissipation drivers 19 are connected to the sides of the two mounting plates 18 that are close to each other. A set of heat dissipation blades 20 is connected to the output end of each heat dissipation driver 19. A network adapter integrated panel 10 is fixedly connected to the left side of the protective cover 1. The heat dissipation and network structure of the protective cover 1 becomes stable. The key to operation is the heat dissipation plate 15, which is made of copper-aluminum composite material with high thermal conductivity. It fits tightly against the back of the server body 21 and quickly conducts the heat generated by the training computing module 23 to the heat dissipation fins 16. The fins accelerate heat dissipation by increasing the heat dissipation area. At the same time, the heat dissipation driver 19 in the heat dissipation slots 11 on the left and right sides of the protective cover 1 drives the heat dissipation blades 20 to rotate at high speed, forming a directional airflow from left to right. This prevents the computing power of the inference computing module 22 from decreasing due to high temperature. The protective filter plate 12 on the outside of the heat dissipation slot 11 is fixed with bolts and can intercept metal debris and dust in the workshop. Only the bolts need to be removed to replace the protective filter plate 12, making maintenance convenient.
[0033] Example 4
[0034] The server main body 21 is electrically connected to the intelligent control module 28 via wires. The intelligent control module 28 is electrically connected to the cache module 26, the large-capacity storage module 27, the inference computing module 22, and the training computing module 23 via wires. The right side of the connection board 33 is connected to the server expansion docking module 35, which is electrically connected to the expansion outdoor unit 30 via wires. The back of the protective cover 1 is embedded with the power interface 17. The intelligent control module 28 connected to the server main body 21 not only realizes resource scheduling, but also automatically cuts off the power supply to the faulty unit through wires when it detects an abnormal GPU load in the inference computing module 22, and diverts the task to other modules. At the same time, it instructs the cache module 26 to back up the current model data to avoid audit interruption. When calling the training computing module 23 at night, it can also link the large-capacity storage module 27 to read data in segments to improve training efficiency.
[0035] Working principle: The server body 21 serves as the core control hub, connecting the inference computing module 22 and the training computing module 23 via wires. These two types of modules can flexibly switch via standardized interfaces, handling large-scale model training and high-concurrency real-time inference respectively. The storage module group 25, in conjunction with the high-speed cache module 26 and the large-capacity storage module 27, provides adaptive storage support for different tasks and interacts with the server body 21. The intelligent control module 28 connects to each core module via wires, realizing task identification, dynamic resource allocation, and status monitoring, and provides feedback on the operating status through the status indicator light 7 on the potentiometer 6. The protective cover 1 and the mounting base 24 construct a protective space. The cooling system operates in both active and passive modes. The cooling driver 19 drives the cooling fins 20 through the cooling slots 11 for ventilation, and heat is conducted to the cooling fins 16 via the heat conduction plate 15 for dissipation. The protective filter plate 12 prevents dust, and the rubber pads 4 on the mounting base 2 provide shock absorption and stability. During operation, first fix the protective cover 1 with the mounting base 2, connect the power interface 17 to supply power, open the protective door 5 to check the module installation, close the door and lock the lock 9, observe the status indicator light 7 to confirm that the initialization is complete, then configure according to the task type. The training task receives data through the network transfer integration panel 10, the system activates the training computing module 23 and the corresponding storage module, the inference task receives the request source through the expansion interface 14, switches to the inference computing module 22, and monitors the operation through the observation window 8 and the status indicator light 7. When the task scale expands, add modules through the expansion interface 14 of the server expansion module 13. In case of failure, open the protective door 5 and use the intelligent control module 28 to locate and hot-swap the module. Regularly remove the protective filter plate 12 to clean the dust. After the training task is completed, switch the module status through the intelligent control module 28. After all tasks are completed, turn off the power and lock the protective door 5 to achieve efficient integrated operation and adapt to multiple AI tasks.
[0036] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any indirect modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A modular inference and training integrated server, comprising a protective shield (1), characterized in that, The inner wall of the protective cover (1) is connected to a mounting base (24), and the inner wall of the mounting base (24) is connected to a server body (21). The upper surface of the server body (21) is connected to a storage module group (25), which is electrically connected to a cache module (26) via wires. The storage module group (25) is also electrically connected to a large-capacity storage module (27) via wires. The front of the server body (21) is connected to an inference computing module (22), and the front of the server body (21) is connected to a training computing module (23). Both the inference computing module (22) and the training computing module (23) are electrically connected to the server body (21) via wires. The right side of the protective cover (1) is connected to a server expansion module. The server expansion module (13) has a set of expansion interfaces (14) connected to its left side and a floating quick connector (29) connected to its right side. The floating quick connector (29) includes two limiting slides (31) connected to the server expansion module (13). The inner wall of each limiting slide (31) is slidably connected to a mounting bracket (34). The inner wall of the mounting bracket (34) is inlaid with a cable tray (36). The inner wall of the mounting bracket (34) is inlaid with two floating parts (32). The right side of each floating part (32) is connected to a connecting plate (33). The side of the server expansion module (13) and the mounting bracket (34) that are close to each other is connected to two electric push rods (37).
2. The modular inference and training integrated server according to claim 1, characterized in that: The bottom surface of the protective cover (1) is connected to the mounting base (2), and the bottom surface of the mounting base (2) is provided with a groove (3), and the inner wall of the groove (3) is connected to a rubber pad (4).
3. The modular inference and training integrated server according to claim 1, characterized in that: The protective cover (1) is hinged to a protective door (5) on the front, and a potential plate (6) is connected to the front of the protective door (5). Two sets of status indicator lights (7) are connected to the front of the potential plate (6).
4. A modular inference and training integrated server according to claim 3, characterized in that: The protective door (5) has a lock (9) attached to its front side, and an observation window (8) is provided on the front side of the protective door (5).
5. A modular inference and training integrated server according to claim 4, characterized in that: A heat-conducting plate (15) is connected to the back of the protective cover (1), and a set of heat dissipation fins (16) is connected to the back of the heat-conducting plate (15). Heat dissipation grooves (11) are provided on the left side and the right side of the protective cover (1). A protective filter plate (12) is connected to the side of the two heat dissipation grooves (11) that are far apart from each other by bolts.
6. A modular inference and training integrated server according to claim 5, characterized in that: The inner walls of the two heat sinks (11) are each connected to a set of mounting plates (18), and the two sets of mounting plates (18) are each connected to a heat sink driver (19) on the side that is close to each other. The output end of each heat sink driver (19) is connected to a set of heat sink blades (20).
7. A modular inference and training integrated server according to claim 6, characterized in that: A network adapter integrated panel (10) is fixedly connected to the left side of the protective cover (1).
8. A modular inference and training integrated server according to claim 1, characterized in that: The server body (21) is electrically connected to an intelligent control module (28) via wires. The intelligent control module (28) is electrically connected to a cache module (26), a large-capacity storage module (27), an inference computing module (22), and a training computing module (23) via wires.
9. A modular inference and training integrated server according to claim 1, characterized in that: The right side of the connecting plate (33) is connected to a server expansion docking module (35), and the server expansion docking module (35) is electrically connected to an expansion outdoor unit (30) via wires.
10. A modular inference and training integrated server according to claim 9, characterized in that: The back of the protective cover (1) is fitted with a power interface (17).