Storage device, system and method based on liquid cooling, storage medium and electronic equipment

By placing memory particles horizontally in the memory module and connecting the cold board horizontally, the problem of low running rate of the memory module is solved, achieving more efficient cooling and faster running rate, suitable for high-performance servers.

CN120045038APending Publication Date: 2025-05-27INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202412000538.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the memory module needs to be installed horizontally in a condensate tube, which increases the spacing between memory sticks, thereby reducing the operating speed of the memory module.

Method used

A liquid cooling-based storage device is designed, where the memory particles in the memory module are placed horizontally, and the cold plate and the memory module are horizontally connected, which avoids the vertical installation of the cold plate, increases the contact area between the cold plate and the memory module, saves space, and improves the heat dissipation efficiency.

Benefits of technology

Through this design, the problem of low running rate of memory modules is solved, the operation efficiency and cooling performance of memory modules are improved, and the memory needs of high-performance servers are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045038A_ABST
    Figure CN120045038A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a storage device, system and method based on liquid cooling, a storage medium and electronic equipment, and relates to the field of hardware design, and the storage device comprises a memory connector which is used for being connected with a mainboard of a server; the memory module is used for storing data, the memory module is connected with the memory connector according to preset pin configuration, the memory module comprises a plurality of memory particles, and the memory particles are horizontally placed on the memory module; the cold plate is used for absorbing heat generated by the memory module, and the cold plate is horizontally connected with the memory module. By the adoption of the scheme, the technical problem that in the related technology, the running speed of the memory module is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of hardware design, and more specifically, to a liquid-cooled storage device, system and method, storage medium, and electronic device. Background Art

[0002] The current rate requirements of servers also increase the memory requirements at the same time. In existing memories, the dual in-line memory module (DIMM) form is usually used. Dual in-line memory modules are usually placed on both sides of the memory in a server and can be arranged in parallel in a certain interface order on the left and right sides of the CPU. With the increase in power consumption, a reasonable heat dissipation design for the memory is required.

[0003] For memories with low power consumption, air cooling can meet the heat dissipation requirements. However, for servers with large memory capacity and high power consumption, liquid cooling is required to achieve the heat dissipation effect. However, in the related art, since the memory modules are placed in parallel on both sides of the CPU, the cold plate needs to be added between the memory strips, that is, the condensation pipe is installed in the horizontal space, which will increase the distance between adjacent memory strips and cause the memory strip operation rate to be low. That is, the memory module in the related art has the problem of low operation rate.

[0004] In view of the technical problem that the current memory module has a low operation rate in the related art, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present application provide a liquid-cooled storage device, system and method, storage medium, and electronic device to at least solve the technical problem that the current memory module has a low operation rate in the related art.

[0006] According to an embodiment of the present application, a liquid-cooled storage device is provided, including: a memory connector for connecting to the motherboard of the server; a memory module for storing data, wherein the memory module is connected to the memory connector according to a preset pin configuration, the memory module includes a plurality of memory particles, and the memory particles are horizontally placed on the memory module; a cold plate for absorbing the heat generated by the memory module, wherein the cold plate is horizontally connected to the memory module.

[0007] In an exemplary embodiment, memory particles are placed on both sides of the front of the memory module, and there are two data channels in one side of the memory module.

[0008] In an exemplary embodiment, the data channels of the memory module transmit data to the processor, and all the memory particles in the data channels are connected by chip select signals.

[0009] According to another embodiment of the present application, a liquid-cooled based storage system is provided, including: a memory device, which is the device of any one of the device embodiments; a motherboard for mounting different modules and implementing data transmission between different modules; a processor for controlling the memory device and reading or writing data in the memory device, wherein the processor is connected to the motherboard and is in point-contact connection with the memory device.

[0010] In an exemplary embodiment, a plurality of memory devices are placed on the motherboard, and the memory devices are evenly distributed on both sides of the processor.

[0011] In an exemplary embodiment, a first memory device and a second memory device are placed on the motherboard, wherein the height of the memory connector in the first memory device is greater than the height of the second memory device.

[0012] In an exemplary embodiment, the first memory device and the second memory device are placed on both the front and back of the motherboard.

[0013] According to yet another embodiment of the present application, a liquid-cooled based storage method is further provided, including: reading the resistance change of the liquid leakage detection line through a management controller on the server motherboard, and initializing the memory module when there is no resistance change; the management controller obtains the temperature of the memory module and adjusts the coolant flow rate of the cold plate connected to the memory module according to the temperature.

[0014] In an exemplary embodiment, when the resistance of the liquid leakage detection line changes, the server is powered off.

[0015] In an exemplary embodiment, the above-mentioned reading the resistance change of the liquid leakage detection line through a management controller on the server motherboard and initializing the memory module when there is no resistance change includes: powering on the management controller when the standby power is turned on; the management controller scans the liquid leakage detection line through the first communication bus and reads the resistance change of the liquid leakage detection line; when there is no resistance change, powering on the motherboard; when the power-on timing is to power on the memory, the memory connector inputs a first voltage to the memory module through the voltage pin; the voltage conversion module in the memory module converts the received first voltage into a second voltage; when all the memory particles in the memory module receive the second voltage, the memory module sends a power-on completion signal to the motherboard; the clock chip in the memory module generates a signal with a first frequency and sends it to the memory particles; the processor on the motherboard obtains the configuration information of the memory module through the second communication bus and determines the presence information of the memory module; when the presence information of the memory module indicates that the memory module is working properly, it is determined that the initialization of the memory module is completed.

[0016] According to another embodiment of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0017] According to another embodiment of the present application, there is also provided an electronic device including a memory and a processor, the memory storing a computer program, and the processor being configured to run the computer program to execute the steps in any one of the above method embodiments.

[0018] According to another embodiment of the present application, there is also provided a computer program product including a computer program, and the steps of the methods in various embodiments of the present application are implemented when the computer program is executed by a processor.

[0019] Through the present application, a memory connector is used to connect to the motherboard of the server; a memory module is used to store data, wherein the memory module is connected to the memory connector according to a preset pin configuration, the memory module includes a plurality of memory chips, and the memory chips are horizontally placed on the memory module; a cold plate is used to absorb the heat generated by the memory module, wherein the cold plate is horizontally connected to the memory module. The memory chips in the memory module are horizontally distributed and connected to the motherboard through the memory connector, and the cold plate can be horizontally placed on the memory module instead of being vertically placed, thereby increasing the contact area between the cold plate and the memory module and saving the space occupied by the vertical arrangement. Thus, the technical problem that the current memory module has a relatively low operating speed in the related art is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a hardware structure block diagram of an optional liquid-cooled based storage method according to an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of an optional liquid-cooled based storage device according to an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of another optional liquid-cooled based storage device according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of an optional two-sided used liquid-cooled based storage device according to an embodiment of the present application;

[0024] Figure 5 is a schematic diagram of an optional liquid-cooled based storage system according to an embodiment of the present application;

[0025] Figure 6 is a schematic diagram of another optional liquid-cooled based storage system according to an embodiment of the present application;

[0026] Figure 7 is a schematic diagram of an optional double - layer structure liquid - cooled based storage system according to an embodiment of the present application;

[0027] Figure 8 is a schematic diagram of another optional double - layer structure liquid - cooled based storage system according to an embodiment of the present application;

[0028] Figure 9 is a flowchart of an optional liquid - cooled based storage system according to an embodiment of the present application. Detailed implementation manners

[0029] In the following, embodiments of the present application will be described in detail with reference to the drawings and in combination with embodiments.

[0030] It should be noted that terms such as "first", "second", etc. in the description and claims of the present application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence.

[0031] Figure 1 is a hardware structure block diagram of an optional liquid - cooled based storage method according to an embodiment of the present application. As Figure 1 shown, according to one aspect of an embodiment of the present application, a liquid - cooled based storage device is provided. As an optional implementation manner, the above - mentioned liquid - cooled based storage device can be, but is not limited to, applied to a liquid - cooled based storage system in a hardware environment as Figure 1 shown. Among them, the liquid - cooled based storage system can include, but is not limited to, a motherboard 102, a processor 104, and a liquid - cooled based storage device 106.

[0032] Among them, the motherboard (Motherboard) 102 is the central platform of the server. It provides physical interfaces and circuit connections for connecting and managing all hardware components in the server. A variety of functions are integrated on the motherboard, including but not limited to a processor socket, memory slots, storage interfaces, network interfaces, power management circuits, and a bus for communication and control. The design and quality of the motherboard directly affect the performance, stability, and scalability of the server.

[0033] The CPU (Central Processing Unit) 104 is the "brain" of the server, responsible for executing instructions and processing data. In a server environment, the CPU usually has higher computing power and a multi - core design to support concurrent processing and high - intensity computing requirements. Server CPUs (such as Intel's Xeon or AMD's EPYC series) are usually equipped with high - speed caches and advanced instruction sets to optimize the performance of multi - tasking, virtualization, and data - intensive applications.

[0034] The liquid-cooled storage device 106 is the hardware in a server for storing running data and programs. It consists of a series of DRAM (Dynamic Random Access Memory) chips, usually in the form of DIMM (Dual In-line Memory Module), RDIMM (Registered DIMM), or LRDIMM (Load Reduced DIMM). The communication speed between the memory module and the CPU directly affects the server's response time and data processing rate. In a server, to improve performance and stability, the memory module usually supports error checking (ECC) and hot-swap capabilities for maintenance or upgrade without shutting down the system.

[0035] In the server architecture, these three components work closely together to jointly determine the server's computing power and data processing efficiency. The motherboard provides the physical platform and communication bridge, enabling the CPU to access the memory module through the memory bus for data reading, writing, and processing. The type, capacity, and speed of the memory module (liquid-cooled storage device) are one of the key factors affecting server performance, while the architecture and core count of the CPU determine the upper limit of the server's computing power. A reasonable motherboard design and a combination of high-performance CPU and memory module are the keys to building a high-performance server.

[0036] The above usage scenarios are just examples. Any scenario related to liquid-cooled storage can utilize the solution of this application. This embodiment does not make any limitations in this regard.

[0037] In this embodiment, a liquid-cooled storage device is provided. Figure 2 It is a schematic diagram of an optional liquid-cooled storage device according to an embodiment of the present application. As Figure 2 shown, the liquid-cooled storage device includes:

[0038] A memory connector 202 for connecting to the motherboard of the server;

[0039] A memory module 204 for storing data. Among them, the memory module is connected to the memory connector according to a preset pin configuration. The memory module includes multiple memory chips, and the memory chips are horizontally placed on the memory module;

[0040] A cold plate 206 for absorbing the heat generated by the memory module. Among them, the cold plate is horizontally connected to the memory module.

[0041] It should be noted that the memory connector 202 is a specially designed connector for physically and electrically connecting the memory module to the motherboard 208 of the server. Different from the traditional DIMM slot, it uses a crimping method, reducing the loss during signal transmission, improving signal integrity and transmission rate. The memory module 204 contains multiple DRAM chips, which are modules for storing data and instructions. In this application, the DRAM chips are placed horizontally on the memory module, different from the traditional vertical placement method. This can make more effective use of space, increase memory density, and facilitate liquid cooling. The cold plate 206 is a heat dissipation device for absorbing the heat generated by the memory module and taking away the heat through liquid circulation. In this design, the cold plate is horizontally connected to the memory module and directly attached above the memory chips, improving the heat dissipation efficiency and avoiding the disadvantages of traditional vertical cold plates, such as occupying too much lateral space and low heat dissipation efficiency.

[0042] The memory module is the core of the storage device. It contains multiple storage units (memory chips), and these units are connected to the connector through a preset pin configuration to ensure the correct reading and writing of data. The horizontally placed memory chip design helps to improve density and heat dissipation effect.

[0043] In an alternative embodiment, the memory connector uses crimping technology, which can reduce signal loss and increase the transmission speed compared with traditional slots. The memory chips on the memory module are placed horizontally, which not only increases the storage density of the memory but also facilitates liquid cooling, because the cold plate can be directly attached above the memory chips to quickly take away the heat. This design not only solves the signal integrity problem of traditional memory technology during high-speed operation but also effectively solves the heat dissipation problem in the high-density memory layout by improving the heat dissipation method, providing the server with more efficient and stable data storage and processing capabilities.

[0044] In an alternative embodiment, assume there is a server based on DDR5 technology that needs to achieve higher memory density and better heat dissipation effect in a limited space. First, the traditional DIMM slot can be replaced with a new type of crimping add-on memory connector 202, which predefines the pin distribution compatible with RDIMM to meet the SI (Signal Integrity) specification and ensure high transmission rate even under 2SPC (Dual In-line Package Configuration). Then, design and install the memory module 204, which can use DRAM chips mounted on both sides (front and back), significantly increasing the memory capacity. Each memory chip is placed horizontally on the module, and this layout allows the cold plate 206 to be seamlessly attached above them. The cold plate is internally designed with a circulating liquid path and can be efficiently connected to an external liquid cooling system to absorb and take away the heat generated during memory operation.

[0045] In an alternative embodiment, the preset pin configuration may include: Data pins (DQ): used to transfer memory data. Data Strobe (DQS): used to synchronize data transfer. Control signals: including Chip Select (CS), Row Address Strobe (RAS), Column Address Strobe (CAS), etc., used to control memory operations. Address signals: used to specify addresses in the memory. Power and ground: provide power to the DIMM module and ensure stable signal transmission.

[0046] Table 1 is a schematic pin configuration table of an alternative liquid-cooled storage device according to an embodiment of the present application; as shown in Table 1, in a Registered DIMM (RDIMM) design, a DIMM module can support multiple physical Banks or Ranks, and each Rank can operate independently. DDR1A and DDR1B here represent two different data channels, which can be addressed and operated separately, thereby improving memory bandwidth and performance. By using two independent data channels, the memory module can perform two data transfers simultaneously, one on the DDR1A channel and the other on the DDR1B channel. This design effectively improves the data transfer bandwidth of the memory module. Designing two channels allows the memory module to support more Banks, thereby increasing the memory capacity. Each channel can be configured with a different number of Banks to meet different memory capacity and performance requirements. Each channel has its own data, command, and control signals, which helps to reduce interference between signals and improve signal integrity. Each channel may have its own power management requirements, especially in high-frequency or high-capacity memory modules, and separate power management can more effectively control power noise and power consumption.

[0047] Table 1 Schematic Pin Configuration Table of Liquid-Cooled Storage Device

[0048]

[0049]

[0050] For the above pin configuration, the following explanations are made:

[0051] DRAM chip pins: DDR1B_DQ3[7], DDR1B_DQ3[5], DDR1B_DQ3[3], etc.: These are the data pins (DQ) of the DRAM chip, used to transfer data. On the connector, these pins need to be connected to the corresponding data pins on the motherboard.

[0052] DDR1B_DQS3_t, DDR1B_DQS3_c: These are the Data Strobe (DQS) signals of the DRAM chip, used to synchronize data transfer. "t" and "c" represent the positive and negative poles of the differential signal respectively.

[0053] Control and Address Pins: DDR1B_CA[0], DDR1B_CA[1], etc.: These are address pins used to specify addresses in memory.

[0054] DDR1B_CS[0], DDR1B_CS[1]: Chip select pins used to select memory chips.

[0055] Power Pin: VIN_BULK: This is the power pin used to supply power to the memory module.

[0056] Ground (VSS): The ground pin is used to provide a reference voltage point to ensure signal stability and reduce noise.

[0057] Special Function Pin: PWR_GOOD_1: Power good signal indicating that the power supply has stabilized.

[0058] RSP_B, PAR_B: May be related to response or parity.

[0059] MICHANICAL KEY: Mechanical key used to ensure correct installation of the memory module.

[0060] Undefined or Reserved Signals (RFU): These pins are not currently in use and may be reserved for future expansion or upgrade.

[0061] When making the actual layout, the following factors need to be considered: Signal path length: Try to shorten the path length of critical signals (such as DQS) to reduce delay and improve signal quality. Differential signal matching: Differential signals (such as DQS) need to be correctly matched to maintain signal integrity. Power and ground layout: Ensure uniform distribution of power and ground pins to provide a stable power supply and reduce power noise. Avoid signal interference: Reduce interference between signals through appropriate layout and shielding measures.

[0062] Figure 3 is a schematic diagram of another optional liquid-cooled storage device according to an embodiment of the present application, as Figure 3As shown, the memory module can include a flash PCB, DRAM, PMIC, etc. Multiple memory chips can be included in the memory module, and the memory chips can be symmetrically distributed. Among them, DRAM (Dynamic Random Access Memory) is the most common type of system memory, used to store data and programs when the computer is running. It is a volatile memory, meaning that the stored data will be lost when the power is cut off. PMIC (Power Management Integrated Circuit) is responsible for managing the power supply of the memory module, including providing stable voltage and current, and regulating the power supply under different working states (such as active, standby, sleep) to optimize energy efficiency and performance. SENSOR (temperature sensor) is used to monitor the temperature of the memory module in real time to ensure the stability and reliability of the memory module in a high-temperature environment. When the temperature exceeds the preset threshold, the system can take cooling measures or reduce performance to protect the hardware. DB (Data Buffer) is used to temporarily store data between the memory module and the memory controller to relieve the data transmission bottleneck and improve the efficiency and reliability of data transmission. RCD (Register Clock Driver) is a component used in the RDIMM (Registered DIMM) module. It is responsible for driving and buffering the clock signal to ensure the stable transmission of the clock signal in the memory module, thus maintaining data synchronization. SPD HUB (Serial Presence Detect Hub) is a hub used in the memory module. It is responsible for collecting the configuration information of the memory module (such as size, speed, timing, etc.) and providing this information to the system BIOS or operating system through the serial presence detect (SPD) technology for correct configuration and use of the memory.

[0063] These components work together to ensure that the memory module can operate stably under various working conditions, while providing efficient data storage and access capabilities. By integrating these functions, the memory module can meet the requirements of modern computer systems for performance, reliability, and energy efficiency.

[0064] Through this application, a memory connector is used to connect to the motherboard of a server; a memory module is used to store data. Among them, the memory module is connected to the memory connector according to a preset pin configuration. The memory module includes multiple memory chips, and the memory chips are horizontally placed on the memory module; a cold plate is used to absorb the heat generated by the memory module, and the cold plate is horizontally connected to the memory module. The memory chips in the memory module are horizontally distributed and connected to the motherboard through the memory connector. The cold plate can be horizontally placed on the memory module instead of being vertically placed, thereby increasing the contact area between the cold plate and the memory module and saving the space occupied by vertical arrangement. Thus, the technical problem of the relatively low operating speed of the current memory module in the related art is solved.

[0065] In an alternative embodiment, memory chips are placed on both sides of the front of the memory module, and there are two data channels in one side of the memory module.

[0066] It should be noted that placing memory chips on both the front and back means that when the memory module is designed, memory chips are placed not only on one side of the circuit board but also on the other side, thereby significantly increasing the storage capacity without increasing the overall thickness of the module. Data channel: In a memory module, a data channel refers to the path for data transmission, usually connected to a memory controller. Each channel can independently transmit data, improving the data transmission efficiency and bandwidth. Such as DDR1A and DDR1B.

[0067] Facing different capacity requirements, different-capacity DRAMs can be selected based on this layout. The capacity ranges from 512M to 2G, and memory chips can be mounted on both the front and back, further expanding the memory capacity. The memory capacity can be more than doubled on this basis, thus forming a basic connection relationship for a memory module supporting 2SPC.

[0068] In an alternative embodiment, the memory module can adopt the technology of double-sided mounting of DRAM chips, which doubles the memory capacity within the same physical space. More importantly, each side is equipped with two data channels, which means that the storage chips on each side can independently exchange data with the memory controller, greatly enhancing the data transmission ability of the memory. This dual-channel design, combined with the double-sided chip layout, significantly improves the data transmission bandwidth of the memory module, provides strong support for high-load applications, and also provides a physical basis for the effective heat dissipation of the liquid cooling system because more chips can be evenly distributed, which is beneficial to the heat absorption and distribution of the cold plate.

[0069] Figure 4 It is a schematic diagram of an alternative two-sided liquid-cooling-based storage device according to an embodiment of the present application. As Figure 4 (a) shows, it is the front of the memory module. As Figure 4(b) shows the reverse side of the memory module. The chips on the front and reverse sides can be similar or symmetrical to reduce interference. When designing a high-performance server for AI computing, the memory module adopting the above technical solution is used. 24 DRAM chips are placed on each side of the memory module, forming two independent data channels. There are a total of 48 chips, providing doubled storage capacity.

[0070] Through the above implementation manners of the present application, since there are two data channels on each side, this enables data transmission not to be restricted even under a high-density layout, maintaining a high transmission rate, ensuring that the memory operating temperature is within a safe range, and preventing the thermal protection mechanism from being triggered, which may affect the system performance. The server can achieve a substantial increase in memory capacity within a limited space, while maintaining a high transmission rate and good heat dissipation performance, meeting the requirements of high-load applications such as AI model training and big data processing, and reducing the system cost and heat dissipation challenges brought about by increasing the memory.

[0071] In an alternative implementation manner, the data channels of the above memory module transmit data to the processor, and all the memory chips in the data channels are connected through chip select signals.

[0072] It should be noted that in the memory module, a data channel refers to a set of wire harnesses and related circuits for data transmission, which connect the memory chips and the processor, enabling data to be transmitted at high speed between the two. A chip select signal is a signal used to select a specific memory chip or group in the memory module, so that the processor can read or write data. The chip select signal provides the processor with access control to multiple memory chips and is the key to implementing multi-channel memory technology.

[0073] Each data channel in the memory module is directly connected to the processor, allowing data to be transmitted bidirectionally between the processor and the memory chips. Through the chip select signal, the processor can precisely control and access any memory chip within the data channel, realizing efficient data reading and writing operations, while managing the access to multiple memory chips, avoiding conflicts, and improving the overall performance.

[0074] In an alternative implementation manner, the data channels play a core role. They are the bridges for data exchange between the processor and the memory chips. Each data channel integrates a series of DRAM chips, and these chips are connected through dedicated chip select signal lines, enabling the processor to precisely select and access specific memory chips as needed. The use of the chip select signal ensures that the processor can still effectively manage the memory resources under a high-density memory layout, avoiding data transmission conflicts, and improving the data processing speed and response ability of the system.

[0075] In an alternative embodiment, assume that a server for high-performance computing (HPC) is being designed. This server needs to process a large amount of data, so it has extremely high requirements for memory bandwidth and capacity. To ensure system stability and data integrity, complex signal integrity (SI) and power management (PMIC) mechanisms can also be designed. Through optimized data channel design, low latency and high signal quality can be maintained even under high-frequency DDR5 or DDR6 operations. The PMIC ensures the power supply of memory chips in different operating states, dynamically adjusts the power supply according to the activity of the chip select signal, reduces power consumption, and extends the service life of the memory module.

[0076] Through the above embodiments of the present application, the memory chips and data channels connected by the chip select signal can achieve efficient management and utilization of memory resources in high-performance servers, improve data processing speed, meet the stringent requirements of HPC applications for memory. At the same time, through the application of liquid cooling technology, it ensures the stable operation and low power consumption characteristics of the system under a high-density memory layout.

[0077] In this embodiment, a liquid-cooled storage system is provided. Figure 5 It is a schematic diagram of an alternative liquid-cooled storage system according to an embodiment of the present application, as Figure 5 shown. The liquid-cooled storage system includes:

[0078] A memory device 502, where the memory device is the device of any device embodiment;

[0079] A main board 504 for carrying different modules and enabling data transmission between different modules;

[0080] A processor 506 for controlling the memory device and reading or writing data in the memory device. The processor is connected to the main board and is in point-to-point contact connection with the memory device.

[0081] It should be noted that the point-to-point contact connection is a connection technology, different from the traditional slot or gold finger connection method. It uses direct point-to-point contact to reduce signal loss and improve signal integrity, and is suitable for highly integrated and high-density designs in liquid-cooled storage systems.

[0082] The memory device 502 may include a memory connector 508, a memory module 510, and a cold plate 512.

[0083] The main board 506 is the foundation of the storage system. It not only supports the installation of multiple modules but also coordinates the data interaction between these modules. It is the center of data transmission for the entire system. The processor 506 is the control core of the storage system. Through its connection with the main board, it further realizes high-efficiency data reading and writing operations with the memory device. This point-contact connection method reduces signal latency and improves system performance.

[0084] In an alternative embodiment, the system consists of three parts: a memory device, a main board, and a processor. Among them, the memory device adopts an innovative liquid cooling technology, which can significantly improve the memory operation speed and density. The main board, as the framework of the system, not only carries different types of modules but also coordinates the communication between the modules. It is the key to data transmission. The processor is closely connected to the main board and directly communicates with the memory device in a point-contact manner. This design reduces the signal latency and loss brought by the traditional connection method, ensuring high speed and efficiency in data processing.

[0085] When designing a high-performance AI training server, a liquid-cooled storage system technical solution can be adopted. First, a memory device is customized. It consists of multiple memory modules. Each module is connected to the main board contact point by point using a compression additional memory connector. This connection method reduces signal loss and improves the signal integrity of data transmission. The DRAM particles on the memory module are mounted on both the front and back sides. Through optimized pin definitions and layouts, it supports two data channels, improving data processing speed and bandwidth.

[0086] The main board is designed as a multi-module support platform. Each memory device has a corresponding compression additional connector at its position to achieve a high-density memory layout. The main board also integrates the processor and other necessary control logics. The processor directly communicates with the memory device through a point-contact connection method. This design reduces the signal path, improves the speed of data reading and writing, and ensures the large data reading and writing requirements during AI training.

[0087] Through this application, a memory connector is used to connect to the main board of the server; a memory module is used to store data. Among them, the memory module is connected to the memory connector according to a preset pin configuration. The memory module includes multiple memory particles, and the memory particles are horizontally placed on the memory module; a cold plate is used to absorb the heat generated by the memory module. Among them, the cold plate is horizontally connected to the memory module. The memory particles in the memory module are horizontally distributed and connected to the main board through the memory connector. The cold plate can be horizontally placed on the memory module without the need for vertical placement, thereby increasing the contact area between the cold plate and the memory module and saving the space occupied due to vertical arrangement. Thus, it solves the technical problem that the current memory module in the related technology has a relatively low operating rate.

[0088] In an alternative embodiment, a plurality of memory devices are placed on the main board, and the memory devices are evenly distributed on both sides of the processor.

[0089] It should be noted that the even distribution is on the main board, and the plurality of memory devices are designed to be placed at equal distances and intervals on both sides of the processor. This layout helps to optimize the signal transmission path, reduce signal latency, and at the same time balance the internal heat distribution of the system, improving the overall heat dissipation efficiency.

[0090] Figure 6 is a schematic diagram of another alternative liquid-cooled storage system according to an embodiment of the present application; as Figure 6 shown, a processor 606 is installed on the main board 604, and there are two memory devices 602 measured on the processor 606. A plurality of memory devices are placed on the main board, and they are symmetrically and evenly distributed on the left and right sides of the processor. This design not only maximizes the use of the space on the main board, but also ensures the balance of data transmission and the uniformity of heat dissipation, which is the key to realizing a high-performance server memory architecture.

[0091] In an alternative embodiment, the main board plays the role of a core platform, which carries a plurality of liquid-cooled memory devices. These devices are carefully arranged on both sides of the processor in an evenly distributed manner. This layout minimizes the signal transmission path between the processor and the memory devices, reduces signal latency, and improves data transmission speed. Secondly, the even distribution is beneficial to heat dissipation. The heat generated by the memory devices can be more evenly absorbed by the cold plate, avoiding local overheating, ensuring that the temperature of the entire system is controlled within a safe range, and enhancing the stability and reliability of the system.

[0092] In an alternative embodiment, during layout, it can be ensured that each memory device is evenly distributed on both sides of the processor. This layout strategy not only considers the shortest signal transmission path, but also fully considers the balance of heat dissipation. The processor is located at the center of the main board, and the memory devices are symmetrically distributed on both sides. Each device is equidistant from the processor, minimizing the latency of data transmission. In addition, the liquid-cooled cold plate is also designed to cover the entire memory area, so that both the memory devices on the front side and the back side of the processor can be effectively cooled, ensuring that the temperature of the memory devices remains within a safe range under high-load operation, avoiding performance degradation or hardware damage caused by overheating.

[0093] Through the above embodiments of the present application, by evenly distributing the memory devices on both sides of the processor and combining the liquid-cooled heat dissipation technology, the optimization of the high-performance server memory architecture is achieved, providing higher data transmission speed and capacity, while ensuring the stable operation and efficient heat dissipation of the system. This design is particularly suitable for application scenarios with extremely high memory requirements such as AI computing and big data processing.

[0094] In an alternative embodiment, a first memory device and a second memory device are placed on the motherboard, wherein the height of the memory connector in the first memory device is greater than the height of the second memory device.

[0095] It should be noted that the first memory device and the second memory device may refer to different memory devices installed on the motherboard, and they may have different designs and specifications, especially in terms of the height of the memory connector, to meet different space requirements and heat dissipation strategies.

[0096] Although both the first memory device and the second memory device are used to store data, the height of their memory connectors can be different. The memory connector of the first memory device is higher, which is usually to solve specific space constraints and heat dissipation requirements. For example, a higher connector may be used in situations that require stronger heat dissipation capabilities, while a shorter connector may be used in space-saving designs.

[0097] Figure 7 It is a schematic diagram of an alternative double-layer liquid-cooled storage system according to an embodiment of the present application; as Figure 7 shown, a processor 706 is installed on the motherboard 704, and a first memory device 702 and a second memory device 708 may also be installed on the motherboard 704. The height of the memory connector of the first memory device 702 may be greater than the height of the second memory device 708. The motherboard is designed to support memory devices of different specifications to adapt to complex and changeable application scenarios. The positions and layouts of the first memory device and the second memory device on the motherboard are carefully planned. The height of the memory connector of the first memory device is greater than that of the second memory device. This design choice is based on a comprehensive consideration of the heat dissipation requirements and space utilization rate of the memory device. A higher memory connector may mean more heat dissipation space, which is suitable for high-power, high-density DRAM particles, while a shorter connector saves more space and is suitable for situations with less heat dissipation requirements.

[0098] Suppose a server motherboard for high-performance computing (HPC) and AI acceleration is being designed. When designing the motherboard, different memory device requirements for heat dissipation and space are considered. The first memory device is designed as a high-density, high-bandwidth DRAM stack module, suitable for scenarios that require a large amount of data storage and fast data access, such as large-scale neural network training. Due to the heat generation problem of high-density DRAM particles, a compressed additional memory connector with a height of 7.5 mm is selected for the first memory device, so that a larger and more effective cold plate can be installed above the device to achieve efficient heat dissipation.

[0099] The second memory device is designed as a memory module with standard capacity and low power consumption, suitable for occasions with relatively low memory requirements but strict requirements for system stability and cost control, such as the operating system storage of servers. Considering space savings and cost control, a compressed additional memory connector with a height of 1.75 mm can be selected for the second memory device. This shorter design makes the entire device more compact, occupies less space, and has relatively lower heat dissipation requirements, allowing for the use of a thinner cold plate design.

[0100] In the motherboard layout, the first memory device is located on one side of the processor, and a large cold plate is installed above it. The second memory device is located on the other side of the processor, and a thin cold plate is used above it. This layout ensures that both memory devices on both sides of the processor can be properly cooled. At the same time, through memory connectors with different heights, the motherboard can support different types of memory devices within a limited space, meeting diverse requirements and improving the overall computing performance and energy efficiency ratio.

[0101] When the server is running, the processor can flexibly select the first memory device to handle high-load computing tasks, taking advantage of its high bandwidth and large capacity. At the same time, the second memory device provides stable background support to maintain the operating efficiency and stability of the system. This design that combines memory connectors of different heights not only optimizes memory performance but also achieves a balance between heat dissipation and space utilization. It is an innovative direction in server memory architecture design, especially suitable for high-performance computing environments that need to balance performance, heat dissipation, and cost.

[0102] In an alternative embodiment, the first memory device and the second memory device are placed on both the front and the back of the motherboard.

[0103] It should be noted that the front and the back refer to the two sides of the motherboard. The front usually installs the processor, chipset, and other main electronic components, while the back can be used for the installation of additional components such as radiators and additional memory devices. In this technical solution, both the front and the back of the motherboard are used to install memory devices to increase the memory capacity and density of the system.

[0104] Utilize the space on both the front and the back of the motherboard to install the first memory device and the second memory device simultaneously. The purpose of this is to significantly increase the memory capacity without increasing the physical size of the server, and at the same time, according to different heat dissipation requirements, reasonably arrange the layout of the devices and the heat dissipation strategy.

[0105] Figure 8 It is a schematic diagram of another alternative double-layer structure liquid-cooled storage system according to an embodiment of the present application; as Figure 8As shown, a processor 806 is installed on the main board 804, and a first memory device 802 and a second memory device 808 are installed on both the front and back sides. In this application, the main board adopts an efficient space utilization strategy to meet the demand for a large amount of memory resources in high-performance computing devices. The first memory device and the second memory device are not only arranged on one side of the main board, but also installed on its back. This double-sided design maximally utilizes all available space on the main board and significantly improves the memory density of the server. By placing memory devices on both the front and back sides of the main board simultaneously, the memory capacity can be doubled without significantly increasing the volume of the server, which is crucial for the construction of high-density computing nodes.

[0106] Suppose a high-performance server is designed for a large data center, which needs to provide an ultra-large memory capacity within limited rack space. According to this application, a strategy of installing memory devices on both sides of the main board is adopted. On the front side of the main board, the first memory devices are installed. These devices adopt a high-density DRAM particle layout, use a 7.5mm high compression additional memory connector, and a large cold plate is installed to cope with the large amount of heat generated by high-speed data exchange. On the back side, the second memory devices are installed. They may adopt lower-power DRAM particles, use a 1.75mm high connector, and are also equipped with a cold plate, but the design is thinner and lighter to save space.

[0107] This double-sided installation layout enables the server to achieve a memory capacity that originally required 4U or more space within the standard 1U or 2U rack space. During installation, first install the second memory devices on the back side of the main board, ensure that the connectors are in close contact with the main board interfaces, then install the cold plates and conduct a leak check. After that, install the first memory devices and their cold plates on the front side of the main board in the same steps. This design not only improves the memory capacity, but also effectively controls the heat through the liquid cooling system, ensuring the stable operation of the memory devices under high load.

[0108] After the server runs, the processor can access the memory devices on both the front and back sides simultaneously. Through an optimized signal transmission path, good signal quality and data transmission speed can be maintained even under a high-density layout. The system management chip continuously monitors the temperature of each side of the memory devices and dynamically adjusts the coolant flow rate of the cold plates to maintain the best cooling effect and prevent overheating. This innovative main board design, combined with an efficient liquid cooling technology, significantly improves the memory capacity and computing performance of the server, while reducing the cooling cost and power consumption. It is an ideal choice for constructing high-performance computing nodes in modern data centers.

[0109] In this embodiment, a liquid-cooled storage method is provided. Figure 9It is a flowchart of an optional liquid-cooled based storage system according to an embodiment of the present application. As shown in the figure, the liquid-cooled based storage method includes:

[0110] S902, read the resistance change of the liquid leakage detection line through the management controller on the server motherboard. In the case of no resistance change, initialize the memory module;

[0111] S904, the management controller obtains the temperature of the memory module and adjusts the coolant flow rate of the cold plate connected to the memory module according to the temperature.

[0112] It should be noted that the management controller on the server motherboard refers to the baseboard management controller (BMC) or a similar system management chip on the server motherboard. It is responsible for monitoring the operating status of the server, including parameters such as temperature, voltage, and fan speed, and can control operations such as starting, restarting, and shutting down the server.

[0113] In the liquid-cooled system, to prevent coolant leakage, liquid leakage detection lines are usually set. These detection lines are usually made of special materials and can sense the presence or leakage of coolant. Under normal circumstances, the resistance value of the liquid leakage detection line remains stable; when there is coolant leakage, the resistance value will change, triggering an alarm to prevent hardware damage.

[0114] In a liquid-cooled environment, a method for managing the operation and maintenance of server memory modules includes initializing the memory module and dynamically monitoring the memory temperature, and adjusting the coolant flow rate of the cold plate to maintain the stable operation of the memory module.

[0115] The liquid-cooled based storage method aims to ensure that the server memory module can operate safely and efficiently in a liquid-cooled environment. Before the server starts, the BMC first reads the resistance change of the liquid leakage detection line to confirm that there is no sign of coolant leakage. This step is a prerequisite for ensuring the safe operation of the liquid cooling system. Once no leakage is detected, the BMC will initialize the memory module to prepare for data transmission.

[0116] During the operation of the server, the BMC will continuously monitor the temperature of the memory module and obtain real-time temperature information through the built-in temperature sensor. According to the temperature of the memory module, the BMC can dynamically adjust the coolant flow rate of the cold plate connected to the memory module. When the temperature rises, increasing the flow rate can more effectively remove heat, preventing the memory module from overheating, thus ensuring the integrity of the memory data and the stability of the system. This temperature-based dynamic adjustment mechanism is the key for the liquid-cooled system to efficiently handle high-power, high-heat-load memory devices.

[0117] Through the present application, a memory connector is used to connect to the motherboard of a server; a memory module is used to store data. Among them, the memory module is connected to the memory connector according to a preset pin configuration. The memory module includes a plurality of memory chips, and the memory chips are horizontally placed on the memory module; a cold plate is used to absorb the heat generated by the memory module. Among them, the cold plate is horizontally connected to the memory module. The memory chips in the memory module are horizontally distributed and are connected to the motherboard through the memory connector. The cold plate can be horizontally placed on the memory module without the need for vertical placement, thereby increasing the contact area between the cold plate and the memory module and saving the space occupied by vertical arrangement. Thus, the technical problem of the relatively low operating speed of the current memory module in the related art is solved.

[0118] In an alternative embodiment, when the resistance of the liquid leakage detection line changes, the server is powered off.

[0119] It should be noted that in a liquid cooling system, when signs of coolant leakage are detected, the system will immediately take power-off measures to prevent the coolant from short-circuiting electronic components and causing damage. This safety mechanism is a key link in the application of liquid cooling technology, ensuring a rapid response in the event of leakage and avoiding greater losses.

[0120] In an alternative embodiment, when implementing a liquid-cooled storage system, safety and reliability are important considerations in the design. To prevent damage that coolant leakage may cause to server hardware, a liquid leakage detection line is integrated into the system, which can continuously monitor whether the coolant leaks into areas where it should not exist. When the BMC detects a change in the resistance value of the liquid leakage detection line, which usually means there is coolant leakage, the system will immediately start an emergency protection program to power off the server. Power-off is the most direct and effective measure to prevent the coolant from further causing system short-circuit or damaging electronic components, providing basic security for the system and data.

[0121] Suppose in a high-performance server equipped with a liquid-cooled storage system in a data center, the system is performing a large-scale data processing task, such as the training of a deep learning model. The BMC of the server continuously monitors the operating status of all liquid-cooled memory modules, including temperature and the change in the resistance of the liquid leakage detection line.

[0122] At a certain moment, the BMC detects that the resistance of the liquid leakage detection line of a certain memory module suddenly decreases, which indicates that the coolant may have leaked near the detection line, meaning there is a risk of coolant leakage. To prevent further damage, the BMC immediately responds and sends a power-off command to the power management module of the server, and the power of the server is quickly cut off, and all computing and data transmission activities immediately stop.

[0123] After a power outage, the operation and maintenance team of the data center quickly intervened, isolated the server, inspected the leakage points of the liquid cooling system, repaired the coolant pipes or replaced the damaged memory modules. After confirming that the system was safe and the leakage problem had been resolved, the server was powered on again. The BMC checked the status of all memory modules once more. After ensuring that everything was normal, the server continued to execute the data processing tasks that had been interrupted before. During the whole process, the integrity of the data and the security of the system were effectively guaranteed.

[0124] This embodiment demonstrates that in a liquid-cooled storage system, by real-time monitoring of the resistance change of the liquid leakage detection line, the system can quickly respond to the leakage risk, implement the server power-off measure, ensure the safe operation of the data center hardware, avoid potential serious losses, and reflect the necessary safety protection mechanism when the liquid-cooling technology is applied in the field of high-performance computing.

[0125] In an alternative embodiment, the above-mentioned method of reading the resistance change of the liquid leakage detection line through the management controller on the server motherboard, and initializing the memory module when there is no change in the resistance, includes: powering on the management controller when the standby power supply is turned on; the management controller scans the liquid leakage detection line through the first communication bus to read the resistance change of the liquid leakage detection line; powering on the motherboard when there is no change in the resistance; when the power-on sequence is to power on the memory, the memory connector inputs a first voltage to the memory module through the voltage pin; the voltage conversion module in the memory module converts the received first voltage into a second voltage; when all the memory chips in the memory module receive the second voltage, the memory module sends a power-on completion signal to the motherboard; the clock chip of the memory module generates a signal with a first frequency and sends it to the memory chips; the processor on the motherboard obtains the configuration information of the memory module through the second communication bus and determines the presence information of the memory module; when the presence information of the memory module indicates that the memory module is working properly, it is determined that the initialization of the memory module is completed.

[0126] It should be noted that the first communication bus and the second communication bus are signal lines on the server motherboard for data communication between different components. The first communication bus usually refers to a low-speed communication line used for scanning the liquid leakage detection line and transmitting temperature information, such as the I3C / I2C bus. The second communication bus is a line for high-speed data transmission between the processor and the memory module, such as the DDR memory bus. The standby power supply provides a minimum amount of power even when the server is not fully powered on, which is used for the operation of key components such as the BMC to facilitate system monitoring and management. The first voltage and the second voltage refer to the voltage initially input by the memory connector (the first voltage) and the power supply voltage adjusted by the internal voltage conversion module of the memory module (the second voltage) during the server power-on process. The second voltage is more suitable for the normal operating voltage of the memory chips.

[0127] The management controller (such as BMC) on the server motherboard plays a crucial role during the system power-on process. Its primary task is to power on first when the standby power is turned on, and then scan the liquid leakage detection line through the first communication bus (such as I3C / I2C bus) to monitor whether the coolant leaks. After confirming no leakage risk, the management controller instructs the motherboard to power all components, entering the next step of the power-on sequence.

[0128] During the power-on sequence, the memory connector inputs the first voltage to the memory module through the voltage pins, which is usually an initial and stable power supply voltage. The voltage conversion module (such as PMIC) inside the memory module is responsible for converting the first voltage into the second voltage suitable for the memory chips to work, ensuring that the memory chips start and operate under the correct voltage conditions. After all memory chips receive the second voltage, the memory module sends a power-on completion signal to the motherboard, indicating that the memory module is ready for data transmission.

[0129] Subsequently, the clock chip (such as RCD) of the memory module starts to work, generating a signal with the first frequency to provide a synchronous clock signal for the memory chips, ensuring the accuracy and integrity of data transmission. Meanwhile, the processor on the motherboard obtains the configuration information of the memory module, including parameters such as capacity and speed, through the second communication bus (such as DDR memory bus), and determines the presence information of the memory module to check whether all memory modules have been correctly initialized and are ready to work.

[0130] After confirming that the presence information of the memory module indicates that it can work normally, the processor finally determines that the initialization of the memory module is completed, marking that the server memory subsystem is ready to start executing data processing and storage tasks. The entire initialization process is to ensure that in a liquid-cooled environment, the memory module starts safely and stably, avoiding potential damage to the hardware caused by coolant leakage, while optimizing the power supply and signal synchronization of the memory chips, improving the overall performance and reliability of the memory system.

[0131] Next, the use of this application in various scenarios of the server will be described:

[0132] Scenario 1: Installation and initialization of the memory module in an air-cooled environment;

[0133] In an alternative embodiment, engineers are preparing to deploy a server with an air-cooling system in a data center. First, they select a Compression Attached Memory Module (CAMM) that complies with the DDR5 standard, which has an optimized Signal Integrity (SI) design and improved heat dissipation performance. On the server motherboard, the engineers find the pre-set CAMM installation location and fix the memory module to the motherboard through screw holes and studs to ensure a stable connection. When the server is powered on, the system follows the established power-on sequence. When it reaches the power-on stage of the memory module, the CAMM connector supplies power to the memory module through the VIN_BULK pin. After the power supply starts, the Power Management Integrated Circuit (PMIC) inside the memory module converts the voltage suitable for DRAM operation according to the power-on sequence and voltage requirements of the memory chips, ensuring the safe startup of the memory chips. After the PMIC completes the voltage conversion and power supply, through the PWR_GOOD pin, the memory module sends a power-on completion signal to the CPLD (Complex Programmable Logic Device) on the motherboard, and the motherboard starts to initialize the memory, including setting the clock frequency and voltage parameters. After the initialization is completed, the Row Command Decoder (RCD) of the memory module starts to work, generating a clock signal with a fixed frequency for synchronizing the read and write operations of the DRAM. The host accesses the SPD HUB chip on the memory module through the I3C bus to query the presence of the memory. After confirmation, the memory module and the host exchange memory addresses, and the system enters the normal data processing state, and the GPU acceleration function is activated.

[0134] Scenario 2: Installation and Monitoring of Memory Modules under a Liquid-Cooling System;

[0135] Prepare to install a server equipped with a liquid-cooling system, select a CAMM memory module that supports liquid cooling, and install a cold plate before installation to improve the heat dissipation efficiency. Before installation, the engineers visually inspected the cold plate and confirmed that there were no obvious gaps or damages to prevent coolant leakage. Subsequently, they powered on the server. First, the standby power (STBY) was powered on to ensure the normal startup of the BMC (Baseboard Management Controller), which is responsible for monitoring the system status. The BMC scans the liquid leakage detection line through the I2C bus to check whether there is coolant leakage in the system. Once a change in resistance is detected, indicating a leakage risk, the system will immediately power off. After confirming no leakage, the engineers continued to power on, initialize, and exchange data for the memory module according to the steps in the air-cooling environment. During system operation, especially under high load, the BMC continuously monitors the memory temperature and reads the temperature sensor data through the I3C bus. Once it detects an abnormal increase in temperature, the BMC will notify the CMC (Central Management Controller) or RMC (Remote Management Controller) to increase the coolant flow rate to keep the memory module operating within a safe temperature range and avoid performance degradation or hardware damage caused by overheating.

[0136] Scenario 3: Increasing Memory Density and Capacity;

[0137] When greater memory capacity and higher density are required, the use of dual-sided mounting and stacked memory modules can be considered. First, power off the server and close the water lock valve of the cold plate to prevent coolant leakage when replacing the memory module. Start installing the CAMM memory module from the back of the chassis, ensuring that the cold plate is properly installed and there is no risk of leakage. Next, install the memory module on the front side, and also perform cold plate installation and leakage inspection. When installing the stacked memory module, different height CAMM connector positions are preset on the motherboard, with the lowest height being 1.75 mm and the highest reaching 7.5 mm. Install the low-height memory module first, then the high-height module, and uniformly check for leakage. After installing the multi-layer memory module, adjust the cold plate design to increase the initial coolant circulation speed to adapt to the additional heat generated by the high-density memory. By dynamically adjusting the coolant flow rate, the BMC ensures the temperature safety of the memory module during high-speed data exchange and improves the system operation efficiency. Finally, power on all the memory modules and the cold plate simultaneously, conduct a dual inspection of liquid leakage and temperature. After confirming that there is no error, the server starts running as a whole, and the memory module exchanges data with the host to support high-performance computing and large data processing requirements.

[0138] Scenario 4: Combining front and back mounting with stacked memory modules;

[0139] In the pursuit of extreme memory density, the engineer decides to use both front and back mounted memory modules and stacked memory modules in the same chassis. The memory module on the back of the chassis can be installed first, followed by the module on the front, starting from the low height to the high height in sequence. During the installation process, a strict liquid leakage inspection is carried out after each memory module is installed to ensure the tightness of the cold plate and the safety of the system. After all the memory modules are installed, adjust the coolant circulation speed of the cold plate to adapt to the heat dissipation requirements brought by the high density and high power consumption of the memory module, ensuring that the memory module can remain stable even when working under overload. After the system is powered on, follow the power-on timing sequence to initialize all the memory modules, confirm the status of the memory particles, adjust the clock signal, and complete the memory address exchange to ensure efficient data transmission between the memory module and the host.

[0140] Through the above embodiments, it can be seen how the Compression Attached Memory Module (CAMM) and its liquid cooling support design improve system performance in different application scenarios. Whether in an air-cooled or liquid-cooled environment, whether increasing the single-layer capacity or increasing the memory density through multi-layer stacking, the CAMM memory module demonstrates strong adaptability and expandability, meeting the high requirements of high-performance computing for the memory system.

[0141] The present application also discloses an extended solution for intelligent temperature control, which is a key component for the efficient maintenance and optimization of servers and data centers, especially for systems adopting high-density memory and liquid cooling technologies. The following is a detailed description of an intelligent temperature control solution based on AI technology;

[0142] A Smart Dynamic Temperature Control and Optimization System (SDTCOS) can achieve precise monitoring and intelligent regulation of the internal temperature of servers, optimize cooling efficiency and energy consumption. Automatically adjust the cooling system according to load changes, including the coolant flow rate in the liquid cooling system and the fan speed of the air cooling system, to maintain the best temperature and performance balance. Through AI predictive analysis, adjust the cooling strategy in advance to prevent performance degradation or hardware damage caused by overheating.

[0143] It may include: a temperature sensor network: install high-precision temperature sensors at key parts of the server (such as CPU, GPU, memory modules, etc.) to collect temperature data in real time and provide input for the intelligent control system. An AI control center: adopt an embedded AI chip or cloud service to analyze the collected temperature data, predict future load and temperature trends, and dynamically adjust the cooling strategy. A cooling system interface: include communication interfaces with the liquid cooling system (such as the coolant pump and flow valve of the cold plate) and the air cooling system (such as the fan controller) to achieve real-time control of the cooling system. An energy consumption monitoring and optimization module: monitor the energy consumption of the entire cooling system, and optimize the cooling efficiency and reduce unnecessary energy consumption according to the AI analysis results.

[0144] The specific implementation methods are as follows: Data collection and initialization: When the server is started, the temperature sensor network begins to collect temperature data of each key part, and transmits the data to the AI control center for initial analysis to establish a temperature baseline. AI model training and deployment: Based on historical temperature data and system load, an AI prediction model is trained. This model can predict future temperature changes according to real-time load to achieve early adjustment. The model is deployed locally on the server or in the cloud, and can be flexibly selected according to the server's network connection status and privacy requirements. Dynamic cooling strategy adjustment: Liquid cooling system control: When it is predicted that the temperature will rise, the AI control center increases the flow rate of the coolant through the cooling system interface, and vice versa. In high-density memory applications, by precisely controlling the coolant flow rate of each cold plate, the temperature balance of each memory module is ensured. Air cooling system control: For air cooling components, the AI system adjusts the fan speed according to temperature prediction to avoid excessive noise and energy consumption caused by too high fan speed. Performance and health monitoring: The AI system continuously monitors system performance indicators and hardware health status. Once a performance decline or hardware damage risk caused by overheating is detected, emergency cooling measures are immediately taken, such as accelerating the coolant circulation or turning on the standby cooling system. Energy consumption optimization and reporting: The AI control center analyzes the energy consumption of the cooling system through the energy consumption monitoring and optimization module, and adjusts the cooling strategy to achieve the best balance between performance and energy efficiency. At the same time, reports are generated regularly to provide energy consumption optimization suggestions and system maintenance guidance.

[0145] In addition, the system can also implement an intelligent early warning mechanism: The AI system can predict the overheating risk in advance according to the temperature trend and system status, and notify the operation and maintenance personnel by means of text messages, emails or notifications to avoid potential system failures. Adaptive learning: The AI control center will continuously optimize the prediction model according to real-time data to improve the accuracy and response speed of temperature control. Remote control and management: Through network connection, the AI control center allows remote operation and maintenance, which is convenient for cross-regional data center management and resource optimization, especially in large-scale clusters, to achieve unified intelligent temperature management.

[0146] Through the above implementation methods, through intelligent dynamic temperature control, it is possible to avoid system performance decline and hardware failures caused by overheating. Intelligently adjust the cooling system, reduce unnecessary consumption of cooling resources, and improve the overall energy efficiency of the data center. Through precise temperature control, reduce the time when the hardware is at extreme temperatures and extend the service life of the hardware. Through remote monitoring and intelligent early warning, reduce the number of on-site maintenance times, optimize the maintenance plan of the cooling system, and reduce the operating cost.

[0147] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, this computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which may be a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present application.

[0148] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and the computer program is configured to execute the steps in any one of the above method embodiments when running.

[0149] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk or optical disk, etc., various media that can store computer programs.

[0150] An embodiment of the present application further provides an electronic device, including a memory and a processor, a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0151] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the above processor, and the input / output device is connected to the above processor.

[0152] An embodiment of the present application further provides a computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores the computer program product, and the computer program realizes the steps of the methods described in various embodiments of the present application when executed by a processor.

[0153] The specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.

[0154] Obviously, those skilled in the art should understand that the various modules or steps of the present application described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0155] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present application shall be included within the protection scope of the present application.

Claims

1. A storage device based on liquid cooling, characterized in that: The device comprises: Memory connector, used to connect to the server's motherboard; A memory module for storing data, wherein the memory module is connected to the memory connector according to a preset pin configuration, the memory module includes a plurality of memory particles, and the memory particles are horizontally placed on the memory module; a cold plate for absorbing heat generated by the memory module, wherein the cold plate is horizontally connected to the memory module.

2. The device according to claim 1, characterized in that The memory chips are placed on both sides of the front of the memory module, and there are two data channels on one side of the memory module.

3. The device according to claim 2, characterized in that The data channel of the memory module transmits data to the processor, and all the memory particles in the data channel are connected via a chip select signal.

4. A storage system based on liquid cooling, characterized in that: The system comprises: A memory device, wherein the memory device is the device according to any one of claims 1 to 3; Mainboard, used to carry different modules and realize data transmission between different modules; A processor is used to control the memory device and read or write data in the memory device, wherein the processor is connected to the mainboard and is connected to the memory device in a contact point manner.

5. The system according to claim 4, characterized in that A plurality of the memory devices are placed on the mainboard, and the memory devices are evenly distributed on both sides of the processor.

6. The system according to claim 5, characterized in that A first memory device and a second memory device are placed on the mainboard, wherein a height of a memory connector in the first memory device is greater than a height of the second memory device.

7. The system according to claim 6, characterized in that The first memory device and the second memory device are placed on the front and back of the mainboard at the same time.

8. A storage method based on liquid cooling, characterized in that: The method comprises: Reading the resistance change of the liquid leakage detection line through the management controller on the server mainboard, and initializing the memory module when the resistance does not change; The management controller obtains the temperature of the memory module and adjusts the flow rate of the cooling liquid of the cold plate connected to the memory module according to the temperature.

9. The method according to claim 8, characterized in that The method further comprises: When the resistance of the liquid leakage detection line changes, the server is powered off.

10. The method according to claim 8, characterized in that The step of reading the resistance change of the liquid leakage detection line by the management controller on the server mainboard and initializing the memory module when the resistance does not change comprises: When the standby power supply is turned on, powering on the management controller; The management controller scans the liquid leakage detection line through the first communication bus to read the resistance change of the liquid leakage detection line; When the resistance does not change, supplying power to the mainboard; When the power-on sequence is memory power-on, the memory connector inputs a first voltage to the memory module through a voltage pin; The voltage conversion module in the memory module converts the received first voltage into a second voltage; When all memory chips in the memory module receive the second voltage, the memory module sends a power-on completion signal to the mainboard; The clock chip of the memory module generates a signal of a first frequency and sends it to the memory particles; The processor on the mainboard obtains the configuration information of the memory module through the second communication bus, and determines the in-place information of the memory module; When the presence information of the memory module indicates that the memory module operates normally, it is determined that the initialization of the memory module is completed.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 8 to 10 when executed by a processor.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 8 to 10 are implemented.

Citation Information

Patent Citations

  • Memory information reading device and method, computing equipment mainboard, equipment and medium

    CN114924998A

  • Cold plate type heat dissipation device for server

    CN115857644A

  • Server liquid leakage processing method, system and device, electronic equipment and medium

    CN117033063A

  • Memory expansion device, server and server cluster

    CN117807013A

  • Memory liquid cooling radiator easy to disassemble memory bank and control method

    CN118550382A

Cited By

  • Server

    CN120723040A

  • Circuit board assembly and electronic equipment

    CN121455300A