Control method of a server

By disabling the eDPC function of CXL devices while enabling the eDPC function on the server, and controlling the server status according to the error type, the problem of insufficient support for CXL devices by the eDPC function is solved, and the reliability and stability of the system are improved.

CN120723522BActive Publication Date: 2025-11-21INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511203059.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-21
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

In scenarios involving large-scale deployment of CXL devices, the eDPC function provides insufficient support for the devices, leading to system downtime in the event of uncorrectable errors, which affects business continuity.

Method used

With the eDPC function enabled on the server, disable the eDPC function on the CXL device and control the server's operating status according to the type of uncorrectable error, including maintaining the operating status unchanged in the case of non-downtime type errors and restarting the server in the case of downtime type errors.

Benefits of technology

This effectively avoids system downtime caused by insufficient support for CXL devices by the eDPC function, improves system reliability and stability, prevents the spread of erroneous data, and maintains the operation of critical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723522B_ABST
    Figure CN120723522B_ABST
Patent Text Reader

Abstract

The application discloses a kind of server control methods, belong to server technical field.The control method of the server includes: control server downstream port isolation function opens;For the first equipment of supporting high-speed interconnection technology connected to the server, disable the downstream port isolation function corresponding to the first equipment;In the case where it is detected that the first equipment occurs uncorrectable error, and the uncorrectable error is non-down type error, keep the current running state of the server unchanged;In the case where it is detected that the first equipment occurs uncorrectable error, and the uncorrectable error is down type error, control the server to be down and restart.The control method of the server of the application can avoid the occurrence of IERR due to the poor support of eDPC function to CXL equipment, and can effectively improve the reliability and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of server technology, and in particular relates to a server control method. Background Technology

[0002] Compute Express Link (CXL) technology, by supporting memory-coherent scaling and efficient data transfer, has become a key technology for building high-performance computing platforms. However, in scenarios with large-scale deployment of CXL devices, system stability issues caused by device failures are becoming increasingly prominent. In particular, when an Uncorrectable Error (UCE) occurs, it can lead to a system-wide IERR (Intra-Intra-Outgoing Error), severely impacting business continuity. While some technologies utilize Enhanced Downstream Port Containment (eDPC) to mitigate downtime caused by uncorrectable errors, even with eDPC enabled, UCE errors in CXL devices can still lead to IERR errors, indicating poor support for eDPC on CXL devices. Summary of the Invention

[0003] This application aims to at least solve one of the technical problems existing in the related art. To this end, this application proposes a server control method that can avoid the occurrence of IERR caused by poor support of eDPC function for CXL devices, and can effectively improve the reliability and stability of the system.

[0004] Firstly, this application provides a method for controlling a server, including:

[0005] Enable downstream port isolation for the control server;

[0006] For the first device that supports high-speed interconnection technology connected to the server, disable the downstream port isolation function corresponding to the first device.

[0007] If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the current operating state of the server shall remain unchanged.

[0008] If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the server is controlled to crash and restart.

[0009] According to the server control method provided in the embodiments of this application, by disabling the eDPC function of the CXL device when the server has the eDPC function enabled, and controlling the server's operating state according to the type of uncorrectable error when the CXL device has an uncorrectable error, it is possible to avoid the occurrence of IERR caused by the poor support of the eDPC function for the CXL device, and effectively improve the reliability and stability of the system.

[0010] One embodiment of the server control method of this application, wherein disabling the downstream port isolation function corresponding to the first device includes:

[0011] Set the target control bit of the downstream port isolation control register corresponding to the first device to the target value to disable the downstream port isolation function of the first device.

[0012] One embodiment of the server control method of this application, after detecting an uncorrectable error in the first device, the method further includes:

[0013] Obtain the first status information of the downstream port isolation status register corresponding to the first device;

[0014] Based on the first status information, determine whether the downstream port isolation function corresponding to the first device has been triggered.

[0015] A server control method according to an embodiment of this application, wherein when an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, maintaining the current operating state of the server unchanged, includes:

[0016] If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the first error information corresponding to the uncorrectable error is obtained, the first error information is sent to the client connected to the server, and the current operating state of the server remains unchanged.

[0017] One embodiment of the server control method of this application, wherein when an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the method controls the server to crash and restart, includes:

[0018] If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the second error information corresponding to the uncorrectable error is obtained, the second error information is sent to the client connected to the server, and the server is controlled to crash and restart.

[0019] One embodiment of the server control method of this application, after the downstream port isolation function of the control server is enabled, the method further includes:

[0020] For the second device connected to the server that does not support high-speed interconnection technology, the downstream port isolation function of the corresponding second device shall be kept enabled.

[0021] If an uncorrectable error is detected in the second device, the link of the second device is disabled based on the downstream port isolation function.

[0022] One embodiment of the server control method of this application, after disabling the link of the second device based on the downstream port isolation function when an uncorrectable error is detected in the second device, the method further includes:

[0023] Uninstall the driver module corresponding to the second device.

[0024] One embodiment of the server control method of this application, after disabling the link of the second device based on the downstream port isolation function when an uncorrectable error is detected in the second device, the method further includes:

[0025] Obtain the third error information corresponding to the uncorrectable error, and send the third error information to the client connected to the server.

[0026] One embodiment of the server control method of this application, after disabling the link of the second device based on the downstream port isolation function when an uncorrectable error is detected in the second device, the method further includes:

[0027] Obtain the status information corresponding to the link status register of the second device;

[0028] Based on the status information corresponding to the link status register, it is determined whether the link of the second device is disabled.

[0029] In a second aspect, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the server control method described in the first aspect above.

[0030] Thirdly, this application provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the server control method as described in the first aspect above.

[0031] The above-described one or more technical solutions in the embodiments of this application have at least one of the following technical effects:

[0032] By disabling the eDPC function on the CXL device when the server has eDPC enabled, and controlling the server's operating status according to the type of uncorrectable error when the CXL device encounters an uncorrectable error, it is possible to avoid IERR caused by poor support of the eDPC function for the CXL device, thereby effectively improving the reliability and stability of the system.

[0033] Furthermore, by accurately identifying CXL devices through manufacturer identification codes and category identifiers, the operation of disabling the eDPC function can be performed only on CXL devices, avoiding accidental operation on non-CXL devices and improving control accuracy.

[0034] Furthermore, if an uncorrectable error of a non-downtime type is detected in the first device, the error information is reported to the client, and the current operating state of the server remains unchanged. This avoids unnecessary downtime caused by non-critical errors, maintains the operation of critical services, and improves system availability.

[0035] Furthermore, if an uncorrectable error of the type of device failure is detected, the control server can crash and restart, which can immediately cut off the error propagation path, prevent erroneous data from spreading to other devices or hosts, and ensure data integrity.

[0036] Furthermore, by maintaining the eDPC function of non-CXL devices, the link of the non-CXL device can be disabled based on the eDPC function when an uncorrectable error is detected in the non-CXL device. This can prevent the error from spreading and ensure system stability and data integrity.

[0037] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0038] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0039] Figure 1 This is one of the flowcharts illustrating the server control method provided in the embodiments of this application;

[0040] Figure 2 This is a second schematic flowchart of the server control method provided in the embodiments of this application;

[0041] Figure 3This is the third flowchart illustrating the server control method provided in the embodiments of this application;

[0042] Figure 4 This is a schematic diagram of the structure of the computer program product provided in the embodiments of this application;

[0043] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0044] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0045] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0046] The following description, in conjunction with the accompanying drawings, details the server control method, server control device, electronic device, and readable storage medium provided in this application through specific embodiments and application scenarios.

[0047] The server control method can be applied to the terminal, and can be executed by the hardware or software in the terminal.

[0048] The terminal includes, but is not limited to, portable communication devices such as mobile phones or tablets with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads). It should also be understood that, in some embodiments, the terminal may not be a portable communication device, but rather a desktop computer with touch-sensitive surfaces (e.g., touchscreen displays and / or touchpads).

[0049] The following embodiments describe a terminal including a display and a touch-sensitive surface. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, mouse, and joystick.

[0050] The server control method provided in this application embodiment can be executed by an electronic device or a functional module or entity in an electronic device that can implement the server control method. The electronic devices mentioned in this application embodiment include, but are not limited to, mobile phones, tablets, computers, cameras, and wearable devices. The server control method provided in this application embodiment will be described below using an electronic device as the execution subject.

[0051] like Figure 1 As shown, the control method of the server includes steps 110, 120, 130 and 140.

[0052] Step 110: Enable the downstream port isolation function of the control server;

[0053] In this step, the downstream port isolation function is a mechanism in PCIe (peripheral component interconnect express, a high-speed serial bus standard) that can be used to detect errors in data transmission and to "poison" corrupted data packets to ensure that the system can safely handle errors and avoid data corruption or system crashes.

[0054] The server-level global downstream port isolation function can be enabled by default, meaning that all ports that support eDPC will have this function enabled by default.

[0055] You can fix the default values ​​of UEFI BIOS (Unified Extensible Firmware Interface Basic InputOutput System, i.e., modern BIOS) options and change the default value of the IIOeDPC Support option (where IIO refers to Intel Input / Output subsystem) to be enabled by default, that is, enable the server's global eDPC function.

[0056] Enabling the IIO eDPC Support option configures the IIO ports of all PCIe devices.

[0057] When the IIO eDPC Support option is enabled, all IIO ports will be configured and iterated during the machine reboot process.

[0058] Step 120: For the first device that supports high-speed interconnection technology connected to the server, disable the downstream port isolation function corresponding to the first device;

[0059] In this step, Compute Express Link (CXL) technology can be used to improve the communication efficiency between the CPU (Central Processing Unit) and accelerators, memory expansion devices, and other computing resources. CXL can provide low-latency, high-bandwidth, and memory-coherent connectivity, supporting collaborative work between the CPU and GPUs (Graphics Processing Units), FPGAs (Field Programmable Gate Arrays), and memory expansion devices. CXL technology has lower latency and higher throughput than traditional PCIe communication. CXL technology supports multiple device types and topologies and can coexist with existing PCIe devices.

[0060] The server can connect to various types of devices. For example, it can connect to first-class devices that support high-speed interconnect technology (i.e., CXL devices), as well as PCIe devices that do not support high-speed interconnect technology (non-CXL devices). For instance, CXL devices can be connected through a dedicated CXL slot, while non-CXL devices can be connected through a regular PCIe slot.

[0061] In some embodiments, prior to step 120, the method may further include:

[0062] The first device is identified from the devices connected to the server based on the manufacturer identification code and category identifier of the device connected to the server.

[0063] In this embodiment, the vendor identifier (VENDOR_ID) is the identifier of the vendor corresponding to the device.

[0064] For example, the manufacturer identification code of the first device can be a fixed value of 0x1E98.

[0065] The category identifier can be the Type field in the PCIe device configuration space, used to identify the type of device.

[0066] For example, the Type type corresponding to the CXL 2.0 version device is Type2.

[0067] The software program determines whether the current device is the first device. It can identify the device manufacturer through the VENDOR_ID field of the PCIe configuration space, and then identify the device type through the Type field of the PCIe configuration space. When both the VENDOR_ID and Type fields meet the characteristics of a CXL device, the device is determined to be the first device (CXL device).

[0068] According to the server control method provided in the embodiments of this application, CXL devices are accurately identified by manufacturer identification code and category identifier, so that the operation of disabling eDPC function is performed only on CXL devices, which can avoid misoperation of non-CXL devices and improve the accuracy of control.

[0069] For the first device that supports high-speed interconnect technology, the corresponding downstream port isolation function can be disabled.

[0070] For example, the downstream port isolation function of the first device can be disabled through methods such as hardware registers, software configuration, firmware device or driver customization.

[0071] For example, the eDPC function can be forcibly disabled by directly modifying the hardware registers of the CXL device; or the eDPC state can be indirectly modified through the software interface or debugging tools provided by the operating system; or the eDPC function of the CXL device can be disabled in any other feasible way. The choice can be made based on user needs, and this application does not limit it.

[0072] Step 130: If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, maintain the current operating state of the server unchanged;

[0073] In this step, an Uncorrectable Error (UCE) is a serious error detected by the hardware in a computer system but which cannot be corrected by an automatic error correction mechanism. It is usually detected by hardware (such as CPU, memory, CXL devices, and PCIe devices) during data transmission or processing.

[0074] Uncorrectable errors can include memory errors, PCIe / CXL device errors, and configuration space errors.

[0075] Among them, PCIe / CXL device errors are caused by the data link layer detecting unrecoverable transmission errors (such as CRC check failures and poisoned TLPs).

[0076] Uncorrectable errors can be categorized into downtime and non-downtime types.

[0077] Uncorrectable errors of the downtime type will directly cause the system or device to malfunction. For example, the device may trigger a link reset, causing the host to access to time out; if the timeout is not recovered, the watchdog timer may trigger a system reset, etc.

[0078] Non-downtime type uncorrectable errors will not immediately cause system crashes, but may affect device performance or data integrity. For example, a CXL device may detect a poison code in the memory consistency protocol, but may not trigger a link termination; device register access may return an error code, and a single device may still operate with some functions.

[0079] If an uncorrectable error of a non-downtime type is detected in the first device, the server's current operating state can be maintained, and the machine will not crash.

[0080] Step 140: If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, control the server to crash and restart.

[0081] In this step, if an uncorrectable error of the type of first device failure is detected, the server can be proactively restarted to ensure rapid isolation of the fault in the event of a serious error, protect data integrity, and restore service through a restart.

[0082] During the research and development process, the inventors discovered that in the relevant technologies, the current IIO eDPC Support option is configured for all IIO ports. When the eDPC setting option is enabled and the value is set to On Fatal and Non-Fatal Errors (or On FatalError), eDPC is enabled and triggered when the downstream port detects an unmasked, uncorrectable error or when the downstream port receives an ERR_NONFATAL or ERR_FATAL message (or ERR_FATAL message). This setting applies to all IIO ports.

[0083] After enabling the IIO eDPC Support option, CXL devices will directly crash and report an IERR error when a UCE error is triggered. Analysis of the crash location points to the CXL memory's base address. This IERR crash is caused by the CXL device, possibly due to a UCE error triggering the eDPC function, which disables the CXL device. When the CXL device is disabled, its memory can no longer be accessed. At this time, the host is still sending access requests to the CXL device's memory. Since there is no signal notifying the host that the memory on the CXL device is inaccessible, the host's access requests continue to receive no data return. The host will consider the access to time out (TOR Timeout), thus causing the IERR error. The eDPC function does not support CXL devices well.

[0084] In this application, when the global eDPC function is enabled on the server, the eDPC function of the CXL device is disabled. When an uncorrectable error occurs on the CXL device, if the error is of the crash type, the server crashes and restarts; if the error is not of the crash type, the server maintains its current operating state. This can effectively avoid the occurrence of IERR caused by the poor support of the eDPC function for the CXL device.

[0085] According to the server control method provided in the embodiments of this application, by disabling the eDPC function of the CXL device when the server has the eDPC function enabled, and controlling the server's operating state according to the type of uncorrectable error when the CXL device has an uncorrectable error, it is possible to avoid the occurrence of IERR caused by the poor support of the eDPC function for the CXL device, and effectively improve the reliability and stability of the system.

[0086] like Figure 2 As shown, in some embodiments, step 120 may include:

[0087] Set the target control bit of the downstream port isolation control register corresponding to the first device to the target value to disable the downstream port isolation function of the first device.

[0088] In this embodiment, the downstream port isolation control register can be the dpcctl register (Downstream Port Containment Control Register), which is used to configure the isolation function.

[0089] A target value can be forcibly written to the target control bit of the downstream port isolation control register to disable the downstream port isolation function of the first device.

[0090] The target control bit can be bit [1-0], and the target value can be 0.

[0091] That is, when the current device is a CXL device, the eDPC function can be disabled by software program, and bits [1-0] of the dpcctl register can be forcibly written to 0.

[0092] If an unmasked, uncorrectable error is detected in the first device, the eDPC function will not be activated. This avoids the situation where, after an uncorrectable error occurs, the host's requests to access memory on the CXL device continuously fail to receive data, resulting in access timeouts and thus causing IERR errors.

[0093] In some embodiments, after detecting an uncorrectable error in the first device, the method may further include:

[0094] Obtain the first status information of the downstream port isolation status register corresponding to the first device;

[0095] Based on the first state information, determine whether the downstream port isolation function corresponding to the first device has been triggered.

[0096] In this embodiment, the downstream port isolation status register can be the dpcsts register (Downstream Port Containment Status Register), used to monitor the isolation status.

[0097] Whether the eDPC function has been triggered can be confirmed by checking bit 0 of the downstream port isolation status register.

[0098] For example, if bit 0 of the downstream port isolation status register is detected to be 0, it can be confirmed that the eDPC function has not been triggered.

[0099] Even if the eDPC function of the CXL device is disabled and an uncorrectable error is detected in the CXL device, the PCI Express below the downlink port is still in communication state, and there is a potential for error data propagation. The corresponding operation can be performed based on the error type of the uncorrectable error to suppress the error propagation.

[0100] In some embodiments, step 130 may include:

[0101] If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the first error information corresponding to the uncorrectable error is obtained, the first error information is sent to the client connected to the server, and the current operating state of the server remains unchanged.

[0102] In this embodiment, such as Figure 3 As shown, in the event of an uncorrectable error of non-crash type in the first device (CXL device), the UEFI BIOS can collect the first error information and report it to the client through the BMC (Baseboard Management Controller) or the operating system layer, or it can send it directly to the client.

[0103] Among them, UEFI BIOS can act as the initiator of error handling, responsible for collecting information and triggering the reporting / crash process.

[0104] As an independent management controller, the BMC can receive error messages reported by the BIOS and forward them to the client via the network.

[0105] The client can be used as a monitoring tool for operations and maintenance personnel to receive and display error information.

[0106] The operating system has already started when the error occurs, and the operating system can act as a secondary transmitter of error information (through the system log).

[0107] The first error information may include error type (such as non-downtime type), device identifier, timestamp, error code and error location. The error location can be obtained through BDF (Bus (bus number), Device (device number), Function (function number)) information to quickly locate the faulty device based on the BDF address.

[0108] According to the server control method provided in the embodiments of this application, when an uncorrectable error of non-downtime type is detected in the first device, the error information is reported to the client, and the current operating state of the server remains unchanged, avoiding unnecessary downtime due to non-serious errors, maintaining the operation of critical services, and improving system availability.

[0109] In some embodiments, step 140 may include:

[0110] If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the second error information corresponding to the uncorrectable error is obtained, the second error information is sent to the client connected to the server, and the server is controlled to crash and restart.

[0111] In this embodiment, the second error information may include information such as the error type (e.g., crash type) and the error location.

[0112] like Figure 3 As shown, UEFI BIOS can collect second error information, report the second error information to the client, and control the server to crash and restart.

[0113] According to the server control method provided in the embodiments of this application, when an uncorrectable error of the first device type of crash is detected, the server is controlled to crash and restart, which can immediately cut off the error propagation path, prevent the erroneous data from spreading to other devices or the host, and ensure data integrity.

[0114] In some embodiments, after step 110, the method may further include:

[0115] For the second device connected to the server that does not support high-speed interconnection technology, keep the downstream port isolation function of the second device enabled.

[0116] If an uncorrectable error is detected in the second device, the link to the second device is disabled based on the downstream port isolation function.

[0117] In this embodiment, the second device that does not support high-speed interconnect technology can be a non-CXL device, such as an NVME device (Non-Volatile Memory Express, a non-volatile storage device accessed via a PCIe interface).

[0118] For non-CXL devices, the corresponding eDPC function can be kept enabled.

[0119] You can determine whether the eDPC function of the second device is enabled by checking the status information of the dpcctl register corresponding to the second device.

[0120] For example, when bit[1-0] of the dpcctl register is 0, it can be determined that the eDPC function of the second device is not enabled; when bit[1-0] is 1 or 2, it can be determined that the eDPC function of the second device is enabled.

[0121] If an uncorrectable error is detected in the second device, the eDPC function will be enabled and triggered. The status information of the dpcsts register can be checked to determine whether the eDPC function has been triggered. If bit 0 of the register is 1, the eDPC function has been triggered. If bit 0 of the register is 0, the eDPC function has not been triggered. If the eDPC function has been triggered, the link of the second device can be disabled according to the eDPC function.

[0122] After an uncorrectable PCIe error occurs, eDPC can automatically disable the link below the downstream port, thereby preventing the erroneous TLP (Transaction Layer Packet) from propagating upstream or downstream. During downstream port control, the LTSSM (Link Training and Status State Machine) associated with the downstream port is guided to a disabled state. When the DPC Trigger Status in the DPC status register is set, the state machine LTSSM will remain disabled, and the transaction layer, i.e., the data link layer, will no longer accept upstream TLPs. If the condition that triggers DPC is related to an upstream TLP, all subsequent upstream TLPs already received from the data link layer can be discarded.

[0123] In some embodiments, after disabling the link of the second device based on the downstream port isolation function, the method may further include:

[0124] Obtain the status information corresponding to the link status register;

[0125] Based on the status information corresponding to the link status register, determine whether the link of the second device is disabled.

[0126] In this embodiment, the link status register can be the linksts register (Link Status Register), used to monitor the physical link connection status.

[0127] The link speed of the second device can be confirmed by checking the linksts register. If the link speed is 0x1, it can be determined that the link of the second device is disabled.

[0128] According to the server control method provided in the embodiments of this application, by maintaining the eDPC function of non-CXL devices, the link of non-CXL devices can be disabled based on the eDPC function when an uncorrectable error is detected in a non-CXL device. This can prevent error propagation and ensure system stability and data integrity.

[0129] In some embodiments, after disabling the link of the second device based on the downstream port isolation function in the event of an uncorrectable error detected in the second device, the method may further include:

[0130] Uninstall the driver module corresponding to the second device.

[0131] In this embodiment, if an uncorrectable error occurs in the second device, the error information will be written to the kernel ring buffer and can be viewed using the dmesg command.

[0132] dmesg can resolve Hardware Errors (indicating a hardware error). After the operating system detects the error, it can trigger the eDPC recovery process to unload the driver module of the second device. This is to prevent the CPU from sending access requests to the second device but continuously failing to receive data, resulting in access timeouts and thus causing a failure.

[0133] In some embodiments, after disabling the link of the second device based on the downstream port isolation function in the event of an uncorrectable error detected in the second device, the method may further include:

[0134] Obtain the third error information corresponding to the uncorrectable error and send the third error information to the client connected to the server.

[0135] In this embodiment, such as Figure 3 As shown, the UEFI BIOS can collect third-party error information and report it to the BMC. The BMC can then send the third-party error information to the client for the user to view, or it can send it directly to the client.

[0136] In some embodiments, prior to step 110, a first device (CXL device) and a second device (other PCIe devices other than CXL) may be installed on the server, and DDR5 memory may be installed on the first device.

[0137] The computer program product provided in this application is described below. The computer program product described below can be referred to in correspondence with the server control method described above.

[0138] The server control method provided in this application can be executed by a computer program product. This application uses the execution of the server control method by a computer program product as an example to illustrate the computer program product provided in this application.

[0139] This application also provides a computer program product.

[0140] like Figure 4 As shown, the computer program product includes: a first processing module 410, a second processing module 420, a third processing module 430, and a fourth processing module 440.

[0141] The first processing module 410 is used to control the activation of the downstream port isolation function of the server;

[0142] The second processing module 420 is used to disable the downstream port isolation function of the first device corresponding to the server connection supporting high-speed interconnection technology.

[0143] The third processing module 430 is used to maintain the current operating state of the server when an uncorrectable error is detected in the first device and the uncorrectable error is a non-downtime type error.

[0144] The fourth processing module 440 is used to control the server to crash and restart when an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error.

[0145] The computer program product provided in the embodiments of this application disables the eDPC function of the CXL device when the server has the eDPC function enabled, and controls the server's operating state according to the type of uncorrectable error when the CXL device has an uncorrectable error. This can avoid the occurrence of IERR caused by poor support of the eDPC function for the CXL device, and effectively improve the reliability and stability of the system.

[0146] In some embodiments, the second processing module 420 may also be used for:

[0147] Set the target control bit of the downstream port isolation control register corresponding to the first device to the target value to disable the downstream port isolation function of the first device.

[0148] In some embodiments, the computer program product may further include a fifth processing module, configured to obtain first status information of the downstream port isolation status register corresponding to the first device after detecting an uncorrectable error in the first device;

[0149] Based on the first state information, determine whether the downstream port isolation function corresponding to the first device has been triggered.

[0150] In some embodiments, the third processing module 430 can also be used for:

[0151] If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the first error information corresponding to the uncorrectable error is obtained, the first error information is sent to the client connected to the server, and the current operating state of the server remains unchanged.

[0152] In some embodiments, the fourth processing module 440 can also be used for:

[0153] If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the second error information corresponding to the uncorrectable error is obtained, the second error information is sent to the client connected to the server, and the server is controlled to crash and restart.

[0154] In some embodiments, the computer program product may further include a sixth processing module, which is used to keep the downstream port isolation function of the second device connected to the server enabled after the downstream port isolation function of the control server is enabled.

[0155] If an uncorrectable error is detected in the second device, the link to the second device is disabled based on the downstream port isolation function.

[0156] In some embodiments, the computer program product may further include a seventh processing module for unloading the driver module corresponding to the second device after disabling the link of the second device based on the downstream port isolation function in the event that an uncorrectable error is detected in the second device.

[0157] In some embodiments, the computer program product may further include an eighth processing module, which, after disabling the link of the second device based on the downstream port isolation function in the event that an uncorrectable error is detected in the second device, obtains third error information corresponding to the uncorrectable error and sends the third error information to a client connected to the server.

[0158] In some embodiments, the computer program product may further include a ninth processing module, configured to further include, after disabling the link of the second device based on the downstream port isolation function in the event that an uncorrectable error has been detected in the second device:

[0159] Obtain the status information corresponding to the link status register of the second device;

[0160] Based on the status information corresponding to the link status register, determine whether the link of the second device is disabled.

[0161] The computer program product in this application embodiment can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope.

[0162] The computer program product in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0163] The computer program product provided in this application embodiment can achieve... Figures 1 to 3 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0164] In some embodiments, such as Figure 5As shown, this application embodiment also provides an electronic device 500, including a processor 501, a memory 502, and a computer program stored in the memory 502 and executable on the processor 501. When the program is executed by the processor 501, it implements the various processes of the above-described server control method embodiment and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0165] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0166] On the other hand, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the various processes of the above-described server control method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0167] In another aspect, this application embodiment provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor, and the processor is used to run programs or instructions to implement the various processes of the above-described server control method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0168] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0169] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0170] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0171] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for controlling a server, characterized in that, include: Enable downstream port isolation for the control server; For the first device that supports high-speed interconnect technology connected to the server, disable the downstream port isolation function corresponding to the first device. If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the current operating state of the server shall remain unchanged. If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the server is controlled to crash and restart.

2. The server control method according to claim 1, characterized in that, Disabling the downstream port isolation function corresponding to the first device includes: Set the target control bit of the downstream port isolation control register corresponding to the first device to the target value to disable the downstream port isolation function of the first device.

3. The server control method according to claim 1, characterized in that, After detecting an uncorrectable error in the first device, the method further includes: Obtain the first status information of the downstream port isolation status register corresponding to the first device; Based on the first status information, determine whether the downstream port isolation function corresponding to the first device has been triggered.

4. The server control method according to any one of claims 1-3, characterized in that, The step of maintaining the current operating state of the server unchanged when an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, includes: If an uncorrectable error is detected in the first device, and the uncorrectable error is a non-downtime type error, the first error information corresponding to the uncorrectable error is obtained, the first error information is sent to the client connected to the server, and the current operating state of the server remains unchanged.

5. The server control method according to any one of claims 1-3, characterized in that, The step of controlling the server to crash and restart when an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, includes: If an uncorrectable error is detected in the first device, and the uncorrectable error is a crash type error, the second error information corresponding to the uncorrectable error is obtained, the second error information is sent to the client connected to the server, and the server is controlled to crash and restart.

6. The server control method according to any one of claims 1-3, characterized in that, After the downstream port isolation function of the control server is enabled, the method further includes: For the second device connected to the server that does not support high-speed interconnection technology, the downstream port isolation function of the corresponding second device shall be kept enabled. If an uncorrectable error is detected in the second device, the link of the second device is disabled based on the downstream port isolation function.

7. The server control method according to claim 6, characterized in that, After disabling the link of the second device based on the downstream port isolation function in the event of an uncorrectable error detected in the second device, the method further includes: Uninstall the driver module corresponding to the second device.

8. The server control method according to claim 6, characterized in that, After disabling the link of the second device based on the downstream port isolation function in the event of an uncorrectable error detected in the second device, the method further includes: Obtain the third error information corresponding to the uncorrectable error, and send the third error information to the client connected to the server.

9. The server control method according to claim 6, characterized in that, After disabling the link of the second device based on the downstream port isolation function in the event of an uncorrectable error detected in the second device, the method further includes: Obtain the status information corresponding to the link status register of the second device; Based on the status information corresponding to the link status register, it is determined whether the link of the second device is disabled.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the server control method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Data processing unit integration

    CN118369648A

  • Correctable error tracking and link recovery

    US20230281080A1