PCIE expansion chassis, server, data transmission control method and apparatus, and product
By introducing multi-channel redundant design and switching units into the PCIE link, the data transmission channel is automatically switched, and the service interruption caused by PCIE link failure is solved, and business stability and reliability are maintained without restarting the system.
Patent Information
- Application Number
- PCT/CN2024/109726
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-17
- Filing Date
- 2024-08-05
- Publication Date
- 2025-07-24
AI Technical Summary
When an existing PCIE link occurs with UCE or CE errors, conventional repair methods require restarting the system or replacing the equipment, resulting in business interruption and troubleshooting takes a long time, affecting business stability and reliability.
The PCIE expansion box adopts a multi-channel redundancy design, which automatically switches to the redundant channel during data transmission through the switching unit, ensuring that the business is maintained stable without shutting down.
It improves the stability and reliability of PCIE links, reduces business interruptions caused by link failures, simplifies the fault repair process, and improves the operation and maintenance flexibility of the system.
Smart Images

Figure CN2024109726_24072025_PF_FP_ABST
Abstract
Description
PCIE expansion box, server, method, device and product for controlling data transmission
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to a Chinese patent application filed with the Patent Office of China on January 17, 2024, with application number 202410064974.X and entitled “PCIE expansion box, server, method, device and product for controlling data transmission”, the entire contents of which are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of PCIE technology, and in particular to a PCIE expansion box, a server, a method for controlling data transmission, a device, and a product. Background Art
[0004] PCIE (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard that utilizes high-speed serial communication technology to significantly improve data transmission speeds. In practical applications, depending on bandwidth requirements, commonly used PCIE slots include X2, X4, X8, and X16. Different slot speeds support a variety of PCIE devices. As PCIE speeds continue to increase, the variety of supported PCIE devices is also expanding. This has led to an increasing number of PCIE link issues, and the scope and severity of these issues are also expanding. While minor error reports, such as CE (Correctable Error), may have minimal impact on services, more serious issues such as UCE (Uncorrectable Error), speed reduction, bandwidth reduction, and even card dropout can severely disrupt normal service operations. For example, fatal uncorrectable errors often cause PCIE link and hardware failures, requiring a link and hardware reset, resulting in service interruption.
[0005] During routine maintenance of PCIE devices and links, if the number of UCE or CE errors in a PCIE link reaches a certain threshold, the fault is usually repaired by locating the error-reporting device and removing it. When the device supports hot-plugging, the device can be replaced by pressing the hot-plug (hot-plugging, hot plugging) button. However, this method can only solve faults caused by the device. When a PCIE link fails, replacing or initializing the device cannot directly solve the problem. It is necessary to restart the system to try to solve the link problem. However, when a problem occurs in the physical link, even restarting the system in time cannot restore the system. In this case, troubleshooting takes a lot of time, seriously affecting the normal operation of the business. Therefore, how to improve the stability and reliability of the PCIE link is a problem that needs to be solved.
[0006] Summary of the Invention
[0007] In view of this, the present application aims to propose a high-speed serial computer expansion bus expansion box, server, method, device and product for controlling data transmission to improve the stability and reliability of the high-speed serial computer expansion bus link.
[0008] To achieve the above objectives, the technical solutions of this application are as follows:
[0009] In a first aspect, some embodiments of the present application provide a high-speed serial computer expansion bus expansion box, the high-speed serial computer expansion bus expansion box comprising: a switching unit and at least one high-speed serial computer expansion bus slot, the switching unit comprising a downstream interface and at least two upstream interfaces; wherein:
[0010] The high-speed serial computer expansion bus slot is used to install a high-speed serial computer expansion bus device;
[0011] The downstream interface is used to connect to the high-speed serial computer expansion bus slot;
[0012] At least two uplink interfaces are used to connect to the central processing unit on the motherboard side, thereby forming a multi-channel data link;
[0013] During the data transmission process, the switching unit selects any channel in the multi-channel data link for data transmission.
[0014] Optionally, the bandwidth of the downstream interface in the switching unit is the sum of the bandwidths of all high-speed serial computer expansion bus slots;
[0015] The bandwidth of any uplink interface in the switching unit shall not be less than the bandwidth of the downlink interface.
[0016] Optionally, the bandwidths of all uplink interfaces in the switching unit are equal.
[0017] Optionally, the high-speed serial computer expansion bus slot is one X16 slot, or two X8 slots;
[0018] The switching unit includes a downstream interface and two upstream interfaces, wherein the bandwidth of each upstream interface is equal to the bandwidth of the downstream interface; the bandwidth of the downstream interface is the sum of the bandwidths of all high-speed serial computer expansion bus slots.
[0019] According to a second aspect of some embodiments of the present application, a server is provided, the server including:
[0020] At least one high-speed serial computer expansion bus expansion box, the high-speed serial computer expansion bus expansion box being the high-speed serial computer expansion bus expansion box provided by the first aspect of some embodiments of the present application, and the high-speed serial computer expansion bus expansion box being a multi-channel high-speed serial computer expansion bus expansion box;
[0021] The central processing unit is connected to the uplink interface of each high-speed serial computer expansion bus expansion box through multiple high-speed serial computer expansion bus interfaces on the mainboard side, and performs data communication with the high-speed serial computer expansion bus device on the high-speed serial computer expansion bus expansion box; when any channel in the multi-channel data link fails, the high-speed serial computer expansion bus interface of the channel on the mainboard side is disabled.
[0022] Optionally, a basic input / output system (BIOS) runs on the central processing unit; the BIOS is configured to perform the following steps: when the server starts, the BIOS allocates corresponding high-speed serial computer expansion bus resources to at least one channel in each multi-channel data link; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box.
[0023] Optionally, the server further includes:
[0024] The baseboard management controller is used to identify the channel of each high-speed serial computer expansion bus expansion box when the server is started; if the high-speed serial computer expansion bus expansion box has multiple channels, a default channel is set for the high-speed serial computer expansion bus expansion box;
[0025] The switching unit is also used to preferentially select the default channel for data transmission;
[0026] The basic input and output system is also used to allocate corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box.
[0027] Optionally, the baseboard management controller is used to determine whether the high-speed serial computer expansion bus expansion box has multiple channels by identifying the bottom unit group FRU of the high-speed serial computer expansion bus expansion box when the server starts; when the default channel is set, the default channel information is stored in the erasable programmable read-only memory.
[0028] Optionally, the basic input and output system is further configured to perform the following steps when the operating system of the server is started:
[0029] Detecting whether there is an uncorrectable error in the current channel; if there is an uncorrectable error, determining that the current channel has failed; or detecting whether the number of correctable errors occurring in the current channel has reached a first threshold; if the number reaches the first threshold, determining that the current channel has failed;
[0030] When it is determined that the current channel fails, the high-speed serial computer expansion bus interface of the current channel on the mainboard is disabled;
[0031] The switching unit is also used to switch the remaining channels of the data link where the current channel is located for data transmission according to the disable information of the high-speed serial computer expansion bus interface on the mainboard side.
[0032] Optionally, the basic input and output system is further configured to perform the following steps during operation of the high-speed serial computer expansion bus device:
[0033] Continuously reading registers of the central processing unit, and if an uncorrectable error occurs in the register, determining that the current channel has failed; or continuously reading registers of the central processing unit, and if the number of correctable errors in the register reaches a second threshold, determining that the current channel has failed;
[0034] When it is determined that the current channel fails, the high-speed serial computer expansion bus interface of the current channel on the mainboard is disabled;
[0035] The switching unit is also used to switch the remaining channels of the data link where the current channel is located for data transmission according to the disable information of the high-speed serial computer expansion bus interface on the mainboard side.
[0036] Optionally, the basic input and output system is further configured to generate fault information based on all currently faulty channels, and send the fault information and information about switching channels to the operating system and baseboard management controller of the server;
[0037] The server's operating system records a fault log based on the received fault information and the information about the switching channel;
[0038] The baseboard management controller records a fault log based on the received fault information and the information of the switching channel.
[0039] Optionally, the baseboard management controller issues an alarm for a link failure based on the received fault information.
[0040] Optionally, the baseboard management controller is further configured to determine that a device failure has occurred and issue an alarm for the device failure when all channels of the data link have failed in the received fault information.
[0041] According to a third aspect of some embodiments of the present application, a method for controlling data transmission is provided, which is applied to the server provided by the second aspect of some embodiments of the present application. The method includes:
[0042] When the server starts, the channel of each high-speed serial computer expansion bus expansion box is identified; if the high-speed serial computer expansion bus expansion box has multiple channels, a default channel is set for the high-speed serial computer expansion bus expansion box;
[0043] Allocating corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box and initializing the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box;
[0044] Check whether the current channel is faulty;
[0045] If the current channel fails, the high-speed serial computer expansion bus interface of the current channel on the mainboard side is disabled, and the remaining channels of the data link to which the current channel belongs are switched for data transmission.
[0046] Optionally, detect whether the current channel has a fault, including:
[0047] When the server's operating system starts, it detects whether there are any uncorrectable errors in the current channel. If there are any uncorrectable errors, it determines that the current channel has failed.
[0048] Alternatively, when the operating system of the server is started, it is detected whether the number of correctable errors occurring in the current channel reaches a first threshold; if it reaches the first threshold, it is determined that the current channel has a fault.
[0049] Optionally, detect whether the current channel has a fault, including:
[0050] During the operation of a high-speed serial computer expansion bus device, the CPU register is continuously read. If an uncorrectable error occurs in the register, the current channel is determined to be faulty.
[0051] Alternatively, during the operation of the high-speed serial computer expansion bus device, the register of the central processing unit is continuously read, and if the number of correctable errors in the register reaches a second threshold, it is determined that the current channel has failed.
[0052] Optionally, the method for controlling data transmission further includes:
[0053] Generate fault information based on all currently faulty channels, and record fault logs based on the fault information and the information about switching channels;
[0054] Based on the fault information, an alarm is issued for link failure.
[0055] Optionally, the method for controlling data transmission further includes:
[0056] If all channels of the data link fail in the fault information, it is determined that a device failure has occurred;
[0057] Based on the fault information, an alarm is issued for equipment failure.
[0058] According to a fourth aspect of some embodiments of the present application, a device for controlling data transmission is provided, for implementing the method for controlling data transmission provided in the third aspect of some embodiments of the present application, the device including:
[0059] The management module is configured to, when the server is started, identify a channel for each high-speed serial computer expansion bus expansion box; if the high-speed serial computer expansion bus expansion box has multiple channels, set a default channel for the high-speed serial computer expansion bus expansion box; allocate corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box, and initialize the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box;
[0060] The control module is configured to detect whether a current channel fails; if the current channel fails, the high-speed serial computer expansion bus interface of the current channel on the mainboard side is disabled, and the remaining channels of the data link to which the current channel belongs are switched for data transmission.
[0061] According to a fifth aspect of some embodiments of the present application, a non-volatile readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for controlling data transmission as in the fourth aspect of some embodiments of the present application are implemented.
[0062] According to the sixth aspect of some embodiments of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps in the method for controlling data transmission as in the fourth aspect of some embodiments of the present application are implemented.
[0063] By using the high-speed serial computer expansion bus (PCIE) expansion box provided in the present application, a downlink interface and multiple uplink interfaces are integrated through a switching unit to realize a data transmission architecture that upgrades a single-channel data link to a multi-channel data link. By adopting a multi-channel redundant design, the data transmission channel can be switched through the switching unit when a PCIE link reports an error, and resources can be automatically switched through the channel for transmission, thereby stabilizing the normal operation of the business without shutting down. Compared with the traditional single transmission link, the stability and reliability of the PCIE link are greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions of some embodiments of the present application, the following briefly introduces the drawings required for use in the description of some embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0065] FIG1 is a schematic diagram of a high-speed serial computer expansion bus (PCIE) expansion box according to some embodiments of the present application;
[0066] FIG2 is a schematic diagram of the structure of a PCIE link in a server proposed in some embodiments of the present application;
[0067] FIG3 is a schematic diagram of the structure of a PCIE link in a server proposed in some embodiments of the present application;
[0068] FIG4 is a flowchart of a method for controlling data transmission proposed in some embodiments of the present application;
[0069] FIG5 is a schematic diagram of the architecture of a switching unit proposed in some embodiments of the present application;
[0070] FIG6 is a flowchart of a dual-channel PCIE link in some embodiments of the present application;
[0071] FIG7 is a schematic diagram of an apparatus for controlling data transmission proposed in some embodiments of the present application. DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in some embodiments of the present application to clearly and completely describe the technical solutions in some embodiments of the present application. Obviously, some of the embodiments described are part of the embodiments of the present application, not all of the embodiments. Based on some embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of this application.
[0073] It should be understood that references throughout this specification to "one embodiment" or "an embodiment" mean that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" throughout this specification do not necessarily refer to the same embodiment. Furthermore, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0074] In some embodiments of the present application, it should be understood that the size of the serial numbers of the following processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of some embodiments of the present application.
[0075] Some exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments are not intended to represent all implementations consistent with the present application. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0076] It should be noted that, in the absence of conflict, some embodiments and features in the embodiments of the present application may be combined with each other.
[0077] The PCIE links in the server come from different CPU (Central Processing Unit) ports. Due to the different CPU positions and the types and locations of the PCIE interfaces on the motherboard, the PCIE Margin (tolerance) value on each PCIE link is also different. Therefore, when using different PCIE expansion boxes or PCIE devices, there is a chance of compatibility issues.
[0078] Currently, PCIE errors are divided into two categories: uncorrectable errors and correctable errors. Correctable errors can be corrected by the hardware through its own logic without software intervention, and the correction will not cause any information loss. The frequency and number of errors can be recorded by software.
[0079] Uncorrectable errors will affect the functionality of the interface. There is no clear mechanism in the protocol to correct such errors. Such errors usually cause the device to slow down, reduce bandwidth, and even worse, cause the PCIE device to be lost. To repair such errors, the entire link and device need to be reset. Uncorrectable errors can be further divided into fatal and non-fatal. Non-fatal uncorrectable errors may cause specific transmissions to become unreliable, but other functions of the link and hardware are not affected. They are usually caused by unreliable transaction layers but the link meets the requirements. The device driver software provides a recovery mechanism and does not affect the operation of the link and other devices. Fatal uncorrectable errors are caused by unreliable links or hardware. The hardware devices on the link need to be reset, which will involve power outages on the device and business interruptions.
[0080] In the related art, when there is a UCE error or CE error in the PCIE link that reaches a threshold, the fault is usually repaired by removing the error device and replacing or restarting the device. For PCIE devices that support hot plugging, when a problem occurs with the PCIE device, while keeping the server powered on, the hot-plug button is used to inform the system to remove the PCIE device. The hot-plug service controls the device driver to deactivate the device and close the PCIE link. At the same time, the device is powered off, and the BIOS (Basic Input Output System) cancels the resource configuration for the slot where the device is located. When the PCIE device is reinserted, the administrator can detect the power-on through the in-place signal and inform the controller on the OS (Operating System) side to reopen the slot and allocate PCIE resources. The system loads the driver to re-initialize the device.
[0081] However, this approach only addresses device-related issues. When a PCIE link fails, replacing or initializing the device doesn't directly resolve the problem; a system restart is required to resolve the link issue. However, even a system restart won't resolve physical link issues. Furthermore, replacing a device or adjusting the link requires shutting down or disconnecting the power supply, severely impacting normal business operations. In this scenario, repeated debugging and troubleshooting consumes significant time, significantly impacting business operations.
[0082] To this end, the present application provides a multi-channel PCIE redundant link design, which combines the firmware in the switching unit to automatically switch channels. When a PCIE error occurs on the server, the device resources are automatically switched to the redundant channel for transmission, thereby achieving the purpose of stabilizing business data transmission without shutting down.
[0083] The present application will be described in detail below with reference to the accompanying drawings and in combination with some embodiments of the present application.
[0084] FIG1 is a schematic diagram of a high-speed serial computer expansion bus (PCIE) expansion box according to some embodiments of the present application. As shown in FIG1 , the PCIE expansion box includes: a switching unit and at least one PCIE slot, wherein the switching unit includes a downstream interface and at least two upstream interfaces; wherein:
[0085] PCIE slot is used to install PCIE devices;
[0086] The downstream interface is used to connect to the PCIE slot;
[0087] At least two uplink interfaces are used to connect to the CPU on the mainboard side, thereby forming a multi-channel data link;
[0088] During the data transmission process, the switching unit selects any channel in the multi-channel data link for data transmission.
[0089] In some embodiments of the present application, a multi-channel data link hardware architecture is used to implement redundant backup of the service transmission link to improve the stability and reliability of service operation. In the PCIE bus architecture, the PCIE expansion box (commonly referred to as a riser) is an important part of the link and is used to connect the data communication between the bottom plate and the top plate.
[0090] The PCIE expansion box of the present application includes a switching unit and at least one PCIE slot. The PCIE slot is used to install PCIE devices. The switching unit includes a downstream interface and multiple upstream interfaces. The downstream interface is connected to the PCIE slot to communicate with the PCIE device; the multiple upstream interfaces are connected to the PCIE interface on the mainboard side to enable the PCIE device in the PCIE slot to communicate with the CPU. Through the upstream interface and downstream interface of the switching unit, a multi-channel data link is formed between the PCIE device and the CPU. During the operation of the device, data is transmitted through one channel in the multi-channel data link. If the current data channel fails, the switching unit can automatically switch to any of the remaining redundant channels for data transmission to ensure the stability and reliability of the service.
[0091] As some embodiments of the present application, the bandwidth of the downstream interface in the switching unit is the sum of the bandwidths of all PCIE slots;
[0092] The bandwidth of any uplink interface in the switching unit shall not be less than the bandwidth of the downlink interface.
[0093] In actual applications, the commonly used PCIE slots are mainly X2, X4, X8, X16 and other specifications according to bandwidth requirements. Different slots can support PCIE devices with different bandwidth rates. For example, a PCIE3.0 X8 slot can support a network card of the same rate (such as X710). In order to adapt to various configuration requirements, server motherboards usually provide multiple PCIE interfaces. By matching different PCIE expansion boxes, users can insert various external PCIE cards to support different configurations. In the traditional PCIE expansion box architecture, it includes an upstream interface and a downstream interface. The upstream bandwidth is consistent with the downstream bandwidth, and the bandwidth ratio of the downstream interface to the slot is 1:1 or 1:2. For example, if the upstream bandwidth corresponds to the X16 bandwidth resource, the slot can be 1 X16 slot or 2 X8 slots. In some embodiments of the present application, in a multi-channel data link, the bandwidth ratio of the uplink interface to the downlink interface is N to 1 (N is not less than 2). To ensure that the bandwidth of the uplink interface can support the normal operation of the PCIE device, the bandwidth of all uplink interfaces in the multi-channel data link is set to be no less than the bandwidth of the downlink interface, so as to ensure that the bandwidth requirements of the different specifications of PCIE devices supported by the slot are met. For example, if the bandwidth of the downlink interface supports a PCIE device with a maximum specification of X8, the bandwidth of each uplink interface is no less than X8 to ensure that the PCIE device does not experience bandwidth reduction when switching channels for service data transmission, thereby ensuring stable operation of the device.
[0094] As some implementation methods of the present application, the bandwidths of all uplink interfaces in the switching unit are equal.
[0095] In some embodiments, in order to ensure the stability of PCIE device operation and save CPU bandwidth resources, when allocating bandwidth resources for multi-channel data links, equal bandwidth resources are allocated to each uplink interface so that the bandwidth rate of business data transmission can be kept stable before and after switching channels.
[0096] In some embodiments of the present application, the PCIE slot is one x16 slot, or two x8 slots;
[0097] The switching unit includes a downstream interface and two upstream interfaces, wherein the bandwidth of each upstream interface is equal to the bandwidth of the downstream interface; the bandwidth of the downstream interface is the sum of the bandwidths of all PCIE slots.
[0098] Figure 5 is a schematic diagram of the architecture of a switching unit proposed in some embodiments of the present application. As shown in Figure 5, the switching unit can construct a dual-channel data link architecture, including two upstream interfaces and one downstream interface, with channels connected via Virtual PCI-PCI Bridge technology. Virtual PCI-PCI Bridge is a virtual bridging technology that can implement conversion between PCI Express (PCIE) and PCI buses within a computer. This technology allows different devices within a computer to communicate with each other through shared memory space, thereby achieving more efficient resource utilization and data transmission. In some embodiments of the present application, the upstream and downstream interfaces of the switching unit are connected to the CPU and PCIE devices respectively via the PCIE bus, forming a dual-channel data link. During server initialization, the BIOS loads the corresponding firmware into the switching unit, enabling the switching unit to automatically switch channels in a multi-channel data link. During service data transmission, the switching unit selects one channel for data transmission. If the current transmission channel fails, the switching unit can automatically identify and switch to the redundant channel to continue transmitting service data, thereby ensuring the stability of service operation.
[0099] In some embodiments of the present application, a dual-channel data link is constructed through a switching unit to achieve maximum compatibility with the PCIE resources of X16 devices. By adding a redundant data transmission channel, the stability of business operations is improved and the impact of PCIE link failures on user usage scenarios is reduced.
[0100] Based on the same inventive concept, some embodiments of the present application provide a server, which includes:
[0101] At least one PCIE expansion box, wherein the PCIE expansion box is any one of the PCIE expansion boxes in the above embodiments, and the PCIE expansion box is a multi-channel PCIE expansion box;
[0102] The CPU is connected to the uplink interface of each PCIE expansion box through multiple PCIE interfaces on the mainboard side, and performs data communication with the PCIE devices on the PCIE expansion box; when any channel in the multi-channel data link fails, the PCIE interface of the channel on the mainboard side is disabled.
[0103] In some embodiments of the present application, the server includes a CPU and multiple PCIE expansion boxes with multiple channels, each PCIE expansion box is connected to the PCIE interface on the mainboard side via the uplink interface of its switching unit. Figure 2 is a schematic diagram of the structure of the PCIE link in the server proposed in some embodiments of the present application. As shown in Figure 2, taking an N-channel PCIE expansion box as an example, the uplink interface of the PCIE expansion box (interface 1, interface 2...interface N on the PCIE expansion box in Figure 2) is connected to the PCIE interface on the mainboard (interface 1, interface 2...interface N on the mainboard in Figure 2) via a PCIE cable, forming a multi-channel data link between the PCIE device in the slot and the CPU. In some embodiments of the present application, the uplink interface of the PCIE expansion box can be set to multiple, forming a multi-channel data link between the CPU and the PCIE device. The more channels in the data link that serve as backup redundancy, the higher the stability and reliability of the link. When any channel that transmits data fails, the switching unit automatically switches to a redundant channel to continue providing service data support, thereby avoiding service interruption and causing a poor user experience.
[0104] When switching data transmission channels, the CPU port on the motherboard is disabled, effectively disabling the current data transmission channel. The switching unit then determines and executes the channel switch based on the CPU port's disablement information. For faulty data transmission channels, administrators can troubleshoot and repair them during their free time. This solution also reduces the urgency of troubleshooting PCIE link faults, allowing administrators to schedule troubleshooting more efficiently and making server operations more flexible.
[0105] As some embodiments of the present application, a BIOS is running on the CPU; the BIOS is used to perform the following steps: when the server starts, the BIOS allocates corresponding PCIE resources to at least one channel in each multi-channel data link; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all PCIE slots in the PCIE expansion box.
[0106] In some embodiments of the present application, resources are allocated for a multi-channel data link via the BIOS. When the server boots, the BIOS allocates PCIE resources to at least one channel in the multi-channel data link. PCIE resources primarily include bus number, PCIE bus bandwidth, memory space, I / O space, and the like. The bandwidth allocated to each channel must be able to support the maximum bandwidth compatible with the slot on the PCIE expansion box. For example, if the slot on the PCIE expansion box is a single X16 slot, the bandwidth allocated to any channel in the multi-channel data link of the PCIE expansion box cannot be less than the bandwidth corresponding to the X16 slot to ensure that the slot-interrupted PCIE device can operate normally on any channel. Only after a channel has been allocated PCIE resources can it be selected by the switching unit during device operation. In some embodiments of the present application, the multi-channel data link is compatible with the traditional single-channel data link configuration method. When PCIE resources are configured for only one channel in the data link, the data link transmits service data as a single-channel data link.
[0107] In some embodiments of the present application, when allocating PCIE resources to a multi-channel data link, corresponding PCIE resources can be allocated to a portion of the channels in the data link according to actual needs. While ensuring that a backup channel is provided for business data transmission, the CPU resource allocation is flexibly controlled to adapt to the different reliability level requirements of different businesses, thereby achieving the purpose of balancing CPU resource usage and ensuring business operation reliability.
[0108] As some implementation methods of the present application, the server further includes:
[0109] The baseboard management controller is used to identify the channel of each PCIE expansion box when the server starts; if the PCIE expansion box has multiple channels, a default channel is set for the PCIE expansion box;
[0110] The switching unit is also used to preferentially select the default channel for data transmission;
[0111] BIOS is also used to allocate corresponding PCIE resources to all channels of each PCIE expansion box.
[0112] In some embodiments of the present application, the server further includes a BMC (Baseboard Management Controller). FIG3 is a schematic diagram of the structure of a PCIE link in a server proposed in some embodiments of the present application. As shown in FIG3 , when the server is started, the baseboard management controller (BMC) first identifies the PCIE expansion box through the basic input and output system (BIOS). If it is identified that the PCIE expansion box supports a multi-channel data link, the baseboard management controller (BMC) provides a link channel configuration function to the administrator. The administrator specifies a channel in the multi-channel data link of the PCIE expansion box as the default channel through the BMC. During the operation of the server, the switching unit preferentially selects the default channel for service data transmission.
[0113] In some embodiments of the present application, when it is identified that the PCIE expansion box supports multi-channel data links, corresponding PCIE resources are allocated to each channel through BIOS to fully utilize the redundant channels of the data link and provide stable data transmission guarantees for application scenarios with high business reliability requirements.
[0114] As some embodiments of the present application, the baseboard management controller is used to determine whether the PCIE expansion box has multiple channels by identifying the underlying unit group FRU (Field Replaceable Unit) of the PCIE expansion box when the server starts; when the default channel is set, the default channel information is stored in an EPROM (Erasable Programmable Read-Only Memory).
[0115] In some embodiments of the present application, the BMC identifies the FRU (Field Replaceable Unit) unit of the PCIE expansion box through the BIOS to identify whether the PCIE expansion box supports a multi-channel data link. If the PCIE expansion box has multiple channels, the link channel configuration function is provided to the administrator. After the administrator sets the default channel, the configuration information is stored in the EPROM (Erasable Programmable Read-Only Memory) to prevent the configuration information from being lost.
[0116] As some embodiments of the present application, the BIOS is further configured to perform the following steps when the server's operating system is started:
[0117] Detecting whether there is an uncorrectable error in the current channel; if there is an uncorrectable error, determining that the current channel has failed; or detecting whether the number of correctable errors occurring in the current channel has reached a first threshold; if the number reaches the first threshold, determining that the current channel has failed;
[0118] When the current channel is determined to be faulty, the PCIE interface of the current channel on the motherboard is disabled;
[0119] The switching unit is further used to switch the remaining channels of the data link where the current channel is located for data transmission according to the disabling information of the PCIE interface on the mainboard side.
[0120] In some embodiments of the present application, when the server starts up, fault detection and link switching control are performed on multi-channel data links through BIOS.
[0121] When the server boots up, the BIOS performs a fault check to determine whether the default data transmission channel has a UCE error. If a UCE error occurs, the channel is considered faulty and data transmission must be switched to a redundant channel to ensure normal service operation. The BIOS disables the PCIe interface on the motherboard for the current channel. After receiving the CPU port disable information, the switching unit automatically switches to a redundant channel to transmit service data.
[0122] Optionally, if the BIOS does not detect a UCE, but the number of CEs in the current channel reaches a first threshold, meaning the number of correctable errors in the current data transmission channel reaches the threshold at which a channel switch is required, the reliability of the current channel is low, and the current channel is determined to have failed, requiring a switch to a redundant channel to run the service. The BIOS disables the PCIE interface of the current channel on the motherboard. After receiving the CPU port disable information, the switching unit automatically switches to a redundant channel to transmit service data. The first threshold can be set to 1000. In actual applications, the first threshold can be flexibly set, and this application does not impose any restrictions on this.
[0123] As some implementation methods of the present application, the BIOS is further configured to perform the following steps during operation of the PCIE device:
[0124] Continuously reading the CPU registers, and if an uncorrectable error occurs in the registers, determining that the current channel has failed; or continuously reading the CPU registers, and if the number of correctable errors in the registers reaches a second threshold, determining that the current channel has failed;
[0125] When the current channel is determined to be faulty, the PCIE interface of the current channel on the motherboard is disabled;
[0126] The switching unit is further used to switch the remaining channels of the data link where the current channel is located for data transmission according to the disabling information of the PCIE interface on the mainboard side.
[0127] In some embodiments of the present application, during the operation of the PCIE device, the BIOS performs periodic and continuous fault detection on the multi-channel data links, and controls the switching of redundant links when a fault occurs.
[0128] During device operation, the BIOS continuously reads the CPU registers to determine whether a UCE error is present. If so, it indicates a channel failure. The BIOS disables the PCIE interface on the motherboard for the current channel. Upon receiving the CPU port disable information, the switching unit automatically switches to a redundant channel for data transmission, ensuring normal service operation.
[0129] Alternatively, if the BIOS does not detect a UCE error in the register but reads that the number of CEs reaches a second threshold, meaning the number of correctable errors in the current data transmission channel reaches the threshold requiring a channel switch, the reliability of the current channel is low, and the current channel is determined to be faulty, disabling the CPU port on the motherboard for the current channel. Upon receiving the CPU port disable information, the switching unit automatically switches to a redundant channel to ensure normal service operation. In some embodiments of the present application, the second threshold can be flexibly set based on the actual link situation, and this application does not impose any restrictions on this.
[0130] As some embodiments of the present application, the BIOS is further configured to generate fault information based on all currently faulty channels, and send the fault information and information about switching channels to the operating system and baseboard management controller of the server;
[0131] The server's operating system records a fault log based on the received fault information and the information about the switching channel;
[0132] The baseboard management controller records a fault log based on the received fault information and the information of the switching channel.
[0133] In some embodiments of the present application, the BIOS monitors multi-channel data links for faults. In some embodiments of the present application, when the BIOS detects the presence of a UCE in the current data transmission channel, or the number of CEs reaches a set threshold, it determines that the data link of the current channel has failed. At this point, the current channel's fault information and the switching channel information are sent to the OS and BMC, respectively. The OS and BMC then log the faults, allowing administrators to easily review historical device faults.
[0134] As some implementation methods of the present application, the baseboard management controller issues an alarm for a link failure based on the received fault information.
[0135] In some embodiments of the present application, the BMC generates an alarm prompt for a link failure based on the received channel fault information and the switching channel information, and provides a prompt on the management interface so that the management personnel can know the current operating status of the PCIE data link, quickly and accurately locate the faulty link, and make corresponding fault repair plans.
[0136] As some embodiments of the present application, the baseboard management controller is further configured to determine that a device failure has occurred and issue an alarm for the device failure when all channels of the data link have failed in the received fault information.
[0137] In some embodiments of the present application, if all channels in the data link connected to the PCIE device fail, it can be determined that the failure is most likely caused by the PCIE device. At this time, the BMC determines that a PCIE device failure has occurred, generates an alarm prompt for the device failure, and prompts it on the management interface so that the management personnel can quickly locate the faulty device and replace the device in time, thereby reducing the service interruption duration and avoiding a major impact on user services.
[0138] Based on the same inventive concept, some embodiments of the present application provide a method for controlling data transmission, which is applied to the server in some of the above embodiments. Figure 4 is a flow chart of the method for controlling data transmission proposed in some embodiments of the present application. As shown in Figure 4, the method includes:
[0139] S21: When the server starts, a channel is identified for each high-speed serial computer expansion bus (PCIE) expansion box; if the high-speed serial computer expansion bus expansion box has multiple channels, a default channel is set for the high-speed serial computer expansion bus expansion box;
[0140] S22: allocating corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box, and initializing the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box;
[0141] S23: Check whether the current channel has a fault;
[0142] S24: If the current channel fails, the high-speed serial computer expansion bus interface of the current channel on the mainboard side is disabled, and the remaining channels of the data link to which the current channel belongs are switched for data transmission.
[0143] FIG6 is a flowchart of a dual-channel PCIE link in some embodiments of the present application. As shown in FIG6 , a dual-channel data link is used as an example for illustration. In some embodiments of the present application, the steps of using a dual-channel data link for service data transmission and fault detection and monitoring are as follows:
[0144] (1) First, you need to configure the firmware for the switching unit in the PCIE expansion box so that the switching unit can perform automatic channel switching.
[0145] (2) Use the BMC to identify the PCIE expansion box through its underlying unit group FRU. If the PCIE expansion box is a single-channel expansion box, it is configured in the traditional way. If the BMC identifies the PCIE expansion box as a dual-channel PCIE expansion box, it provides a link channel setting interface, and the administrator can set the default channel through the BMC. For example, set channel 1 as the default channel and channel 2 as the redundant channel. The BMC stores the channel setting information in the EPROM.
[0146] (3) When the server starts, the BIOS allocates resources for the two PCIE channels of the PCIE expansion box. The switch unit prioritizes providing resources to the PCIE device based on the default channel 1 set by the BMC. At this stage, the device completes driver loading and initialization.
[0147] (4) When the OS starts, the BIOS checks whether the default channel 1 is normal. If channel 1 is normal, channel 1 will be selected for data transmission during device operation. If there is a UCE error in the default channel 1 or the number of CEs exceeds the threshold, the BIOS will disable the CPU port of the default channel. At this time, the PCIE resources of the upstream channel 1 will be disconnected due to the disabling of the CPU port. At this time, since the remaining channel 2 in the data link can continue to provide resources, the switching unit on the PCIE expansion box will automatically switch the PCIE resources to channel 2 to ensure that resources continue to be provided to the PCIE device and ensure stable operation of the device. At this time, the BMC will record the fault log of the default channel 1 for the operator to remotely view the device operation status.
[0148] When a PCIE device is operating, the BIOS continuously reads CPU register information. If it detects a UCE in default channel 1 or the number of CEs exceeds the threshold, the BIOS disables the CPU port in channel 1 and the switch automatically switches the PCIE resources to channel 2. Furthermore, the BIOS reports the current channel fault to the BMC, which logs the channel switching action and the channel 1 fault log.
[0149] (5) When the link switches to channel 2, channel 2 is tested. If it is detected that UCE error or CE exceeds the threshold in channel 2, it can be determined that the fault is most likely a PCIE device problem. The BMC records the device fault log and can issue a device fault alarm as needed.
[0150] As some implementation methods of the present application, detecting whether a current channel fails includes:
[0151] When the server's operating system starts, it detects whether there are any uncorrectable errors in the current channel. If there are any uncorrectable errors, it determines that the current channel has failed.
[0152] Alternatively, when the operating system of the server is started, it is detected whether the number of correctable errors occurring in the current channel reaches a first threshold; if it reaches the first threshold, it is determined that the current channel has a fault.
[0153] As some implementation methods of the present application, detecting whether a current channel fails includes:
[0154] During the operation of the PCIE device, the CPU register is continuously read. If an uncorrectable error occurs in the register, the current channel is determined to be faulty.
[0155] Alternatively, during the operation of the PCIE device, the CPU register is continuously read, and if the number of correctable errors in the register reaches a second threshold, it is determined that a current channel fails.
[0156] As some implementation methods of the present application, the method for controlling data transmission further includes:
[0157] Generate fault information based on all currently faulty channels, and record fault logs based on the fault information and the information about switching channels;
[0158] Based on the fault information, an alarm is issued for link failure.
[0159] As some implementation methods of the present application, the method for controlling data transmission further includes:
[0160] If all channels of the data link fail in the fault information, it is determined that a device failure has occurred;
[0161] Based on the fault information, an alarm is issued for equipment failure.
[0162] Based on the same inventive concept, some embodiments of the present application provide a device for controlling data transmission. Referring to FIG. 7 , FIG. 7 is a schematic diagram of a device 300 for controlling data transmission proposed in some embodiments of the present application. As shown in FIG. 7 , the device includes:
[0163] The management module 301 is configured to, when the server is started, identify a channel for each PCIE expansion box; if the PCIE expansion box has multiple channels, set a default channel for the PCIE expansion box; allocate corresponding PCIE resources to all channels of each PCIE expansion box, and initialize the PCIE devices on the PCIE expansion box; wherein the bandwidth of each channel is not less than the sum of the bandwidths of all PCIE slots in the PCIE expansion box;
[0164] The control module 302 is configured to detect whether the current channel fails; if the current channel fails, the PCIE interface of the current channel on the motherboard side is disabled, and the remaining channels of the data link to which the current channel belongs are switched for data transmission.
[0165] As some embodiments of the present application, the control module 302 is configured to perform the following steps:
[0166] When the server's operating system starts, it detects whether there are any uncorrectable errors in the current channel. If there are any uncorrectable errors, it determines that the current channel has failed.
[0167] Alternatively, when the operating system of the server is started, it is detected whether the number of correctable errors occurring in the current channel reaches a first threshold; if it reaches the first threshold, it is determined that the current channel has a fault.
[0168] As some embodiments of the present application, the control module 302 is further configured to perform the following steps:
[0169] During the operation of the PCIE device, the CPU register is continuously read. If an uncorrectable error occurs in the register, the current channel is determined to be faulty.
[0170] Alternatively, during the operation of the PCIE device, the CPU register is continuously read, and if the number of correctable errors in the register reaches a second threshold, it is determined that a current channel fails.
[0171] As some embodiments of the present application, the apparatus 300 for controlling data transmission further includes a monitoring module configured to perform the following steps:
[0172] Generate fault information based on all currently faulty channels, and record fault logs based on the fault information and the information about switching channels;
[0173] Based on the fault information, an alarm is issued for link failure.
[0174] As some embodiments of the present application, the monitoring module is further configured to perform the following steps:
[0175] If all channels of the data link fail in the fault information, it is determined that a device failure has occurred;
[0176] Based on the fault information, an alarm is issued for equipment failure.
[0177] Based on the same inventive concept, some embodiments of the present application provide a non-volatile readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the method for controlling data transmission in some embodiments of the present application.
[0178] Based on the same inventive concept, some embodiments of the present application provide an electronic device, which includes a memory, a processor, and a computer program stored in the memory and runnable on the processor. When executed by the processor, the steps in the method for controlling data transmission of some embodiments of the present application are implemented.
[0179] Regarding the apparatus in some of the above embodiments, the specific manner in which each module performs operations has been described in detail in some embodiments of the method and will not be elaborated on here.
[0180] The above are only some preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
[0181] For the sake of simplicity, some method embodiments are described as a series of actions. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that some embodiments described in this specification are preferred embodiments, and the actions and components involved are not necessarily required by this application.
[0182] Those skilled in the art will appreciate that some embodiments of the present application may be provided as methods, devices, or computer program products. Therefore, some embodiments of the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, some embodiments of the present application may take the form of a computer program product implemented on one or more non-volatile readable storage media (including but not limited to magnetic disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code.
[0183] Some embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to some embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0184] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0185] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0186] Although preferred embodiments of some embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including all changes and modifications that fall within the scope of the preferred embodiments and some embodiments of the present application.
[0187] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0188] The above is a detailed introduction to the PCIE expansion box, server, method, device and product for controlling data transmission provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core idea of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A server, characterized in that, Including: At least one high-speed serial computer expansion bus expansion box, and the high-speed serial computer expansion bus expansion box is a multi-channel high-speed serial computer expansion bus expansion box; The high-speed serial computer expansion bus expansion box includes a switching unit and at least one high-speed serial computer expansion bus slot. The switching unit includes a downlink interface and at least two uplink interfaces. Among them: The high-speed serial computer expansion bus slot is used to install high-speed serial computer expansion bus devices; the downlink interface is used to connect the high-speed serial computer expansion bus slot; the at least two uplink interfaces are used to connect the central processor at the motherboard end, so as to form a multi-channel data link between the central processor and any high-speed serial computer expansion bus slot. The connection to the central processor at the motherboard end specifically includes: the at least two uplink interfaces are connected one-to-one with the corresponding number of interfaces on the motherboard end; during data transmission, the switching unit selects any channel in the multi-channel data link for data transmission; if the current data channel fails, it automatically switches to any channel in the remaining redundant channels for data transmission; The central processor is respectively connected to the uplink interfaces of each high-speed serial computer expansion bus expansion box through multiple high-speed serial computer expansion bus interfaces at the motherboard end, and performs data communication with the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; when any channel in the multi-channel data link fails, the high-speed serial computer expansion bus interface of the channel at the motherboard end is disabled.
2. The server according to claim 1, wherein The bandwidth of the downlink interface in the switching unit is the sum of the bandwidths of all high-speed serial computer expansion bus slots; The bandwidth of any uplink interface in the switching unit is not less than the bandwidth of the downlink interface.
3. The server according to claim 1, characterized in that, The bandwidths of all uplink interfaces in the switching unit are equal.
4. The server according to claim 1, characterized in that, The high-speed serial computer expansion bus slot is an X16 slot or two X8 slots; The switching unit includes a downlink interface and two uplink interfaces. Among them, the bandwidth of each uplink interface is equal to the bandwidth of the downlink interface; the bandwidth of the downlink interface is the sum of the bandwidths of all high-speed serial computer expansion bus slots.
5. The server according to claim 1, wherein The basic input / output system runs on the central processor; the basic input / output system is used to perform the following steps: when the server starts, the basic input / output system allocates corresponding high-speed serial computer expansion bus resources to at least one channel in each multi-channel data link; among them, the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box.
6. The server according to claim 5, characterized in that Also including: The baseboard management controller is used to identify channels for each high-speed serial computer expansion bus expansion box when the server starts; If the high-speed serial computer expansion bus expansion box has multiple channels, set a default channel for the high-speed serial computer expansion bus expansion box; The switching unit is also used to preferentially select the default channel for data transmission; The basic input / output system is further configured to allocate corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box.
7. The server according to claim 6, wherein The baseboard management controller is configured to, when the server starts up, determine whether the high-speed serial computer expansion bus expansion box has multiple channels by identifying the field replaceable unit (FRU) of the underlying unit group of the high-speed serial computer expansion bus expansion box; and store the information of the default channel in the erasable programmable read-only memory after the default channel setting is completed.
8. The server according to claim 5, characterized in that, The basic input / output system is further configured to perform the following steps when the operating system of the server starts up: Detect whether there is an uncorrectable error in the current channel; if there is an uncorrectable error, determine that the current channel has failed; or, detect whether the number of correctable errors occurring in the current channel reaches a first threshold; If the first threshold is reached, determine that the current channel has failed; When it is determined that the current channel has failed, disable the high-speed serial computer expansion bus interface of the current channel at the motherboard end; The switching unit is further configured to switch the remaining channels of the data link where the current channel is located for data transmission according to the disabling information of the high-speed serial computer expansion bus interface at the motherboard end.
9. The server according to claim 5, characterized in that, The basic input / output system is further configured to perform the following steps during the operation of the high-speed serial computer expansion bus device: Continuously read the registers of the central processing unit; if an uncorrectable error appears in the registers, determine that the current channel has failed; or, continuously read the registers of the central processing unit; if the number of correctable errors in the registers reaches a second threshold, determine that the current channel has failed; When it is determined that the current channel has failed, disable the high-speed serial computer expansion bus interface of the current channel at the motherboard end; The switching unit is further configured to switch the remaining channels of the data link where the current channel is located for data transmission according to the disabling information of the high-speed serial computer expansion bus interface at the motherboard end.
10. The server according to claim 8 or 9, characterized in that The basic input / output system is further configured to generate fault information based on all the channels that have failed currently, and send the fault information and the information of the switched channels to the operating system and the baseboard management controller of the server; The operating system of the server records a fault log according to the received fault information and the information of the switched channels; The baseboard management controller records a fault log according to the received fault information and the information of the switched channels.
11. The server according to claim 10, wherein The baseboard management controller issues an alarm for a link fault according to the received fault information.
12. The server according to claim 10, wherein The baseboard management controller is further configured to, when all channels of the data link have failed in the received fault information, determine that a device fault has occurred and issue an alarm for the device fault.
13. A method for controlling data transmission, characterized in that, Applied to the server according to any one of claims 1-12, including: When the server starts up, perform channel identification on each high-speed serial computer expansion bus expansion box; if the high-speed serial computer expansion bus expansion box has multiple channels, set a default channel for the high-speed serial computer expansion bus expansion box; Allocate corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box, and initialize the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; wherein, the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box; Detect whether the current channel has a fault; If the current channel has a fault, disable the high-speed serial computer expansion bus interface of the current channel at the motherboard end, and switch the remaining channels of the data link to which the current channel belongs for data transmission.
14. The method for controlling data transmission according to claim 13, wherein Detecting whether the current channel has a fault includes: When the operating system of the server starts, detect whether there are uncorrectable errors in the current channel; if there are uncorrectable errors, determine that the current channel has a fault; Or, when the operating system of the server starts, detect whether the number of correctable errors that appear in the current channel reaches a first threshold; if the first threshold is reached, determine that the current channel has a fault.
15. The method for controlling data transmission according to claim 13, wherein Detecting whether the current channel has a fault includes: During the operation of the high-speed serial computer expansion bus device, continuously read the register of the central processing unit If there are uncorrectable errors in the register, determine that the current channel has a fault; Or, during the operation of the high-speed serial computer expansion bus device, continuously read the register of the central processing unit, if the number of correctable errors in the register reaches a second threshold, determine that the current channel has a fault.
16. The method for controlling data transmission according to claim 14 or 15, characterized in that, Further includes: Generate fault information based on all the channels that currently have faults, and record a fault log according to the fault information and the information of the switched channels; Perform an alarm for the link fault according to the fault information.
17. The method for controlling data transmission according to claim 16, wherein Further includes: If all channels of the data link have faults in the fault information, determine that a device fault has occurred; Perform an alarm for the device fault according to the fault information.
18. A device for controlling data transmission, characterized in that, For implementing the method according to any one of claims 13-17, includes: A management module, configured to identify channels for each high-speed serial computer expansion bus expansion box when the server starts; if the high-speed serial computer expansion bus expansion box has multiple channels, set a default channel for the high-speed serial computer expansion bus expansion box; allocate corresponding high-speed serial computer expansion bus resources to all channels of each high-speed serial computer expansion bus expansion box, and initialize the high-speed serial computer expansion bus devices on the high-speed serial computer expansion bus expansion box; wherein, the bandwidth of each channel is not less than the sum of the bandwidths of all high-speed serial computer expansion bus slots in the high-speed serial computer expansion bus expansion box; A control module, configured to detect whether the current channel has a fault; if the current channel has a fault, disable the high-speed serial computer expansion bus interface of the current channel at the motherboard end, and switch the remaining channels of the data link to which the current channel belongs for data transmission.
19. A non-volatile readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, the steps in the method according to any one of claims 13-17 are implemented.
20. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the computer program, the steps in the method according to any one of claims 13-17 are implemented.
Citation Information
Patent Citations
PCIE equipment
CN110362511A
Server using double-slot CPU
CN113177018A
PCIE (Peripheral Component Interface Express) data exchange device
CN114968873A
Converged architecture system, nonvolatile storage system and storage resource acquisition method
CN116185641A
PCIE (Peripheral Component Interface Express) expansion box, server, method, device and product for controlling data transmission
CN117591457A
Cited By
Server and control method
CN120723521A