An AI server with hardware fault self-detection

CN122817010APending Publication Date: 2026-09-25GUANGZHOU DAOQIN ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610978142.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

当前行业内硬件自检主要存在两种主流实现方案:第一种为硬件独立轮询自检方案,各硬件部件对应独立自检任务,所有任务均执行完整硬件参数采集与故障校验,该方式不存在任务分组,能够精准采集单台硬件运行状态,但大批量硬件同步全量自检时会占用大量CPU、内存与传输带宽资源,极易挤占AI模型训练、推理业务所需算力;且同源硬件批量故障时,多任务逐条上报告警,产生海量重复告警信息,运维人员梳理故障根因耗时久;

Benefits of technology

本发明通过设置自检存储单元更新存储若干同源自检组,设置任务执行单元正常仅执行主部件的自检任务,大幅减少了全量自检频率;只有当主部件检出异常时,才触发所属组内从属部件的自检,避免了正常状态下所有硬件部件同步全量自检造成的CPU、内存及带宽资源挤占,从而保障AI训练、推理业务的算力资源;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817010A_ABST
    Figure CN122817010A_ABST
Patent Text Reader

Abstract

The application discloses an AI server with hardware fault self-detection, and relates to the technical field of servers.The application sets a self-checking storage unit to store a plurality of homologous self-checking groups, and sets a task execution unit to normally execute only the self-checking task of the main component, thereby greatly reducing the full self-checking frequency;only when the main component detects an abnormality, the self-checking of the slave components in the group is triggered, thereby avoiding the CPU, memory and bandwidth resource squeezing caused by the synchronous full self-checking of all hardware components in the normal state;the homologous coupling determination unit is set to periodically analyze and determine the fault state correlation coefficient and normal working condition correlation coefficient of each homologous self-checking group, automatically identify the hardware components with low coupling degree with other components in the group, and migrate the hardware components to a more matched homologous self-checking group or a newly-built group by the changing unit in the group, so that the whole process does not need manual intervention, and the dynamic alignment of the self-checking grouping and the physical coupling relationship is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of server technology, specifically to an AI server with self-detection of hardware faults. Background Technology

[0002] With the large-scale deployment of AI server computing power clusters, the number of hardware components such as GPUs, power supplies, cooling fans, and backplane buses inside the clusters is enormous. Periodic self-checks of hardware are a necessary means to ensure stable operation of equipment and to detect component aging and hardware failures in a timely manner. Currently, there are two main implementation schemes for hardware self-testing in the industry: The first is the hardware independent polling self-testing scheme, where each hardware component corresponds to an independent self-testing task. All tasks perform complete hardware parameter collection and fault verification. This method does not have task grouping and can accurately collect the operating status of a single hardware unit. However, when a large number of hardware units perform full self-testing simultaneously, it will consume a lot of CPU, memory and transmission bandwidth resources, which can easily squeeze the computing power required for AI model training and inference. Moreover, when a batch of hardware from the same source fails, multiple tasks report alarms one by one, generating a large number of duplicate alarm messages, and it takes a long time for maintenance personnel to sort out the root cause of the failure. The second type is the static group self-test scheme. R&D personnel pre-divide fixed hardware groups based on the hardware wiring, power supply, and heat dissipation layout at the factory stage. A main test task is set in the group, and the other subordinate tasks reuse the test data of the main task to save resources. At the same time, this type of scheme is generally equipped with a periodic sampling full inspection mechanism to verify the actual working condition of the hardware through low-frequency full sampling inspection. Some operation and maintenance systems also support automatic adjustment of alarm classification rules based on sampling anomalies. However, the group boundaries, root subordinate task configuration, and self-test execution logic of this scheme are all fixed configurations. It can only be automatically optimized at the upper-level alarm display level and cannot be linked to modify the underlying hardware self-test execution strategy. During the entire lifecycle of a server, factors such as data center line upgrades, hardware replacement and relocation, component aging, and rack load adjustments gradually alter the original power supply, heat dissipation, and bus coupling relationships between hardware components, rendering the previously statically defined same-source groups ineffective. Existing static grouping schemes, when random checks reveal asynchronous hardware failures within a group, can only alert the system to abnormal group configurations or automatically split alarm entries. Adjusting task groups, changing root dependency relationships, or modifying hardware testing rules requires manual configuration by maintenance personnel. If maintenance personnel fail to optimize groupings promptly, two drawbacks occur: First, hardware that has broken free of the same-source relationship may still use data reuse mechanisms, leading to missed fault reports when subordinate hardware experiences independent failures due to data reuse. Second, newly added hardware may be incorporated into existing coupling links but not assigned to the same-source group, resulting in continuous independent full-scale self-testing and unnecessary waste of computing resources. To address the above problems, this invention proposes a solution. Summary of the Invention

[0003] The purpose of this invention is to provide an AI server with self-detection of hardware faults, in order to solve the problems mentioned in the background art.

[0004] This invention provides an AI server with self-detection of hardware faults, comprising: The self-test storage unit is used to update the storage of several identical self-test groups. Each identical self-test group contains a hardware component of the main component type and several hardware components of the subordinate component type. Each hardware component corresponds to a preset self-test task. If any hardware component is detected as abnormal during the execution process, a preset execution intervention operation is performed to intervene in the execution of the self-test tasks of all hardware components in the same self-test group to which the hardware component belongs and to generate fault intervention data for the same self-test group. The task execution unit is also used to perform a full inspection task on the same source inspection group containing the hardware component after the number of self-inspection tasks of the hardware component with any component type as the main component reaches a preset fixed amount, and to obtain a full inspection record data of the same source inspection group after the full inspection task is completed. The same-source coupling determination unit is used to immediately store each full inspection record data or fault intervention data received from any same-source inspection group. The same source coupling determination unit is also used to analyze the full inspection record data and fault intervention data of each same source inspection group stored in the coupling determination period at intervals of a coupling determination period, determine the hardware components that need to be split in several same source inspection groups in the coupling determination period based on the analysis results, and generate hardware split groups of several same source inspection groups in the coupling determination period based on them. The unit within the group changes the position of each hardware component within a group for each coupling determination cycle received.

[0005] Furthermore, during the execution of the self-test task of any hardware component of a main component type, if an abnormality is detected in the hardware component, the execution of the self-test tasks of all hardware components within the same self-test group to which the hardware component belongs is intervened, and the following fault intervention data for the same self-test group is generated: Record the detection time and generate a fault code; obtain all hardware components of the same self-test group to which the hardware component belongs that are of the subordinate component type; and execute the self-test task of all the obtained hardware components. After all the hardware components have completed their self-test tasks, obtain all the hardware components that detected abnormalities, along with their detection time and fault codes, and generate a fault intervention data for the same self-test group based on these data.

[0006] Furthermore, the content of a single full inspection task performed by any originating inspection team is as follows: Obtain all hardware components and their self-test tasks included in the same self-test group; For each hardware component obtained, a self-test task of the hardware component is immediately executed, and after the self-test task is completed, the parameter values ​​of several relevant parameters during the execution of the self-test task are obtained and the self-test data of the hardware component is generated based on them. After all the self-test data of the hardware components have been generated, a full inspection record data of the same self-test group is generated based on it.

[0007] Furthermore, the contents of several hardware split groups originating from the same source detection group for any coupling determination period are as follows: S11: Obtain all identical self-test groups stored in the self-test storage unit at the current moment and select one identical self-test group as the coupling determination group; mark all hardware components contained in the coupling determination group as A1, A2, ..., Aa, a≥1, where the hardware component whose component type is the main component is marked as A1; S12: All fault intervention data of the coupling determination group stored within the coupling determination period are respectively labeled as B1, B2, ..., Bb, where b≥1; S13: Obtain all fault intervention data with the same fault code in fault intervention data B1, B2, ..., Bb, extract all detection times of hardware components A1 and A2 from them and mark them as t1, t2, ..., tn respectively, where n is the total number of detection times of hardware components A1 and A2 extracted; Using formula Calculate the correlation coefficient E1 of the fault states of hardware components A1 and A2. In the formula, Ct and Dt represent the fault states of hardware components A1 and A2 at detection time t, respectively, and C and D are the average values ​​of the fault states of hardware components A1 and A2 at detection times t1, t2, ..., tn, respectively. S14: Obtain all self-test data of hardware components A1 and A2 from all full-test record data of the coupling determination group stored in the coupling determination period; mark all relevant parameters selected by the operation and maintenance personnel as F1, F2, ..., Ff respectively; S15: Calculate and obtain the correlation coefficient J1 between hardware components A1 and A2 under normal operating conditions; S16: Based on the fault state correlation coefficient E1 and the normal operating condition correlation coefficient J1, determine whether hardware components A1 and A2 are mutually coupled from the same source. If E1≥P1 and J1>P2 during the determination process, hardware components A1 and A2 are determined to be mutually coupled from the same source; otherwise, hardware components A1 and A2 are determined not to be mutually coupled from the same source, and hardware components A1 and A2 are marked as pending statistics. S17: According to S11 to S16, determine whether any two hardware components A1, A2, ..., Aa are mutually coupled from the same source. After all determinations are completed, obtain the number of times each hardware component is marked with a statistical marker. If the number of times several hardware components are marked with a statistical marker reaches a preset fixed frequency, then these hardware components are determined to be hardware components that need to be split within the coupling determination group of the coupling determination period. Based on this, generate the hardware splitting group of the coupling determination group of the coupling determination period. The preset fixed frequency is based on the total number of hardware components. If no hardware component is marked with a statistical marker and the number of times it reaches the preset fixed frequency, then no processing is performed. S18: Select all the same source test groups as coupling determination groups respectively, and generate hardware split groups of several same source test groups in the coupling determination period according to S12 to S17.

[0008] Furthermore, in S13, the principle for determining the fault state of hardware component A1 at detection time t is as follows: if hardware component A1 detects an abnormality at detection time t, its fault state at retrieval time t is 1; if hardware component A1 does not detect an abnormality at detection time t, its fault state at retrieval time t is 0. The principle for determining the fault state of hardware component A2 is the same.

[0009] Furthermore, in S15, the calculation of the correlation coefficient J1 between hardware components A1 and A2 under normal operating conditions is as follows: SS11: Extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A1, and sort all the extracted parameter values ​​from the order in which the self-test data was generated; similarly, extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A2, and sort all the extracted parameter values ​​from the order in which the self-test data was generated. SS12: Calculate and obtain the Pearson correlation coefficient G1 of hardware components A1 and A2 based on the relevant parameter F1. In the calculation process, firstly, all parameter values ​​of the relevant parameter F1 are extracted from all self-test data of hardware components A1 and A2 to generate the first, second, ..., h group point sets (H1, I1), (H2, I2), ..., (Hh, Ih), where h refers to the total number of full test record data of the coupling determination group stored in the coupling determination period; SS13: Pearson correlation coefficients G2, G3, ..., Gf of hardware components A1 and A2 generated according to SS11 to SS12, based on relevant parameters F2, F3, ..., Ff; SS14: Calculate the correlation coefficient J1 of hardware components A1 and A2 under normal operating conditions using the formula J1=G1×ɑ1+G2×ɑ2+...+Gf×ɑf. In the formula, ɑ1, ɑ2, ..., ɑf are the preset weights of the relevant parameters F1, F2, ..., Ff, respectively.

[0010] Furthermore, after receiving the hardware splitting of any identical self-test group in any coupling determination period, the changes made to the group positions of several hardware components currently stored in the identical self-test group in the self-test storage unit are as follows: S21: Select any hardware component from the hardware splitting group as the changed component; S22: Calculate the fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in each of the same self-inspection groups except for the optimization group, according to S13 to S15 respectively. S23: Based on the calculated fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in any same self-test group, determine whether the changed component and each hardware component in the same self-test group are mutually coupled according to S16. Select the same self-test group with the most hardware components that are determined to be mutually coupled with the changed component as the change target group of the changed component. At this time, find the stored same self-test group including the changed component in the self-test storage unit, and move the changed component in it into the change target group to complete the change of the changed component's position within the group. If the number of hardware components in any same self-test group that are coupled to the changed component is less than the preset same-source quantity, a new same self-test group is created in the self-test storage unit. The same self-test group is used as the change target group for the changed component. At this time, the same self-test group containing the changed component is found in the self-test storage unit, and the changed component is moved to the change target group. After the movement, the changed component is no longer in the same self-test group that originally contained the changed component, thus completing the position update of the changed component. S24: Select each hardware component in the hardware splitting group of all the same source inspection groups in the received coupling determination period as the changed component, and change the position of each hardware component in the group according to S21 to S23.

[0011] Compared with existing technologies, it has the following advantages: This invention significantly reduces the frequency of full self-tests by setting up a self-test storage unit to update and store several self-test groups of the same type, and by setting the task execution unit to execute only the self-test task of the main component under normal conditions. Only when the main component detects an anomaly will the self-test of the subordinate components in the same group be triggered, thus avoiding the CPU, memory, and bandwidth resource squeezing caused by the simultaneous full self-test of all hardware components under normal conditions, thereby ensuring the computing power resources for AI training and inference operations. This invention, by setting up a co-source coupling determination unit to periodically analyze the full inspection record data and fault intervention data of each co-source self-inspection group, determines the fault state correlation coefficient and normal operating condition correlation coefficient of each co-source self-inspection group, automatically identifies hardware components with low coupling with other components in the group, and has the group's change unit migrate them to a more suitable co-source self-inspection group or a newly created group. The entire process requires no manual intervention, realizing the dynamic alignment of self-inspection grouping and physical coupling relationship, avoiding the situation where hardware components are separated from the co-source relationship but still use the data reuse mechanism, and the occurrence of fault omissions due to the reuse of normal data when subordinate hardware experiences independent faults; This invention automatically separates hardware components that have broken away from their original source relationship, preventing them from continuing to reuse the normal self-test data of the main component and thus masking their own independent faults. At the same time, newly added hardware components that are integrated into the coupling link are automatically assigned to the corresponding original source self-test group, so that they are transformed from full self-test to on-demand subordinate self-test, saving unnecessary computing resources. Attached Figure Description

[0012] Figure 1 This is a functional structure diagram of the present invention. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] Please see Figure 1 This application provides an AI server with hardware fault self-detection, including a self-test storage unit, a task execution unit, a homogeneous coupling determination unit, and an intra-group change unit; The self-test storage unit is used to update and store several identical self-test groups. Each identical self-test group contains a hardware component of type master and several hardware components of type slave. Each hardware component corresponds to a preset self-test task. The self-test task refers to a set of preset detection operations for that hardware component, used to determine whether the hardware component is in normal working condition. The content of the detection operations is determined according to the type of hardware component, including but not limited to: reading the component's internal status register, executing the built-in self-test BIST, measuring key power supply voltage or temperature, verifying communication link connectivity, and running diagnostic instructions. The execution result of the self-test task includes at least two states: "normal" or "abnormal". When the result is abnormal, the detection time and fault code are also recorded. In this application, the initial self-test group is constructed and pre-stored in the self-test storage unit by the administrator based on the physical coupling relationship of the target server's hardware power supply link, heat dissipation structure, and interconnect bus; The administrator selects one hardware component from all the hardware components in the group as the master component, and the rest as slave components. The selection is based on the fact that the self-test results of the master component can represent the normal or abnormal working condition of the common links in the group, such as power supply, heat dissipation, and bus. Usually, the hardware component that undertakes core computing, control, or is located at a critical node of the link is selected as the master component. For example, in an AI server chassis, the same power output branch on the backplane supplies power to three GPU cards and their associated power supply chips. The three GPU cards are located in the same heat dissipation duct and share a set of cooling fans. The three GPUs are interconnected through the same PCIe high-speed backplane bus. The self-test tasks corresponding to the three GPUs are divided into a common self-test group. In this group, the GPU with the strongest direct association with the common link and the widest self-test coverage is usually selected as the master component, and the other GPUs are slave components. The task execution unit is used to execute the self-test task of the main hardware components of each self-test group at a preset self-test cycle, and reuse the execution results. The specific reuse method is as follows: The execution results of the self-test task of the main component are obtained, including but not limited to normal or abnormal status, fault codes, and several operating parameters such as voltage and temperature. These are directly used as the self-test data of all hardware components of the same self-test group whose component type is a subordinate component in the self-test cycle. The subordinate hardware components no longer execute their own self-test tasks separately, but inherit the self-test data of the main component by default. If the self-test result of the main component is normal, the subordinate components are all considered normal and do not need to execute additional self-test tasks. If any hardware component is detected as abnormal during execution, a preset execution intervention operation is performed to intervene in the execution of the self-test tasks of all hardware components in the same self-test group to which the hardware component belongs, and to generate fault intervention data for the same self-test group. The fault intervention data is then transmitted to the same-source coupling determination unit. Specifically, during the execution of the self-test task of any hardware component of the main component type, if an abnormality is detected in the hardware component, the detection time is recorded and a fault code is generated. All hardware components of the same self-test group to which the hardware component belongs, whose component type is a subordinate component, are obtained, and the self-test tasks of all the obtained hardware components are executed. After all the hardware components have completed their self-test tasks, obtain all the hardware components that detected abnormalities, their detection time, and fault codes, and generate a fault intervention data for the same self-test group based on them. The task execution unit is also used to perform a full inspection task on the same source inspection group containing the hardware component after the number of self-inspection tasks of the hardware component with any component type as the main component reaches a preset fixed amount, and to obtain a full inspection record data of the same source inspection group after the full inspection task is completed. The content of a single full inspection task performed by any originating inspection team is as follows: Obtain all hardware components and their self-test tasks included in the same self-test group; For each hardware component acquired, a self-test task is immediately executed for that hardware component. After the self-test task is completed, the parameter values ​​of several related parameters during the execution of the self-test task are acquired. These related parameters are selected as needed by the operation and maintenance personnel and are used to measure the degree of correlation between hardware components, including but not limited to execution time, CPU utilization, memory utilization, bandwidth consumption, average temperature, average voltage, and average power consumption during the execution of the self-test task, and self-test data of the hardware component is generated based on these parameters. After all the self-test data of the hardware components have been generated, a full inspection record data of the same source self-test group is generated based on it, and the full inspection record data is transmitted to the same source coupling determination unit. The same-source coupling determination unit is used to store each full inspection record data or fault intervention data received from any same-source inspection group immediately after receiving it. The same-source coupling determination unit is also used to analyze the full inspection record data and fault intervention data of each same-source inspection group stored in the coupling determination period at every coupling determination period, determine the hardware components that need to be split in several same-source inspection groups in the coupling determination period based on the analysis results, generate hardware split groups of several same-source inspection groups in the coupling determination period according to them, and transmit the generated hardware split groups of several same-source inspection groups in the coupling determination period to the group change unit. The hardware splitting group for generating several identical source detection groups for any coupling determination period is as follows: S11: Obtain all identical self-test groups stored in the self-test storage unit at the current moment and select one identical self-test group as the coupling determination group; mark all hardware components contained in the coupling determination group as A1, A2, ..., Aa, a≥1, where the hardware component whose component type is the main component is marked as A1; S12: All fault intervention data of the coupling determination group stored within the coupling determination period are respectively labeled as B1, B2, ..., Bb, where b≥1; S13: Obtain all fault intervention data with the same fault code in fault intervention data B1, B2, ..., Bb, extract all detection times of hardware components A1 and A2 from them and mark them as t1, t2, ..., tn respectively, where n is the total number of detection times of hardware components A1 and A2 extracted; Using formula Calculate the correlation coefficient E1 of the fault states of hardware components A1 and A2. In the formula, Ct and Dt represent the fault states of hardware components A1 and A2 at detection time t, respectively, and C and D are the average values ​​of the fault states of hardware components A1 and A2 at detection times t1, t2, ..., tn, respectively. The principle for determining the fault state of hardware component A1 at detection time t is as follows: if hardware component A1 detects an abnormality at detection time t, its fault state at retrieval time t is 1; if hardware component A1 does not detect an abnormality at detection time t, its fault state at retrieval time t is 0. The principle for determining the fault state of hardware component A2 is the same. It should be noted that the fault state correlation coefficient is a human-defined parameter used to characterize the degree of coupling between hardware components, such as power supply, heat dissipation, and bus. The closer the fault state correlation coefficient is to 1, the stronger the coupling relationship between the two corresponding hardware components and the higher the degree of homogeneity. S14: Obtain all self-test data of hardware components A1 and A2 from all full-test record data of the coupling determination group stored in the coupling determination period; mark all relevant parameters selected by the operation and maintenance personnel as F1, F2, ..., Ff respectively; S15: Calculate the correlation coefficient J1 between hardware components A1 and A2 under normal operating conditions. The calculation is as follows: SS11: Extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A1, and sort all the extracted parameter values ​​from the self-test data in the order in which they were generated. Similarly, extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A2, and sort all the extracted parameter values ​​from the order in which the self-test data was generated. SS12: Calculate and obtain the Pearson correlation coefficient G1 of hardware components A1 and A2 based on the relevant parameter F1. In the calculation process, firstly, all parameter values ​​of the relevant parameter F1 are extracted from all self-test data of hardware components A1 and A2 to generate the first, second, ..., h group point sets (H1, I1), (H2, I2), ..., (Hh, Ih), where h refers to the total number of full test record data of the coupling determination group stored in the coupling determination period; In the first set of points, H1 and I1 are the parameter values ​​of the relevant parameter F1 in the self-test data of hardware components A1 and A2 in the first full inspection record data stored in the coupling determination period, according to the order of storage. The second, third, ..., h sets of points follow the same pattern. SS13: Pearson correlation coefficients G2, G3, ..., Gf of hardware components A1 and A2 generated according to SS11 to SS12, based on relevant parameters F2, F3, ..., Ff; SS14: The correlation coefficient J1 of hardware components A1 and A2 under normal operating conditions is calculated using the formula J1=G1×ɑ1+G2×ɑ2+...+Gf×ɑf. In the formula, ɑ1, ɑ2, ..., ɑf are the preset weights of the relevant parameters F1, F2, ..., Ff, which are preset by the operation and maintenance management personnel based on the importance of each relevant parameter in measuring the degree of coupling correlation. S16: Determine whether hardware components A1 and A2 are mutually coupled based on the fault state correlation coefficient E1 and the normal operating condition correlation coefficient J1. Based on the determination result, select whether to mark hardware components A1 and A2 with a statistical label. The judgment is as follows: If E1≥P1 and J1>P2, it indicates that the power supply, heat dissipation, bus and other coupling relationships between hardware components A1 and A2 are strong and the degree of co-coupling is high. In this case, hardware components A1 and A2 are determined to be co-coupled. Conversely, if E1≥P1 and J1>P2, it indicates that the power supply, heat dissipation, bus and other coupling relationships between hardware components A1 and A2 are weak and the degree of co-coupling is low. In this case, hardware components A1 and A2 are determined not to be co-coupled. Hardware components A1 and A2 are marked with a statistical mark, and P1 and P2 are the preset fault state correlation coefficient threshold and normal working condition correlation coefficient threshold, respectively. S17: According to S11 to S16, determine whether any two hardware components A1, A2, ..., Aa are mutually coupled from the same source. After all determinations are completed, obtain the number of times each hardware component is marked with a statistical marker. If the number of times several hardware components are marked with a statistical marker reaches a preset fixed frequency, then these hardware components are determined to be hardware components that need to be split within the coupling determination group of the coupling determination period. Based on this, generate the hardware splitting group of the coupling determination group of the coupling determination period. The preset fixed frequency is based on the total number of hardware components. If no hardware component is marked with a statistical marker and the number of times it reaches the preset fixed frequency, then no processing is performed. S18: Select all the same source inspection groups as coupling determination groups respectively, generate several hardware splitting groups of the same source inspection groups in the coupling determination period according to S12 to S17 and transmit them to the group change unit. The group-based change unit changes the group-based position of each hardware component within a group for each coupling determination cycle received by several hardware split groups from the same source inspection group. Upon receiving a hardware split from any identical self-test group during any coupling determination period, the following changes are made to the group positions of several hardware components currently stored in the self-test storage unit within the identical self-test group: S21: Select any hardware component from the hardware splitting group as the changed component; S22: Calculate the fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in each of the other self-inspection groups except the aforementioned self-inspection group, according to S13 to S15 respectively. S23: Change the position of the component within the group. The changes are as follows: For the calculated fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in any same self-test group, according to S16, it is determined whether the changed component and each hardware component in the same self-test group are mutually coupled from the same source. The same self-test group with the most hardware components that are determined to be mutually coupled from the same source with the changed component is selected as the change target group of the changed component. At this time, the same self-test group including the changed component is found in the self-test storage unit, and the changed component is moved to the change target group to complete the change of the changed component's position within the group. If the number of hardware components in any same self-test group that are coupled to the changed component is less than the preset same-source quantity, a new same self-test group is created in the self-test storage unit. The same self-test group is used as the change target group for the changed component. At this time, the same self-test group containing the changed component is found in the self-test storage unit, and the changed component is moved to the change target group. After the movement, the changed component is no longer in the same self-test group that originally contained the changed component, thus completing the position update of the changed component. S24: Select each hardware component in the hardware splitting group of all the same source inspection groups in the received coupling determination period as the changed component, and change the position of each hardware component in the group according to S21 to S23. It should be noted here that if a hardware component was a main component in its own source inspection group before the position change within the group, then after the position change within the group, a hardware component is selected for the same source inspection group and the target group to which the hardware component is moved, and the component type is changed accordingly. The changes to the component type of a selected hardware component within the same source inspection group are as follows: Since there are no hardware components of the same type as the main component in the same self-testing group, for the remaining hardware components in the same self-testing group, according to the same principle as the initial selection of the main component, that is, the self-testing result of the main component can represent the working condition of the common link in this group, a hardware component that undertakes the core function or is in the critical node of the link is usually selected to re-select a hardware component as the main component. The changes to the component type of a selected hardware component within the target group are as follows: Based on the same principle as the initial selection of the main component, i.e., the self-test result of the main component can represent the working condition of the common link in this group, one hardware component is selected as the main component from all hardware components in the target group, and the remaining hardware components are subordinate components; if there is already a main component in the target group, the original main component is changed to a subordinate component, or a new selection is made according to the above principle, and the newly selected result shall prevail.

[0015] Some of the data in the above formulas are numerical calculations with dimensions removed, and the contents not described in detail in this specification are all prior art known to those skilled in the art.

[0016] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. An AI server with self-detection of hardware faults, characterized in that, include: The self-test storage unit is used to update the storage of several identical self-test groups. Each identical self-test group contains a hardware component of the main component type and several hardware components of the subordinate component type. Each hardware component corresponds to a preset self-test task. If any hardware component is detected as abnormal during the execution process, a preset execution intervention operation is performed to intervene in the execution of the self-test tasks of all hardware components in the same self-test group to which the hardware component belongs and to generate fault intervention data for the same self-test group. The task execution unit is also used to perform a full inspection task on the same source inspection group containing the hardware component after the number of self-inspection tasks of the hardware component with any component type as the main component reaches a preset fixed amount, and to obtain a full inspection record data of the same source inspection group after the full inspection task is completed. The same-source coupling determination unit is used to immediately store each full inspection record data or fault intervention data received from any same-source inspection group. The same source coupling determination unit is also used to analyze the full inspection record data and fault intervention data of each same source inspection group stored in the coupling determination period at intervals of a coupling determination period, determine the hardware components that need to be split in several same source inspection groups in the coupling determination period based on the analysis results, and generate hardware split groups of several same source inspection groups in the coupling determination period based on them. The unit within the group changes the position of each hardware component within a group for each coupling determination cycle received.

2. An AI server with hardware fault self-detection according to claim 1, characterized in that, If an abnormality is detected in the self-test task of a hardware component of any component type that is the main component, the execution of the self-test tasks of all hardware components in the same self-test group to which the hardware component belongs will be intervened, and the following fault intervention data for the same self-test group will be generated: Record the detection time and generate a fault code; obtain all hardware components of the same self-test group to which the hardware component belongs that are of the subordinate component type; and execute the self-test task of all the obtained hardware components. After all the hardware components have completed their self-test tasks, obtain all the hardware components that detected abnormalities, along with their detection time and fault codes, and generate a fault intervention data for the same self-test group based on these data.

3. An AI server with hardware fault self-detection according to claim 1, characterized in that, The content of a single full inspection task performed by any originating inspection team is as follows: Obtain all hardware components and their self-test tasks included in the same self-test group; For each hardware component obtained, a self-test task of the hardware component is immediately executed, and after the self-test task is completed, the parameter values ​​of several relevant parameters during the execution of the self-test task are obtained and the self-test data of the hardware component is generated based on them. After all the self-test data of the hardware components have been generated, a full inspection record data of the same self-test group is generated based on it.

4. An AI server with hardware fault self-detection according to claim 1, characterized in that, The hardware splitting group for generating several identical source detection groups for any coupling determination period is as follows: S11: Obtain all identical self-test groups stored in the self-test storage unit at the current moment and select one identical self-test group as the coupling determination group; mark all hardware components contained in the coupling determination group as A1, A2, ..., Aa, a≥1, where the hardware component whose component type is the main component is marked as A1; S12: All fault intervention data of the coupling determination group stored within the coupling determination period are respectively labeled as B1, B2, ..., Bb, where b≥1; S13: Obtain all fault intervention data with the same fault code in fault intervention data B1, B2, ..., Bb, extract all detection times of hardware components A1 and A2 from them and mark them as t1, t2, ..., tn respectively, where n is the total number of detection times of hardware components A1 and A2 extracted; Using formula Calculate the correlation coefficient E1 of the fault states of hardware components A1 and A2. In the formula, Ct and Dt represent the fault states of hardware components A1 and A2 at detection time t, respectively, and C and D are the average values ​​of the fault states of hardware components A1 and A2 at detection times t1, t2, ..., tn, respectively. S14: Obtain all self-test data of hardware components A1 and A2 from all full-test record data of the coupling determination group stored in the coupling determination period; mark all relevant parameters selected by the operation and maintenance personnel as F1, F2, ..., Ff respectively; S15: Calculate and obtain the correlation coefficient J1 between hardware components A1 and A2 under normal operating conditions; S16: Based on the fault state correlation coefficient E1 and the normal operating condition correlation coefficient J1, determine whether hardware components A1 and A2 are mutually coupled from the same source. If E1≥P1 and J1>P2 during the determination process, hardware components A1 and A2 are determined to be mutually coupled from the same source; otherwise, hardware components A1 and A2 are determined not to be mutually coupled from the same source, and hardware components A1 and A2 are marked as pending statistics. S17: According to S11 to S16, determine whether any two hardware components A1, A2, ..., Aa are mutually coupled from the same source. After all determinations are completed, obtain the number of times each hardware component is marked with a statistical marker. If the number of times several hardware components are marked with a statistical marker reaches a preset fixed frequency, then these hardware components are determined to be hardware components that need to be split within the coupling determination group of the coupling determination period. Based on this, generate the hardware splitting group of the coupling determination group of the coupling determination period. The preset fixed frequency is based on the total number of hardware components. If no hardware component is marked with a statistical marker and the number of times it reaches the preset fixed frequency, then no processing is performed. S18: Select all the same source test groups as coupling determination groups respectively, and generate hardware split groups of several same source test groups in the coupling determination period according to S12 to S17.

5. An AI server with hardware fault self-detection according to claim 4, characterized in that, In S13, the principle for determining the fault state of hardware component A1 at detection time t is as follows: if hardware component A1 detects an abnormality at detection time t, its fault state at retrieval time t is 1; if hardware component A1 does not detect an abnormality at detection time t, its fault state at retrieval time t is 0. The principle for determining the fault state of hardware component A2 is the same.

6. An AI server with hardware fault self-detection according to claim 4, characterized in that, S15, The calculation of the correlation coefficient J1 between hardware components A1 and A2 under normal operating conditions is as follows: SS11: Extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A1, and sort all the extracted parameter values ​​from the order in which the self-test data was generated; similarly, extract all parameter values ​​of relevant parameter F1 from all self-test data of hardware component A2, and sort all the extracted parameter values ​​from the order in which the self-test data was generated. SS12: Calculate and obtain the Pearson correlation coefficient G1 of hardware components A1 and A2 based on the relevant parameter F1. In the calculation process, firstly, all parameter values ​​of the relevant parameter F1 are extracted from all self-test data of hardware components A1 and A2 to generate the first, second, ..., h group point sets (H1, I1), (H2, I2), ..., (Hh, Ih), where h refers to the total number of full test record data of the coupling determination group stored in the coupling determination period; SS13: Pearson correlation coefficients G2, G3, ..., Gf of hardware components A1 and A2 generated according to SS11 to SS12, based on relevant parameters F2, F3, ..., Ff; SS14: Calculate the correlation coefficient J1 of hardware components A1 and A2 under normal operating conditions using the formula J1=G1×ɑ1+G2×ɑ2+...+Gf×ɑf. In the formula, ɑ1, ɑ2, ..., ɑf are the preset weights of the relevant parameters F1, F2, ..., Ff, respectively.

7. An AI server with hardware fault self-detection according to claim 1, characterized in that, Upon receiving a hardware split from any identical self-test group during any coupling determination period, the following changes are made to the group positions of several hardware components currently stored in the self-test storage unit within the identical self-test group: S21: Select any hardware component from the hardware splitting group as the changed component; S22: Calculate the fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in each of the same self-inspection groups except for the optimization group, according to S13 to S15 respectively. S23: Based on the calculated fault state correlation coefficient and normal operating condition correlation coefficient between the changed component and each hardware component in any same self-test group, determine whether the changed component and each hardware component in the same self-test group are mutually coupled according to S16. Select the same self-test group with the most hardware components that are determined to be mutually coupled with the changed component as the change target group of the changed component. At this time, find the stored same self-test group including the changed component in the self-test storage unit, and move the changed component in it into the change target group to complete the change of the changed component's position within the group. If the number of hardware components in any same self-test group that are coupled to the changed component is less than the preset same-source quantity, a new same self-test group is created in the self-test storage unit. The same self-test group is used as the change target group for the changed component. At this time, the same self-test group containing the changed component is found in the self-test storage unit, and the changed component is moved to the change target group. After the movement, the changed component is no longer in the same self-test group that originally contained the changed component, thus completing the position update of the changed component. S24: Select each hardware component in the hardware splitting group of all the same source inspection groups in the received coupling determination period as the changed component, and change the position of each hardware component in the group according to S21 to S23.