Link fault processing method and electronic equipment
By stress testing and performance analysis of the target link, combined with problem information, the faulty object is identified and handled, which solves the problem of insufficient fault diagnosis in the existing technology and improves the efficiency and accuracy of fault location and handling.
Patent Information
- Application Number
- CN202511318186.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2025-12-12
AI Technical Summary
Existing technologies for link fault detection and diagnosis suffer from insufficient comprehensiveness in fault investigation and inaccurate fault location, resulting in low fault handling efficiency and impacting production progress and system reliability.
By stress testing the target link, performance and problem information are obtained. Combined with changes in performance indicators and related changes, the target objects to be processed are identified, and targeted processing information is generated to handle the faults.
It improves the accuracy of fault location and the flexibility of handling, meets the fault handling needs in different scenarios, and reduces business losses caused by excessively long fault troubleshooting time.
Smart Images

Figure CN121125448A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a link fault processing method and an electronic device. BACKGROUND
[0002] Multi-device cooperation becomes a key way to improve computing power and meet data processing needs. Each independent device or system can be communicated through a link to meet the needs of multi-device cooperation.
[0003] For the detection and diagnosis of the link, most of them focus on theoretical analysis or single-dimensional detection, which has problems such as not comprehensive fault troubleshooting and not accurate fault positioning, resulting in low fault processing efficiency and affecting production progress and system reliability. SUMMARY
[0004] In view of the above problems, the present application provides a link fault processing method and an electronic device.
[0005] According to a first aspect of the present application, a link fault processing method is provided, comprising: performing stress testing on a target link formed by connecting a plurality of processors to obtain performance information of the target link under a predetermined stress condition; wherein the processor includes a plurality of functional modules, and the performance information includes at least one performance indicator corresponding to at least one of the data transmission dimension, the request processing dimension, and the resource consumption dimension; obtaining problem information generated by the target link under the predetermined stress condition, the problem information indicating a fault occurring in the stress testing process of the target link; based on the problem information, determining at least one target performance indicator causing the fault from the performance information, and determining a target object to be processed in the target link based on at least one of the change of the target performance indicator and the associated change between a plurality of target performance indicators; the target object includes at least one of any processor in the target link and any functional module in any processor; and generating processing information for the target object according to the target performance indicator and the problem information, so as to process the fault based on the processing information.
[0006] According to a second aspect of the present application, an electronic device is provided, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0007] In the embodiments of the present application, the target object is determined based on the performance information and the problem information, which helps to improve the accuracy of target object positioning, provides a clear direction for subsequent fault processing, and improves the efficiency and pertinence of fault processing. The processing information for processing the fault is determined according to the target performance indicator and the problem information, which can improve the flexibility of fault processing to meet the fault processing needs in different scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0008] The above content of the present application and other purposes, features and advantages will be more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0009] Figure 1 The application scenario of the link fault processing method and the electronic device according to the embodiments of the present application is schematically shown;
[0010] Figure 2 The flow chart of the link fault processing method according to the embodiments of the present application is schematically shown;
[0011] Figure 3 The principle diagram of obtaining the problem information of the target link under the predetermined pressure condition according to the embodiments of the present application is schematically shown;
[0012] Figure 4 The principle diagram of generating the processing information for the target object according to the target performance index and the problem information according to the embodiments of the present application is schematically shown;
[0013] Figure 5 The principle diagram of obtaining the predefined information of the multiple processing steps respectively and determining the execution order between the multiple processing steps according to the predefined information according to the embodiments of the present application is schematically shown;
[0014] Figure 6 The structural block diagram of the computer system of the electronic device according to the embodiments of the present application is schematically shown. DETAILED DESCRIPTION
[0015] In order to make the people in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person of ordinary skill in the art without creative labor should belong to the scope of protection of the present application.
[0016] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0017] The embodiments of the present application provide a link fault processing method and an electronic device. Before introducing the technical solutions provided by the embodiments of the present application, the related technologies involved in the present application are described.
[0018] In some scenarios, different devices or systems may need to work together to achieve complementary functions, improve overall performance, meet complex requirements, etc. A link can serve as a bridge for mutual communication and information exchange between independent devices or systems, meeting the needs of collaborative work between multiple devices.
[0019] For example, in high-performance computing and data center scenarios, the computing power can be improved by the way of collaborative work of multiple processors. External global memory interconnect technology (xGMI) can be used as a high-speed interconnection bridge between processors, storage devices and other high-speed components to realize high-speed data transmission and stable communication between devices on the link, so as to ensure efficient collaborative operation of the multi-processor system.
[0020] For the detection and problem diagnosis of the link, most of them focus on theoretical analysis or single-dimensional detection, and there are problems that the fault troubleshooting is not comprehensive and the fault positioning is not accurate. For example, only the theoretical model based on idealized assumptions is used to analyze the link fault, without considering the actual deviation and real-time data in the running process, resulting in a large deviation between the fault diagnosis result and the actual cause. Or, only the fault information of the physical layer or the link layer is concerned, and the root cause of the fault information is ignored, and only the surface fault object is repaired without solving the root cause of the fault, so that the fault occurs repeatedly. When troubleshooting, it may mainly rely on the subjective experience of technicians or general processing strategies, without combining the performance data of the specific fault scene, which cannot adapt to the dynamically changing running environment, resulting in non-standard processing flow, poor fault handling effect and other problems.
[0021] Figure 1 An application scenario of the link fault processing method and the electronic device according to the embodiment of the present application is schematically shown.
[0022] As shown in Figure 1 , the application scenario 100 according to the embodiment can include a target link 101, a network 102 and a computing device 103. The target link 101 can include, for example, a first device 201, a second device 201 and a third device 201, and the target link 101 can be a heterogeneous computing community composed of the first device 201, the second device 202 and the third device 203 for achieving a specific business target. The network 102 provides a medium for a communication link between the first device 201, the second device 201 and the third device 201, the target link 101 and the computing device 103. The network 102 can include various connection types, such as wired, wireless communication links or optical fiber cables, etc. The computing device 103 can be a physical or virtual device directly connected or located around the target link 101, or can be a remote server or a cloud computing platform.
[0023] The first device 201, the second device 201 and the third device 201 can be, for example, a computing device, such as a server (physical machine, virtual machine, container), a personal computer, a mobile terminal, etc., providing computing power; a network device, such as a router, a switch, a gateway, a load balancer, a firewall, etc., responsible for data routing, switching, distribution and security control; a storage device, such as a hard disk array (NAS / SAN), a database server, a distributed file system, etc., responsible for persistent storage and efficient read / write of data; a special-purpose device, such as an encryption card, a GPU acceleration card, an Internet of Things sensor, etc., providing specific functions.
[0024] It should be noted that the link fault processing method provided by the embodiment of the present application can generally be executed by the computing device 103. Accordingly, the link fault processing apparatus provided by the embodiment of the present application can generally be arranged in the computing device 103. The link fault processing method provided by the embodiment of the present application can also be executed by a computing device or a computing device cluster different from the computing device 103 and capable of communicating with the target link 101 and / or the computing device 103. Accordingly, the link fault processing apparatus provided by the embodiment of the present application can also be arranged in a computing device or a computing device cluster different from the computing device 103 and capable of communicating with the target link 101 and / or the computing device 103.
[0025] It should be understood that Figure 1 the number of target links, networks and computing devices in the above-mentioned application scenario is only illustrative. According to the implementation needs, there can be any number of target links, networks and computing devices.
[0026] Figure 2A flowchart of a link fault processing method according to an embodiment of the present application is shown.
[0027] As shown in Figure 2 The link fault processing method of this embodiment includes operations S210-S240.
[0028] In operation S210, a stress test is performed on a target link formed by connecting a plurality of processors, to obtain performance information of the target link under a predetermined stress condition.
[0029] In some embodiments, the target link can be formed based on External Global Memory Interconnect (xGMI) technology, and cooperative communication of multiple processors is achieved through a unified high-speed channel. The processors can include multiple functional modules, such as a computing module, a storage module, an interconnection module, a control module, etc. The performance information can include at least one performance indicator corresponding to at least one of a data transmission dimension, a request processing dimension, and a resource consumption dimension, such as response time (time required for the link to receive a request and return a response), throughput (amount of data successfully transmitted per unit time), resource utilization rate (proportion of CPU, memory, cache, etc. in the test), concurrency capability, etc.
[0030] In some embodiments, the predetermined stress condition can be set according to the expected use scenario and business requirements of the target link. The predetermined stress condition can include parameters such as the number of concurrent users, data transmission rate, request frequency, etc. A stress test environment is constructed through the predetermined stress condition to perform a stress test on the target link.
[0031] For example, a continuous test request can be initiated on the target link according to the set predetermined stress condition through a suitable stress test tool. During the test, the running state of the target link is monitored in real time, and performance information including multiple performance indicators is collected. The performance indicators can include, for example, response time, throughput, resource utilization rate, etc.
[0032] In operation S220, problem information generated by the target link under the predetermined stress condition is obtained; the problem information indicates a fault occurring in the stress test process of the target link.
[0033] In some embodiments, during the stress test process, the running state of the target link (such as log output, error prompt information, etc. of the system) can be monitored in real time through the monitoring function of the monitoring system or test tool. When the target link fails, such as system crash, response timeout, data transmission error, etc., the time of failure occurrence, specific performance, impact range, etc. related information are recorded in time to form the problem information.
[0034] At operation S230, at least one target performance indicator causing the fault is determined from the plurality of performance indicators based on the problem information, and at least one of a change in the target performance indicator and a change in an association between the plurality of target performance indicators is determined to determine a target object to be processed in the target link.
[0035] In some embodiments, the acquired problem information is analyzed in association with the performance indicators, and by comparing changes in the performance indicators before and after the fault, a performance indicator having an association with the fault is found and determined as the target performance indicator.
[0036] According to the determined target performance indicator, the influence of each device or hardware or software in the device in the target link on the performance indicator can be further analyzed to determine a target object causing the abnormality of the target performance indicator, and the target object includes at least one of any processor in the target link and any functional module in any processor, such as a certain processor, network device, software module, etc. For example, if it is found that the long response time is caused by slow query response of a certain database server, the database server is the target object.
[0037] The performance information is a quantitative mapping of the behavior of the object in the target link, and can be determined by the interaction behavior of each object in the link. For example, the throughput decrease can be caused by insufficient processing capacity of the processor or by network bandwidth bottleneck, and the target object can be determined by analyzing the change in the target performance indicator. There can be a certain association between different performance indicators, and the target object to be processed can also be determined by cross-verification of the plurality of performance indicators. For example, if the throughput decreases and the response time increases, it indicates that the system is in a resource occupation state, and the fault can be generated based on the resource occupation state, and the functional module of the specific occupation resource can be located and determined as the target object.
[0038] Only by the problem information to determine the target object can cause the problem of inaccurate positioning of the target object. For example, the problem information indicates that a software error is found, but it cannot be determined whether the software runs slowly due to insufficient hardware performance of the device to cause the error or the software itself has a logical vulnerability. By combining the device response time and other indicators in the performance information, the root cause of the fault and the target object to be processed can be more accurately determined, the efficiency and accuracy of fault positioning are improved, and the business loss caused by too long fault troubleshooting time is reduced.
[0039] At operation S240, processing information for the target object is generated according to the target performance indicator and the problem information, so as to process the fault based on the processing information.
[0040] In some embodiments, processing information for the target object can be generated according to the determined target performance indicators and problem information. The processing information may, for example, include a detailed description of the fault, abnormal conditions of the target performance indicators, specific processing suggestions and steps. For example, for the problem of slow response of the database server query, the processing information can suggest optimizing the database query statement, increasing the database index, adjusting the database server configuration parameters, etc.
[0041] The generated processing information is conveyed to the relevant operation and maintenance personnel or developers, and the operation and maintenance personnel or developers repair and optimize the target object according to the processing suggestions and steps, so as to solve the fault of the target link in the stress test process and restore the normal operation of the system.
[0042] The performance information provides objective operation data of the target link under stress test, and the problem information records abnormal conditions of the target link in the stress test. By correlating and analyzing the performance information and the problem information, the root cause of the fault can be more accurately determined, the target object to be processed can be accurately located, the deviation of fault location caused by relying on a single information source can be avoided, and the accuracy of fault location can be improved. After the target performance and the target object are determined, the processing information for the target object is generated based on the target performance indicators and the problem information, which can improve the flexibility and accuracy of fault processing to meet the fault processing requirements in different fault scenarios.
[0043] The process of stress testing the target link is described below in combination with specific embodiments.
[0044] In some embodiments, the target link formed by a plurality of processors is stress tested to obtain performance information of the target link under a predetermined stress condition, including: stress testing the target link according to a plurality of test periods to obtain a plurality of test results, wherein each test period includes at least one stress test phase, and the stress test phases in the same test period have different stress values, the stress value representing the amount of data applied to the target link per unit time. According to the plurality of stress test results, the performance information of each object in the target link is determined, the performance information representing the running state of the plurality of objects in the target link, and the object including at least one of any processor in the target link and any functional module in any processor.
[0045] In some embodiments, a reasonable number of test cycles can be determined according to the complexity of the target link, available test time, and other factors. In each test cycle, at least one stress test phase is divided, and the number of stress test phases can be determined according to the characteristics of the link and test requirements. A corresponding stress value is set for each stress test phase, and the stress value represents the number applied to the target link per unit time. The stress value can be set following the principle of from low to high to comprehensively evaluate the performance of the target link under different loads.
[0046] For example, the present application can include 3 test cycles, each of which can include a 1-hour high-load stress test phase and a 20-minute low-load test phase. It should be noted that "high" and "low" here are relative concepts.
[0047] The stress values of the same stress test phase in different test cycles can be the same or different. Stress testing of multiple test cycles based on different stress values can improve the comprehensiveness of stress testing. Stress testing of multiple test cycles based on the same stress value can capture possible failures at different stages. For example, in the initial test phase, there may be some failures due to incomplete hardware initialization, unstable software driver loading, and other reasons. As the test continues, the device enters a stable running phase, and problems such as overheating and component aging due to long-term high-load operation may occur.
[0048] Different stress test phases simulate different data transmission intensities to diagnose the performance, stability, and potential defects of the target link. High-load stress test phases can expose the extreme capacity and extreme risk of the link, and low-load stress test phases can determine the basic stability and benchmark of the link. Combining high-load stress testing and low-load stress testing can improve the comprehensiveness of stress testing and expose potential problems of the target link under different loads.
[0049] In some embodiments, the target link is sequentially stress tested according to the pre-set stress test phases and stress values. In each stress test phase, the corresponding stress is continuously applied, and relevant data such as throughput and response time during the test process are recorded.
[0050] In some embodiments, the test results of each stress test phase in multiple test cycles are sorted and summarized to determine the performance of the target link under different stresses, thereby determining the performance information of each object. For example, the test results can include test cycle, stress test phase, stress value, and statistical information such as the value of each performance indicator at each time.
[0051] The application sets multiple test periods to test the target link from different angles and time periods, fully considers the performance of the target link in different stages and different load environments, and avoids the contingency and limitations of a single test period. Each test period includes multiple pressure value different pressure test stages, which can simulate the running of the target link under different loads, comprehensively evaluate the performance indicators of the target link under different load conditions, and more accurately understand the performance characteristics of the target link.
[0052] The following will introduce the process of obtaining problem information of the target link under predetermined pressure conditions in combination with specific embodiments. Figure 3
[0053] Figure 3 The schematic diagram of obtaining problem information of the target link under predetermined pressure conditions according to the embodiment of the application is shown.
[0054] As shown in Figure 3 , the embodiment of obtaining problem information of the target link under predetermined pressure conditions includes: determining a first problem record of the target link according to the state information of each object in the target link in any pressure test stage; analyzing the log generated in the running process of the target link to obtain a second problem record of the target link; and correlating the first problem record and the second problem record based on time correlation and / or event correlation to obtain the problem information of the target link in any pressure test stage.
[0055] In some embodiments, the state information of each object in the target link collected in any pressure test stage is analyzed in depth, and the abnormal state can be identified through the set threshold and judgment rules. For example, for network delay, a normal range (such as less than 100 ms) can be set, and when the actual delay exceeds the range, it is judged as abnormal. According to the analysis result of the state information, the problems found are recorded in a certain format to form the first problem record. The content of the first problem record can include the time of the problem, the involved object, the specific performance of the problem, the possible impact range, etc.
[0056] In some embodiments, during the running of the target link, a large amount of log information is generated by various objects, including system logs, application logs, error logs, etc. The log collection tool can be used to collect the logs generated during the running of the target link and analyze the collected logs to extract key information (such as timestamps, log levels, error codes, related objects, etc. in the logs). According to the log content, the problems recorded therein are identified, such as the "database connection timeout" error in the application log, the "disk space shortage" warning in the system log, etc. The problems found by analyzing the logs are recorded in a format similar to the first problem record, forming a second problem record. The record content should include the time of the problem occurrence, the log source object, the problem details recorded in the log, etc. For example: "[2024-01-01 10:01:00] The log of application server B records 'database connection pool is exhausted, cannot get new connection', causing part of the business request to fail".
[0057] In some embodiments, the related first problem record and second problem record can be integrated by time correlation or event correlation to form the complete problem information of the target link in any stress test phase. The problem information may, for example, include a comprehensive description of the problem, a list of objects involved, and the correlation between problems. For example: "During the stress test phase of 2024-01-01 10:00 - 10:05, the target link has the following problems: server A CPU usage is too high (95%), causing the response time to be prolonged; at the same time, network device C bandwidth utilization reaches 100%, causing server D connected to C to be unable to normally receive data".
[0058] The first problem record and the second problem record can be associated with time as a clue. For example, by checking the problem records that occur in the same time period, it can be analyzed whether there is a time sequence or simultaneous occurrence between them. For example, the CPU usage of the server is recorded as too high (first problem record) at 10:00, and the database connection problem of the application server is recorded (second problem record) at 10:01. Through time correlation, it can be inferred that the CPU usage is too high may be related to the decrease in database connection processing capacity.
[0059] The logical relationship between them can also be found by analyzing the events and objects involved in the problem records. For example, the first problem record mentions that the bandwidth utilization of network device C reaches 100%, and the second problem record mentions that server D connected to network device C cannot normally receive data. Through event correlation, it can be judged that the bandwidth congestion of network device C may cause server D to have data reception abnormalities.
[0060] The embodiments of the present application comprehensively monitor the target link from two dimensions of system running state and business logic execution by combining the state information and running log of each object in the target link, avoiding the omission and misjudgment caused by a single information source. The problem information obtained by associating the first problem record and the second problem record can solve the problem of isolation between the first problem record and the second problem record, effectively improve the management efficiency of the problem, reduce the amount of data to be processed, and improve the efficiency of subsequent processing. For example, only through the state information, the server performance decline can be found, but by combining the log for association analysis, the server performance decline can be associated with the database access anomaly. The completeness and accuracy of the problem information are improved, and the problem troubleshooting efficiency is improved.
[0061] In some embodiments, the target link can include a plurality of parallel sub-links, and in an ideal state, the load and error rate of the plurality of sub-links should be close to uniform distribution. According to the characteristics of the parallel sub-links, the embodiments of determining the first problem record of the target link are proposed.
[0062] The process of determining the first problem record of the target link is introduced below in combination with specific embodiments.
[0063] In some embodiments, the first problem record of the target link is determined according to the state information of each object in the target link, including: obtaining the state information of each of the plurality of sub-links in the stress test phase respectively, the state information indicating the retraining times of the sub-link, the retraining being an operation triggered by the target link in the test process to cope with a specific situation, the specific situation including at least one of signal quality decline, hardware failure or change, and environmental factor change; in response to the retraining times of the plurality of sub-links satisfying a preset condition, generating the first problem record. The preset condition includes at least one of the retraining times of the plurality of sub-links existing in the sub-link greater than a first preset threshold, and the difference between the retraining times of any two sub-links greater than a second preset threshold, the first problem record indicating the sub-link with a fault.
[0064] In some embodiments, state information collection points can be deployed in each sub-link of the target link. These collection points can be monitoring modules integrated in sub-link related components (such as servers, databases, middleware, etc.), or independent monitoring agents. For example, monitoring software is installed on the key server of each sub-link to collect the running state data of the server in real time. The running information can be system response time, throughput, retraining times, etc., and the person skilled in the art can select the corresponding state information according to the actual situation.
[0065] The state information in the embodiments of the present application can be the retraining number of the sub-link. Link retraining is a key mechanism in high-speed serial communication links, used to dynamically adjust physical layer parameters during link operation to restore or maintain stable signal transmission quality. Link retraining refers to the devices (transmission end and reception end) at both ends of the link re-negotiate and optimize the physical layer parameters (such as signal level, equalization setting, clock recovery, etc.) to solve the problem of communication performance degradation caused by changes in link state (such as signal attenuation, interference, temperature fluctuation). The essence is to dynamically calibrate the link to adapt to changes in the link state, thereby ensuring the reliability and stability of data transmission. The triggering conditions of retraining may, for example, include: signal quality degradation (such as exceeding the error rate), hardware failure or change (such as component aging, poor contact), environmental factor change (such as temperature fluctuation, electromagnetic interference), etc.
[0066] In the pressure test phase, each collection point can collect the retraining number of the sub-link in real time at a predetermined time interval (such as every 10 seconds). The retraining number can reflect the adaptability and stability of the sub-link under the current pressure test scenario.
[0067] The retraining numbers of the plurality of sub-links in a certain time period are counted respectively, and the retraining numbers of the plurality of sub-links after counting are judged to check whether there is a sub-link that meets a preset condition. The preset condition can include that the retraining number of the sub-link is greater than a first preset value, and the difference between the retraining numbers of any two sub-links is greater than a second preset threshold. The first preset threshold can be set according to the normal performance range of the target link and the business tolerance. If the retraining number of a sub-link exceeds the first preset threshold, the sub-link may have a serious fault or continuous performance degradation. The second preset threshold is used to measure the balance of the retraining numbers between the sub-links. If the difference between the retraining numbers of two sub-links is greater than the second preset threshold, it means that the performance of the sub-link is significantly lower than that of other sub-links, and there may be a local fault or an imbalance in configuration. As long as the retraining number of the sub-link meets at least one of the above two conditions, it is determined that the retraining number of the sub-link meets the preset condition, and a corresponding first problem record is generated. The first problem record may, for example, include the time of the problem, the sub-link identifier involved, the specific value of the retraining number, and the type of the preset condition met, etc.
[0068] In the embodiments of the present application, the first preset threshold is set to 10 times. If it is found in the pressure test process that the retraining number of a certain sub-link reaches 13 times, it means that the sub-link may have a problem. The first preset threshold is determined from three aspects of target link reliability boundary, engineering experience, and preventive maintenance. Taking the xGMI link as an example, the link should rarely need to be retrained in a stable state. Under normal load, the ideal case is that the retraining number per hour tends to 0. In large-scale deployment in a data center, when the daily retraining number of the link exceeds 10 times, the failure rate significantly increases (such as memory access error, PCIe device drop), and the first preset threshold is set to 1 time in the present application, so that early warning can be given before the link has a serious failure, thereby avoiding sudden downtime.
[0069] The second preset threshold is set to 5 times, that is, the difference between the maximum retraining number and the minimum retraining number exceeds 5, that is, it is determined that there is a faulty sub-link. For example, the retraining number of sub-link A is 2 times, and the retraining number of sub-link B is 8 times. The difference between them is 6 times, which is greater than the second preset threshold, indicating that the performance difference of the two sub-links when facing pressure is large, and sub-link B may have a physical layer problem (such as PCB trace defect, connector contact failure) or driver configuration error.
[0070] The retraining number is one of the important indicators reflecting the stability of each sub-link in the target link. The present application determines the faulty sub-link through the retraining number of the sub-link itself and the difference between the retraining numbers of different sub-links, so that fast and accurate fault link positioning can be achieved. By focusing on the retraining number of the sub-link in the pressure test stage and generating the first problem record, the stability of the target link under different pressure test stages can be more comprehensively evaluated, which helps to find deeper stability problems.
[0071] The process of obtaining the second problem record of the target link is introduced below in combination with specific embodiments.
[0072] In some embodiments, the log generated in the running process of the target link is analyzed to obtain the second problem record of the target link, including: identifying the log generated in the running process of the target link based on a preset blacklist, determining an object corresponding to a fault string as an abnormal object in the target link, the preset blacklist including a fault string in the log for indicating a fault; generating the second problem record of the target link based on fault information recorded in the log of the abnormal object.
[0073] In some embodiments, the preset blacklist can include a plurality of fault strings. The fault string may, for example, be a fault description keyword of a pilot extracted from historical link fault cases, an error code, a specific log mode, etc., or a fault string supplemented and input by a technician according to experience.
[0074] The log generated by the target link during the test process can be acquired in real time by the log collection module, and the collected log is preprocessed (such as filtering irrelevant log content, merging repeated log content). The preprocessed log is matched with the fault strings in the blacklist one by one, and according to the matching result, the object identification (such as sub-link ID, device MAC address, port number, etc.) involved in the log is extracted, and the object is determined as an abnormal object. The matching method can be exact matching or regular expression fuzzy matching.
[0075] The log generated by the target link during the operation process may include, for example, dmesg / message log, sel log of BMC, etc. The preset blacklist and fault strings corresponding to different logs can be different. For example, the fault strings contained in the dmesg / message log blacklist are shown in Table 1.
[0076] Table 1 dmesg / message log blacklist
[0077]
[0078] The fault strings contained in the sel log blacklist are shown in Table 2.
[0079] Table 2 sel log blacklist
[0080]
[0081] After determining the abnormal object, the log segments of the abnormal object before and after the fault occurs (such as 1 minute before the fault occurs to 1 minute after the fault occurs) can be screened from the log according to the abnormal object identification, and the fault information of the abnormal object can be extracted from the log segment. The fault information may include, for example, fault time, fault duration, association time (such as retraining trigger, hardware reset), etc. The log (such as ADDC log) specially used to record the diagnostic information of the hardware system specific error can also be acquired, and the fault information of the abnormal object can be extracted from the log to generate a second problem record. The second problem record may include, for example, abnormal object representation, fault string, time when the fault object appears, fault type, etc.
[0082] The present application realizes accurate matching of fault strings through a preset blacklist, quickly locates abnormal objects in combination with log analysis, accurately captures known fault patterns, avoids irrelevant log interference, and improves the accuracy of abnormal object positioning.
[0083] The following describes the process of determining the target object to be processed in the target link based on the target performance index in combination with specific embodiments.
[0084] In some embodiments, the target performance indicator is determined from the corresponding performance information that causes the fault from the plurality of performance indicators, and the target object to be processed in the target link is determined based on at least one of a change in the target performance indicator and a change in a correlation between the plurality of target performance indicators, including: taking the performance information corresponding to the time when the problem information occurs as the target performance information according to the time when the problem information occurs; determining a candidate performance indicator from the plurality of performance indicators based on a change in each performance indicator in the target performance information in a predetermined period; the predetermined period includes the time when the problem information occurs; determining at least one target performance indicator from the candidate performance indicators based on analysis of the candidate performance indicators according to a target analysis dimension; the target analysis dimension includes at least one of comparative analysis and trend analysis; analyzing the change in a single performance indicator and / or performing combined analysis on the change in each of the plurality of target performance indicators; and determining the target object that causes the abnormality of the target performance indicator according to the change analysis result and / or the combined analysis result.
[0085] In some embodiments, the performance information of the target link is recorded in chronological order, and the problem information is often generated based on the fault occurring in a specific time point or time period, and it is necessary to align the time to correspond the problem information and the performance information generated at the same time or in the same time period, so as to perform subsequent correlation analysis.
[0086] The performance information corresponding to the time when the problem information occurs can be selected from the collected performance information according to the time when the problem information occurs, and the performance information is taken as the target performance information. For example, if fault 1 occurs in the target link at time point a, the performance information corresponding to time point a is taken as the target performance information. For each performance indicator in the target performance information, the change in the performance indicator in a predetermined period is analyzed, for example, whether the signal strength fluctuates significantly in the period, whether the bit error rate suddenly rises, etc. According to the change of each performance indicator, those indicators that change significantly in the predetermined period or that may be associated with the fault are selected from the plurality of performance indicators as candidate performance indicators. For example, if it is found that the signal strength suddenly starts to fluctuate greatly in the problem occurrence period, the signal strength can be taken as a candidate performance indicator.
[0087] The candidate indicators are analyzed by the target analysis dimension, and at least one performance indicator that has a greater impact on the fault is determined from the candidate indicators as the target performance indicator according to the analysis result, and the object that causes the abnormality of the indicator is determined according to the determined target performance indicator, and the object is determined as the target object.
[0088] Exemplarily, the target analysis dimension can include comparative analysis, trend analysis, etc. The comparative analysis can include horizontal comparison and vertical comparison. Exemplarily, the current value of an index can be compared with its own historical same period, historical average. For example, the current disk I / O is 5 times of the same time in the past. The index of a problematic host can also be compared with the same index of other normal hosts. For example, only the CPU of host A in a cluster is abnormal, and other hosts are normal, so the problem is likely to be in host A itself. The trend analysis refers to analyzing whether the change trend of an index is sudden rise / fall or slow growth / leakage. Different change trends can be caused by different reasons. For example, sudden change is usually related to sudden traffic, code release, restart, etc., and slow change is usually related to resource leakage, data accumulation, etc.
[0089] In the link system, different performance indexes are usually associated with specific hardware components or software modules. By further technical analysis and troubleshooting on the objects related to the performance index, the specific object causing the target performance anomaly is determined, and the object is taken as the target object to be processed. For example, the target performance index is the abnormal rise of the disk I / O read / write delay of the storage array, which can be related to the back-end disk (HDD / SSD), the cache module of the array controller, or the cable connecting the disk backplane. Determining the target object based on the target performance index can include at least one of the following: analyzing the change of a single target performance index; and combining and analyzing the changes of multiple target performance indexes. According to the change analysis result and / or the combined analysis result, the target object causing the target performance index anomaly is determined.
[0090] Analyzing a single target performance index can focus on the dynamic change of a single index, quickly identify related objects and narrow down the troubleshooting range, and be directly associated with the objects related to the change of the index. Multi-index combined analysis can explain the complex causal relationship hidden behind a single index by correlating multiple single indexes, identify indirect root causes, and further improve the accuracy and comprehensiveness of target object determination.
[0091] A large amount of performance information will be generated during the stress test phase. The application can effectively reduce the data analysis range by preliminarily screening the performance information based on the time of the problem information, and narrow down the analysis range from the entire test process to the time range closely related to the problem information, effectively reducing the amount of data to be processed. During the stress test process, some performance indexes may appear temporary abnormal fluctuations due to accidental factors or external interference. The candidate performance indexes can be further analyzed based on the target analysis dimension to exclude performance indexes that are abnormal in the problem period but are not key factors, exclude accidental factors, and improve the accuracy of target performance index determination.
[0092] The following will be described in combination withFigure 4 And in specific embodiments, a process of generating processing information for a target object according to a target performance indicator and problem information is introduced.
[0093] Figure 4 A schematic diagram of a process of generating processing information for a target object according to a target performance indicator and problem information according to embodiments of the present application is shown.
[0094] As shown in Figure 4 the embodiment, the process of generating processing information for a target object according to a target performance indicator and problem information includes: determining a plurality of processing steps for processing the fault based on the target performance indicator and the problem information; obtaining predefined information of the plurality of processing steps respectively, and determining an execution order of the plurality of processing steps according to the predefined information, the predefined information being determined based on a maintenance success rate and a maintenance effect of the processing steps in historical maintenance data, and used to measure the processing effectiveness of the processing steps on the fault; combining the plurality of processing steps based on the execution order to obtain the processing information for the target object; and wherein the processing information is used to guide a user to execute the plurality of processing steps based on the execution order.
[0095] In some embodiments, the processing steps that can be used to solve the fault can be found in a preset step library according to the target performance indicator and the problem information respectively. The processing steps that can be used to solve the fault can also be found in the preset step library based on a problem scenario obtained by combining the target performance indicator and the problem information.
[0096] For example, the processing steps related to improving or restoring the target performance indicator can be filtered out from a preset database according to the target performance indicator. The processing steps suitable for the type of fault or having corresponding characteristics can be found in the preset step library according to the classification result of the problem information. The processing steps found based on the target performance indicator and the problem information are comprehensively filtered to remove repeated or irrelevant steps, and a set of processing steps for processing the fault is obtained. Alternatively, a problem scenario can be constructed according to the combination of the target performance indicator and the problem information, the constructed problem scenario is matched with the scenario tags in the preset step library, and the coarse-grained steps suitable for the current problem scenario are found.
[0097] In some embodiments, for each determined processing step, predefined information of the processing step is obtained, the predefined information being determined based on a maintenance success rate and a maintenance effect of the processing step in historical maintenance data, and used to measure the processing effectiveness of the processing step on the fault. For example, the maintenance information of the processing step can be determined according to the performance of the processing step in the historical maintenance data. The processing success rate can be determined by counting the proportion of similar faults that are successfully solved by the processing step. The maintenance effect can be evaluated from the speed of solving the fault, the degree of improvement of system performance, etc.
[0098] The effectiveness of each processing step for fault processing is measured according to predefined information, and a processing step with a high repair success rate and good repair effect is usually selected first. The predefined information can be presented in a numerical manner, and the larger the value of the predefined information of a processing step, the higher the effectiveness of the processing step for fault processing, and the earlier the execution order in the processing information.
[0099] The multiple processing steps are sequentially arranged according to the execution order, and the processing information is generated based on the arranged processing steps. The processing information may, for example, include the specific operation method of each processing step, the required tools, matters needing attention, etc., so as to guide the user to execute the multiple processing steps based on the execution order.
[0100] The embodiments of the present application determine the processing steps through the target performance indicators and the problem information, and determine the execution order according to the predefined information of the historical repair data, thereby avoiding blind attempts of the user when processing faults, helping to realize the standardization of fault processing, and improving the accuracy and reliability of fault processing.
[0101] In some embodiments, the predefined information value can be obtained based on historical data, expert experience or a machine learning model, and can represent the effectiveness of the step in solving a specific fault under a specific fault scenario (i.e., a combination of target performance indicators and problem information). For example, the predefined information value of processing step A under indicator 1 and problem information X is 0.8, and the value under indicator 2 and problem information Y is 0.5, etc.
[0102] The determination method of the predefined information value can include: obtaining multiple historical processing data; performing modeling analysis on the multiple historical processing data to obtain initial information corresponding to each processing step under multiple fault scenarios; adjusting the initial information corresponding to each processing step under the multiple fault scenarios based on engineering experience to obtain multiple predefined information of each processing step under the multiple fault scenarios; periodically collecting newly added processing data within a specified time range, and updating the multiple predefined information of each processing step under the multiple fault scenarios based on the newly added processing data.
[0103] In some embodiments, the historical processing data can include fault information (such as target performance indicators, problem information, etc.), processing information (processing steps taken, processing step execution order, execution result of each step, final solution state of the fault, etc.), etc. A machine learning model or a statistical model can be used to analyze the historical processing data to calculate the initial information of each processing step under a specific fault scenario, wherein the fault scenario can be determined by the combination of the target performance indicators and the problem information.
[0104] After determining the initial information for each step, the initial information is adjusted based on engineering experience to obtain the predefined information for each step. Engineering experience refers to the tacit knowledge about technical decisions and risk trade-offs accumulated by engineers (or operations and maintenance experts) in the long-term process of solving practical problems. Engineering experience cannot be fully quantified by data, but it can effectively make up for the blind spots and defects of pure data models and improve the accuracy of predefined information.
[0105] During the operation of the target link, new processing data may be continuously generated. When a fixed period (such as weekly or monthly) is reached or when the amount of new processing data reaches a certain threshold, the update process is automatically triggered. Based on the newly added processing data, the predefined information of each processing step under different fault scenarios is updated so that fault handling can be carried out in a timely manner based on the updated predefined information, thereby improving the adaptability of the fault handling method.
[0106] For example, updating the predefined information of each processing step under different fault scenarios based on the newly added processing data can include two methods: incremental update or full update. Incremental update refers to adjusting the original predefined information value only based on the new maintenance data, while full update refers to merging the newly added maintenance data with the historical maintenance data and regenerating the predefined value of the processing step based on the merged maintenance data.
[0107] This application's embodiments adjust the initial predefined information based on engineering experience. While ensuring the objectivity of the predefined information values, it also references domain knowledge and risk control, effectively improving the accuracy of the predefined information. The introduction of periodically updated knowledge enables the determination of predefined information to have dynamic adaptability, thereby ensuring that fault handling decisions are always synchronized with the dynamically changing external environment.
[0108] In some embodiments, the processing steps may include at least one of reinstalling specified hardware of the target object, replacing specified hardware of the target object, or updating specified firmware of the target object. Specified hardware may include, for example, a hard drive, a motherboard, etc., and specified firmware may include, for example, the motherboard's BIOS, the router's firmware, the graphics card's firmware, etc.
[0109] Taking the specified hardware as hard drive and the specified firmware as hard drive backplane firmware as an example, for example... Figure 4 The process of generating the processing information shown is further explained. Let's assume a server (target object) in the target link experiences a failure scenario where the target performance information indicates "disk I / O latency" and the problem information indicates "disk SMART error alarm". By querying the direct database, we find that the processing steps for this failure scenario (high I / O latency, SMART error) are: reinstalling the hard drive, replacing the hard drive, and updating the hard drive backplane firmware. The predefined information value corresponding to each processing step is N. (重新安装硬盘) =0.6, N (更换硬盘) =0.9, N(更新硬盘背板固件) =0.2. The execution order of the processing steps is: replacing the hard disk, reinstalling the hard disk, and updating the hard disk backplane firmware. Processing information for the server is generated based on the execution order of the processing steps and the specific execution operations of each processing step, so as to guide the field engineer to process the target object according to the execution order of the processing steps and the execution operations in the processing information.
[0110] In some embodiments, the target object can involve multiple faults or one fault involves multiple target performance indicators and multiple problem information. The following describes the process of determining the execution order between multiple processing steps according to pre-defined information in combination with specific embodiments. Figure 5 and specific embodiments.
[0111] Figure 5 The schematic diagram of acquiring pre-defined information of multiple processing steps respectively and determining the execution order between multiple processing steps according to the pre-defined information according to the embodiments of the present application is shown.
[0112] As shown in Figure 5 , the process of determining the execution order between multiple processing steps includes: in the case where a fault involves multiple target performance indicators and multiple problem information, acquiring pre-defined information of multiple processing steps for combinations of multiple target performance indicators and problem information respectively; summing up multiple pre-defined information values of the same processing step, and determining the execution order between multiple processing steps based on the sum results of each processing step; wherein the execution step with the larger sum result has a higher execution order.
[0113] In some embodiments, for each processing step, the pre-defined information value of the step for each combination of target performance indicators and problem information is acquired, and the pre-defined information values of different processing steps under different combinations are different. For each processing step, sum up all the pre-defined information values of the step for different combinations (i.e. fault scenarios), and sort all the processing steps according to the sum results of each processing step. The processing step with the larger sum result has a higher priority and a higher execution order.
[0114] Taking the specified hardware as CPU and motherboard and the specified firmware as BIOS as an example, the process of determining the execution order between multiple processing steps as shown in Figure 5 is further described.
[0115] Suppose the fault involves fault scenario z1 (combination of performance indicator x1 and fault information y1) and fault scenario z2 (combination of performance indicator x2 and fault information y2), then the pre-defined information values of each processing step under the two scenarios are acquired respectively, as shown in Table 3.
[0116] Table 3 Correspondence between fault scenarios and processing steps
[0117]
[0118] The sum of the predefined information values of the processing steps under the fault scenario z1 and the fault scenario z2 is summed up respectively, and the sum of the predefined information values of the four processing steps is 1.2, 0.9, 1, and 0.1 in turn. The execution order between the processing steps is: reinstalling the CPU, replacing the motherboard, replacing the CPU, and updating the BIOS version.
[0119] Suppose that the fault scenario of a server (target object) in the target link is that the target performance information indicates "hard disk IO delay" and the problem information indicates "hard disk SMART error alarm". By querying the direct library, it is found that the processing steps under the fault scenario (high IO delay, SMART error) are reinstalling the hard disk, replacing the hard disk, and updating the hard disk backplane firmware. The predefined information values corresponding to the processing steps are N (重新安装硬盘) =0.6, N (更换硬盘) =0.9, and N (更新硬盘背板固件) =0.2. Therefore, the execution order of the processing steps is: replacing the hard disk, reinstalling the hard disk, and updating the hard disk backplane firmware. Based on the execution order, processing information for the server is generated, which includes the execution order of the processing steps and the specific execution operations of each processing step, so as to guide the field engineers to process the target object according to the execution order of the processing steps and the execution operations in the processing information.
[0120] In a complex fault, a single performance indicator or problem information may not reflect the cause of the fault. By integrating the predefined information of the processing steps under multiple target performance indicators and problem information, the limitation of relying only on a single performance indicator or problem information is effectively avoided, the comprehensiveness and accuracy of the execution order determination are effectively improved, and the processing information generation demand under the complex fault is met. Steps with high priority usually have a greater impact on solving the fault, and preferentially executing these steps can faster alleviate the problem, reduce system downtime or performance decline period, and improve operation efficiency and service reliability.
[0121] In some embodiments, the fault processing method of the present application can further include: in response to the execution of at least one processing step in the processing information being completed, detecting the target object; in the case that the detection result of the target object meets the preset condition, terminating the process of processing the fault based on the processing information; in the case that the detection result of the target object does not meet the preset condition, continuing to execute the next processing step in the processing information until the detection result of the target object meets the preset condition.
[0122] For example, in response to the completion of one processing step in the processing information, a detection procedure is triggered to detect the target object (e.g., a database service) determined previously. The detection can be to obtain performance indicators and status information that can reflect whether the fault has been recovered. For example, the database connection status is checked, a test query statement is executed, and the response time indicator is checked to see if it has returned to normal. The detection result is compared with a preset condition, which is a predefined explicit criterion indicating that the fault has been resolved. If the preset condition is met, it means that the fault has been successfully resolved, and the execution of the processing information is terminated, and all subsequent steps are not executed. If the preset condition is not met, the next processing step in the processing information is executed.
[0123] According to another aspect of the embodiments of the present application, an electronic device is provided, which includes a memory and a processor. The memory stores a computer program. The processor is configured to execute the steps of any of the method embodiments described above by using the computer program.
[0124] In an example embodiment, the electronic device described above can further include a transmission device and an input / output device. The transmission device is connected to the processor. The input / output device is connected to the processor.
[0125] The specific examples in the embodiments described above can refer to the examples described in the embodiments and example implementations described above, and will not be described herein again.
[0126] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, which includes a stored program. When the program is executed, the steps of any of the method embodiments described above are performed.
[0127] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0128] According to another aspect of the embodiments of the present application, a computer program product is provided, which includes a computer program / instruction containing program codes for performing the method shown in the flowchart. In such an embodiment, reference is made to the embodiments and example implementations described above. Figure 6The computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit 601, it performs various functions provided in the embodiments of this application. The sequence numbers of the embodiments of this application above are merely for description and do not represent the superiority or inferiority of the embodiments.
[0129] Figure 6 A schematic block diagram of a computer system for an electronic device according to an embodiment of this application is shown. Figure 6 As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 602 or programs loaded from storage section 908 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output interface 605 (I / O interface) is also connected to the bus 604.
[0130] The following components are connected to the input / output interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a local area network card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 610 as needed so that computer programs read from it can be installed into the storage section 608 as needed.
[0131] In particular, according to embodiments of the present application, the processes described in the various method flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 609, and / or installed from the removable media 611. When the computer program is executed by the central processing unit 601, various functions defined in the system of the present application are executed.
[0132] It should be noted that, Figure 6 The computer system 600 of the electronic device shown is only an example and should not impose any limitation on the functions and use range of embodiments of the present application.
[0133] According to still another aspect of embodiments of the present application, a link fault processing apparatus is also provided, comprising: a testing module configured to perform stress testing on a target link formed by a plurality of processors connected together to obtain performance information of the target link under a predetermined stress condition; the performance information comprising at least one performance indicator corresponding to at least one of a data transmission dimension, a request processing dimension, and a resource consumption dimension; an obtaining module configured to obtain problem information generated by the target link under the predetermined stress condition, the problem information indicating a fault occurring in the stress testing process of the target link; a determining module configured to determine at least one target performance indicator causing the fault from the performance information based on the problem information, and determine a target object to be processed in the target link based on at least one of a change in the target performance indicator and an associated change between a plurality of target performance indicators; the target object comprising at least one of any processor in the target link and any functional module in any processor; and a generating module configured to generate processing information for the target object according to the target performance indicator and the problem information, so as to process the fault based on the processing information.
[0134] It should be noted that the testing module in this embodiment can be configured to perform the above operation S210, the obtaining module in this embodiment can be configured to perform the above operation S220, the determining module in this embodiment can be configured to perform the above operation S230, and the generating module in this embodiment can be configured to perform the above operation S240.
[0135] Optionally, the test module comprises: a test submodule, configured to perform stress testing on the target link according to a plurality of test periods to obtain a plurality of test results, wherein each test period comprises at least one stress testing phase, and stress testing phases in the same test period have different stress values, the stress value representing an amount of data applied to the target link per unit time; and a first determination submodule, configured to determine performance information of each object in the target link according to the plurality of stress test results, the performance information representing running states of the plurality of objects in the target link, and the object including at least one of any processor in the target link or any functional module in any processor.
[0136] Optionally, the acquisition module comprises: a second determination submodule, configured to determine a first problem record of the target link according to state information of each object in the target link in any stress testing phase; an analysis submodule, configured to analyze logs generated in the running process of the target link to obtain a second problem record of the target link; and an association submodule, configured to associate the first problem record and the second problem record based on time association and / or event association to obtain problem information of the target link in any stress testing phase.
[0137] Optionally, the second determination submodule further comprises: an acquisition unit, configured to acquire state information of each of the plurality of sublinks in the stress testing phase respectively, the state information representing a number of retraining of the sublink, the retraining being an operation triggered by the target link in the test process to cope with a specific situation, and the specific situation including at least one of a decrease in signal quality, a hardware fault or change, or an environmental factor change; and a generation unit, configured to generate the first problem record in response to the number of retraining of the plurality of sublinks satisfying a preset condition, the preset condition including at least one of the plurality of sublinks having a sublink with a number of retraining greater than a first preset threshold and a difference between the number of retraining of any two sublinks being greater than a second preset threshold, and the first problem record indicating a sublink with a fault.
[0138] Optionally, the analysis submodule further comprises: an identification unit, configured to identify the logs generated in the running process of the target link based on a preset blacklist, and determine an abnormal object in the target link corresponding to a fault string in the logs, the preset blacklist including the fault string in the logs for indicating a fault; and a generation unit, configured to generate the second problem record of the target link based on fault information recorded in the logs of the abnormal object.
[0139] Optionally, the determining module further comprises: a target performance information determining sub-module, configured to take performance information corresponding to the time when the problem information occurs as target performance information according to the time when the problem information occurs; a candidate performance index determining sub-module, configured to determine candidate performance indexes from the plurality of performance indexes based on changes of each performance index in the target performance information in a predetermined time period; the predetermined time period comprises the time when the problem information occurs; an analyzing sub-module, configured to analyze the candidate performance indexes based on a target analysis dimension, and determine at least one target performance index from the candidate performance indexes; the target analysis dimension comprises at least one of contrast analysis and trend analysis; and a target object determining sub-module, configured to analyze changes of a single target performance index and / or to perform combined analysis on changes of a plurality of target performance indexes respectively, and determine a target object causing abnormality of the target performance index according to the change analysis result and / or the combined analysis result.
[0140] Optionally, the generating module further comprises: a processing step determining sub-module, configured to determine a plurality of processing steps for processing the fault based on the target performance index and the problem information; an execution order determining sub-module, configured to acquire predefined information of the plurality of processing steps respectively, and determine an execution order of the plurality of processing steps according to the predefined information; the predefined information is determined based on a maintenance success rate and a maintenance effect of the processing steps in historical maintenance data, and is used to measure processing effectiveness of the processing steps on the fault; and a combining sub-module, configured to combine the plurality of processing steps based on the execution order, and obtain processing information for the target object; wherein the processing information is used to guide a user to execute the plurality of processing steps based on the execution order.
[0141] Optionally, the execution order determining sub-module further comprises: an acquisition unit, configured to acquire predefined information of combinations of the plurality of processing steps for the plurality of target performance indexes and the problem information respectively in a case where the fault involves the plurality of target performance indexes and the plurality of problem information; and a summation unit, configured to sum a plurality of predefined information values of the same processing step, and determine the execution order between the plurality of execution steps based on summation results of the respective processing steps; wherein the execution step with a larger summation result has a more advanced execution order.
[0142] Obviously, those skilled in the art should understand that each module or each step of the above-mentioned embodiments of the present application can be realized by a general computing device, which can be concentrated on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in an order different from here, or they can be manufactured into each integrated circuit module respectively, or multiple modules or steps among them can be manufactured into a single integrated circuit module to realize. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.
[0143] The above merely provides preferred embodiments of the present application, and is not used to limit the present application. Those skilled in the art can make various modifications and variations without departing from the spirit of the present application. Any modifications, equivalent replacements, improvements, etc., made within the principles of the present application should be included in the protection scope of the present application.
Claims
1. A link failure handling method, characterized in that, include: A stress test is performed on a target link formed by multiple processors to obtain the performance information of the target link under predetermined stress conditions; wherein, the processor includes multiple functional modules, and the performance information includes at least one performance index corresponding to at least one of the following dimensions: data transmission dimension, request processing dimension, and resource consumption dimension. Obtain problem information generated by the target link under the predetermined pressure condition, the problem information indicating the failure of the target link during the pressure test; Based on the problem information, at least one target performance indicator causing the fault is determined from the performance information, and the target object to be processed in the target link is determined based on at least one of the changes in the target performance indicator and the correlation changes among multiple target performance indicators; the target object includes any processor in the target link and at least one of any functional modules in any processor; Based on the target performance indicators and the problem information, processing information is generated for the target object so as to process the fault based on the processing information.
2. The method according to claim 1, characterized in that, The stress test performed on the target link formed by connecting multiple processors to obtain the performance information of the target link under a predetermined stress includes: The target link is stress tested in multiple test cycles to obtain multiple test results. Each test cycle includes at least one stress test phase. The stress values of the stress test phases in the same test cycle are different. The stress value represents the amount of data applied to the target link per unit time. Based on the results of multiple stress tests, the performance information of each object in the target link is determined. The performance information represents the running status of multiple objects in the target link. The objects include any processor in the target link and at least one of any functional modules in any processor.
3. The method according to claim 2, characterized in that, The step of obtaining problem information generated by the target link under predetermined pressure conditions includes: In any stress test phase, the first problem record of the target link is determined based on the status information corresponding to each object in the target link; The logs generated during the operation of the target link are analyzed to obtain the second problem record of the target link; Based on time correlation and / or event correlation, the first problem record and the second problem record are correlated to obtain the problem information of the target link in any stress test phase.
4. The method according to claim 3, characterized in that, The target link includes multiple parallel sub-links. Based on the status information of each object in the target link, the first problem record of the target link is determined, including: The status information of each of the multiple sub-links during the stress test phase is obtained respectively. The status information indicates the number of retraining times of the sub-link. The retraining is an operation triggered by the target link during the test to cope with specific situations. The specific situations include at least one of signal quality degradation, hardware failure or change, and environmental factor change. In response to the retraining count of multiple sub-links meeting a preset condition, a first problem record is generated; The preset conditions include that among the multiple sub-links, there is a sub-link whose number of retraining attempts is greater than a first preset threshold, and the difference between the number of retraining attempts of any two sub-links is greater than at least one of the second preset thresholds. The first problem record indicates the sub-link with a fault.
5. The method according to claim 3, characterized in that, The parsing of logs generated during the operation of the target link yields a second problem record for the target link, including: Based on a preset blacklist, the logs generated during the operation of the target link are identified, and the objects corresponding to the fault strings are identified as abnormal objects in the target link. The preset blacklist includes fault strings in the logs used to indicate faults. Based on the fault information recorded in the log of the abnormal object, a second problem record of the target link is generated.
6. The method according to claim 1, characterized in that, Based on the problem information, at least one target performance indicator causing the fault is determined from the performance information, and the target object to be processed in the target link is determined based on at least one of the changes in the target performance indicator and the correlation changes among multiple target performance indicators, including: Based on the time when the problem information appears, the performance information corresponding to that time is used as the target performance information; Based on the changes of each performance indicator in the target performance information over a predetermined period, candidate performance indicators are determined from the plurality of performance indicators; the predetermined period includes the time when the problem information occurs. The candidate performance indicators are analyzed based on the target analysis dimensions to determine at least one target performance indicator; the target analysis dimensions include at least one of comparative analysis and trend analysis. Analyze the changes in a single target performance index and / or combine the changes in multiple target performance indicators; Based on the analysis results of the changes and / or the combined analysis results, the target object causing the abnormality of the target performance index is determined.
7. The method according to claim 1, characterized in that, The step of generating processing information for the target object based on the target performance indicators and the problem information includes: Based on the target performance indicators and the problem information, multiple processing steps for handling the fault are determined; Predefined information for each of the multiple processing steps is obtained, and the execution order of the multiple processing steps is determined based on the predefined information. The predefined information is determined based on the repair success rate and repair effect of the processing steps in historical repair data, and is used to measure the effectiveness of the processing steps in handling the fault. The multiple processing steps are combined based on the execution order to obtain processing information for the target object; wherein, the processing information is used to guide the user to execute the multiple processing steps based on the execution order.
8. The method according to claim 7, characterized in that, The processing steps include at least one of reinstalling the specified hardware of the target object, replacing the specified hardware of the target object, and updating the specified firmware of the target object.
9. The method according to claim 7, characterized in that, The step of acquiring predefined information for the plurality of processing steps and determining the execution order of the plurality of processing steps based on the predefined information includes: In cases where the fault involves multiple target performance indicators and multiple problem information, predefined information for the combination of multiple target performance indicators and problem information is obtained for each of the multiple processing steps. Multiple predefined information values for the same processing step are summed, and the execution order among the multiple execution steps is determined based on the summation results of each processing step; wherein, the execution step with the larger summation result is executed earlier.
10. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 9.