Node inspection method and device, computer equipment and storage medium
By preprocessing node performance data and generating inspection instructions, the lag and inefficiency of node inspection are solved, efficient and intelligent fault detection and repair are achieved, and the stability and reliability of the system are ensured.
Patent Information
- Application Number
- CN202510672270.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
The existing node inspection mechanism has lag and inefficiency problems, and environmental abnormalities cannot be discovered in time, resulting in further deterioration of failures and affecting business continuity.
By obtaining the original performance data of the node for preprocessing, identifying abnormal performance items and associated data, extracting feature information of alarm information, and performing feature vector conversion to generate inspection instructions to achieve accurate node inspection.
Improve fault response efficiency, enhance preventive maintenance capabilities, ensure the stability and reliability of the system, and reduce the need for manual intervention.
Smart Images

Figure CN120498959A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer network operation and maintenance, and in particular to a node inspection method, device, computer equipment and storage medium. Background Art
[0002] With the continuous expansion of server clusters and the rapid development of related storage technologies, the connection topology between nodes has become increasingly complex, forming a vast network. This network not only includes a large number of hardware devices but also a complex interweaving of software modules. While this highly integrated environment greatly improves data processing capabilities and system flexibility, it also significantly increases potential points of failure. The larger the cluster, the more components requiring maintenance multiply. A small failure in any link can trigger a chain reaction throughout the cluster, resulting in service interruption or performance degradation.
[0003] Currently, node health inspection mechanisms rely heavily on user-triggered initiation, but this triggering mode exhibits significant lag. Cluster health status is typically only obtained after an environmental problem has occurred, making it difficult to detect environmental anomalies in a timely manner. Furthermore, current node health inspections are often comprehensive, resulting in unnecessary inspections of many healthy modules, reducing the efficiency of node health inspections. In distributed storage systems, traditional inspection methods are typically based on fixed time intervals or manual initiation. This approach lacks accurate perception of the system's real-time status and the ability to proactively respond. This can lead to excessive and unnecessary inspections when the system is operating normally, wasting system resources. Furthermore, when performance issues arise, problems are often not discovered in a timely manner because the inspection cycle has not yet expired, further exacerbating the problem and impacting business continuity. Summary of the Invention
[0004] The present application provides a node inspection method to at least solve the problem of improving inspection efficiency in related technologies.
[0005] This application provides a node inspection method, including:
[0006] Obtain the original performance data of the node, pre-process the original performance data, and obtain the data to be analyzed;
[0007] Determine the abnormal performance items of the node based on the data to be analyzed and obtain associated data corresponding to the abnormal performance items;
[0008] Obtain node alarm information and extract preset feature information in the node alarm information;
[0009] By converting the abnormal performance items, associated data and preset characteristic information into characteristic vectors, the corresponding characteristic vectors are obtained and then spliced together to form a comprehensive characteristic vector. In response to the abnormality of the comprehensive characteristic vector, a node inspection instruction is triggered.
[0010] Inspect nodes according to node inspection instructions.
[0011] This application also provides a node inspection device, including:
[0012] The acquisition module is used to obtain the original performance data of the node, pre-process the original performance data of the node, and obtain the data to be analyzed;
[0013] The abnormal data calculation module is used to perform statistical calculations on the data to be analyzed, compare the calculation results with the preset threshold value, and obtain the node performance abnormality data and the node performance data corresponding to the node performance abnormality data;
[0014] The alarm information acquisition module is used to obtain node alarm information and extract preset feature information in the node alarm information;
[0015] An instruction generation module is used to convert abnormal performance items, associated data, and preset characteristic information into feature vectors to obtain corresponding feature vectors, and then concatenate them to form a comprehensive feature vector. In response to an abnormality in the comprehensive feature vector, a node inspection instruction is triggered.
[0016] The execution module is used to inspect the corresponding nodes according to the node inspection instruction.
[0017] The present application also provides a computer device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned node inspection methods when executing the computer program.
[0018] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned node inspection methods are implemented.
[0019] This application obtains and preprocesses raw node performance data, ensuring data quality and accuracy, providing a reliable foundation for subsequent analysis. The system identifies abnormal performance items based on the data to be analyzed and further obtains associated data related to these abnormal performance items, helping to accurately locate the root cause of the problem. Node alarm information is extracted and pre-defined feature information is identified from it, allowing the alarm information to be combined with the node's performance data, thereby enhancing the depth and breadth of analysis. By embedding and splicing the pre-defined feature information, abnormal performance items, and associated data, node-specific inspection instructions are generated. These instructions are not only targeted and effective, but also help quickly identify and resolve potential performance issues through precise inspection operations, ensuring the healthy operation of the system. This integrated process significantly improves the automation and intelligence level of the system, reduces the need for manual intervention, and enables timely and effective inspection and repair when performance anomalies occur, thereby ensuring system stability and reliability. This integrated data analysis and inspection instruction generation method not only improves fault response efficiency but also enhances preventive maintenance capabilities, helping to improve the long-term operation and maintenance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of a node inspection method provided in an embodiment of the present application.
[0022] Figure 2 This is a structural block diagram of a passenger counting device provided in an embodiment of the present application.
[0023] Figure 3 A timing diagram of a node inspection method provided in an embodiment of the present application.
[0024] Figure 4 This is a diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0025] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0027] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0028] In one embodiment, Figure 1 As shown, a node inspection method is provided, comprising the following steps:
[0029] S100: Obtain the original performance data of the node, pre-process the original performance data, and obtain the data to be analyzed.
[0030] Raw node performance data refers to performance metrics collected from the node without any processing or analysis. This data can include CPU utilization, memory usage, disk I / O performance, network traffic, and more. It is a direct source of data reflecting the node's operating status. Preprocessing involves cleaning, converting, and organizing raw data to make it more suitable for subsequent analysis and processing. Preprocessing methods include standardization and normalization. The node is a server node, meaning a node in the network. This node is a server.
[0031] Specifically, by obtaining the raw performance data of the nodes and performing preprocessing, the quality of the data and the accuracy of subsequent analysis can be significantly improved. Raw performance data often contains noise, missing values, or inconsistent information, which can lead to erroneous conclusions or system misjudgments when used directly in analysis. By cleaning, denoising, and standardizing this data, the system can ensure data consistency and reliability, providing a solid foundation for subsequent performance analysis, anomaly detection, and fault diagnosis. Preprocessed data is more suitable for efficient analysis, reducing unnecessary redundant information, avoiding delays or errors caused by data issues, and also improving the speed and efficiency of data processing.
[0032] S200: According to the data to be analyzed, determine the abnormal performance items of the node and obtain the associated data corresponding to the abnormal performance items.
[0033] Abnormal performance items refer to performance indicators or system behaviors identified during analysis that deviate from the normal, preset range. Associated data refers to other data related to the abnormal performance items. This data can help us better understand the root cause of the problem and provide further diagnostic clues. For example, if a node's CPU usage is abnormal, associated data may include information about the node's load, running processes, and network traffic.
[0034] Specifically, by analyzing the data to be analyzed, the system can promptly identify nodes experiencing performance deviations and accurately determine which performance indicators are outside normal ranges. By acquiring relevant data related to abnormal performance items, the system can fully understand the background of the problem from multiple dimensions. For example, when a node's CPU is abnormal, the relevant data may indicate that it is due to excessive resource consumption by a specific process or excessive system load. This detailed analysis enables more accurate problem diagnosis, thereby improving troubleshooting efficiency.
[0035] S300: Obtain node alarm information and extract preset feature information in the node alarm information.
[0036] Node alarm information refers to warning signals sent by nodes in the system, indicating that a performance indicator or system status is abnormal or deviates from normal values. Preset feature information refers to key feature data extracted from node alarm information and derived according to pre-set standards or rules.
[0037] Specifically, by acquiring node alarm information and extracting pre-set feature information, the system's ability to identify and respond to potential problems can be effectively enhanced. This accelerates fault response and enhances the system's ability to predict and prevent faults, helping to improve the system's overall stability and maintainability.
[0038] S400: By performing feature vector conversion on abnormal performance items, associated data and preset feature information, corresponding feature vectors are obtained, and they are spliced together to form a comprehensive feature vector. In response to the abnormality of the comprehensive feature vector, a node inspection instruction is triggered.
[0039] Specifically, by converting abnormal performance items, associated data, and preset feature information into feature vectors, a comprehensive feature vector is formed and compared with the preset risk threshold, enabling accurate judgment of the node's health status. When the node status exceeds the preset threshold, the system automatically triggers an inspection instruction to conduct troubleshooting and repair. By effectively integrating information from multiple data sources, the accuracy of anomaly detection is improved, reducing the possibility of false positives and missed reports. At the same time, intelligent automatic inspection triggering reduces the pressure of manual monitoring and improves operation and maintenance efficiency.
[0040] S500: Inspect the node according to the node inspection instruction.
[0041] Among them, inspections include: hardware inspections, network inspections, and software inspections.
[0042] Specifically, inspecting nodes according to node inspection instructions can significantly improve system operation and maintenance efficiency and problem response speed. This not only speeds up fault detection and resolution, but also improves system stability and reliability, reducing system operation and maintenance costs and the need for manual intervention.
[0043] By acquiring and preprocessing raw node performance data, the system ensures data quality and accuracy, providing a reliable foundation for subsequent analysis. Based on the data to be analyzed, the system identifies abnormal performance items and further obtains related data, helping to accurately locate the root cause of the problem. Node alarm information is extracted and pre-defined signature information is identified, allowing the alarm information to be combined with the node's performance data, thereby enhancing the depth and breadth of analysis. By embedding and combining the pre-defined signature information, abnormal performance items, and related data, node-specific inspection instructions are generated. These instructions are not only targeted and effective, but also enable precise inspection operations to quickly identify and resolve potential performance issues, ensuring healthy system operation. This integrated process significantly enhances the system's automation and intelligence, reduces the need for manual intervention, and enables timely and effective inspection and remediation when performance anomalies occur, thereby ensuring system stability and reliability. This integrated data analysis and inspection instruction generation approach not only improves fault response efficiency but also enhances preventive maintenance capabilities, contributing to improved long-term system operation and maintenance effectiveness.
[0044] In one embodiment, obtaining raw performance data of a node and preprocessing the raw performance data to obtain data to be analyzed include:
[0045] At each sampling interval, the raw performance data of the node is collected;
[0046] identifying an original format of the original performance data and converting the original format of the original performance data into a preset format;
[0047] The original unit of the raw performance data is identified and the original unit of the raw performance data is converted into a preset unit.
[0048] The preset formats include at least one of the following: JSON format, CSV format, and XML format; and the preset units include at least one of the following: bytes or milliseconds.
[0049] Specifically, the preprocessing of the original performance data of the node can greatly improve the efficiency and accuracy of data analysis. First, the setting of the sampling interval enables the system to regularly collect the performance data of the node to ensure that the status of the node can be reflected in real time. For example, if the sampling interval is 1 minute, the system can collect the latest performance status of the node within every minute, helping administrators to detect performance problems in a timely manner. Since different monitoring tools or systems may use different data formats, directly using the data in the original format for analysis will bring compatibility issues. By converting the data into a preset unified format, the system can ensure the consistency of the data and facilitate subsequent processing and analysis. For example, the system uses XML format to record data, and the analysis tool only supports CSV format. The preprocessing step can convert the data into CSV format, making subsequent analysis smooth. Similarly, identifying and converting the original units to preset units helps ensure that various performance indicators can be compared and analyzed under a unified measurement unit when conducting large-scale analysis.
[0050] In one embodiment, obtaining node alarm information and extracting preset characteristic information from the node alarm information include:
[0051] Determine the alarm type based on the node alarm information;
[0052] Determine the alarm level corresponding to the alarm type based on the node alarm information;
[0053] Summarize the alarm types and alarm levels as preset feature information.
[0054] Among them, alarm types include: hardware alarms, software alarms and network data alarms. Hardware alarms include: disk failure and memory errors; software alarms include: service interruption and file system errors; network data alarms include: data loss and data inconsistency; alarm levels include: severe alarms and general alarms. Hardware alarms, service interruption and data loss are severe alarms; file system errors and data inconsistency are general alarms.
[0055] Specifically, by summarizing pre-set characteristic information based on alarm type and level, the system can quickly identify the nature and severity of the problem upon receiving an alarm, helping operations personnel prioritize critical issues. Based on the level and type of the alarm, the system can more efficiently allocate operations resources.
[0056] In one embodiment, feature vector conversion is performed on abnormal performance items, associated data, and preset feature information to obtain corresponding feature vectors, which are then concatenated to form a comprehensive feature vector. In response to an abnormality in the comprehensive feature vector, a node inspection instruction is triggered, including:
[0057] Obtain historical data of the node, which includes at least one of the following: historical performance data, number of historical alarm information, and number of historical fault cases;
[0058] Determining a risk threshold vector based on the historical data, wherein the risk threshold vector includes at least one risk element, and the at least one risk element represents at least one of the following: a historical performance threshold corresponding to the historical performance data, an alarm quantity threshold corresponding to the number of historical alarm information, and a case quantity threshold corresponding to the number of historical fault cases;
[0059] Performing feature vector conversion on the abnormal performance item, the associated data, and the preset feature information to obtain a feature vector of the abnormal performance item, a feature vector of the associated data, and a feature vector of the preset feature information;
[0060] The abnormal performance item feature vector, the associated data feature vector and the preset feature information feature vector are concatenated to obtain a comprehensive feature vector;
[0061] Compare the elements in the comprehensive feature vector with the corresponding risk elements to determine whether the node is abnormal;
[0062] In response to a node abnormality, a node inspection instruction is triggered.
[0063] Risk elements are indicators used to assess node health. Each element represents a specific risk metric. Historical performance data reflects the node's performance over a period of time (such as CPU utilization, memory usage, network bandwidth, disk I / O, etc.). Statistical analysis of historical performance data can determine the node's normal performance range and calculate corresponding performance thresholds based on historical data. These thresholds, as part of the risk element, are used to assess whether a node is experiencing performance anomalies.
[0064] Specifically, by analyzing the node's historical performance data, historical alarm information, and historical failure cases and setting risk thresholds, the node's health status can be comprehensively and accurately assessed. By converting abnormal performance items, related data, and preset feature information into feature vectors and splicing them into comprehensive feature vectors, the system can comprehensively consider multiple influencing factors and provide a more comprehensive and accurate node status assessment. After comparing with the risk threshold, the system can intelligently determine whether there is an anomaly. Once an anomaly is detected, it can automatically trigger an inspection instruction to ensure that the node problem is handled in a timely manner. This process not only greatly improves operation and maintenance efficiency and reduces the need for manual intervention, but also reduces the probability of system failures by early detection and processing of node anomalies, thereby enhancing the stability and reliability of the system.
[0065] In one embodiment, after inspecting the node according to the node inspection instruction, the method further includes:
[0066] Get current inspection data;
[0067] Classify the current inspection data and compare the classified current inspection data with the corresponding historical inspection data to obtain the health trend of the corresponding node;
[0068] Generate inspection reports based on the health trends of nodes.
[0069] Among them, the node inspection instruction refers to the instruction issued by the system or administrator, requiring an inspection of a certain node (such as a storage server, switch, computer and other hardware or software components). The instruction usually includes inspection items, inspection frequency, inspection method, etc. Current inspection data refers to the data collected when executing the node inspection instruction. Historical inspection data refers to all data collected when inspecting the node over a period of time in the past. It records the node's health status, performance trends, fault conditions, etc. The inspection report is a detailed document or report generated based on the inspection data and health trends, including: hardware health change trends, software health change trends, alarm data change trends, potential risks, recommended improvement measures, etc.
[0070] Specifically, by comparing current inspection data with historical inspection data, the system can promptly detect changes in node health. For example, a drop in hard drive read / write speed might not cause a serious problem in the short term, but if a trend of worsening performance is detected by comparing historical data, proactive measures can be taken, such as replacing the hard drive or adjusting the load, to prevent a more severe failure.
[0071] When the inspection command is executed, the system collects data such as the CPU utilization, memory usage, and disk health status of the current node. The system then compares the current data with the data from the past few inspections and finds that the disk I / O rate of a certain storage node has gradually become lower, while historical data shows that the disk I / O rate of this node has not changed much in the past year. Through health trend analysis, the system can conclude that the disk of this node is experiencing performance degradation, which may be due to the gradual aging of the hard disk or other reasons. Based on this trend, the system automatically generates an inspection report, pointing out that the health trend of the node is gradually declining, and recommends that the administrator replace the hard disk in advance to avoid data loss or service interruption. Based on the report, the administrator can decide whether to take immediate action or arrange for hard disk replacement according to actual business needs. This automated health trend analysis and report generation enables administrators to detect problems in advance and take preventive measures, reducing the risk of failure and improving the reliability of the entire system.
[0072] In one embodiment, the method further comprises:
[0073] In response to triggering a node inspection instruction, determining whether there is a node inspection task currently;
[0074] In response to the existence of a node inspection task, the current node inspection instruction is canceled.
[0075] Specifically, repeated inspection tasks may generate redundant alarms or system logs, increasing the system's burden. By determining whether an inspection task is already in progress, the system avoids duplication and ensures that each inspection task does not conflict with other tasks. This reduces unnecessary interference in the system and helps improve task execution efficiency. Operations and maintenance personnel no longer need to manually intervene in repeatedly triggered inspection tasks; the system automatically processes inspection instructions and makes judgments, ensuring the orderly execution of inspection tasks. This system not only reduces the workload of administrators but also improves the system's autonomous operation and maintenance capabilities.
[0076] In one embodiment, Figure 3 As shown, according to the data to be analyzed, determining the abnormal performance items of the node and obtaining the associated data corresponding to the abnormal performance items include:
[0077] according to:
[0078]
[0079] The data to be analyzed is smoothed to obtain low-frequency features, where t(u) represents the transformation result, f(x) represents the original performance input sequence, N represents the number of nodes, and s(x, u) is expressed as:
[0080]
[0081] Indicates mapping the frequency range between [0, π], u represents the translation parameter, and α(u) is expressed as:
[0082]
[0083] Performing statistical calculations on the low-frequency features to determine the statistical characteristic value distribution of the low-frequency features;
[0084] Compare the statistical characteristic value distribution with the preset threshold range to determine whether there is performance anomaly in the node;
[0085] In response to a performance anomaly on a node, the anomaly item and associated data corresponding to the anomaly item are recorded.
[0086] Among them, statistical characteristic values include: mean, standard deviation, maximum and minimum values, median, skewness, kurtosis, periodicity, trend and fluctuation range.
[0087] Specifically, in a large-scale distributed storage system, the system regularly monitors the performance of each node to ensure stable operation. Performance data (such as CPU usage, memory usage, and network bandwidth) is collected from each node. To reduce the impact of short-term fluctuations and noise on the analysis results, the system smoothes this performance data. For example, a sliding average or low-pass filter is used to extract low-frequency features from the data and remove high-frequency noise. Low-frequency features, such as long-term performance trends and cyclical fluctuations, are extracted from the smoothed data. These features help the system understand the health of the node over a longer period of time. The system performs statistical calculations on these low-frequency features, such as the mean and standard deviation, to obtain a distribution of their statistical features. This distribution provides basic data for subsequent performance analysis and anomaly detection. The system compares the calculated low-frequency feature statistics with a preset threshold. If the statistical value exceeds the preset threshold, the system determines that the node has a performance anomaly. For example, if CPU usage remains high for a long time and exceeds the preset threshold, the node may be considered to have a performance bottleneck or hardware failure. When the system identifies a performance anomaly, it automatically records the relevant anomaly items (such as the anomaly type, severity, etc.) and the associated data (such as the specific performance indicator value, the time when the anomaly occurred, the relevant nodes, etc.). This information helps operation and maintenance personnel locate the problem and take timely measures to repair it. Through smoothing and low-frequency feature extraction, the system can identify the true performance trend from massive short-term data fluctuations. This enables the system to detect performance anomalies more accurately, rather than relying solely on instantaneous high-frequency fluctuations. Through noise removal and smoothing, the system can reduce false alarms caused by short-term fluctuations. This method can effectively reduce misjudgments caused by temporary anomalies caused by system load, network latency, etc. By analyzing low-frequency features, the system can monitor the long-term health of the node and promptly detect potential performance bottlenecks or hardware failures, so that measures can be taken in advance to avoid system failures and reduce the risk of service interruptions.
[0088] For example, during health monitoring of a storage node, the system might detect a gradual decrease in the node's disk I / O rate. Through smoothing, the system removes transient fluctuations and extracts the long-term trend of disk performance. Based on this trend, the system determines that the node's disk health is deteriorating and compares this anomaly with a preset disk performance threshold. Because the statistical characteristics exceed the normal range, the system determines a performance anomaly and records detailed data related to the anomaly (such as the magnitude of the I / O rate drop and the affected nodes).
[0089] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0090] In one embodiment, it includes:
[0091] Predict potential node failures based on historical performance data, alarm information, and failure cases, and obtain prediction results;
[0092] Analyze the abnormal items of the current node based on the predicted results and the current node data to see if they are the same as the predicted results;
[0093] In response to the prediction result being different from the abnormal item of the current node, the potential cause is identified through the associated data;
[0094] In response to the prediction result being the same as the abnormal item of the current node, fault processing is performed.
[0095] Specifically, the system can form a more accurate fault prediction model based on accumulated historical data, identifying potential failure modes in advance, significantly reducing the risk of system downtime and business interruption caused by sudden failures. If the predicted result matches the abnormal item in the current node, the system has effectively predicted the impending failure and can promptly implement remedial measures. This mechanism can quickly initiate fault handling procedures, reduce manual intervention time, and ensure high system availability and business continuity. If the predicted result differs from the abnormal item in the current node, in-depth analysis of the associated data can identify the potential root cause of the failure. This process effectively addresses abnormal situations that existing prediction models may not cover, flexibly responding to complex or unforeseen fault types, and avoiding the significant risks of misdiagnosis or missed faults. This data-driven anomaly identification and root cause analysis can provide real-time diagnosis and remediation solutions for new and unknown failure modes, further enhancing the system's intelligent operation and maintenance capabilities.
[0096] The embodiment of the present application also provides a passenger flow counting device, such as Figure 2 As shown, it includes: an acquisition module 210, an abnormal data calculation module 220, an alarm information acquisition module 230, an instruction generation module 240 and an execution module 250.
[0097] The acquisition module 210 is used to obtain the original performance data of the node and pre-process the original performance data of the node to obtain the data to be analyzed;
[0098] The abnormal data calculation module 220 is used to perform statistical calculations on the data to be analyzed, and compare the calculation results with a preset threshold value to obtain node performance abnormality data and node performance data corresponding to the node performance abnormality data;
[0099] The alarm information acquisition module 230 is used to obtain node alarm information and extract preset feature information in the node alarm information;
[0100] The instruction generation module 240 is used to convert the abnormal performance items, associated data and preset characteristic information into feature vectors to obtain corresponding feature vectors, and then combine them to form a comprehensive feature vector. In response to the abnormality of the comprehensive feature vector, it triggers the generation of node inspection instructions;
[0101] The execution module 250 is used to inspect the corresponding node according to the node inspection instruction. The description of the features in the embodiment corresponding to the node inspection device can be found in the relevant description of the embodiment corresponding to the node inspection method, which will not be repeated here.
[0102] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned node inspection method embodiments when running.
[0103] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0104] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned node inspection method embodiments are implemented.
[0105] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two, such as Figure 4 As shown, in order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to function. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0106] The above is a detailed introduction to a node inspection method provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A node inspection method, characterized in that: include: Obtaining raw performance data of the node, and preprocessing the raw performance data to obtain data to be analyzed; Determine, based on the data to be analyzed, abnormal performance items of the node and obtain associated data corresponding to the abnormal performance items; Obtaining node alarm information and extracting preset feature information of the node alarm information; By performing feature vector conversion on the abnormal performance item, the associated data and the preset feature information, a corresponding feature vector is obtained, and the feature vectors are spliced to form a comprehensive feature vector. In response to an abnormality in the comprehensive feature vector, a node inspection instruction is triggered; The node is inspected according to the node inspection instruction.
2. A node inspection method according to claim 1, characterized in that: The acquiring of raw performance data of the node and preprocessing of the raw performance data to obtain data to be analyzed includes: Collecting the raw performance data at every sampling interval; Identifying an original format of the original performance data and converting the original format into a preset format; The original unit of the original performance data is identified and the original unit is converted into a preset unit.
3. A node inspection method according to claim 1, characterized in that: Obtaining node alarm information and extracting preset feature information from the node alarm information includes: Determine the alarm type according to the node alarm information; Determining an alarm level corresponding to the alarm type according to the node alarm information; The alarm types and the alarm levels are summarized as preset feature information.
4. A node inspection method according to claim 1, characterized in that: By performing feature vector conversion on the abnormal performance item, the associated data, and the preset feature information, a corresponding feature vector is obtained, and the feature vectors are spliced to form a comprehensive feature vector. In response to an abnormality in the comprehensive feature vector, a node inspection instruction is triggered, including: Acquire historical data of the node, the historical data including at least one of the following: historical performance data, number of historical alarm information, and number of historical fault cases; Determining a risk threshold vector based on the historical data, wherein the risk threshold vector includes at least one risk element, and the at least one risk element represents at least one of the following: a historical performance threshold corresponding to the historical performance data, an alarm quantity threshold corresponding to the number of historical alarm information, and a case quantity threshold corresponding to the number of historical fault cases; Performing feature vector conversion on the abnormal performance item, the associated data, and the preset feature information to obtain an abnormal performance item feature vector, an associated data feature vector, and a preset feature information feature vector; Concatenate the abnormal performance item feature vector, the associated data feature vector, and the preset feature information feature vector to obtain a comprehensive feature vector; Comparing the elements in the comprehensive feature vector with the corresponding risk elements to determine whether the node is abnormal; In response to the node being abnormal, a node inspection instruction is triggered to be generated.
5. A node inspection method according to claim 1, characterized in that: After inspecting the node according to the node inspection instruction, the method further includes: Get current inspection data; Classifying the current inspection data, and comparing the classified current inspection data with corresponding historical inspection data to obtain a health trend of the node; The inspection report is generated according to the health trend of the node.
6. A node inspection method according to claim 5, characterized in that: The method further comprises: In response to triggering the node inspection instruction, determining whether an inspection task for the node is currently being executed; In response to the existence of the inspection task, the current node inspection instruction is canceled.
7. A node inspection method according to claim 1, characterized in that: Determining abnormal performance items of a node according to the data to be analyzed and obtaining associated data corresponding to the abnormal performance items includes: according to: The data to be analyzed is smoothed to obtain low-frequency features, where t(u) represents the transformation result, f(x) represents the original performance input sequence, N represents the number of nodes, and s(x, u) is expressed as: Indicates mapping the frequency range between [0, π], u represents the translation parameter, and α(u) is expressed as: Performing statistical calculations on the low-frequency features to determine a distribution of statistical characteristic values of the low-frequency features; Comparing the statistical characteristic value distribution with a preset threshold interval to determine whether the node has performance abnormalities; In response to a performance abnormality of the node, the abnormal item and associated data corresponding to the abnormal item are recorded.
8. A node inspection device, characterized in that: include: An acquisition module is used to acquire original performance data of a node and pre-process the original performance data of the node to obtain data to be analyzed; An abnormal data calculation module is used to perform statistical calculations on the data to be analyzed, and compare the calculation results with a preset threshold to obtain node performance abnormality data and node performance data corresponding to the node performance abnormality data; An alarm information acquisition module is used to obtain node alarm information and extract preset feature information from the node alarm information; An instruction generation module is configured to perform feature vector conversion on the abnormal performance item, the associated data, and the preset feature information to obtain corresponding feature vectors, and to concatenate the feature vectors to form a comprehensive feature vector. In response to an abnormality in the comprehensive feature vector, a node inspection instruction is triggered; An execution module is used to inspect the corresponding node according to the node inspection instruction.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.