A visual operation and maintenance data analysis system
By employing modules such as node data aggregation, indicator sorting, and multi-parameter linkage identification in the visualized operation and maintenance data analysis system, the problem of low operation and maintenance efficiency in existing operation and maintenance data analysis systems under scenarios of multi-parameter trend linkage and uneven resource utilization has been solved, realizing full-link linkage analysis and rapid response.
Patent Information
- Application Number
- CN202510913770.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-03
AI Technical Summary
Existing visual operation and maintenance data analysis systems struggle to automatically identify abnormal nodes and resource redundancy when faced with complex scenarios involving multi-parameter trend linkages and uneven resource usage. This forces operation and maintenance personnel to repeatedly search and manually compare data, limiting response speed and systemic management.
The system employs a node data collection module, an indicator sorting and filtering module, a multi-parameter linkage identification module, and a resource node screening module. By analyzing node operation data through the intelligent computing center backbone network, it groups and sorts the data by time label, identifies linkage anomaly trend links, and filters resource control number groups, thereby achieving full-link linkage analysis of operation and maintenance data.
It has improved the response speed and systematization of operation and maintenance data, realized the visualization, archiving and traceability of operation and maintenance events, promoted the transformation of operation and maintenance data from single-point monitoring to full-link linkage analysis, and improved operation and maintenance efficiency and management level.
Smart Images

Figure CN120407350B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated operation and maintenance technology, and in particular to a visual operation and maintenance data analysis system. Background Technology
[0002] Automated operations and maintenance (O&M) is an important branch of information technology management. Its core aspects include unified management of IT system assets, automated execution of O&M tasks, automated deployment and upgrades of server software, configuration file management, performance data collection and processing, real-time status monitoring, dynamic scheduling of resource utilization, fault diagnosis and automatic recovery, user access control and security, and multi-dimensional log management. This technology aims to improve the operational efficiency and management level of large-scale enterprise IT infrastructure through automation. Traditional visual O&M data analysis systems involve collecting and organizing IT O&M data within an automated O&M platform, using indicator calculation methods to statistically analyze basic data such as server CPU and GPU utilization, memory usage, disk read / write performance, network bandwidth utilization, application response time, error rate, and business traffic. The analysis results are then output in the form of charts and reports to allow O&M personnel to observe the system's operational status.
[0003] Current operational data analysis models mainly rely on individual statistics and static distribution of indicators, lacking automatic identification and priority screening of multi-parameter trend linkages. Node classification is merely a superficial archiving, making it difficult to cope with complex scenarios with frequent dynamic changes in nodes or uneven distribution of resource usage. Source tracing analysis relies on breakpoint data comparison, making it difficult to trace event chains and the impact paths formed by high-risk nodes. This results in operational personnel having to repeatedly search and manually compare data, limiting both response speed and the systematic nature of management. Summary of the Invention
[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a visual operation and maintenance data analysis system.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a visual operation and maintenance data analysis system, the system comprising:
[0006] The node data collection module is based on the backbone network of the intelligent computing center. It analyzes the operation data of each node, groups the nodes by time label, arranges the changes in CPU and GPU usage ratios and memory usage ratios in order, and archives disk read / write time, total inbound / outbound traffic and business response time to obtain a set of node indicator sequences.
[0007] The indicator sorting and filtering module compares the CPU and GPU usage ratios and total inbound and outbound traffic of each node under the same time tag based on the node indicator sequence set, summarizes the differences, sorts them by type, identifies the nodes with the highest ranking, and obtains the operation and maintenance sorted node group.
[0008] Based on the operation and maintenance sorting node group, the multi-parameter linkage identification module analyzes the synchronous changes of CPU and GPU usage ratios, memory usage ratios and business response times of the leading nodes in the current and adjacent time tags, determines the time-series curves of each parameter, and obtains the linkage anomaly trend link.
[0009] Based on the aforementioned linkage anomaly trend link, the resource node screening module filters unmarked nodes in the associated resource segment, compares the memory usage ratio trend and the fluctuation range of inbound and outbound traffic, and determines the node number by combining the sorting intersection to obtain the resource control number group.
[0010] The present invention is improved in that the node indicator sequence set includes node performance data, archived time series, and group classification information; the operation and maintenance sorting node group includes priority identifier, node identification code, and sorting index number; the linkage abnormal trend link includes linkage feature sequence, abnormal fluctuation segment, and trend matching identifier; and the resource control number group includes control node list, allocation node code, and redundant resource identifier.
[0011] The present invention is improved in that the node data collection module includes:
[0012] The data acquisition submodule is based on the backbone network of the intelligent computing center. It analyzes the real-time data collected on the CPU and GPU usage ratio, memory usage ratio, disk read and write time, total inbound and outbound traffic and business response time of various nodes. The data is organized according to time tags to generate a set of operating indicators.
[0013] The joint arrangement submodule compares the correlation between the CPU and GPU usage ratios and memory usage ratios of each node under the same time tag based on the set of operating indicators, and arranges them in sequence according to the changing trend of each indicator to obtain a joint changing trend group.
[0014] The indicator archiving submodule filters the disk read / write time, total inbound / outbound traffic, and business response time of each type of node under continuous time labels based on the joint change trend group, and archives the parameter trends to obtain a set of node indicator sequences.
[0015] The present invention is improved in that the index sorting and filtering module includes:
[0016] The difference summarization submodule analyzes the CPU and GPU usage ratios and total inbound and outbound traffic of the same type of nodes under each time tag based on the node indicator sequence set, compares the fluctuation range of the corresponding parameters of each node in the group, calculates the difference magnitude of parameters between different nodes, and obtains the performance change magnitude.
[0017] The group comparison submodule calls the performance change magnitude, groups nodes according to node type, compares the difference distribution of CPU and GPU usage ratio and total inbound and outbound traffic of each group of nodes under the same time label, judges the distribution breadth of parameters among nodes in the same group, filters out nodes with key performance differences, and obtains the difference distribution range.
[0018] The sorting and extraction submodule analyzes the CPU and GPU usage ratios, total inbound and outbound traffic, and group parameter benchmarks of nodes under the current time tag based on the difference distribution interval. It adjusts the impact of historical performance changes and linkage fluctuations within the group, identifies nodes with outstanding sorting performance, and uses them as key operation and maintenance analysis objects to obtain the operation and maintenance sorting node group.
[0019] The present invention is improved in that the multi-parameter linkage recognition module includes:
[0020] The time-series parameter extraction submodule analyzes the nodes ranked first in the operation and maintenance sorting node group, compares the CPU and GPU usage ratio, memory usage ratio and business response time in chronological order at multiple time points, and integrates the time-series information of each node's indicators to obtain indicator time-series curve data.
[0021] The synchronization trend identification submodule calls the time series curve data of the indicators, compares the direction of change of each indicator at each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator within the synchronization interval to obtain the synchronization trend interval table.
[0022] The linkage link construction submodule determines the trend fluctuation of nodes on the CPU and GPU load panel and response latency layer based on the synchronization trend interval table, calculates the correlation between each group of trends, obtains the total linkage trend intensity, determines the time series curve of each parameter, and obtains the linkage abnormal trend link.
[0023] The present invention is improved in that the resource node screening module includes:
[0024] The unmarked node submodule of the section, based on the linked abnormal trend link, calls the nodes it covers, determines the resource section associated with the node, filters the nodes in the same section that have not been marked as abnormal, optimizes the node status screening process, and obtains the resource test number set.
[0025] The memory traffic linkage submodule compares the memory usage ratio and total inflow / outflow of each node at consecutive time points based on the resource test number set, obtains the joint trend difference of each node, and then performs number mapping according to the resource segment affiliation to obtain the segment trend difference sequence.
[0026] The number intersection extraction submodule identifies the top-ranked node numbers based on the segment trend difference sequence, analyzes the intersection of the nodes with the ranked nodes, determines the distribution of the intersection numbers in the node set, and obtains the resource regulation number group.
[0027] The present invention has an improvement, wherein the system further includes:
[0028] The event trajectory tracing module retrieves time-series data on node CPU and GPU usage ratios, disk read / write times, and business response times based on the resource control number group, and compares it with the linked fluctuating node curves to obtain the trajectory information of the associated nodes.
[0029] The associated node trajectory information includes historical trajectory sequences, node association features, and event archive identifiers.
[0030] The present invention is improved in that the event trajectory tracing module includes:
[0031] The trajectory data extraction submodule analyzes the recent CPU and GPU usage ratios, disk read / write activity, and business response time records based on the resource control number group, and organizes them into continuous node operation performance to obtain the node operation sequence.
[0032] The fluctuation curve comparison submodule compares the changing trends of each indicator of each node in the node operation sequence with the curve of the linked fluctuation node, judges the synchronous changes of each node indicator curve in the same time sequence, identifies the node performance with highly consistent trends, and obtains the synchronous characteristic node group.
[0033] The trajectory information generation submodule screens nodes with key synchronous changes based on the synchronous feature node group, optimizes the abnormal segments and change characteristics of each node in the time series curve, and obtains the trajectory information of the associated nodes.
[0034] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0035] In this invention, the operational data of different nodes are hierarchically grouped and archived, and various operation and maintenance parameters are dynamically linked and merged by combining time tags. Under the changing trends of multiple indicators, operation and maintenance focus objects are automatically screened, and abnormal nodes and resource redundancy nodes are accurately screened based on comprehensive feature trajectories. This promotes the transformation of operation and maintenance data from single-point monitoring to full-link linkage analysis. In the process of visualization archiving and tracing, the intuitiveness and globality of operation and maintenance event response are improved, so as to promote the systematic closed-loop monitoring of node status, resource allocation and operation and maintenance risks. Attached Figure Description
[0036] Figure 1 This is a system flowchart of the present invention;
[0037] Figure 2 This is a flowchart of the node data collection module in this invention;
[0038] Figure 3 This is a flowchart of the index sorting and filtering module in this invention;
[0039] Figure 4 This is a flowchart of the multi-parameter linkage recognition module in this invention;
[0040] Figure 5 This is a flowchart of the resource node screening module in this invention;
[0041] Figure 6 This is a flowchart of the event trajectory tracing module in this invention. Detailed Implementation
[0042] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0043] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0044] Example: Please refer to Figure 1 This invention provides a technical solution: a visual operation and maintenance data analysis system comprising:
[0045] The node data collection module is based on the backbone network of the intelligent computing center. It collects the operation data of computing nodes, storage nodes and transmission nodes, groups the nodes according to time tags, arranges the joint changes of CPU and GPU usage ratio and memory usage ratio in chronological order, and archives disk read and write time, total inbound and outbound traffic and business response time according to node type to obtain a node indicator sequence set.
[0046] The indicator ranking and filtering module is based on the node indicator sequence set. It compares the CPU and GPU usage ratios and total inbound and outbound traffic of each node under the same time tag, summarizes the differences of each data item, groups and sorts them according to node type, identifies the top-ranked nodes in each group as key operation and maintenance analysis objects, and obtains the operation and maintenance ranked node group.
[0047] The multi-parameter linkage identification module is based on the operation and maintenance sorting node group. It analyzes the synchronous changes of CPU and GPU usage ratio, memory usage ratio and business response time of the top five nodes in the current and adjacent time tags. Through the CPU and GPU load panel of the server group and the response latency layer of the storage array, the time-series curve of each parameter is determined, and the linkage anomaly trend link is obtained.
[0048] The resource node screening module is based on the linkage anomaly trend link, filters the unmarked nodes in the resource segment to which the associated nodes belong, compares the change trend of their memory usage ratio with the fluctuation range of the total inbound and outbound traffic, and retains the number of the first node by combining the sorting intersection to obtain the resource control number group.
[0049] The event trajectory tracing module retrieves continuous time-series data of the CPU and GPU usage ratio, disk read / write time, and business response time of nodes over the past three days based on resource control number groups. It then compares the time-series data with the curves of interconnected fluctuating nodes to obtain the trajectory information of related nodes for the visualization, archiving, and source analysis of operation and maintenance events.
[0050] The node indicator sequence set includes node performance data, archived time series, and grouping and classification information. The operation and maintenance sorting node group includes priority identifier, node identification code, and sorting index number. The linkage anomaly trend link includes linkage feature sequence, abnormal fluctuation segment, and trend matching identifier. The resource control number group includes control node list, allocation node code, and redundant resource identifier. The associated node trajectory information includes historical trajectory sequence, node association characteristics, and event archive identifier.
[0051] In Module 1, the intelligent computing center backbone network refers to the high-bandwidth, high-reliability network infrastructure built within the intelligent computing center, connecting all computing nodes, storage nodes, and transmission nodes. It serves as the main channel for all data flow, collection, and scheduling. Computing nodes are server resource units deployed with processors and memory, used for data processing and business operations. Storage nodes exist in the form of disk arrays or distributed storage servers, primarily responsible for data persistence and read / write operations. Transmission nodes are network devices such as switches and routers, primarily responsible for the flow of data between nodes in the intelligent computing center. Time stamps record the time points of data collection, used for time-series alignment and trend analysis of data from different nodes during subsequent analysis. Node types refer to three categories: "computing nodes," "storage nodes," or "transmission nodes," categorized by hardware function and role. Joint changes refer to the trend analysis of the linkage between two indicators, such as the CPU or GPU usage ratio and the memory usage ratio, of the same node under the same time stamp (e.g., whether they rise or fall simultaneously). Business response time refers to the total time from initiation of a business request to receiving a response, typically reflecting the response efficiency and user experience of application services.
[0052] In Module 2, each data difference refers to the magnitude of the difference between the same type of parameters of each node under the same time label. For example, the difference in the CPU or GPU usage ratio of different computing nodes, or the difference in the total inbound and outbound traffic; the nodes ranked higher refer to the nodes with the most outstanding performance indicators after the difference is summarized and ranked (such as high CPU or GPU load or large traffic, etc.), and these nodes are usually the focus of operation and maintenance.
[0053] In Module 3, the top five nodes refer to the top five critical nodes, which are usually objects with high load, abnormal status, or significant business impact; synchronous changes refer to multiple indicators (such as CPU or GPU usage ratio, memory usage ratio, and business response time) showing changes in the same direction in the time series (such as synchronous increase or synchronous decrease), reflecting the linkage effect of the system or business; the response latency layer refers to a graphical interface or data layer that visualizes the response time distribution of devices such as storage arrays, used to identify storage bottlenecks or anomalies; the time series curve refers to a trend line that reflects the change of a certain parameter (such as CPU or GPU usage ratio, response time, etc.) over time, used to observe the changing patterns and fluctuation characteristics of the indicator.
[0054] In Module 4, "resource segment" refers to a resource partition (such as a cabinet, server room, or service pool) in the intelligent computing center, divided according to physical location or business function, used for local resource management and scheduling; "unmarked node" refers to a node that is not classified as abnormal or under special attention during the anomaly identification and linkage identification process, i.e., a node in a "normal" state; "change trend" refers to analyzing the continuous change direction (rising, falling, or remaining stable) of a certain indicator (such as memory usage ratio) over a period of time to determine the trend of resource load; "fluctuation amplitude" refers to the maximum range of change of a certain indicator within the observation period, used to identify nodes with drastic load changes; "ranking intersection" refers to cross-comparing the ranking results of different indicators to screen nodes that perform well in multiple dimensions.
[0055] In Module 5, continuous time-series data refers to the historical data sequence collected under multiple consecutive time tags of a node, such as the CPU or GPU usage ratio change curve of a node within three days; linked fluctuation nodes refer to the nodes marked as needing attention in the aforementioned linked identification module due to the synchronous fluctuation of multiple indicators.
[0056] Please see Figure 2 The node data collection module includes:
[0057] The data acquisition submodule is based on the backbone network of the intelligent computing center. It analyzes the real-time data collected on the CPU and GPU usage ratio, memory usage ratio, disk read and write time, total inbound and outbound traffic and business response time of various nodes. The data is organized according to time tags to generate a set of operating indicators.
[0058] Define the category attributes of each node and determine the collection scope. For example, node A is a computing node, node B is a storage node, and node C is a transmission node. Collect corresponding operational metrics such as CPU or GPU utilization, memory usage, disk operation time, network traffic, and service response time. Within a minute-by-minute collection cycle, monitor the CPU or GPU utilization of node A. By reading the real-time feedback of node operations on the processor's running status, calculate the ratio of its active processing time to its total runtime within one minute. If node A's active processing time is 2825 milliseconds and the total processing time per minute is 10,000 milliseconds, then the CPU or GPU utilization is 28.25%. Within the same cycle, memory usage is calculated based on the ratio of used memory recorded by the operating system to total memory. If node A has a total memory of 32GB and actually uses 12GB, then the memory usage is 37.5%. Disk read / write time is collected by statistically analyzing the total time spent on disk interface read / write operations. If node B performs 50 read / write operations within one minute, with an average time of 7 milliseconds per operation, the total read / write time is 350 milliseconds. The inbound and outbound traffic of transmission node C is obtained by statistically analyzing the number of input and output bytes of the network interface. If the received traffic is 2.4GB and the sent traffic is 1.2GB within one minute, the total traffic is 3.6GB. The business response time is collected through log analysis, recording the time difference between the receipt of a request and the return of a response. If five business calls occur within one minute, with response times of 140, 180, 210, 170, and 200 milliseconds respectively, the average business response time is 180 milliseconds. All the above data are tagged according to the sampling time to form an indicator data package based on time tags, and categorized under their respective node numbers to form a set of operating indicators under the node type classification.
[0059] The joint permutation submodule compares the correlation between the CPU and GPU usage ratios and memory usage ratios of each node under the same time tag based on the set of operating indicators, and arranges them in order according to the changing trend of each indicator to obtain the joint changing trend group.
[0060] Data from all nodes is merged according to time labels. Comparisons are then performed sequentially between computing nodes with the same time label. For example, at time T5, node X has a CPU or GPU utilization of 65% and a memory utilization of 58%, node Y has 82% and 71%, and node Z has 78% and 64%. These constitute the metric combinations for each node at that time. Trend comparisons are then performed on each pair of these combinations. First, it is determined whether the CPU or GPU utilization and memory utilization increased or decreased simultaneously in the previous time period. For example, if node Y's CPU or GPU utilization was 72% and memory utilization was 63% at time T4, and changed to 82% and 71% at T5, it indicates that its metrics show a combined upward trend. Process nodes Z and X, identify all nodes that change in the same direction, and calculate the magnitude of change of two indicators for each node. If node Y's CPU or GPU increases by 10% and memory increases by 8%, and node Z's CPU or GPU increases by 9% and memory increases by 7%, then node Y can be determined to have the most dramatic change in this period. Nodes are sorted according to the magnitude of change to form a sorting sequence. In actual operation, when objects at the top of the node sorting consistently show similar trends, they can be identified as objects with significant joint change trends in this round. Each record entry in this sorting group contains a node number, time label, indicator change direction, and joint change magnitude, which are used to form a joint change trend group.
[0061] The indicator archiving submodule filters the disk read / write time, total inbound / outbound traffic, and business response time of each type of node under continuous time labels based on the joint change trend group, archives the parameter trends, and obtains a set of node indicator sequences.
[0062] The indicator archiving submodule extracts the top five nodes from the joint trend group and monitors and archives their indicator trends over several subsequent time tags. For example, selecting node Y in the time period from T5 to T9, records its disk read / write times as 300, 340, 370, 410, and 460 milliseconds, inbound / outbound traffic as 3.1, 3.5, 3.9, 4.0, and 4.2 GB, and business response time as 190, 205, 218, 230, and 240 milliseconds. A data list is created for each indicator of this node at each time point, and a trend curve is plotted. The growth rate is then manually analyzed; for example, disk write time shows a periodic increasing trend. The total increase was from 300 milliseconds to 460 milliseconds, the traffic increased from 3.1GB to 4.2GB, and the response time increased from 190 milliseconds to 240 milliseconds. This indicates that the node's load increased and its performance decreased during this period. The values of the above three indicators under each time label were recorded to form trend data and marked as the indicator sequence of node Y in T5~T9. This data was further archived into the indicator sequence set corresponding to the node to form the basic dataset for continuous trend analysis. The structure of the indicator sequence set is based on time label and indicator items, recording the absolute value and direction of change of each indicator in a continuous time period for subsequent node screening and event identification processing.
[0063] Please see Figure 3 The indicator sorting and filtering module includes:
[0064] The difference summarization submodule analyzes the CPU and GPU usage ratios and total inbound and outbound traffic of the same type of nodes under each time tag based on the node indicator sequence set, compares the fluctuation range of the corresponding parameters of each node in the group, calculates the difference magnitude of parameters between different nodes, and obtains the performance change magnitude.
[0065] First, extract all compute nodes, storage nodes, and transmission nodes under the same time label from the dataset. After classifying the nodes by type, obtain the CPU or GPU usage ratio and total inbound / outbound traffic recorded for each node under that time label. Then, perform a difference analysis on the parameters within each node category. For example, under the T10 time label, if the CPU or GPU usage rates of compute nodes A, B, and C are 55%, 78%, and 62% respectively, and the total inbound / outbound traffic is 1.2GB, 3.0GB, and 1.5GB respectively, then by subtracting the minimum value from the maximum value for the CPU or GPU usage rate, the fluctuation range of the CPU or GPU usage rate within this group is found to be 78% - 55% = 23%. Similarly, the fluctuation range of the total inbound / outbound traffic is 3.0GB - 1.2GB = 1.8GB. Then, by analyzing every pair of nodes... The absolute differences between the values are summed. The difference in CPU or GPU utilization between nodes A and B is 23%, between nodes A and C is 7%, between nodes B and C is 16%, and so on to complete the sum of the combined differences. If the differences in total inbound and outbound traffic are 1.8GB, 0.3GB, and 1.5GB respectively, the difference magnitude of the corresponding combination is recorded. By statistically analyzing the maximum fluctuation value of each type of parameter and the average value of the combined differences, the overall performance change within the group is measured. If the fluctuation value of CPU or GPU utilization is greater than 20% and the fluctuation value of inbound and outbound traffic exceeds 1.5GB, it is considered that there is a node group with a large performance change magnitude under this time tag. The node number, time tag, parameter type, fluctuation range value, and node pair combination information are further organized into a change magnitude dataset for subsequent group comparison.
[0066] The performance variation of submodule calls is compared by grouping nodes according to node type. The distribution of CPU and GPU usage ratios and total inbound and outbound traffic in each group of nodes is compared under the same time label. The distribution breadth of parameters among nodes in the same group is determined, and nodes with key performance differences are selected to obtain the difference distribution range.
[0067] Based on node type, all nodes are independently grouped according to computation, storage, and transmission. Within each group, a differential distribution analysis is performed based on the performance data corresponding to the same time label. For each node in each group, the CPU or GPU utilization rate and total inbound / outbound traffic at a given time label are taken, and a one-dimensional sequence is constructed and sorted. After sorting, the difference between the first and last parts of the sequence is compared. For example, at time T12 in the computation node group, the CPU or GPU utilization rates of nodes X, Y, Z, and W are 45%, 68%, 53%, and 80%, respectively. The constructed sequence is 45%, 53%, 68%, and 80%, with a difference of 35%. If a difference in CPU or GPU utilization rate greater than 30% is considered too large, then the node group is marked as widely distributed. A sorting analysis is also performed on the inbound / outbound traffic. If nodes X, Y, Z, and W have CPU or GPU utilization rates of 1.1GB, 68%, 53%, and 80%, respectively, then the distribution is considered too large. 1.4GB, 2.2GB, 3.8GB, the difference after sorting is 2.7GB. Determine if the distribution difference of inbound and outbound traffic exceeds the preset threshold. If it is set to 2.5GB, then this group meets the requirement of abnormal distribution breadth. Then, find two key nodes within this group, namely the nodes near the extreme values in the distribution of each parameter, namely the node X with the lowest CPU or GPU utilization and the node W with the highest utilization, and the node X with the smallest inbound and outbound traffic and the node W with the largest utilization. When the extreme value nodes of the two parameters overlap, that is, nodes X and W are both critical performance objects, they can be selected as key nodes of performance difference and included in the difference distribution interval record. The record content is marked with node number, type, corresponding time tag, CPU or GPU and traffic distribution sorting position, difference threshold judgment result, whether it constitutes overlapping extreme value nodes, and stored in the difference distribution interval set.
[0068] The sorting and extraction submodule analyzes the CPU and GPU usage ratios, total inbound and outbound traffic, and grouping parameter benchmarks of nodes under the current time label based on the difference distribution interval. It then adjusts the impact of historical performance changes and related fluctuations within the group, using the following formula:
[0069] ;
[0070] Nodes with outstanding ranking performance are identified as key operational analysis objects, resulting in an operational ranking node group. Representative node The range of sorting performance, For nodes CPU or GPU usage percentage under the current time tag For nodes Total inbound and outbound traffic under the current time tag. and The groups to which the nodes belong are respectively The baseline for balancing CPU or GPU usage ratio and total inbound / outbound traffic. This represents the discrete level of the performance variation in this group. Indicates the node in the consecutive Performance changes of each time tag Indicates the node at the 1st Traffic linkage intensity under each time tag This indicates the number of time tags for continuous observation, i.e., the total number of time points for continuous tracking of a node.
[0071] Ranking performance amplitude refers to the degree of deviation of each node's current CPU or GPU usage ratio from its group baseline (group average level) among nodes of the same type. It is a comprehensive measurement that combines the historical performance changes and traffic linkage effects over a period of time. Ranking performance amplitude is used to reflect the strength and degree of change of a node's operating characteristics among nodes of the same type. This indicator can identify key nodes that deserve special attention in the current operation and maintenance analysis. The larger the ranking performance amplitude, the more prominent the difference and impact of the node in the indicator performance.
[0072] Based on the differential distribution intervals, analyze the CPU or GPU usage ratio and total inbound / outbound traffic of the candidate nodes in the current time label. Unify the parameter dimension processing method and use a normalization method to eliminate dimensional differences. Normalize the original data according to the maximum and minimum value ranges of the groups belonging to the nodes, using the following formula:
[0073] ;
[0074] Let the candidate node number be Its CPU or GPU usage ratio is The minimum CPU or GPU usage ratio in the corresponding group is The maximum is After normalization:
[0075] ;
[0076] Similarly, its inflow and outflow are The minimum and maximum values of the group are respectively and Then the normalized inflow and outflow rates are:
[0077] ;
[0078] The equilibrium benchmark for the group to which the node belongs is:
[0079] ;
[0080] The molecular part is calculated as follows:
[0081] ;
[0082] The performance changes of the node over three consecutive time tags are as follows:
[0083] ;
[0084] The corresponding traffic linkage intensity is:
[0085] ;
[0086] The product of the three terms is:
[0087] ;
[0088] The sum is as follows:
[0089] ;
[0090] The normalized discrete level of the group performance variation is:
[0091] ;
[0092] The denominator is:
[0093] ;
[0094] Substitute into the sorting performance range formula:
[0095] ;
[0096] This result indicates the node Currently, there are significant deviations compared to other nodes in the same group in two key metrics: CPU or GPU usage ratio and total inbound / outbound traffic. Furthermore, it exhibits strong performance fluctuations and traffic correlation characteristics over continuous time periods. (Numerical results...) The significantly higher performance of this node compared to most other nodes in the same group indicates that it is an unstable entity with high resource consumption and a high potential for abnormal linkage during the current operation and maintenance cycle. This numerical result serves as the direct criterion for identifying nodes with outstanding ranking performance in the ranking extraction submodule. This is achieved by calculating the performance of all nodes... By sorting the nodes in descending order and selecting the top-ranked nodes, the operation and maintenance sorting node group can be derived.
[0097] Please see Figure 4 The multi-parameter linkage recognition module includes:
[0098] The time-series parameter extraction submodule is based on the operation and maintenance sorting node group. It analyzes the nodes with the highest sorting, compares the CPU and GPU usage ratio, memory usage ratio and business response time in chronological order at multiple time points, and integrates the time-series information of each node's indicators to obtain the indicator time-series curve data.
[0099] First, read the CPU or GPU usage ratio, memory usage ratio, and service response time records for each node across multiple consecutive time tags. Arrange the data chronologically and create a timeline structure. For example, for node A from T1 to T5, the records are: CPU or GPU usage ratio: 72%, 76%, 80%, 75%, 78%; memory usage ratio: 68%, 72%, 75%, 73%, 74%; service response time: 180ms, 200ms, 230ms, 215ms, 225ms. Create a snapshot of these three data items for each time tag. Then, compare the snapshots of the same node across multiple time tags sequentially, recording the direction of change for each indicator. For example, from T1 to T2, the CPU or GPU usage ratio might be 72%, 76%, 80%, 75%, 78%, 68%, 72%, 75%, 73%, 74%, 180ms, 200ms, 230ms, 215ms, 225ms. A rise from 2% to 76% is recorded as an increase; a rise from T2 to T3 to 80% is also marked as an increase; a fall from T3 to T4 to 75% is marked as a decrease, and so on, to complete the trend extraction of each indicator. The changing trends of the three indicators are matched within the same time period, that is, within each continuous label interval, it is compared whether the three indicators change in the same direction. If they are consistent, the current time period is marked as a consistent trend segment; if they are inconsistent, the identification is interrupted. All consistent trend segments are summarized, and the start and end time labels, corresponding node numbers, changing directions of the three indicators, and numerical change magnitudes of each segment are recorded. The above information is collected into indicator time series curve data, where each curve structure contains elements such as node ID, time period start and end, indicator item, trend direction, value list, and trend duration period.
[0100] The synchronization trend identification submodule calls the time series curve data of the indicators, compares the direction of change of each indicator at each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator within the synchronization interval to obtain the synchronization trend interval table.
[0101] The curve content is traversed node by node, and the direction of change is compared from front to back according to the time period. For each time label interval, trend markers for CPU or GPU usage ratio, memory usage ratio, and service response time are extracted. If node B shows an increase of 1, 2, and 3 respectively in the interval from T6 to T9, then this interval is considered to have a consistent trend direction. If, during the interval from T9 to T10, CPU or GPU usage decreases, memory usage decreases, and response time increases, then it is considered inconsistent. All segments with consistent trends are filtered and marked, and consecutive segments with consistent trends are grouped into candidate intervals for synchronous trends. The numerical differences of the changes in the three indicators within each candidate interval are calculated and recorded. The system checks whether the changes in the indicators are within a set range. If the changes in the three indicators within this segment are +8%, +6%, and +50ms respectively, the synchronization threshold is set as follows: the change in any indicator must not be less than 50% of the average value of the overall trend. For example, if the average trend is +6%, then each indicator must exceed +3% or an equivalent change. If this condition is met, the segment is confirmed as a synchronized trend segment. Within the synchronized trend segment, the indicator change pattern during this time period is further recorded, such as whether it is continuously rising, continuously falling, or changing repeatedly in cycles. The changes in each synchronized segment are organized into a record by four items: node, time period, change value, and trend direction, thus constructing a synchronized trend segment table.
[0102] The linkage link construction submodule determines the trend fluctuation of nodes on the CPU and GPU load panels and response latency layers based on the synchronization trend interval table, and calculates the correlation between each set of trends using the following formula:
[0103] ;
[0104] Obtain the total intensity of the linkage trend, determine the time-series curves of each parameter, and obtain the linkage anomaly trend chain, where, This represents the total strength of the linkage trend formed by the node group under continuous time labels. This represents the total number of nodes to be sorted, i.e., the number of nodes being analyzed. Indicates the first The change in CPU or GPU usage ratio of a node across different time tags represents the difference in CPU or GPU usage ratio between adjacent data points in the time series. Indicates the first The change in memory usage ratio of a node under different time labels is the difference in memory usage ratio between adjacent data points in the time series. Indicates the first The number of service response times changes for each node under different time labels represents the difference in service response times between adjacent data points in the time series. Indicates the first The trend item for each node in the storage array response latency layer reflects the response changes related to the storage device. Indicates the first The trend item for each node in the CPU or GPU load panel of the server group reflects the trend of CPU or GPU load changes for that node.
[0105] The total strength of the linkage trend refers to a comprehensive quantitative result of the trend strength obtained by performing joint trend calculation on three core operation and maintenance indicators (CPU or GPU usage ratio, memory usage ratio, and business response time) of the top-ranked nodes under multiple time tags during the analysis process. This result is used to measure whether there is a phenomenon of strong consistency in trend fluctuations and obvious linkage characteristics among multiple nodes within a certain time range.
[0106] Based on the synchronization trend interval table, determine the trend fluctuation data of each node in the CPU or GPU load panel and storage array response latency layer of the server group. For each node marked with the same direction, collect the change amplitude of three types of indicators within its corresponding time period, namely, the change in CPU or GPU usage ratio. Changes in memory usage Changes in business response time Because the three types of indicators mentioned above have different units, among which and In percentage terms, while Measured in milliseconds, all participating items need to be normalized before combined calculations. The maximum value normalization method is used, dividing the change in each indicator by its maximum value in the sample node. If the three maximum changes in the sample node are in the following order: Based on the synchronization trend interval table, determine the trend fluctuation data of each node in the CPU or GPU load panel and storage array response latency layer of the server group. For each node marked with the same direction, collect the change amplitude of three types of indicators within its corresponding time interval, namely the change in CPU or GPU usage ratio. Changes in memory usage Changes in business response time Because the three types of indicators mentioned above have different units, among which and In percentage terms, while Therefore, since the calculations are in milliseconds, all participating items need to be normalized before the combined calculations. The maximum value normalization method is used, which divides the change in each indicator by its maximum value in the sample node. If the maximum changes of the three items in the sample node are, in order... , , If the monitored value of a certain node (such as node 01) is , , The normalization result is: , , Then read the trend data item of that node in the response delay layer. and CPU or GPU load panel trend items Assuming , The difference term is calculated as follows: Substituting the above data into the formula, the calculation for node 01 is as follows:
[0107] The joint squared terms are: ;
[0108] The response adjustment item is: ;
[0109] The trend difference item is: ;
[0110] The final sub-item is: ;
[0111] If node 02 and node 03 are calculated respectively Given values of 2.415 and 2.536, the overall trend strength of the node group is:
[0112] ;
[0113] This value represents the overall level of linkage intensity formed by the node group within the current time label segment. It can serve as an important basis for subsequent plotting of node time-series curves and extraction of linkage link segments, thereby obtaining the linkage anomaly trend links. The formula is obtained through... This item strengthens the feedback of the joint trend of computing node load changes, through The impact of drastic changes in response time should be reasonably amplified and further combined with The differences in the fluctuation coupling between storage and computing nodes are quantified to generate an indicator that reflects the cross-system trend resonance relationship. It is used to assist in identifying aggregation regions of node-coordinated fluctuations.
[0114] Please see Figure 5 The resource node screening module includes:
[0115] The unmarked node submodule is based on the linkage anomaly trend link. It calls the nodes it covers, determines the resource segment associated with the node, filters the nodes in the same segment that have not been marked as anomaly, optimizes the node status screening process, and obtains the resource test number set.
[0116] Extract the node numbers of the nodes involved in the linkage anomaly from the link record, confirm that each node is currently in the anomaly linkage sequence, and obtain its resource segment information in the node archive table. Assuming nodes A, B, and C are identified as anomalous nodes, and their resource segments are segment 1, segment 2, and segment 1 respectively, then segment 1 and segment 2 are identified as the currently involved resource segments. Then, read the complete resource topology diagram and retrieve all node sets contained under segment 1 and segment 2. Assuming segment 1 contains nodes A, C, D, and E, and segment 2 contains nodes B, F, G, and H, where nodes D, E, F, G, and H are nodes not currently participating in the linkage trend, their node status field is used to determine if they are not marked as anomalous, i.e., they do not appear in the linkage trend link record. Subsequently, a screening condition is constructed: nodes not marked as anomalous in the current resource segment. Node numbers meeting the condition are incremented. The nodes are added to the list to be screened. Then, the unmarked nodes are initially assessed for their status. The CPU or GPU utilization, memory usage, and total inbound and outbound traffic changes for the last three time tags are read. If node E's CPU or GPU usage changes from 62% to 70% to 77%, memory usage from 60% to 66% to 72%, and inbound and outbound traffic from 2.1GB to 2.4GB to 2.8GB during the time period T11 to T13, it is determined that it has a growth pattern similar to an abnormal trend and is marked as a potential target of attention. If the change of node F is lower than the set minimum fluctuation threshold (e.g., CPU or GPU difference less than 10%, traffic change less than 0.5GB), it is determined that the node does not have any abnormal signs and is not included in the scope of attention. Finally, the node numbers of all unmarked but significantly fluctuating nodes in all resource segments are integrated to form a set of resource numbers to be tested.
[0117] The memory traffic linkage submodule compares the changes in memory usage ratio and total inbound / outbound traffic of each node at consecutive time points based on the resource test set, using the following formula:
[0118] ;
[0119] The joint trend difference of each node is obtained, and then numbered and mapped according to the resource segment affiliation to obtain the segment trend difference sequence, where, Indicates the first The joint trend difference of each node Represents a node The change in memory usage percentage between adjacent time tags reflects the trend of node memory load. Represents a node The magnitude of change in total inbound and outbound traffic reflects the trend of network traffic changes at a node. This refers to the number of nodes participating in this round of analysis;
[0120] Compare the memory usage ratio and total inbound / outbound traffic of each node under two consecutive time labels, and calculate the memory change rate of each node. Variation in inflow and outflow rates Then, normalize the two types of participating terms separately, and substitute them into the formula. If there are five nodes with the following numbers: to ;
[0121] The change in the node's memory usage percentage between the two time tags is as follows: ;
[0122] The corresponding change in total inflow and outflow volume is ;
[0123] Normalized They are respectively , They are respectively ;
[0124] Substitute it into the denominator and calculate:
[0125] ;
[0126] Substituting the whole term into the denominator, we get:
[0127] ;
[0128] Calculate each node sequentially. value:
[0129] ;
[0130] ;
[0131] ;
[0132] ;
[0133] ;
[0134] This result indicates that the joint trend difference among nodes... The larger the value, the more significant the change in the memory usage ratio of that node. Changes in total inflow and outflow There is a higher degree of synchronous fluctuation between them, that is, the two indicators show the same direction and amplitude linkage characteristics under continuous time labels. This indicator reflects the coupling behavior strength of the node in multiple operation and maintenance data.
[0135] The numbering intersection extraction submodule identifies the top-ranked node numbers based on the segment trend difference sequence, analyzes the intersection of the nodes with the ranked nodes, determines the distribution of the intersection numbers in the node set, and obtains the resource regulation number group.
[0136] The joint trend difference refers to the strength of the correlation between the change in the memory usage ratio of the same node and the change in the total inbound and outbound traffic under continuous time labels. This quantity is used to measure the synchronicity and coupling of nodes during changes in resource pressure (memory) and network load (traffic).
[0137] Extract the node IDs with the highest fluctuation amplitude in each resource segment. Let's say nodes D and E in segment 1 rank first and second in the comprehensive evaluation of memory usage and traffic changes, with differences of 16% and 1.3GB, and 14% and 1.2GB, respectively. Add the IDs of D and E to the list of top-ranked nodes according to the difference ranking rules. Simultaneously, read the set of node IDs previously marked as having abnormal linkage in the ranked node list. Compare the IDs of the top-ranked node ID set with the set of nodes with abnormal linkage, identifying duplicate IDs. For example, if node E appears in both sets, it is identified as an intersection node. Further analyze the distribution of all intersection IDs within their respective resource segments. The specific analysis method is as follows: The method involves statistically analyzing the proportion of intersection numbers in each segment and the proportion of each number in the total number of original nodes. For example, if segment 1 has 6 nodes and the intersection numbers are nodes E and A, accounting for 2 / 6 (33.3%), then a significant number intersection phenomenon exists in segment 1. Based on this result, a list of node intersection structures is compiled. Then, a comprehensive judgment is performed based on the resource location, indicator trend value, and sorting order of each intersection node to identify the most representative intersection numbers. Priority is set to filter the nodes ranked higher in the intersection. If the indicator trend value of node E in the intersection nodes is higher than that of A, then E is added to the resource control number group as the final retained node. At the same time, its sorting value, resource ownership, number of intersections, trend direction, and other information are marked, and the resource control number group is output.
[0138] Please see Figure 6 The event tracking module includes:
[0139] The trajectory data extraction submodule analyzes the recent CPU and GPU usage ratios, disk read / write activity, and business response time-series records based on resource control number groups, and organizes them into continuous node operation performance to obtain the node operation sequence.
[0140] The system reads the operational data records of each included node over the past three days. Specifically, it retrieves the CPU or GPU usage ratio, disk read / write time, and service response time data for each node at time tags T1 to T24 (with a sampling interval of 3 hours) from storage. A time series table structure is created for each data item, grouped by node number. Within each group, the sampled data is arranged chronologically and filled into an indicator matrix to form a continuous performance curve for each node. For example, node A's CPU or GPU usage at time points T1 to T5 is 58%, 63%, 68%, 64%, and 70%, respectively; its disk read / write times are 290ms, 310ms, 345ms, 330ms, and 375ms, respectively; and its service response time is 180ms. The data entries for this node are 195ms, 210ms, 200ms, and 215ms. The continuous time series data entries are complete and consist of the hourly values of three indicators. The trend direction of each indicator is identified and the slope of change is marked. For example, if the CPU or GPU generally increases from T1 to T5, the disk read / write time increases, and the response time also increases periodically, this set of data is regarded as a node with a continuous positive growth trend. If the CPU or GPU sequence of another node B fluctuates at 65%, 60%, 67%, 62%, and 68%, and the direction of change is unstable, it is marked as an unstable node. The time series of the three types of indicators of all nodes are written into their node operation sequence, and after being uniformly numbered based on the time label, they are stored in the indicator trajectory database.
[0141] The fluctuation curve comparison submodule compares the changing trends of each indicator of each node with the linked fluctuation node curve in the node operation sequence, judges the synchronous changes of each node indicator curve under the same time sequence, identifies the node performance with highly consistent trends, and obtains the synchronous characteristic node group.
[0142] For each node, a single metric curve is constructed for CPU or GPU, disk read / write time, and response time. The direction of change sequence is calculated for two adjacent time stamps, generating a direction vector sequence. For example, if the direction vector of node A's CPU or GPU curve from T1 to T5 is [+1, +1, -1, +1], and the corresponding direction vector of the linked node L is [+1, +1, -1, +1], then node A is completely consistent with this metric. If its disk and response metric direction vectors are also completely consistent with node L, or if the difference is only within a single time point, this node is initially identified as a trend-synchronized node. Further analysis is then performed on the slope of the metric changes for each node. The similarity score within the rate of change interval is calculated separately. If the slope error of node A and node L is less than 5% in at least 4 out of 5 consecutive segments across the three indicators, the node is judged to have high consistency with the linkage curve. The node is then included in the candidate set of synchronous feature nodes. Subsequently, all nodes that maintain trend synchronization with the linkage curve across the three indicators and whose errors are all within the acceptable threshold range are counted. If nodes A, C, and E meet this condition while B and D are only partially consistent, then A, C, and E are identified as the synchronous feature node group. Each node in this group records the length of the synchronous segment of the indicator curve, the proportion of consistent change direction, and the average trend offset value for use in the next stage of screening.
[0143] The trajectory information generation submodule screens key nodes with synchronous changes based on the synchronous feature node group, optimizes the abnormal segments and change characteristics of each node in the time series curve, and obtains the trajectory information of the associated nodes.
[0144] First, extract the synchronous change segments from the CPU or GPU, disk I / O, and response time series of each node. These segments are those where the three indicators show consistent changes in a continuous time range. For example, if node C shows a continuous increase in CPU or GPU, disk I / O, and response values from T6 to T10, with a direction vector of [+1, +1, +1, +1], then this segment is marked as a candidate segment for anomalies. Next, evaluate the magnitude of changes within this anomaly segment. If the growth rate of any two of the three indicators exceeds a preset threshold (e.g., CPU or GPU increase exceeding 12%, response time increase exceeding 40ms), then this segment is confirmed as an anomaly. This time range is then marked as a high-fluctuation anomaly zone in node C's curve. All indicator values, time labels, direction vectors, and magnitude data for this node within this segment are included in the trajectory information item. Repeat this process to traverse all synchronous feature nodes, forming a complete list of trajectory anomaly segments. Finally, summarize the indicator change characteristics, maximum and minimum value points, volatility, and comparison with adjacent nodes exhibited by each node in its respective anomaly segment to complete the generation of the associated node trajectory information.
[0145] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A visual operation and maintenance data analysis system, characterized in that, The system includes: The node data collection module is based on the backbone network of the intelligent computing center. It analyzes the operation data of each node, groups the nodes by time label, arranges the changes in CPU and GPU usage ratios and memory usage ratios in sequence, and archives disk read / write time, total inbound / outbound traffic and business response time to obtain a set of node indicator sequences. The indicator sorting and filtering module compares the CPU and GPU usage ratios and total inbound and outbound traffic of each node under the same time tag based on the node indicator sequence set, summarizes the differences, sorts them by type, identifies the nodes with the highest ranking, and obtains the operation and maintenance sorted node group. Based on the operation and maintenance sorting node group, the multi-parameter linkage identification module analyzes the synchronous changes of CPU and GPU usage ratios, memory usage ratios and business response times of the leading nodes in the current and adjacent time tags, determines the time-series curves of each parameter, and obtains the linkage anomaly trend link. Based on the aforementioned linkage anomaly trend link, the resource node screening module filters unmarked nodes in the associated resource segment, compares the memory usage ratio trend and the fluctuation range of inbound and outbound traffic, and determines the node number by combining the sorting intersection to obtain the resource regulation number group.
2. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The node indicator sequence set includes node performance data, archived time series, and group classification information. The operation and maintenance sorting node group includes priority identifier, node identification code, and sorting index number. The linkage anomaly trend link includes linkage feature sequence, abnormal fluctuation segment, and trend matching identifier. The resource control number group includes control node list, allocation node code, and redundant resource identifier.
3. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The node data collection module includes: The data acquisition submodule is based on the backbone network of the intelligent computing center. It analyzes the real-time data collected on the CPU and GPU usage ratio, memory usage ratio, disk read and write time, total inbound and outbound traffic and business response time of various nodes. The data is organized according to time tags to generate a set of operating indicators. The joint arrangement submodule compares the correlation between the CPU and GPU usage ratios and memory usage ratios of each node under the same time tag based on the set of operating indicators, and arranges them in sequence according to the changing trend of each indicator to obtain a joint changing trend group. The indicator archiving submodule filters the disk read / write time, total inbound / outbound traffic, and business response time of each type of node under continuous time labels based on the joint change trend group, and archives the parameter trends to obtain a set of node indicator sequences.
4. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The index sorting and filtering module includes: The difference summarization submodule analyzes the CPU and GPU usage ratios and total inbound and outbound traffic of the same type of nodes under each time tag based on the node indicator sequence set, compares the fluctuation range of the corresponding parameters of each node in the group, calculates the difference magnitude of parameters between different nodes, and obtains the performance change magnitude. The group comparison submodule calls the performance change magnitude, groups nodes according to node type, compares the difference distribution of CPU and GPU usage ratio and total inbound and outbound traffic of each group of nodes under the same time label, judges the distribution breadth of parameters among nodes in the same group, filters out nodes with key performance differences, and obtains the difference distribution range. The sorting and extraction submodule analyzes the CPU and GPU usage ratios, total inbound and outbound traffic, and group parameter benchmarks of nodes under the current time tag based on the difference distribution interval. It adjusts the impact of historical performance changes and linkage fluctuations within the group, identifies nodes with outstanding sorting performance, and uses them as key operation and maintenance analysis objects to obtain the operation and maintenance sorting node group.
5. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The multi-parameter linkage recognition module includes: The time-series parameter extraction submodule analyzes the nodes ranked first in the operation and maintenance sorting node group, compares the CPU and GPU usage ratio, memory usage ratio and business response time in chronological order at multiple time points, and integrates the time-series information of each node's indicators to obtain indicator time-series curve data. The synchronization trend identification submodule calls the time series curve data of the indicators, compares the direction of change of each indicator at each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator within the synchronization interval to obtain the synchronization trend interval table. The linkage link construction submodule determines the trend fluctuation of nodes on the CPU and GPU load panel and response latency layer based on the synchronization trend interval table, calculates the correlation between each group of trends, obtains the total linkage trend intensity, determines the time series curve of each parameter, and obtains the linkage abnormal trend link.
6. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The resource node screening module includes: The unmarked node submodule of the section, based on the linked abnormal trend link, calls the nodes it covers, determines the resource section associated with the node, filters the nodes in the same section that have not been marked as abnormal, optimizes the node status screening process, and obtains the resource test number set. The memory traffic linkage submodule compares the memory usage ratio and total inflow / outflow of each node at consecutive time points based on the resource test number set, obtains the joint trend difference of each node, and then performs number mapping according to the resource segment affiliation to obtain the segment trend difference sequence. The number intersection extraction submodule identifies the top-ranked node numbers based on the segment trend difference sequence, analyzes the intersection of the nodes with the ranked nodes, determines the distribution of the intersection numbers in the node set, and obtains the resource regulation number group.
7. The visual operation and maintenance data analysis system according to claim 1, characterized in that, The system also includes: The event trajectory tracing module retrieves time-series data on node CPU and GPU usage ratios, disk read / write times, and business response times based on the resource control number group, and compares it with the linked fluctuating node curves to obtain the trajectory information of the associated nodes. The associated node trajectory information includes historical trajectory sequences, node association features, and event archive identifiers.
8. The visual operation and maintenance data analysis system according to claim 7, characterized in that, The event trajectory tracing module includes: The trajectory data extraction submodule analyzes the recent CPU and GPU usage ratios, disk read / write activity, and business response time records based on the resource control number group, and organizes them into continuous node operation performance to obtain the node operation sequence. The fluctuation curve comparison submodule compares the changing trends of each indicator of each node in the node operation sequence with the curve of the linked fluctuation node, judges the synchronous changes of each node indicator curve in the same time sequence, identifies the node performance with highly consistent trends, and obtains the synchronous characteristic node group. The trajectory information generation submodule screens nodes with key synchronous changes based on the synchronous feature node group, optimizes the abnormal segments and change characteristics of each node in the time series curve, and obtains the trajectory information of the associated nodes.
Citation Information
Patent Citations
Data analysis method and device, electronic equipment and storage medium
CN118394603A
Management and control method, device and equipment of intelligent operation and maintenance terminal and storage medium
CN119065946A