Visual operation and maintenance data analysis system
By adopting modules such as node data collection, index sorting and multi-particle linkage identification in the visual operation and maintenance data analysis system, the problem of multi-parameter trend linkage identification of operation and maintenance data analysis in the existing technology is solved, and dynamic linkage merger and traceability analysis of operation and maintenance data is realized, improving operation and maintenance response speed and systemativity.
Patent Information
- Application Number
- CN202510913770.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-07-03
AI Technical Summary
When facing multi-parameter trend linkage and complex scenarios, existing visual operation and maintenance data analysis systems are difficult to automatically identify and prioritize screening, resulting in repeated search and manual comparison of operation and maintenance personnel, and systematic response speed and management are limited.
The node data collection module, index sorting screening module, multi-account linkage identification module and resource node screening module are adopted to analyze node data through the backbone network of the Intelligent Computing Center, group by time tag, identify linkage abnormal trend links, filter resource regulation number groups, and realize dynamic linkage merger and traceability analysis of operation and maintenance data.
It improves the intuitiveness and globality of operation and maintenance data analysis, realizes systematic closed-loop monitoring of operation and maintenance incident response, and promotes the accuracy of node status and resource allocation.
Smart Images

Figure CN120407350A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of automated operation and maintenance technology, and in particular to a visual operation and maintenance data analysis system. Background Art
[0002] The field of automated operations and maintenance is a key branch of information technology management. Its core areas include unified management of IT system assets, automated execution of operations and maintenance tasks, automated deployment and upgrades of server software, configuration file management, performance data collection and processing, real-time status monitoring, dynamic scheduling of resource utilization, fault diagnosis and automatic recovery, user authority control and security assurance, and multi-dimensional log management. This technical field is dedicated to improving the operational efficiency and management level of large-scale enterprise IT infrastructure through automated means. Traditional visual operations and maintenance data analysis systems collect and organize IT operations and maintenance data within an automated operations and maintenance platform, using indicator calculation methods to perform statistical analysis on basic data such as server CPU and GPU utilization, memory usage, disk read and write performance, network bandwidth utilization, application response time, error rate, and business traffic. The analysis results are then output in the form of charts and reports to facilitate operations and maintenance personnel in observing the system's operating status.
[0003] The current operation and maintenance data analysis model of existing technologies is mainly based on the individual statistics and static distribution display of indicators. It lacks automatic identification and priority screening of multi-parameter trend linkage. Node classification is superficial archiving, which makes it difficult to cope with complex scenarios with frequent dynamic changes of nodes or uneven distribution of resource usage. Tracing the source analysis relies on breakpoint data comparison, which makes it difficult to trace the event chain and the impact path formed by high-risk nodes. As a result, operation and maintenance personnel need to repeatedly search and manually compare, and the response speed and management system are limited. Summary of the Invention
[0004] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a visual operation and maintenance data analysis system.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a visual operation and maintenance data analysis system, the system comprising: The node data collection module, based on the backbone network of the intelligent computing center, analyzes the operating data of each node, groups node types by time tags, arranges the changes in the CPU and GPU usage ratios and memory usage ratios, and archives disk read and write time, total inbound and outbound traffic, and service response time to obtain a node indicator sequence set. The indicator sorting and screening module compares the CPU and GPU usage ratios and the total inbound and outbound traffic of each node under the same time tag based on the node indicator sequence set, summarizes the differences, groups and sorts them by type, identifies the top-ranked nodes, and obtains the operation and maintenance ranked node group; Based on the operation and maintenance sorting node group, the multi-parameter linkage recognition module analyzes the synchronous changes in the CPU and GPU usage ratios, memory occupancy ratio, and service response duration of the top-ranked nodes at the current and adjacent time tags, determines the time series curves of each parameter, and obtains the linkage abnormal trend link; Based on the linkage abnormal trend link, the resource node screening module screens the unlabeled nodes in the associated resource section, compares the memory occupancy ratio trend and the fluctuation range of the in-out traffic, and combines the sorting intersection to determine the node numbers, obtaining the resource regulation number group.
[0006] The improvements of the present invention are that the node index sequence set includes node performance data, archived time series, and grouping and classification information, the operation and maintenance sorting node group includes priority identification, node identification code, and sorting index number, the linkage abnormal trend link includes linkage feature sequence, abnormal fluctuation section, and trend matching identification, and the resource regulation number group includes regulation node list, allocation node code, and redundant resource identification.
[0007] The improvements of the present invention are that the node data collection module includes: The data collection sub-module analyzes the real-time collected content of the CPU and GPU usage ratios, memory occupancy ratio, disk read-write time, total in-out traffic, and service response duration of various nodes based on the backbone network of the intelligent computing center, sorts the data based on the time tag as a benchmark, and generates an operation index set; The joint arrangement sub-module compares the correlation between the CPU and GPU usage ratios and the memory occupancy ratio of each node under the same time tag according to the operation index set, and arranges them in sequence in combination with the change trend of each index, obtaining the joint change trend group; The index archiving sub-module screens the disk read-write time, total in-out traffic, and service response duration of each type of node under continuous time tags according to the joint change trend group, archives the parameter trends, and obtains the node index sequence set.
[0008] The improvements of the present invention are that the index sorting and screening module includes: The difference induction sub-module analyzes the CPU and GPU usage ratios and the total in-out traffic of the same type of nodes under each time tag based on the node index sequence set, compares the fluctuation ranges of the corresponding parameters of each node in the group, calculates the difference amplitude of the parameters between the different nodes, and obtains the performance change amplitude; The grouping comparison sub-module calls the performance change amplitude, groups according to the node type, compares the difference distributions of the CPU and GPU usage ratios and the total in-out traffic of each group of nodes under the same time tag, judges the distribution breadth between the parameters of the nodes in the same group, screens the nodes with key performance differences, and obtains the difference distribution interval; The sorting and extraction submodule analyzes the CPU and GPU usage ratios, total inbound and outbound traffic, and grouping parameter benchmarks of the nodes under the current time tag based on the difference distribution interval, adjusts the historical performance changes and linkage fluctuations within the group, identifies nodes with outstanding sorting performance as key operation and maintenance analysis objects, and obtains the operation and maintenance sorting node group.
[0009] The present invention is improved in that the multi-parameter linkage identification module includes: The timing parameter extraction submodule analyzes the top-ranked nodes based on the operation and maintenance ranking node group, compares the CPU and GPU usage ratios, memory usage ratios, and service response times at multiple time points in chronological order, and integrates the timing information of each node indicator to obtain indicator timing curve data; The synchronization trend identification submodule calls the indicator time series curve data, compares the change direction of each indicator of each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator in the synchronization interval to obtain a synchronization trend interval table; The linkage link construction submodule determines the trend fluctuation of the node on the CPU and GPU load panels and the response delay layer according to the synchronization trend interval table, calculates the correlation degree between each group of trends, obtains the total linkage trend strength, determines the timing curve of each parameter, and obtains the linkage abnormal trend link.
[0010] The present invention is improved in that the resource node screening module includes: The unmarked node submodule of the segment is based on the linkage abnormal trend link, calls the nodes covered by it, determines the resource segment associated with the node, filters the nodes in the same segment that are not marked as abnormal, optimizes the node status screening process, and obtains the resource number set to be tested; The memory-flow linkage submodule compares the memory usage ratio and the total inflow and outflow changes of each node at consecutive time points based on the resource number set to be tested, obtains the joint trend difference of each node, and then performs number mapping based on the resource segment affiliation to obtain the segment trend difference sequence; The number intersection extraction submodule identifies the top-ranked node numbers according to the segment trend difference sequence, analyzes the intersection of the nodes and the sorted nodes, determines the distribution of the intersection numbers in the node set, and obtains the resource control number group.
[0011] The present invention is improved in that the system further comprises: The event trajectory tracing module retrieves the time series data of the node CPU and GPU usage ratio, disk read and write time, and business response time based on the resource control number group, and compares it with the linkage type fluctuation node curve to obtain the associated node trajectory information; The associated node trajectory information includes historical trajectory sequence, node association characteristics, and event archive identification.
[0012] The present invention is improved in that the event trajectory tracing module includes: Based on the resource regulation number group, the trajectory data extraction sub-module analyzes the time series records of the recent CPU and GPU usage ratios, disk read and write conditions, and service response conditions, organizes them into the continuous operation performance of nodes, and obtains the node operation sequence; The fluctuation curve comparison sub-module compares the change trends of each index of each node in the node operation sequence with the curve of the linkage type fluctuation node, judges the synchronous change situation of each node index curve at the same time series, identifies the node performance with highly consistent trends, and obtains the synchronous feature node group; The trajectory information generation sub-module screens the nodes that are key to synchronous changes according to the synchronous feature node group, optimizes the abnormal sections and change characteristics of each node in the time series curve, and obtains the associated node trajectory information.
[0013] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In the present invention, by hierarchically grouping and archiving the operation data of different nodes, and combining time tags to realize the dynamic linkage merging of various operation and maintenance parameters, the objects of operation and maintenance attention are automatically screened under the multi-index change trend, and the abnormal nodes and resource redundant nodes are accurately screened according to the comprehensive feature trajectory, promoting the transformation of operation and maintenance data from single-point monitoring to full-linkage analysis, improving the intuitiveness and overall situation of operation and maintenance event response in the process of visual archiving and traceability, so as to promote the systematic closed-loop monitoring of node status, resource allocation and operation and maintenance risks. Brief Description of the Drawings
[0014] Figure 1 is the system flow chart of the present invention; Figure 2 is the flow chart of the node data collection module in the present invention; Figure 3 is the flow chart of the index sorting and screening module in the present invention; Figure 4 is the flow chart of the multi-parameter linkage recognition module in the present invention; Figure 5 is the flow chart of the resource node screening module in the present invention; Figure 6 is the flow chart of the event trajectory tracing module in the present invention. Detailed Embodiments
[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0016] In the description of the present invention, it should be understood that the terms "length", "width", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings and are only for the convenience of describing the present invention and simplifying the description. They do not indicate or imply that the devices or elements referred to must have a specific direction, be constructed and operate in a specific direction, and therefore should not be understood as limiting the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0017] Example: See Figure 1 The present invention provides a technical solution: a visual operation and maintenance data analysis system comprising: The node data collection module, based on the backbone network of the intelligent computing center, collects operational data from computing nodes, storage nodes, and transmission nodes. It groups node types according to time tags and arranges the combined changes in CPU and GPU usage ratios and memory usage ratios in chronological order. It also archives disk read and write time, total inbound and outbound traffic, and service response time by node type to obtain a node indicator sequence set. The indicator sorting and filtering module compares the CPU and GPU usage ratios and the total inbound and outbound traffic of each node under the same time tag based on the node indicator sequence set. It summarizes the differences in each data item, groups and sorts them by node type, and identifies the top-ranked nodes in each group as key operation and maintenance analysis objects to obtain the operation and maintenance ranked node group. The multi-parameter linkage identification module analyzes the synchronous changes in the CPU and GPU usage ratios, memory usage ratios, and service response times of the top five nodes under the current and adjacent time tags based on the operation and maintenance ranking node group. Using the CPU and GPU load panels of the server group and the response delay layer of the storage array, it determines the timing curves of each parameter and obtains linkage anomaly trend links. The resource node screening module uses the linkage abnormal trend link to screen unmarked nodes in the resource segment to which the associated node belongs. It compares the trend of changes in their memory usage ratio with the fluctuation range of the total inflow and outflow volume, and retains the top node numbers based on the sorted intersection to obtain the resource control number group. The event trajectory tracing module retrieves the continuous time series data of the node's CPU and GPU usage ratio, disk read and write time, and business response time in the past three days based on the resource control number group, and synchronously compares the time series data with the linkage-type fluctuation node curve for visual archiving and traceability analysis of operation and maintenance events to obtain related node trajectory information.
[0018] The node indicator sequence set includes node performance data, archive time series, and group classification information. The operation and maintenance sorting node group includes priority identification, node identification code, and sorting index number. The linkage abnormal trend link includes linkage feature sequence, abnormal fluctuation segment, and trend matching identification. The resource control number group includes the control node list, allocation node code, and redundant resource identification. The associated node trajectory information includes historical trajectory sequence, node association characteristics, and event archiving identification.
[0019] In Module 1, the backbone network of the intelligent computing center refers to the high-bandwidth, high-reliability network infrastructure built within the intelligent computing center, connecting all computing nodes, storage nodes, and transmission nodes, and serving as the main channel for all data flow, collection, and scheduling. Computing nodes are server resource units equipped with processors and memory, responsible for data processing and business operations. Storage nodes exist in the form of disk arrays or distributed storage servers, primarily responsible for data persistence and read and write operations. Transmission nodes are network devices such as switches and routers, primarily responsible for the flow of data between nodes in the intelligent computing center. Time tags record the time point of data collection and are used for subsequent time series alignment and trend analysis of data from different nodes. Node types refer to three categories: "computing nodes," "storage nodes," or "transmission nodes," divided by hardware function and role. Joint changes refer to the linkage trend analysis of two indicators, such as the CPU or GPU usage ratio and the memory usage ratio, for the same node at the same time tag (for example, whether they increase or decrease simultaneously). Business response time refers to the total time it takes for a business request to be initiated and responded to, which usually reflects the response efficiency of the application service and the user experience.
[0020] In Module 2, each data difference refers to the difference between the same type of parameters across nodes at the same time tag. For example, this could be the difference in CPU or GPU usage or the difference in total inbound and outbound traffic between different compute nodes. The top-ranked nodes are those with the most prominent performance (e.g., high CPU or GPU load or high traffic volume) after summarizing and ranking the differences. These nodes are typically the focus of operations and maintenance.
[0021] In Module 3, the top five nodes refer to the key nodes that are ranked in the top five, usually objects with high load, abnormal status, or large business impact; synchronous changes refer to multiple indicators (such as CPU or GPU usage ratio, memory usage ratio, and business response time) showing changes in the same direction in the time series (such as synchronous increase or decrease), reflecting the linkage effect of the system or business; the response delay layer refers to a graphical interface or data layer that visually displays the response time distribution of devices such as storage arrays, which is used to identify storage bottlenecks or anomalies; the timing curve refers to a trend line that reflects the change of a certain parameter (such as CPU or GPU usage ratio, response time, etc.) over time, which is used to observe the change pattern and fluctuation characteristics of the indicator.
[0022] In Module 4, the resource section refers to the resource partition (such as a certain cabinet, computer room or service pool) in the intelligent computing center divided by physical location or business function, which is used for local resource management and scheduling; the unmarked nodes refer to the nodes that are not classified as abnormal or key - concerned nodes during the abnormal recognition and linkage recognition processes, that is, the nodes in the "normal" state; the change trend refers to the continuous change direction (rising, falling or stable) of an indicator (such as the memory occupancy ratio) over a period of time to judge the resource load trend; the fluctuation range refers to the maximum change range of an indicator within the observation period, which is used to identify the nodes with drastic load changes; the sorting intersection refers to cross - comparing the sorting results of different indicators to screen the nodes that are outstanding in multiple dimensions.
[0023] In Module 5, the continuous time - series data refers to the historical acquisition data sequence of a node under multiple consecutive time tags, such as the CPU or GPU usage ratio change curve of a node within three days; the linkage - type fluctuating nodes refer to the nodes marked as needing attention in the aforementioned linkage recognition module due to the synchronous fluctuation of multiple indicators.
[0024] Please refer to Figure 2 , the node data collection module includes: The data acquisition sub - module, based on the backbone network of the intelligent computing center, analyzes the real - time acquisition content of the CPU and GPU usage ratios, memory occupancy ratios, disk read - write times, total in - out traffic and service response durations of various nodes, sorts the data based on time tags, and generates a set of operation indicators; Define the category attributes of each node and determine the collection scope. For example, node A is a computing node, node B is a storage node, and node C is a transmission node. Corresponding running metrics such as CPU or GPU usage rate, memory occupancy rate, disk operation time, network traffic, and business response duration are obtained. During the one-minute collection cycle, monitor the CPU or GPU usage rate of node A. By reading the real-time feedback of the node operation on the processor running state, calculate the ratio of the active processing duration to the total running duration within one minute. If the active processing time of node A is 2,825 milliseconds and the total processing time for the whole minute is 10,000 milliseconds, then the CPU or GPU usage rate is 28.25%. During the same cycle, the memory occupancy rate is calculated based on the ratio between the used memory and the total memory recorded by the operating system. If the total memory of node A is 32 GB and the actual usage is 12 GB, then the memory occupancy rate is 37.5%. The collection method for disk read and write time is to count the total elapsed time of disk interface read and write operations. If node B performs 50 read and write operations on the disk within one minute and the average elapsed time for each operation is 7 milliseconds, then the total read and write time is 350 milliseconds. The in-and-out traffic of transmission node C is obtained by counting the input and output bytes read from the network interface. If the received traffic is 2.4 GB and the sent traffic is 1.2 GB within one minute, then the total traffic is 3.6 GB. The collection of business response duration is achieved through log analysis, recording the time difference between when the request is received and when the response is returned. If there are five business calls within one minute and the response durations are 140, 180, 210, 170, and 200 milliseconds respectively, then the average business response is 180 milliseconds. All the above data are tagged according to the sampling time to form an index data packet based on time tags and classified under their respective node numbers to form a set of running metrics under the node type classification.
[0025] According to the set of running metrics, the combined permutation sub-module compares the correlation between the CPU and GPU usage ratios and the memory occupancy ratios of each node under the same time tag, and arranges them in sequence in combination with the change trend of each metric to obtain a combined change trend group; Merge the data of all nodes according to the time tags, and compare the computing nodes with the same time tags in sequence. For example, at time point T5, the CPU or GPU utilization rate of node X is 65%, and the memory occupancy rate is 58%. For node Y, they are 82% and 71%, and for node Z, they are 78% and 64%, respectively forming the index combinations of each node at this moment. Conduct trend comparison for each pair of the above combinations. First, judge whether the CPU or GPU and memory occupancy rate rise or fall simultaneously in the previous time period. For example, if the CPU or GPU of node Y is 72% and the memory is 63% at time point T4, and it becomes 82% and 71% at T5, it indicates that its indicators show a combined upward trend. Similarly process node Z and X, identify all nodes with the same direction of change, and calculate the change range of the two indicators of each node. If the CPU or GPU of node Y increases by 10% and the memory increases by 8%, and the CPU or GPU of node Z increases by 9% and the memory increases by 7%, it can be determined that the change of node Y is the most drastic in this cycle. Sort the nodes according to the change range to form a sorting sequence. In actual operation, when the objects ranked at the forefront in the node sorting continuously show a similar trend, they can be identified as the objects with a significant combined change trend in this round. Each record entry in this sorting group contains the node number, time tag, index change direction, and combined change range, which are used to form the combined change trend group.
[0026] According to the combined change trend group, the index archiving sub-module filters the disk read / write time, total in / out traffic volume, and service response duration of each type of node under consecutive time tags, archives the parameter trends, and obtains the node index sequence set. The index archiving sub-module extracts the top five nodes sorted from the combined change trend group, monitors and archives their index trends within several subsequent time tags. For example, select node Y in the time period from T5 to T9, and record its disk read / write time as 300, 340, 370, 410, 460 milliseconds in sequence, the in / out traffic as 3.1, 3.5, 3.9, 4.0, 4.2 GB, and the service response duration as 190, 205, 218, 230, 240 milliseconds. Establish a data list for each index of this node at each time point and draw a change trend curve, and manually analyze its growth rate. For example, the disk write time shows an increasing trend in each cycle, and the total increase ranges from 300 milliseconds to 460 milliseconds, the traffic increases from 3.1 GB to 4.2 GB, and the response time increases from 190 milliseconds to 240 milliseconds. This indicates that the load of this node increases and the performance decreases in this cycle. Record the values of the above three indicators at each time tag to form trend data, and mark it as the index sequence of node Y from T5 to T9, and further archive it into the corresponding index sequence set of this node to form a basic data set for continuous trend analysis. The structure of the index sequence set is in the order of time tags and columns of index items, recording the absolute value and change direction of each index within a continuous time period, which is used for subsequent node screening and event identification processing.
[0027] Please refer to Figure 3 , the index sorting and filtering module includes: The difference induction sub-module analyzes the CPU and GPU usage ratios and the total in-out traffic volume of nodes of the same type under each time label based on the node index sequence set, compares the fluctuation ranges of the corresponding parameters of each node in the group, calculates the difference amplitude between the parameters of different nodes, and obtains the performance change amplitude; First, extract all computing nodes, storage nodes, and transmission nodes with the same time label from the dataset. After classifying the nodes by type, sequentially obtain the CPU or GPU usage ratio and the total in-out traffic volume recorded by each node under this time label. Then, perform a difference analysis on the parameters within each type of node. For example, under the time label T10, if the CPU or GPU usage rates of computing nodes A, B, and C are 55%, 78%, and 62% respectively, and the total in-out traffic volumes are 1.2GB, 3.0GB, and 1.5GB, then for the CPU or GPU usage rate, perform the method of subtracting the minimum value from the maximum value to obtain the fluctuation range of the CPU or GPU usage rate within this group as 78% - 55% = 23%. Similarly, the fluctuation range of the total in-out traffic volume is 3.0GB - 1.2GB = 1.8GB. Then, by summing the absolute differences between the values of each pair of nodes, calculate the difference in the CPU or GPU usage rate between nodes A and B as 23%, between nodes A and C as 7%, and between nodes B and C as 16%. And so on to complete the sum of the combined differences. If the differences in the total in-out traffic volume are 1.8GB, 0.3GB, and 1.5GB respectively, record the difference amplitude of the corresponding combination. By statistically analyzing the maximum fluctuation value and the average value of the combined differences of each type of parameter above, measure the overall performance change situation within the group. If the fluctuation value of the CPU or GPU usage rate is greater than 20% and the fluctuation value of the in-out traffic exceeds 1.5GB, then it is considered that there is a group of difference nodes with a relatively large performance change amplitude under this time label. Further, organize the node number, time label, parameter type, fluctuation range value, and node pair combination information into a change amplitude dataset for subsequent execution of group comparison.
[0028] The grouping comparison sub-module calls the performance change amplitude, groups according to the node type, compares the difference distributions of the CPU and GPU usage ratios and the total in-out traffic volume of each group of nodes under the same time label, judges the distribution breadth between the parameters of the nodes in the same group, screens the nodes that are key to the performance differences, and obtains the difference distribution interval; All nodes are independently grouped into computing, storage, and transmission categories according to the node type. Within each group, differential distribution analysis is performed based on the performance data corresponding to the same time tag. After obtaining the CPU or GPU usage ratio and the total inbound and outbound traffic volume of all nodes in a certain group at a certain time tag, a one-dimensional sequence is constructed and sorted. After the sorting is completed, the difference between the head and the tail of the sequence is compared. For example, at time T12 in the computing node group, the CPU or GPU usage rates of nodes X, Y, Z, and W are 45%, 68%, 53%, and 80% respectively. Then, the sorted sequence is constructed as 45%, 53%, 68%, 80%, and the difference is 35%. If it is set that when the difference in CPU or GPU usage rate is greater than 30%, it is considered that the distribution breadth is too large, then this node group is marked as widely distributed. The same sorting analysis is performed on the inbound and outbound traffic. If the inbound and outbound traffic of nodes X, Y, Z, and W are 1.1GB, 1.4GB, 2.2GB, and 3.8GB respectively, then the difference after sorting is 2.7GB. It is judged whether the differential distribution of the inbound and outbound traffic exceeds the preset threshold. If it is set to 2.5GB, then this group meets the requirement of abnormal distribution breadth. Subsequently, two key nodes are searched within this group, that is, the nodes near the extreme ends in each parameter distribution, which are node X with the lowest CPU or GPU usage rate and node W with the highest CPU or GPU usage rate, and node X with the smallest inbound and outbound traffic and node W with the largest inbound and outbound traffic. When the extreme nodes of the two parameters overlap, that is, both node X and W are critical performance objects, then they can be screened into the differential distribution interval record as the key nodes with significant performance differences. The record content includes the node number, the type it belongs to, the corresponding time tag, the sorting positions of the CPU or GPU and traffic distributions, the determination result of the differential threshold, and whether it constitutes an overlapping extreme node, and is stored in the differential distribution interval set.
[0029] Based on the differential distribution interval, the sorting extraction sub-module analyzes the CPU and GPU usage ratios, the total inbound and outbound traffic volume of the nodes at the current time tag, and the grouping parameter benchmark, and adjusts the historical performance changes and the impact of linkage fluctuations within the group, using the formula: ; Identify the nodes with outstanding sorting performance as the key objects for operation and maintenance analysis, and obtain the operation and maintenance sorting node group. Among them, represents the sorting performance amplitude of node , is the CPU or GPU usage ratio of node at the current time tag, is the total inbound and outbound traffic volume of node at the current time tag, and are respectively the CPU or GPU usage ratio and the inbound and outbound traffic volume balance benchmark of the group to which the node belongs , is the discrete level of the performance change amplitude of this group, Indicates the performance change of the node at the th consecutive time tag, indicating the traffic linkage strength of the node at the th time tag, indicating the number of continuously observed time tags, that is, the total number of time points for continuously tracking the node.
[0030] The sorting performance amplitude refers to the degree of deviation of the current CPU or GPU usage ratio and the total in-out traffic volume of each node from its grouping benchmark (group average level) among nodes of the same type, and is a comprehensive measure value combined with its historical performance changes and traffic linkage effects over a period of time. The sorting performance amplitude is used to reflect the strength and abnormal degree of the operation characteristics of the node among nodes of the same type. Through this indicator, key nodes worthy of key attention in the current operation and maintenance analysis can be identified. The larger the sorting performance amplitude, the more prominent the differences and impacts of the node in terms of indicator performance.
[0031] According to the difference distribution interval, analyze the CPU or GPU usage ratio and the total in-out traffic volume of the sorting candidate nodes at the current time tag, unify the parameter dimension processing method, and use the normalization method to eliminate the dimensional differences. Normalize the original data according to the maximum and minimum ranges of the nodes to which they belong. The formula is: ; Let the candidate node number be , its CPU or GPU usage ratio be , the minimum CPU or GPU usage ratio in the corresponding group be , the maximum be , after normalization: ; Similarly, its in-out traffic is , the minimum and maximum in the group are and respectively, then the normalized in-out traffic is: ; The balance benchmark of the group where the node is located is: ; The numerator part is calculated as: ; The performance change of the node under three consecutive time tags is: ; The corresponding traffic linkage strength is: ; The product of the three items is: ; The sum result is: ; The normalized discrete level of the grouped performance variation is: ; The denominator part is: ; Substitute into the sorting performance amplitude formula: ; This result indicates that node currently has a significant deviation from the nodes in the same group in terms of the two key indicators of CPU or GPU usage ratio and the total in-out traffic volume. At the same time, it shows strong performance fluctuations and traffic linkage characteristics within a continuous time period. The numerical result is significantly higher than the performance amplitude of most nodes in the same group, indicating that it is an object with unstable status, high resource consumption, and great potential for abnormal linkage in the current operation and maintenance cycle. This numerical result is the direct judgment basis used in the sorting extraction sub-module to identify outstanding sorting performance. By arranging all the calculated for each node in descending order and selecting the node set at the top of the sorting, the operation and maintenance sorting node group can be derived.
[0032] Please refer to Figure 4 , the multi-parameter linkage recognition module includes: The timing parameter extraction sub-module analyzes the nodes at the top of the sorting based on the operation and maintenance sorting node group, makes multi-time-point comparisons of the CPU and GPU usage ratios, memory occupancy ratio, and business response duration in chronological order, and integrates the timing information of each node's indicators to obtain the indicator timing curve data; First, read the value records of the CPU or GPU usage ratio, memory occupancy ratio and business response time corresponding to the node under multiple consecutive time labels, arrange the data in chronological order and establish a timeline structure. For example, the records of node A from T1 to T5 are CPU or GPU usage ratio: 72%, 76%, 80%, 75%, 78%, memory occupancy ratio: 68%, 72%, 75%, 73%, 74%, business response time: 180ms, 200ms, 230ms, 215ms, 225ms. The above three data items form a group of indicator snapshots under each time label, and then compare the indicator snapshots of the same node under multiple time labels in turn, and record the direction of change of each indicator. For example, the CPU or GPU from T1 to T2 increases from 7 If the change direction of the three indicators is consistent in each time period, that is, in each continuous label interval, the three indicators are compared to see whether the change direction is consistent. If they are consistent, the current time period is marked as a trend-consistent segment. If they are inconsistent, the identification is interrupted. All trend-consistent segments are summarized, and the start and end time labels of each segment, the corresponding node number, the change direction of the three indicators, and the magnitude of the numerical change are recorded. The above information is aggregated into indicator time series curve data, where each curve structure contains elements such as node ID, start and end time period, indicator item, trend direction, value list, and trend duration period.
[0033] The synchronization trend identification submodule calls the indicator time series curve data, compares the change direction of each indicator of each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator in the synchronization interval to obtain the synchronization trend interval table; Traverse the curve content node by node, perform the change direction comparison operation from front to back in the order of time periods, extract the trend markers of CPU or GPU usage ratio, memory occupancy ratio, and business response duration in each time label interval. Suppose node B is in the interval from T6 to T9, and the three indicators show an upward, upward, and upward trend respectively, then it is determined that the interval has a consistent trend direction. If the CPU or GPU decreases, the memory decreases, and the response time increases during the period from T9 to T10, then it is determined to be inconsistent. Screen and mark all the sections with consistent directions, and form the consecutive paragraphs with consistent directions into synchronous trend candidate intervals. Calculate the numerical differences of the changes in the three indicators within each candidate interval respectively, and calculate whether the change amplitudes between the indicators are within the set range. If the change amplitudes of the three indicators in this section are +8%, +6%, and +50ms respectively, and the synchronous determination threshold is set as: the change amplitude of any indicator shall not be less than 50% of the average total trend. For example, if the average trend is +6%, then each indicator needs to exceed +3% or the equivalent change amount. If it is satisfied, this interval is confirmed and registered as a synchronous trend interval. In the synchronous trend interval, further record the indicator change patterns during this time period, such as whether it is continuously rising, continuously falling, or cyclically fluctuating, and organize the change situations of each item in each synchronous section into a record according to the four items of node, time period, change value, and trend direction, and construct a synchronous trend interval table.
[0034] The linked link construction sub-module determines the trend fluctuations of the nodes on the CPU and GPU load panels and the response delay layer according to the synchronous trend interval table, calculates the correlation degree between each group of trends, and uses the formula: ; Obtain the total amount of linked trend intensity, determine the timing curves of each parameter, and obtain the linked abnormal trend link. Among them, represents the total amount of linked trend intensity formed by the node group under consecutive time labels, represents the total number of sorted nodes, that is, the number of nodes to be analyzed, represents the th node's change number of CPU or GPU usage ratio under different time labels, which is the difference between the CPU or GPU usage ratios of adjacent data points in the time series, represents the th node's change number of memory occupancy ratio under different time labels, which is the difference between the memory occupancy ratios of adjacent data points in the time series, represents the th node's change number of business response duration under different time labels, which is the difference between the business response durations of adjacent data points in the time series, represents the th node's trend item in the storage array response delay layer, reflecting the response changes related to the storage device, Indicates the trend item of the -th node in the CPU or GPU load panel of the server group, reflecting the trend of CPU or GPU load change of the node.
[0035] The total amount of linked trend intensity refers to a comprehensive trend intensity quantification result obtained by performing joint trend calculation on the three core operation and maintenance indicators (CPU or GPU usage ratio, memory occupancy ratio, business response time) of the nodes with higher rankings under multiple time tags during the analysis process. This result is used to measure whether there is a strong consistency in trend fluctuations and obvious linkage characteristics among multiple nodes within a certain time range.
[0036] According to the synchronous trend interval table, judge the trend fluctuation data of each node in the CPU or GPU load panel of the server group and the storage array response delay layer. For each node marked with the same direction, collect the change amplitudes of the three types of indicators within its corresponding time period, namely the change in CPU or GPU usage ratio , the change in memory occupancy ratio and the change in business response time . Since the units of the above three types of indicators are different, among them and are in percentage, while is in milliseconds. Therefore, before the combined calculation, all participating items need to be normalized. The maximum normalization method is adopted, and the change of each indicator is divided by its maximum value in the sample nodes. If the three largest changes in the sample nodes are in turn according to the synchronous trend interval table, judge the trend fluctuation data of each node in the CPU or GPU load panel of the server group and the storage array response delay layer. For each node marked with the same direction, collect the change amplitudes of the three types of indicators within its corresponding time period, namely the change in CPU or GPU usage ratio , the change in memory occupancy ratio and the change in business response time . Since the units of the above three types of indicators are different, among them and are in percentage, while is in milliseconds. Therefore, before the combined calculation, all participating items need to be normalized. The maximum normalization method is adopted, and the change of each indicator is divided by its maximum value in the sample nodes. If the three largest changes in the sample nodes are in turn , , , if the monitoring values of a certain node (such as node 01) are , , , then the normalization result is: , , ; Subsequently, read the trend data item of this node in the response delay layer and the trend item of the CPU or GPU load panel , assuming , , calculate the difference item as , substitute the above data into the formula, and substitute the data for node 01 for calculation as follows: The combined square term is: ; The response adjustment term is: ; The trend difference term is: ; The final sub-item is: ; If node 02 and node 03 are calculated respectively to obtain as 2.415 and 2.536, then the total trend intensity of the node group is: ; This value represents the overall level of the linkage intensity formed by the node group within the current time tag segment, and can be used as an important basis for subsequent drawing of the node time series curve and extraction of the linkage link section, and then obtain the linkage abnormal trend link. The formula strengthens the feedback of calculating the combined trend of node load changes through the term, reasonably amplifies the impact of drastic changes in response duration through the term, and further combines to quantify the coupling difference between the fluctuations of storage and computing nodes, and comprehensively generate an index for assisting in identifying the aggregation area of node collaborative fluctuations.
[0037] Please refer to Figure 5 , the resource node screening module includes: The section unlabeled node sub-module, based on the linkage abnormal trend link, calls the nodes it covers, judges the resource sections associated with the nodes, screens the nodes that are not marked as abnormal in the same section, optimizes the node status screening process, and obtains the resource to-be-tested number set; Extract the node numbers involved in the abnormal linkage from the link record, confirm that each node is currently in the abnormal linkage sequence and obtain its resource section information in the node archive table. Suppose nodes A, B, and C are identified as abnormal nodes, and their resource sections are section 1, section 2, and section 1 respectively. Then identify section 1 and section 2 as the currently involved resource sections. Then read the complete resource topology structure diagram and retrieve all the node sets included under section 1 and section 2. Suppose section 1 contains nodes A, C, D, and E, and section 2 contains nodes B, F, G, and H, where nodes D, E, F, G, and H are currently not participating in the linkage trend. Determine that they are not marked as abnormal through the node marking status field, that is, they do not appear in the linkage trend link record. Subsequently, construct a screening condition: nodes that are not marked as abnormal in the current resource section. The node numbers that meet the conditions are added to the list to be screened. Then, conduct a preliminary status determination on the unmarked nodes, and read the CPU or GPU usage rate, memory occupancy rate, and the change range of the total in-out traffic volume under their last 3 time tags. If the CPU or GPU of node E changes from 62% → 70% → 77%, the memory changes from 60% → 66% → 72%, and the in-out traffic changes from 2.1GB → 2.4GB → 2.8GB during the time period from T11 to T13, it is determined that it has a growth pattern similar to the abnormal trend, and this node is marked as a potential object of concern. If the change range of node F is lower than the set minimum fluctuation threshold (such as the CPU or GPU difference is less than 10%, and the traffic change is less than 0.5GB), then it is determined that this node does not have abnormal signs and is not included in the scope of concern. Finally, integrate the node numbers that are not marked but have obvious fluctuations in all the current resource sections to form a set of resource numbers to be tested.
[0038] Based on the set of resource numbers to be tested, the memory flow linkage sub-module compares the memory occupancy ratio and the change in the total in-out traffic volume of each node at consecutive time points, using the formula: ; [[ID=�]] Obtain the combined trend difference amount of each node, and then perform number mapping according to the resource section attribution to obtain the section trend difference sequence, where represents the combined trend difference amount of the th node, represents the change range of the memory occupancy ratio of node [[ID=I5]] under adjacent time tags, which is the change trend of the node's memory load, represents the change range of the total in-out traffic volume of node which is the change trend of the node's network traffic, is the number of nodes participating in this round of analysis; ` Compare the memory occupancy ratio and the total in-out traffic volume of each node at two consecutive time tags, and calculate the memory change range of each node and the change range , respectively normalize the two types of participation items, substitute them into the formula. If there are five node numbers respectively to ; The change range of the memory occupancy ratio of the node between two time tags is ; The corresponding change range of the total in-out flow volume is ; After normalization, are respectively , are respectively ; Substitute them into the denominator part for calculation: ; Substitute into the overall denominator term to get: ; Calculate the value of each node in turn: ; ; ; ; ; This result shows that the joint trend difference amount The larger the numerical value, the higher the degree of synchronous volatility between the change of the memory occupancy ratio of the node and the change of the total in-out flow volume . That is, the two indicators show the linkage characteristics of the same direction and the same amplitude under continuous time tags. This indicator reflects the coupling behavior intensity of the node in multiple operation and maintenance data.
[0039] The number intersection extraction sub-module identifies the node numbers with the top rankings according to the section trend difference sequence, analyzes the intersection of the nodes and the sorted nodes, judges the distribution of the intersection numbers in the node set, and obtains the resource regulation number group.
[0040] The joint trend difference amount refers to the intensity of the linkage trend performance between the change range of the memory occupancy ratio and the change range of the total in-out flow volume of the same node under continuous time tags. This quantity is used to measure the synchronization and coupling degree of the node during the change process of resource pressure (memory) and network load (flow).
[0041] Extract the node numbers with the top fluctuation amplitudes in each resource section. Suppose that in section 1, nodes D and E rank among the top two in the comprehensive evaluation of memory occupancy rate and traffic change amplitude, and their difference values are 16%, 1.3GB and 14%, 1.2GB respectively. According to the difference sorting rule, add the numbers of D and E to the list of nodes with top rankings. At the same time, read the set of node numbers that were previously marked as linkage anomalies in the sorted node list, and perform a number comparison operation on the set of node numbers with top rankings and the set of linkage anomaly nodes to identify the numbers that appear repeatedly. For example, if node E appears in both sets, it is identified as an intersection node. Further judge the distribution of all intersection numbers in the resource sections to which the nodes belong. The specific analysis method is as follows: count the proportion of the number of intersection numbers in each section and the proportion of this number in the total number of original nodes. Suppose there are 6 nodes in section 1, and the intersection numbers are node E and A, with a proportion of 2 / 6, that is, 33.3%. Then record that there is a significant number intersection phenomenon in section 1, and sort out the node intersection structure list according to this result. Subsequently, perform a comprehensive judgment based on the resource location, index trend value, and sorting order of each intersection node to identify the most representative intersection number. It is set to preferentially screen the nodes with top rankings in the intersection. If the index trend value of node E in the intersection nodes is higher than that of A, then E is added to the resource regulation number group as the final retained node, and at the same time, information such as its sorting value, resource attribution, intersection times, and trend direction is marked, and output this resource regulation number group.
[0042] Please refer to Figure 6 , the event traceback module includes: Based on the resource regulation number group, the trace data extraction sub-module analyzes the time series records of the recent CPU and GPU usage ratios, disk read and write conditions, and service response conditions, sorts them into the continuous running performance of the nodes, and obtains the node running sequence; Read the operation data records of the included nodes one by one in the past three days. Specifically, obtain the CPU or GPU usage ratio, disk read and write time, and business response duration data of each node corresponding to the time tags T1 to T24 (with a sampling interval of every 3 hours) from the storage. Establish a time series table structure for each data item, group by node number, arrange the sampling data in chronological order within each group and fill it into the index matrix to form a continuous operation performance curve of the node. For example, the CPU or GPU usage rates of node A at time points T1 to T5 are 58%, 63%, 68%, 64%, and 70%, the disk read and write times are 290ms, 310ms, 345ms, 330ms, and 375ms respectively, and the business response times are 180ms, 195ms, 210ms, 200ms, and 215ms. The continuous time series data record of this node is complete and consists of the hourly values of the three indicators. Identify the trend direction of each indicator and mark the change slope. For example, it is identified that the CPU or GPU generally rises from T1 to T5, the disk read and write time shows an increase, and the response duration also rises in each cycle. This group of data is regarded as a node with a continuous positive growth trend. If the CPU or GPU sequence of another node B fluctuates as 65%, 60%, 67%, 62%, 68% and the change direction is unstable, it is marked as an unstable node. Write the time series of the three indicators of all nodes into their node operation sequences respectively, and store them in the index trajectory database after uniformly numbering based on the time tags.
[0043] The fluctuation curve comparison sub-module compares the change trends of each indicator of each node in the node operation sequence with the curve of the linked fluctuation node, judges the synchronous change situation of each node indicator curve at the same time series, identifies the node performance with highly consistent trends, and obtains the synchronous feature node group; Construct single - metric curves for the CPU or GPU, disk read - write time, and response duration of each node respectively. Calculate the change - direction identification sequence by taking adjacent two time tags and generate a direction - vector sequence. For example, the CPU or GPU curve direction vector of node A from T1 to T5 is [+1, +1, -1, +1], and the corresponding direction vector of the coupled node L is [+1, +1, -1, +1]. Then node A is completely consistent with node L in this metric. If the direction vectors of its disk and response metrics are also completely consistent with those of node L or deviate only within one time point, preliminarily judge this node as a trend - synchronous node. Further, normalize the change slopes of each node's metrics, and calculate the similarity scores within the change - rate intervals respectively. If the slope errors of node A and node L are less than 5% in at least 4 out of 5 consecutive segments for the three metrics, it is judged that this node has a high consistency with the coupled curve. Incorporate the node number into the candidate set of synchronous - feature nodes. Subsequently, count the objects among all nodes that are trend - synchronous with the coupled curve in all three metrics and whose errors are within the acceptable threshold range. For example, nodes A, C, and E meet this condition while B and D are only partially consistent. Then A, C, and E are marked as a synchronous - feature node group. Each node in this group records the synchronous - segment length, the proportion of consistent change directions, and the average trend - deviation value of the metric curves for the next - stage screening.
[0044] The trajectory - information generation sub - module screens the nodes that are crucial for synchronous changes according to the synchronous - feature node group, optimizes the abnormal sections and change characteristics of each node in the time - series curve, and obtains the associated - node trajectory information. First, extract the synchronous - change sections from the CPU or GPU, disk I / O, and response - time series of each node, that is, the change sections where the three metrics show consistent directions within a certain continuous time range. For example, for node C from T6 to T10, the CPU or GPU, disk, and response values all show continuous growth, and the direction vector is [+1, +1, +1, +1]. Then mark this section as a candidate abnormal section. Subsequently, evaluate the change amplitude within this abnormal section. If the growth amplitudes of any two of the three metrics exceed the preset change threshold, such as the CPU or GPU growth amplitude exceeds 12% and the response - time increase exceeds 40 ms, then confirm this section as an abnormal section. Further, mark this time range as a high - volatility abnormal area in the curve of node C, and include all the metric values, time tags, direction vectors, and amplitude data of this node within this section into the trajectory - information item. Repeat the above process to traverse all synchronous - feature nodes to form a complete list of trajectory - abnormal - section records. Finally, summarize the metric - change characteristics, maximum and minimum value points, volatility, and comparison with adjacent nodes shown by each node in its respective abnormal section to complete the generation of associated - node trajectory information.
[0045] The above are only the preferred embodiments of the present invention, and do not limit the present invention in other forms. Any person skilled in the relevant art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A visualization operation and maintenance data analysis system, characterized in that The system comprises: The node data collection module, based on the backbone network of the intelligent computing center, analyzes the operating data of each node, groups node types by time tags, arranges the changes in the CPU and GPU usage ratios and memory usage ratios, and archives disk read and write time, total inbound and outbound traffic, and service response time to obtain a node indicator sequence set. The indicator sorting and screening module compares the CPU and GPU usage ratios and the total inbound and outbound traffic of each node under the same time tag based on the node indicator sequence set, summarizes the differences, groups and sorts them by type, identifies the top-ranked nodes, and obtains the operation and maintenance ranked node group; The multi-parameter linkage identification module analyzes the synchronous changes in the CPU and GPU usage ratios, memory usage ratios, and service response times of the top nodes in the current and adjacent time tags based on the operation and maintenance sorting node group, determines the timing curves of each parameter, and obtains the linkage abnormality trend link; The resource node screening module screens unmarked nodes in the associated resource segment based on the linkage abnormal trend link, compares the memory usage ratio trend and the inflow and outflow fluctuation amplitude, determines the node number based on the sorted intersection, and obtains the resource control number group.
2. The visualization operation and maintenance data analysis system according to claim 1, wherein The node indicator sequence set includes node performance data, archive time series, and group classification information; the operation and maintenance sorting node group includes priority identification, node identification code, and sorting index number; the linkage abnormal trend link includes linkage feature sequence, abnormal fluctuation segment, and trend matching identification; the resource control number group includes control node list, allocation node code, and redundant resource identification.
3. The visualization operation and maintenance data analysis system according to claim 1, wherein The node data collection module includes: The data collection submodule, based on the backbone network of the intelligent computing center, analyzes the real-time collection of CPU and GPU usage ratios, memory usage ratios, disk read and write times, total inbound and outbound traffic, and service response times of various nodes. It organizes the data based on time tags and generates a set of operating indicators. The joint arrangement submodule compares the correlation between the CPU and GPU usage ratios and the memory usage ratios of each node under the same time label based on the set of operating indicators, and arranges them in sequence based on the change trend of each indicator to obtain a joint change trend group; The indicator archiving submodule filters the disk read and write time, total inflow and outflow volume and service response time of each type of node under continuous time labels according to the joint change trend group, archives parameter trends, and obtains a node indicator sequence set.
4. The visualization operation and maintenance data analysis system according to claim 1, wherein The indicator sorting and screening module includes: The difference summarization submodule analyzes the CPU and GPU usage ratios and the total inbound and outbound traffic of the same type of nodes under each time tag based on the node indicator sequence set, compares the fluctuation range of the corresponding parameters of each node in the group, calculates the difference amplitude of the parameters between the difference nodes, and obtains the performance variation amplitude; The group comparison submodule calls the performance variation range, groups nodes according to node type, compares the difference distribution of CPU and GPU usage ratios and total inbound and outbound traffic of each group of nodes under the same time label, determines the distribution breadth of parameters between nodes in the same group, selects nodes with key performance differences, and obtains the difference distribution range; The sorting and extraction submodule analyzes the CPU and GPU usage ratios, total inbound and outbound traffic, and grouping parameter benchmarks of the nodes under the current time tag based on the difference distribution interval, adjusts the historical performance changes and linkage fluctuations within the group, identifies nodes with outstanding sorting performance as key operation and maintenance analysis objects, and obtains the operation and maintenance sorting node group.
5. The visualization operation and maintenance data analysis system according to claim 1, wherein The multi-parameter linkage identification module includes: The timing parameter extraction submodule analyzes the top-ranked nodes based on the operation and maintenance ranking node group, compares the CPU and GPU usage ratios, memory usage ratios, and service response times at multiple time points in chronological order, and integrates the timing information of each node indicator to obtain indicator timing curve data; The synchronization trend identification submodule calls the indicator time series curve data, compares the change direction of each indicator of each node, identifies the interval with the same direction, and calculates the joint change trend of each indicator in the synchronization interval to obtain a synchronization trend interval table; The linkage link construction submodule determines the trend fluctuation of the node on the CPU and GPU load panels and the response delay layer according to the synchronization trend interval table, calculates the correlation degree between each group of trends, obtains the total linkage trend strength, determines the timing curve of each parameter, and obtains the linkage abnormal trend link.
6. The visualization operation and maintenance data analysis system according to claim 1, wherein The resource node screening module includes: The unmarked node submodule of the segment is based on the linkage abnormal trend link, calls the nodes covered by it, determines the resource segment associated with the node, filters the nodes in the same segment that are not marked as abnormal, optimizes the node status screening process, and obtains the resource number set to be tested; The memory-flow linkage submodule compares the memory usage ratio and the total inflow and outflow changes of each node at consecutive time points based on the resource number set to be tested, obtains the joint trend difference of each node, and then performs number mapping based on the resource segment affiliation to obtain the segment trend difference sequence; The number intersection extraction submodule identifies the top-ranked node numbers according to the segment trend difference sequence, analyzes the intersection of the nodes and the sorted nodes, determines the distribution of the intersection numbers in the node set, and obtains the resource control number group.
7. The visualization operation and maintenance data analysis system according to claim 1, wherein The system further comprises: The event trajectory tracing module retrieves the time series data of the node CPU and GPU usage ratio, disk read and write time, and business response time based on the resource control number group, and compares it with the linkage type fluctuation node curve to obtain the associated node trajectory information; The associated node trajectory information includes historical trajectory sequence, node association characteristics, and event archive identification.
8. The visualization operation and maintenance data analysis system according to claim 7, wherein The event trace tracing module includes: The trajectory data extraction submodule analyzes the recent CPU and GPU usage ratios, disk read and write conditions, and business response time series records based on the resource control number group, organizes them into node continuous operation performance, and obtains the node operation sequence; The fluctuation curve comparison submodule compares the change trend of each indicator of each node in the node operation sequence with the linkage fluctuation node curve, determines the synchronous change of the indicator curves of each node in the same time sequence, identifies the node performance with highly consistent trends, and obtains the synchronous feature node group; The trajectory information generation sub-module screens the nodes that are critical for synchronization changes based on the synchronization feature node group, optimizes the abnormal sections and change characteristics of each node in the timing curve, and obtains the associated node trajectory information.
Citation Information
Patent Citations
Modeling method and device for operation and maintenance system
CN113268891A
Data center intelligent operation and maintenance method and device based on digital twinning
CN116595756A
Data analysis method and device, electronic equipment and storage medium
CN118394603A
Management and control method, device and equipment of intelligent operation and maintenance terminal and storage medium
CN119065946A
Management method and device for stable operation of system and server
CN119201630A
Cited By
Server load balancing method and system based on edge computing
CN121301008A
Server load balancing method and system based on edge computing
CN121301008B
Real-time monitoring system and method based on server state
CN121301130A
Remote monitoring method and system for state of pressure-bearing equipment
CN121558117A
User social interaction data management system and method based on cloud computing
CN122173507A