Intelligent operation and maintenance monitoring system and method for ARM computing power node
Through node time synchronization, adaptive alarm rules and communication status detection, the monitoring method of ARM computing power nodes is optimized, solving the problems of sampling timestamp misalignment, insufficient alarm mechanism and misjudgment of communication status, and achieving efficient operation and maintenance monitoring and stability assurance.
Patent Information
- Application Number
- CN202511188457.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-25
AI Technical Summary
The existing ARM computing power node monitoring technology has problems such as sampling timestamp misalignment, lack of dynamic perception of the alarm mechanism, misjudgment of communication status and excessive system resource overhead, which threatens service stability and security.
Through node time synchronization processing, adaptive alarm rule optimization, abnormal response identification and communication status detection, an adaptive node operation response strategy is generated, the alarm mechanism and communication detection are optimized, and the response accuracy and system stability are improved.
It achieves a response capability within seconds, reduces false alarms and missed alarms, improves the accuracy of locating communication anomalies, avoids unnecessary overhead of system resources, and improves operation and maintenance adaptability and service continuity.
Smart Images

Figure CN120729702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data monitoring technology, and more specifically, to an intelligent operation and maintenance monitoring system and method for ARM computing power nodes. Background Art
[0002] With the development of new-generation information technologies such as cloud computing and computing networks, the operating architecture composed of ARM computing nodes has been widely deployed in various business scenarios such as cloud gaming, cloud phones, and mobile offices due to its high energy efficiency and strong scalability. To ensure service continuity and stability, an efficient and intelligent operation and maintenance and monitoring system is becoming an increasingly necessary component, so that the operating architecture composed of ARM computing nodes has the ability to perceive and respond to various abnormal situations of computing nodes. However, there are still some problems with existing monitoring technologies.
[0003] Since ARM computing nodes use computing boards as physical carriers to run instances, there are large differences in hardware structures. The time bases of different boards mostly rely on their own local crystal oscillators or clock modules. Even if a unified sampling trigger mechanism is used when collecting node operation data, it is difficult to eliminate the misalignment of sampling timestamps between multiple nodes, which is not suitable for application to cloud platforms, which require operating systems that respond in seconds. In terms of alarm mechanisms, existing technologies often lack dynamic perception of the current running task type, resulting in misjudgment or missed judgments. For example, high CPU or GPU utilization may be normal in some task types, but when running another task type, it may indicate a computing bottleneck or delay risk. In addition, if a node reports a large number of status indicators at the same time during operation, it is difficult for the receiving terminal to identify the core abnormal indicators that really need to be processed quickly. This causes key signals to be drowned out or even fail, leaving potential security risks. At the same time, in terms of communication detection, traditional technical solutions often use whether the ping is successful or the port response occurs to determine whether the communication status is normal. However, for example, the failure of the semantic content in the response packet or the resource writing defect can easily lead to pseudo communication failure problems. At this time, the content is abnormal, but the communication link structure may not be abnormal. The system directly determines this situation as a communication failure due to recognition timeout or abnormal response format, resulting in a misjudgment of the communication status. When the system determines that an abnormality has occurred, it will execute the corresponding exception handling strategy or restart the system to correct the abnormality. However, if the status indicator repeatedly jitters around the system's default abnormality threshold, the system may be repeatedly restarted or the exception handling strategy may be executed, resulting in a substantial increase in system resource overhead, which seriously threatens the stability of the service.
[0004] In view of this, the present invention proposes an intelligent operation and maintenance monitoring system and method for ARM computing power nodes to solve the above problems. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art and achieve the above-mentioned objectives, the present invention provides the following technical solution: an intelligent operation and maintenance monitoring method for ARM computing power nodes, comprising: S1. Collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; S2. Perform node time synchronization on the node operation status data to generate node synchronization operation data; S3. Optimize the alarm mechanism based on node synchronization data and obtain adaptive alarm rules; S4. Combined with adaptive alarm rules, perform abnormal response identification on the node synchronization data and output abnormal response data queue; S5 abnormal response data queue communication status detection, get the communication status mark data; S6. Generate a node operation response strategy based on the communication status mark data, and send the node operation response strategy to the preset ARM computing power node control terminal.
[0006] Furthermore, the node time synchronization processing method includes: Obtain the board number of the ARM computing power node and group the node operation status data, and output the grouped node data; set the global sampling time standard of the node unit; calculate the difference between the timestamp of the grouped node data and the global sampling time standard of the node unit to obtain the time offset of the corresponding board node unit; compare the time offset with the preset time error tolerance threshold, and divide the node operation status data into synchronous node data subsets and asynchronous node data subsets based on the comparison result; Feature extraction is performed on the asynchronous node data subset to obtain the node time feature; time series clustering is performed on the asynchronous node data subset based on the node time feature to generate a hysteresis node cluster set and a drift trend node cluster set; historical sampling response data of different boards are obtained, and time compensation mapping is performed on the hysteresis node cluster set based on the historical sampling response data to output a time correction node cluster set; the drift period of the drift trend node cluster set is extracted and the drift rate is calculated, and the drift trend node cluster set is time-axis re-parameterized based on the drift period and drift rate to obtain a time reconstruction node cluster set; the synchronous node data subset, the time correction node cluster set and the time reconstruction node cluster set are integrated to obtain the node synchronous operation data.
[0007] Furthermore, the method of optimizing the alarm mechanism includes: Obtain the real-time running task identifier corresponding to the node synchronization running data, and match the running task identifier with the node number to obtain the task node index set; determine the task type based on the task node index set, and construct the resource allocation index for each task type; extract the historical node data of the corresponding task type, calculate the multi-dimensional parameter difference between the resource allocation index of the historical node data and the resource allocation index of the current node, assign sensitivity weights to the resource allocation index based on the multi-dimensional parameter difference, and output the weighted resource allocation index; Perform interactive correlation analysis on weighted resource allocation indicators and output indicator fluctuation correlation characteristics; construct task type anomaly judgment conditions by combining indicator fluctuation correlation characteristics and task types; establish a mapping relationship between task type anomaly judgment conditions and weighted resource allocation indicators, and construct anomaly correlation data sets based on the mapping relationship; identify anomaly subsets with overlapping tasks in the anomaly correlation data sets, and generate alarm intensity sequences based on the trigger response sequence of relevant indicators in the anomaly subsets; and construct adaptive alarm rules by combining relevant indicators, task type anomaly judgment conditions, and alarm intensity sequences.
[0008] Furthermore, the method of constructing the task type abnormality determination condition includes: Task types include cloud gaming, mobile office, and cloud mobile phone; When the task type is cloud gaming, the abnormality criteria include the short-term GPU utilization change, the number of consecutive frame rate jitters, and the image rendering delay. When the task type is mobile office, the task type abnormality determination criteria include CPU mutation change rate, memory allocation, and work task switching frequency; When the task type is cloud phone, the task type abnormality judgment conditions include system access delay, network delay jump frequency and data interruption number.
[0009] Furthermore, the method of performing abnormal response identification includes: Using adaptive alarm rules as a basis for judgment, identify abnormal responses in node synchronization operation data and output preliminary screening of abnormal data; identify the operation behavior corresponding to the preliminary screening abnormal data, and filter the abnormal processing subset based on the operation behavior; set an abnormal risk level for each abnormal processing subset, and determine the relevant abnormal indicators with abnormal risk levels higher than the preset risk level threshold as the preferred abnormal indicators; Construct a dynamic sliding window to traverse all exception processing subsets, obtain all preferred exception indicators within the dynamic sliding window at the same time, and determine the preferred exception indicator as the highest processing priority; For non-preferred abnormal indicators, the distribution density and duration of non-preferred abnormal indicators are obtained using a dynamic sliding window at the same time; the load status characteristics of the corresponding ARM computing power node are collected in real time, and the suboptimal abnormality score is calculated based on the distribution density, duration and load status characteristics. The non-preferred abnormal indicators are sorted based on the suboptimal abnormality score to obtain the suboptimal abnormality queue; the preferred abnormal indicators and the suboptimal abnormality queue are integrated to obtain the abnormal response data queue.
[0010] Furthermore, the communication status detection method includes: Obtain the communication path set of the ARM computing nodes corresponding to the abnormal response data queue; extract the node response data of each communication path in the communication path set, identify the semantic content of the node response data, and construct a node semantic vector matrix based on the semantic content; perform structural integrity analysis on the node semantic vector matrix, and filter the initial abnormal response data packet based on the structural integrity analysis results; Identify the request instructions corresponding to the response data of the remaining nodes, and query the original included fields based on the request instructions; calculate the mismatch rate by comparing the number of existing fields with the number of originally included fields, and perform a second-class screening of the response data of the remaining nodes based on the mismatch rate and a preset threshold to obtain the second-class abnormal response data packet; The resource occupancy rate of the response data of the remaining nodes is extracted to construct an abnormal field occupancy vector; the standard deviation of all elements in the abnormal field occupancy vector is calculated, and the skewness analysis of the abnormal field occupancy vector is performed based on the standard deviation to output the skewness coefficient; if the skewness coefficient does not meet the preset standard coefficient interval, the response data of the remaining nodes are screened into three categories to obtain three categories of abnormal response data packets; the abnormal response data packets screened each time are marked with a communication abnormality status to obtain communication status mark data.
[0011] Furthermore, the method of marking the abnormal communication state includes: If the structural integrity analysis result shows that the structure is incomplete, the initial abnormal response data packet at this time will be judged as a structural destruction abnormal state; if the mismatch rate is higher than the preset threshold, the second-class abnormal response data packet at this time will be judged as a semantic content abnormal state; if the skewness coefficient does not meet the preset standard coefficient range, the third-class abnormal response data packet at this time will be judged as a resource write abnormal state.
[0012] Furthermore, the generation node executes the response strategy in the following manner: Extract the communication status tag of the communication status mark data, match the communication status tag and the relevant abnormal indicators of the corresponding data with the preset abnormality resolution strategy database, and generate appropriate strategy instructions; Extract the relevant abnormal indicator fluctuation value of any node within the continuous preset monitoring period, and construct a state fluctuation matrix based on this value; set a reasonable fluctuation range, and count the number of times all the relevant abnormal indicator fluctuation values of each node in the state fluctuation matrix exceed the reasonable fluctuation range; if the number is greater than the preset frequency threshold, directly execute the abnormal handling action; otherwise, temporarily mark the corresponding node, and record the trend change of the relevant abnormal indicators of the temporarily marked node in the subsequent continuous period to construct a trend offset function; if the accumulation of the trend offset function value in the continuous period does not exceed the preset accumulation threshold and the change direction is the correct direction, then remove the temporary mark and abandon the execution of the abnormal handling action; if the trend offset function value is continuously greater than the preset function threshold in the continuous period and the change direction is an upward trend, then directly execute the abnormal handling action; integrate appropriate strategy instructions and the encoded abnormal handling action execution process to construct a node operation response strategy.
[0013] Furthermore, the method of constructing the state fluctuation matrix includes: All relevant abnormal indicators belonging to the same node are normalized to obtain normalized abnormal indicators; the normalized abnormal indicators are spliced into row vectors, and the normalized abnormal indicator row vectors corresponding to all nodes are superimposed to obtain the state fluctuation matrix.
[0014] An intelligent operation and maintenance monitoring system for ARM computing power nodes, which is used to implement an intelligent operation and maintenance monitoring method for ARM computing power nodes, is characterized by comprising: The data acquisition module is used to collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; The node synchronization module is used to perform node time synchronization processing on the node operation status data and generate node synchronization operation data; The alarm optimization module is used to optimize the alarm mechanism based on the node synchronization operation data and obtain adaptive alarm rules; The anomaly recognition module is used to perform abnormal response recognition on the node synchronization operation data in combination with the adaptive alarm rules and output the abnormal response data queue; The communication detection module is used to detect the communication status of the abnormal response data queue and obtain communication status mark data; The strategy generation module is used to generate a node operation response strategy based on the communication status mark data and send the node operation response strategy to the preset ARM computing power node control terminal; each module is connected by wired and / or wireless means.
[0015] The technical effects and advantages of the intelligent operation and maintenance monitoring system and method of ARM computing power nodes of the present invention are as follows: By collecting the running status data of ARM nodes, we optimize the node synchronization, alarm mechanism, abnormal response, communication detection and policy generation, and realize an intelligent operation and maintenance monitoring method for ARM computing nodes. Compared with the existing experience, by introducing the multi-node time synchronization process, we effectively solve the problem of clock inconsistency caused by different hardware structures of boards, so that the system can adapt to the second-level response in application scenarios. By integrating task type and resource sensitivity, the system can adjust the alarm mechanism threshold according to the actual business impact of the task, avoiding the false alarm or missed alarm situation that is easy to cause by traditional alarm rules. By introducing semantic content detection and skew analysis, the system can effectively solve the problem of clock inconsistency caused by different hardware structures of boards, so that the system can adapt to the second-level response in application scenarios. The analysis breaks through the judgment method of "response means communication success, otherwise communication failure" in traditional communication detection technology, effectively alleviates the situation of abnormal misjudgment, and improves the accuracy of communication anomaly positioning; by detecting the fluctuation trend of node indicators, anti-shake control is achieved during the operation of computing power nodes, effectively avoiding the high-overhead operation of the system executing abnormal response strategies due to slight fluctuations in related node indicators, and improving the overall operation stability; in summary, this intelligent operation and maintenance monitoring method for ARM computing power nodes significantly improves the adaptability and processing accuracy of node abnormal information in the computing power network, and provides stronger operation and maintenance adaptability and service continuity guarantee. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 A schematic diagram of an intelligent operation and maintenance monitoring method for an ARM computing power node of the present invention; Figure 2 This is a schematic diagram of an intelligent operation and maintenance monitoring system for an ARM computing power node of the present invention. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0018] Example 1
[0019] See also Figure 1 As shown, the intelligent operation and maintenance monitoring method of an ARM computing power node described in this embodiment includes: S1. Collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; S2. Perform node time synchronization on the node operation status data to generate node synchronization operation data; S3. Optimize the alarm mechanism based on node synchronization data and obtain adaptive alarm rules; S4. Combined with adaptive alarm rules, perform abnormal response identification on the node synchronization data and output abnormal response data queue; S5 abnormal response data queue communication status detection, get the communication status mark data; S6. Generate a node operation response strategy based on the communication status mark data, and send the node operation response strategy to the preset ARM computing power node control terminal.
[0020] The general application scenario of this embodiment is a cluster of 100 M114 Android cloud servers, and each M114 contains 84 ARM computing boards. Each board corresponds to an independent virtualized Android instance, which is used to provide ARM computing power services. The ARM computing board is the physical carrier output by the ARM computing power node; the operating status data includes physical parameters such as CPU usage, memory occupancy, process survival status and GPU rendering queue length that reflect the node operation status; at the same time, this embodiment fills missing values and compresses redundant data on the operating status data in turn, realizes data cleaning, and obtains higher-quality node operating status data.
[0021] Since each server cluster has 84 ARM computing boards, a high-density heterogeneous physical structure is formed. Each board is essentially a small computing subsystem, and its time base relies on its own local crystal oscillator or BMC clock module. Moreover, in different specific application scenarios such as cloud gaming or mobile office, different nodes will be subject to corresponding business behavior restrictions. Therefore, even if a unified sampling time is used during collection, slight time deviations cannot be avoided, which has a significant impact on businesses that require second-level response.
[0022] Methods for node time synchronization include: Obtain the board number of the ARM computing power node and group the node operation status data, and output the grouped node data. The running instance in the ARM computing power node is extracted and mapped to the ARM board number on which it is deployed, and grouped based on the number. Each group of grouped node data belongs to the same ARM computing board, ensuring clear division of data.
[0023] Set the global sampling time standard of the node unit, where the global sampling time standard of the node unit is a unified time reference line preset by the control terminal based on the ARM computing power node in actual application, which is used for subsequent time synchronization processing operations; calculate the difference between the timestamp of the grouped node data and the global sampling time standard of the node unit to obtain the time offset of the corresponding board node unit, where the time offset is used to measure the time difference between the sampling timestamp of the grouped node data corresponding to each board and the set global sampling time standard of the node unit, and use the gap between the sampling timestamps related to different boards and the unified time standard to reflect the degree of time offset.
[0024] The time offset is compared with a preset time error tolerance threshold, and the node operation status data is divided into a synchronous node data subset and an asynchronous node data subset according to the comparison result, wherein the time error tolerance threshold is a threshold set based on the historical experience of technical personnel in this field, and is used to determine whether the time offset meets the threshold requirements and perform node classification; when the time offset exceeds the preset time error tolerance threshold, the corresponding node unit is determined to be an asynchronous node and combined into an asynchronous node data subset; otherwise, the node unit is determined to be a synchronous node and combined into a synchronous node data subset.
[0025] Feature extraction is performed on the asynchronous node data subset to obtain node time features, where the node time features include information such as sampling interval fluctuations, sampling sequence stability, and offset trends of corresponding nodes, which are used to reflect the changing behavior trajectory of the asynchronous node data subset.
[0026] Based on the node time features, the asynchronous node data subset is subjected to time series clustering processing to generate a hysteresis node cluster set and a drift trend node cluster set. In this embodiment, the DTW dynamic time warping algorithm is used to identify the data with similar time lag behavior in the asynchronous node data subset based on the node time features and cluster them into a hysteresis node cluster set. The data with similar continuous offset increase or decrease trends are clustered into a drift trend node cluster set, which is used to reflect the two abnormal sampling modes in the asynchronous node data subset.
[0027] The historical sampling response data of different boards are obtained, and the time compensation mapping of the hysteresis node cluster set is performed based on the historical sampling response data, and the time correction node cluster set is output, wherein the historical sampling response data of a board card in multiple previous sampling cycles is extracted to obtain the historical sampling response data of the board card; the nodes with response lag in the historical sampling response data are identified, and it is determined whether these lag nodes exist in continuous sampling cycles. If so, it means that these nodes are stable lags and enter the time compensation mapping process; the compensation value is obtained by calculating the average lag time of the response lag nodes in the historical sampling response data, and the value is used to perform time compensation mapping on the current lag node corresponding to the board card; if it is not a stable lag, it is determined to be a temporary lag, and the sampling timestamp of the node is aligned with the global sampling time standard of the node unit to achieve compensation; all compensated lag nodes are integrated to obtain the time correction node cluster set.
[0028] The drift period of the drift trend node cluster set is extracted and the drift rate is calculated. The drift trend node cluster set is time-axis reparameterized based on the drift period and the drift rate to obtain a time-reconstructed node cluster set. The drift trend node cluster set is different from the hysteresis node cluster set. The offset of the hysteresis node cluster set is less variable and can be compensated, while the time offset value of the drift trend node cluster set is in a rapid change with a certain trend. Therefore, the time axis reparameterization method is used for processing in this embodiment.
[0029] Time axis reparameterization is a technique widely used in multiple fields for processing time series data. First, the drift period of a drift trend node cluster set is obtained and the drift rate is calculated. The drift period and drift rate are then used as the basic variables of the reparameterization function in the time axis reparameterization method. It should be noted that the drift period is obtained by plotting a waveform diagram of the node cluster's chronological changes based on the node values of each node cluster in the drift trend node cluster set, where the horizontal axis represents time and the vertical axis represents the specific values of the nodes in the node cluster. The time interval between adjacent repeated offset waveforms extracted from the waveform diagram is the drift period. The drift rate refers to the slope of the sampling point in the offset waveform during any drift period. The drift period guides the period fitting and stability setting of the reparameterization function, while the drift rate guides the slope direction of the sampling time adjustment of the reparameterization function. Finally, the reparameterization function is used to calculate the offset that needs to be compensated for the drift trend node cluster to obtain the time reconstruction node cluster set. The node synchronization operation data is obtained by integrating the synchronization node data subset, the time correction node cluster set, and the time reconstruction node cluster set.
[0030] Ways to optimize the alert mechanism include: Obtain the real-time running task identifier corresponding to the node synchronization running data, and match the running task identifier with the node number to obtain a task node index set, where the running task identifier refers to the task that the relevant ARM computing power node of the node synchronization running data is currently performing when the user is using it, such as running games, browsing the web, or establishing communications and other specific running tasks; the ARM computing power nodes carried by each board have numbers, so each specific running task is bound to the corresponding node number to obtain a task node index set.
[0031] The task type is determined based on the task node index set, and resource allocation indicators are constructed for each task type. The general types of tasks are divided into three categories: cloud games, mobile office, and cloud phones based on the specific running tasks in the task node index set. For each task type, the corresponding resource parameters consumed under the historical stable running state are extracted, and the corresponding resource parameters consumed by each task type are used as the resource allocation indicators of the nodes currently belonging to this task type.
[0032] Extract historical node data of corresponding task types, calculate the multi-dimensional parameter difference between the resource allocation index of historical node data and the resource allocation index of current node, assign sensitivity weights to resource allocation index based on multi-dimensional parameter difference, and output weighted resource allocation index, where multi-dimensional parameter difference refers to the difference between the resource allocation index of node data in historical stable state and each dimensional parameter in resource allocation index of current node, which is used to measure the degree of deviation between current running state and historical stable state; first normalize the calculated multi-dimensional parameter difference to eliminate dimension difference and obtain dimensionless eigenvalue of resource allocation index of this type; then set corresponding initial sensitivity weight for each task type, where initial sensitivity weight is set by querying historical records, and then adjust the initial sensitivity weight based on specific task situation; multiply the sensitivity weight by dimensionless eigenvalue to obtain weighted resource allocation index of a certain task type, avoiding false alarms easily caused by fixed thresholds.
[0033] An interactive correlation analysis is performed on the weighted resource allocation indicators to output indicator fluctuation correlation features. The interactive correlation analysis is used to identify the possible coupling behavior relationships between individual indicator parameters in each weighted resource allocation indicator. The indicator fluctuation correlation feature refers to the correlation strength of multiple coupling behaviors, such as the frame rate jitter amplitude associated with increased GPU occupancy, and the change in system transmission efficiency caused by memory criticality. The indicator fluctuation correlation feature and task type are combined to construct task type anomaly judgment conditions. By matching the indicator fluctuation correlation feature with the task type, the anomaly judgment conditions for the corresponding task type are constructed, that is, a certain type of indicator combination can constitute an effective alarm trigger condition under a certain task type.
[0034] A mapping relationship between task type anomaly determination conditions and weighted resource allocation indicators is established, and an anomaly association data set is constructed based on the mapping relationship, wherein the anomaly determination conditions are bound to the weighted resource allocation indicators of the task type corresponding to the conditions to establish a corresponding mapping relationship; combined with the mapping relationship, an anomaly association data set with mutual correlation of multiple anomaly events is constructed.
[0035] Identify abnormal subsets with overlapping tasks in the abnormal association data set, and generate an alarm intensity sequence based on the trigger response order of related indicators in the abnormal subset, wherein several abnormal events with correlation are extracted as abnormal subsets, and the trigger response order refers to which related indicator enters the abnormal judgment state first, and these related indicators are sorted in descending order based on this order; it should be noted that if there are several related indicators entering the abnormal judgment state at the same time, then at this time, they are sorted in descending order based on the sensitivity weights of several related indicators, and finally all the sorted related indicators are integrated to obtain the alarm intensity sequence; combine the related indicators, task type abnormal judgment conditions and alarm intensity sequence to form an adaptive alarm rule, correspond the related indicators to the task type abnormal judgment conditions, and at the same time combine the alarm intensity sequence to form a sequence rule that reflects the related indicators of multiple task types when making abnormal judgments.
[0036] The methods for constructing task type exception judgment conditions include: Task types include cloud gaming, mobile office, and cloud mobile phones. Different types of tasks have significantly different characteristics in their use of system resources during operation, and different task types have different sensitivities to resource changes.
[0037] When the task type is cloud gaming, the conditions for determining task type abnormalities include the change in GPU short-term utilization, the number of continuous frame rate jitters, and image rendering delay. Since the essential logic of cloud gaming is an application that is highly dependent on the real-time nature of graphics rendering and interactive feedback, it is sensitive to changes in GPU-related indicators. Therefore, the change in GPU short-term utilization, the number of continuous frame rate jitters, and image rendering delay are used as abnormality determination conditions and corresponding thresholds are set; the change in GPU short-term utilization can reflect instantaneous rendering congestion caused by a surge in rendering load, and the number of continuous frame rate jitters can reflect whether the frame rate is stable over a continuous period of time; the image rendering delay can reflect whether the output display status has a delay that affects the gaming experience.
[0038] When the task type is mobile office, the task type abnormality judgment conditions include the CPU mutation change rate, memory allocation and work task switching frequency. When in the mobile office task type, the system mainly processes documents and light calculations, and generally does not involve the GPU level. Therefore, the judgment conditions are concentrated on the CPU and memory levels and corresponding thresholds are set to determine whether office tasks can respond smoothly; the CPU mutation change rate is used to detect whether there is unexpected CPU high-speed idling or sudden scheduling failure; the memory allocation is used to reflect whether the distribution of system memory is reasonable during office work; the work task switching frequency is used to detect whether there is task jamming or multi-task congestion.
[0039] When the task type is cloud phone, the criteria for determining task type abnormalities include system access delay, network delay jump frequency, and number of data interruptions. The operating environment of the cloud phone task type is more complex and may be accompanied by a large number of behaviors such as simulated communications, background tasks, and virtual events. The most sensitive core indicators of this task type are the system's response accessibility and communication stability. Therefore, the system access delay, network delay jump frequency, and number of data interruptions are used as judgment criteria and corresponding thresholds are set. The system access delay is used to determine whether the system is operating normally; the network delay jump frequency is used to determine network stability; and the number of data interruptions is used to measure whether there are abnormal interruptions or response timeouts in the communication link.
[0040] Ways to perform abnormal response identification include: Adaptive alarm rules are used as the judgment basis to identify abnormal responses in the node synchronization operation data, and output preliminary screening of abnormal data. The data in the node synchronization operation data is matched with the adaptive alarm rules. If it meets the conditions of the adaptive alarm rules, it is determined to be abnormal data. All abnormal data are integrated to obtain preliminary screening of abnormal data. The adaptive alarm rules provide early warning of possible abnormal data. The abnormal response identification performed at this time is the real identification of the existing abnormal data.
[0041] Identify the running behavior corresponding to the preliminary screened abnormal data, and filter the exception processing subset based on the running behavior. By performing behavioral path analysis on the running task to which the preliminary screened abnormal data belongs, first determine the task type, and then combine the running behavior to filter out the subset that needs exception processing from the preliminary screened abnormal data; the running behavior refers to the characteristics of parameters that can reflect system resource consumption under the current task type, including process scheduling status and thread load.
[0042] An abnormal risk level is set for each abnormal processing subset, and the relevant abnormal indicators whose abnormal risk level is higher than the preset risk level threshold are determined as preferred abnormal indicators, where the preset risk level threshold refers to the level threshold used to distinguish between general behaviors and high-risk behaviors based on expert experience and known theoretical knowledge. If the abnormal risk level of a certain abnormal processing subset is higher than the threshold, it means that the operating behavior corresponding to the abnormal processing subset is a high-risk behavior; taking the indicators in each abnormal processing subset as a unit, the risk level of each indicator is determined by querying historical records. If there is a relevant abnormal indicator that is higher than the preset risk level threshold, it is determined as a preferred abnormal indicator.
[0043] A dynamic sliding window is constructed to traverse all exception processing subsets, obtain all preferred exception indicators in the dynamic sliding window at the same time, and determine the preferred exception indicator as the highest processing priority. The dynamic sliding window is used to divide the complete time range into several time periods through traversal, and identify the preferred exception indicators in different exception processing subsets in the same time period; the preferred exception indicators in the time period are all determined to be the highest processing priority.
[0044] For non-preferred abnormal indicators, the dynamic sliding window of the same time is used to obtain the distribution density and duration of non-preferred abnormal indicators, among which the remaining indicators that do not belong to the preferred abnormal indicators are non-preferred abnormal indicators; the distribution density and duration of non-preferred abnormal indicators in the same time period as the preferred abnormal indicators are obtained, among which the distribution density refers to the number of occurrences of a non-preferred abnormal indicator in the time period, and the duration refers to the cumulative duration of the intermittent or continuous appearance of the indicator.
[0045] The load status characteristics of the corresponding ARM computing power nodes are collected in real time. The suboptimal anomaly score is calculated by combining the distribution density, duration, and load status characteristics. The non-preferred anomaly indicators are sorted based on the suboptimal anomaly score to obtain the suboptimal anomaly queue. The calculation formula for the suboptimal anomaly score is: ;in, represents the suboptimal anomaly score of any non-optimal anomaly indicator; Indicates the distribution density of the non-preferred abnormal indicator; Indicates the duration of the non-preferred abnormal indicator; The number of characteristics representing the load state; represents the mean value of the load state characteristics; Indicates the standard deviation of the load state characteristics; Indicates the The load state characteristics include parameters such as CPU utilization, GPU rendering utilization, and memory write blocking rate; 、 and They represent the weights of the sum of distribution density, duration and normalized load characteristic value respectively. In this embodiment, , , , those skilled in the art can adjust it according to specific circumstances; based on the suboptimal anomaly score, the suboptimal anomaly queue is obtained by descending sorting.
[0046] The optimal anomaly indicator and the suboptimal anomaly queue are integrated to obtain the anomaly response data queue, in which the optimal anomaly indicator is marked as the highest processing priority, and the remaining indicators are arranged according to the suboptimal anomaly queue to obtain the anomaly response data queue for a certain time period.
[0047] The communication status detection methods include: Get the communication path set of the ARM computing node corresponding to the abnormal response data queue, where the communication path refers to the communication link connection established by the ARM computing node in the network connection during the specific operation process. Each communication path includes information such as the transmission protocol data link, session channel and response port.
[0048] The node response data of each communication path in the communication path set is extracted, and the semantic content of the node response data is identified. At the same time, a node semantic vector matrix is constructed based on the semantic content, where the node response data refers to the response data returned by the ARM computing power node along the communication path after receiving the system instruction, including information such as message structure, field content and status code; the semantic content of the node response data is extracted by using the preset semantic large model, and the semantic content is encoded into vector form to construct a semantic vector matrix.
[0049] A structural integrity analysis is performed on the node semantic vector matrix, and the initial exception response data packet is screened based on the structural integrity analysis results. This is done by judging whether the fields of each data packet in the node semantic vector matrix are complete, whether the field order complies with the communication protocol specification, and whether the data type can be identified. If any one or more of the above requirements are not met, the corresponding node response data is judged as the initial exception response data packet.
[0050] Identify the request instructions corresponding to the response data of the remaining nodes, and query the original included fields based on the request instructions, where the response data of the remaining nodes refers to the data that has passed the structural integrity analysis except for the initial abnormal response data packet; trace back the corresponding request instructions of the response data of the remaining nodes in the initial request, and query the field content that the instruction should theoretically return through the request instruction index, which is the original included field.
[0051] The mismatch rate is calculated by comparing the existing fields with the number of originally included fields. Based on the mismatch rate and a preset threshold, the remaining node response data is subjected to a second-category screening to obtain a second-category abnormal response packet. The field match rate is calculated by calculating the ratio of the actual number of returned fields to the theoretical standard number of returned fields. The mismatch rate is calculated by subtracting the field match rate from 1. If the mismatch rate exceeds the preset threshold, the packet is determined to be a second-category abnormal response packet. The preset threshold is set by those skilled in the art based on actual needs.
[0052] The resource occupancy rates of the remaining node response data are extracted to construct an abnormal field occupancy vector. The remaining node response data is the same as the remaining node response data mentioned above, and is also the data that has passed the structural integrity analysis except for the initial abnormal response data packet. The resource occupancy rates of the nodes corresponding to all the remaining node response data are combined into an array and converted into a vector form to obtain the abnormal field occupancy vector.
[0053] Calculate the standard deviation of all elements in the abnormal field occupancy vector, perform skewness analysis on the abnormal field occupancy vector based on the standard deviation, and output the skewness coefficient. Skewness analysis is a statistical method used to measure the asymmetry of data distribution. This method uses the standard deviation to construct the third-order standard moment formula to calculate the skewness coefficient. The skewness coefficient is used to measure the degree of asymmetry in the data distribution.
[0054] If the skewness coefficient does not meet the preset standard coefficient interval, the remaining node response data will be screened into three categories to obtain three types of abnormal response data packets. The preset standard coefficient interval refers to the theoretical numerical interval of the skewness coefficient that should be obtained by using skewness analysis in the current application scenario, which is set based on the theoretical knowledge of skewness analysis.
[0055] In this embodiment, if the skewness coefficient is greater than 0 and greater than or equal to the upper boundary of the preset standard coefficient interval, it means that the resource occupancy rate of most fields is normal at this time, and a few fields may be in 0 or NULL state, and there is a risk of resource loss, such as an empty value of the interface field or an unwritten resource field; if the skewness coefficient is less than 0 and less than or equal to the lower boundary of the preset standard coefficient interval, it means that the resource occupancy rate of most fields is abnormal at this time, and a few fields are normal, and there is a one-sided resource abnormality trend, such as high occupancy of most fields; the abnormal response data packet screened each time is marked with a communication abnormality status to obtain communication status marking data, and the communication abnormality corresponding to the abnormal response data packet screened each time is marked as a label to obtain communication status marking data including the communication abnormality status label.
[0056] The abnormal response data packets screened each time are: the initial abnormal response data packet with incomplete structure screened out after the initial structural integrity analysis results; the second type of abnormal response data packet with abnormal semantic content obtained by the nodes other than the initial abnormal response data packet after the second type of screening; the third type of abnormal response data packet with abnormal resource writing obtained by the nodes other than the initial abnormal response data packet after the third type of screening; It should be noted that in the second and third type screening processes, there may be nodes that are screened out every time. In this case, when marking the communication abnormality status of these nodes, a state superposition marking method is adopted, that is, a node may have two communication abnormality states.
[0057] Methods for marking abnormal communication status include: If the result of the structural integrity analysis is that the structure is incomplete, the initial abnormal response data packet at this time will be judged as a structural destruction abnormal state. If the result of the structural integrity analysis is that the structure is incomplete, it means that there are structural level deficiencies or other abnormal conditions, and a structural destruction abnormal state label is generated to mark this type of data; if the mismatch rate is higher than the preset threshold, the second type of abnormal response data packet at this time will be judged as a semantic content abnormal state. If the mismatch rate is higher than the preset threshold, it means that there is an error in the field content, and a semantic content abnormal state label is generated to mark this type of data; if the skewness coefficient does not meet the preset standard coefficient range, the third type of abnormal response data packet at this time will be judged as a resource write abnormal state. If the skewness coefficient cannot fall within the preset standard coefficient range, it means that the resource occupancy is unbalanced, and a resource write abnormal state label is generated to mark this type of data.
[0058] The ways to generate node operation response strategies include: The communication status tag of the communication status marking data is extracted, and the communication status tag and the relevant abnormal indicators of the corresponding data are matched with the preset abnormality resolution strategy database to generate appropriate strategy instructions, wherein the abnormal state corresponding to the communication status tag and the abnormal state corresponding to the relevant abnormal indicators are used as preconditions and matched with the preset abnormality resolution strategy database to generate an instruction combination that can resolve the above abnormal state, which is the appropriate strategy instruction.
[0059] The relevant abnormal indicator fluctuation value of any node within the continuous preset monitoring period is extracted, and a state fluctuation matrix is constructed based on the value, where the relevant abnormal indicator fluctuation value refers to the change amplitude of several relevant abnormal indicators corresponding to a certain node within the continuous preset monitoring period; the abnormal state fluctuation matrix is used to horizontally reflect the fluctuation of each indicator of a certain node.
[0060] Set a reasonable fluctuation range and count the number of times the fluctuation values of all relevant abnormal indicators of each node in the state fluctuation matrix exceed the reasonable fluctuation range. The reasonable fluctuation range refers to the maximum deviation threshold that each indicator can accept under normal fluctuation conditions.
[0061] If the number is greater than the preset frequency threshold, the exception handling action is directly executed, where the preset frequency threshold refers to a reasonable number limit value set by querying existing data and expert experience; if the change range of all relevant abnormal indicators of a node exceeds the reasonable fluctuation range, and the number is greater than the preset frequency threshold, the abnormality at this time is judged as high abnormal fluctuation, and it is necessary to select an abnormal handling action that can handle the abnormal situation at this time from the preset abnormal resolution strategy database to handle it quickly.
[0062] Otherwise, the corresponding node will be temporarily marked, and the trend changes of the relevant abnormal indicators of the temporarily marked node in the subsequent continuous cycles will be recorded to construct a trend offset function. If the number of times is less than the preset frequency threshold, the corresponding node may be in a state of slight fluctuation. The system does not directly perform the abnormal processing action, but temporarily marks it, indicating that the node has potential abnormalities but has not yet caused serious fluctuations and out of control, to prevent the system's abnormal processing actions from being mistakenly triggered and affecting normal operation; the value changes of the corresponding relevant abnormal indicators of the temporarily marked node in the subsequent continuous cycles are recorded, and the trend offset function is constructed based on this; the calculation formula of the trend offset function is: ;in, The function value of the temporarily marked node in any of the subsequent consecutive periods, where any period refers to the target period selected arbitrarily from the subsequent consecutive periods and calculated using the trend shift function; Indicates the number of abnormal indicators related to the suspended marking node in any of the subsequent consecutive cycles. At the same time, any of the cycles refers to the number of abnormal indicators related to the suspended marking node in any of the subsequent consecutive cycles. Refers to the target period with the same period, which needs to be calculated using the trend shift function; Indicates the first Related abnormal indicators; Indicates the target period The value of the relevant abnormal indicator, Indicates that in the previous cycle of the target cycle The value of the relevant abnormal indicator.
[0063] If the accumulated value of the trend offset function does not exceed the preset accumulation threshold in consecutive periods and the change direction is reasonable, the suspension mark is removed and the execution of the exception handling action is abandoned. The preset accumulation threshold refers to the upper limit of the sum of the function values in consecutive periods set by querying existing data and expert experience; if the accumulated value of the trend offset function does not exceed the preset accumulation threshold in several consecutive periods, and the change direction is the correct direction, that is, it shows stabilization or oscillation, this situation indicates that the node cannot cause serious fluctuation out of control in this change state, so the suspension mark is removed and the execution of the exception handling action is abandoned.
[0064] If the trend deviation function value is continuously greater than the preset function threshold value over a continuous period and the direction of change is an upward trend, the exception handling action is directly executed. If the trend deviation function value is continuously greater than the preset function threshold value over a continuous period, the accumulated amount at this time has already exceeded the preset accumulation amount threshold value, and the direction is consistent with the direction of worsening. At this time, the corresponding node is considered to have truly entered an abnormal state, and the corresponding exception handling action is directly triggered for processing. The node operation response strategy is constructed by integrating appropriate policy instructions and the encoded exception handling action execution process. The preset function threshold value is set by those skilled in the art according to actual needs.
[0065] The ways to construct the state fluctuation matrix include: All relevant abnormal indicators belonging to the same node are normalized to obtain normalized abnormal indicators, where the dimensional differences of different indicators are eliminated through normalization to facilitate reasonable mathematical calculations; the normalized abnormal indicators are spliced into row vectors, and the normalized abnormal indicator row vectors corresponding to all nodes are superimposed to obtain a state fluctuation matrix, where each column of the state fluctuation matrix represents a dimension parameter of the normalized abnormal indicator, and each row represents all normalized abnormal indicators corresponding to a node.
[0066] This embodiment collects the running status data of ARM nodes and optimizes them from the aspects of node synchronization, alarm mechanism, abnormal response, communication detection and policy generation, thus realizing an intelligent operation and maintenance monitoring method for ARM computing nodes. Compared with the existing experience, by introducing the time synchronization process of multiple nodes, the problem of clock inconsistency caused by different hardware structures of boards is effectively solved, so that the system can adapt to the second-level response in application scenarios. By integrating task type and resource sensitivity, the system can adjust the alarm mechanism threshold according to the actual business impact of the task, avoiding the false alarm or missed alarm situation that is easy to cause by traditional alarm rules. By introducing semantic content detection and bias detection, the system can effectively solve the problem of clock inconsistency caused by different hardware structures of boards and cards, so that the system can adapt to the second-level response in application scenarios. State analysis breaks through the judgment method of "response means communication success, otherwise communication failure" in traditional communication detection technology, effectively alleviates the situation of abnormal misjudgment, and improves the accuracy of communication anomaly positioning; by detecting the fluctuation trend of node indicators, anti-shake control is achieved during the operation of computing power nodes, effectively avoiding the high-overhead operation of the system executing abnormal response strategies due to slight fluctuations in the relevant indicators of the nodes, and improving the overall operation stability; in summary, this intelligent operation and maintenance monitoring method for ARM computing power nodes significantly improves the adaptability and processing accuracy of node abnormal information in the computing power network, and provides stronger operation and maintenance adaptability and service continuity guarantee.
[0067] Example 2
[0068] See also Figure 2 As shown, for the parts not described in detail in this embodiment, please refer to the description of Example 1. An intelligent operation and maintenance monitoring system for ARM computing power nodes is provided, including: The data acquisition module is used to collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; The node synchronization module is used to perform node time synchronization processing on the node operation status data and generate node synchronization operation data; The alarm optimization module is used to optimize the alarm mechanism based on the node synchronization operation data and obtain adaptive alarm rules; The anomaly recognition module is used to perform abnormal response recognition on the node synchronization operation data in combination with the adaptive alarm rules and output the abnormal response data queue; The communication detection module is used to detect the communication status of the abnormal response data queue and obtain communication status mark data; The strategy generation module is used to generate a node operation response strategy based on the communication status mark data and send the node operation response strategy to the preset ARM computing power node control terminal; each module is connected by wired and / or wireless means.
[0069] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art will be able to modify the technical solutions described in the foregoing embodiments or to substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
[0070] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0071] In the description of the present invention, it should be understood that the terms "first", "second", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0072] In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0073] In the description of the present invention, “several” means one or more, and “a large number” means two or more.
[0074] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0075] The formulas in this manual are all dimensionless and calculated using numerical values. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters and thresholds in the formulas are set by technicians in this field based on actual conditions.
[0076] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. An intelligent operation and maintenance monitoring method for ARM computing nodes, characterized in that: include: S1. Collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; S2. Perform node time synchronization on the node operation status data to generate node synchronization operation data; S3. Optimize the alarm mechanism based on node synchronization data and obtain adaptive alarm rules; S4. Combined with adaptive alarm rules, perform abnormal response identification on the node synchronization data and output abnormal response data queue; S5 abnormal response data queue communication status detection, get the communication status mark data; S6. Generate a node operation response strategy based on the communication status mark data, and send the node operation response strategy to the preset ARM computing power node control terminal.
2. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 1, characterized in that: The method for performing node time synchronization processing includes: Obtain the board number of the ARM computing power node and group the node operation status data, and output the grouped node data; set the global sampling time standard of the node unit; calculate the difference between the timestamp of the grouped node data and the global sampling time standard of the node unit to obtain the time offset of the corresponding board node unit; compare the time offset with the preset time error tolerance threshold, and divide the node operation status data into synchronous node data subsets and asynchronous node data subsets based on the comparison result; Feature extraction is performed on the asynchronous node data subset to obtain the node time feature; time series clustering is performed on the asynchronous node data subset based on the node time feature to generate a hysteresis node cluster set and a drift trend node cluster set; historical sampling response data of different boards are obtained, and time compensation mapping is performed on the hysteresis node cluster set based on the historical sampling response data to output a time correction node cluster set; the drift period of the drift trend node cluster set is extracted and the drift rate is calculated, and the drift trend node cluster set is time-axis re-parameterized based on the drift period and drift rate to obtain a time reconstruction node cluster set; the synchronous node data subset, the time correction node cluster set and the time reconstruction node cluster set are integrated to obtain the node synchronous operation data.
3. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 2, characterized in that: The method of optimizing the alarm mechanism includes: Obtain the real-time running task identifier corresponding to the node synchronization running data, and match the running task identifier with the node number to obtain the task node index set; determine the task type based on the task node index set, and construct the resource allocation index for each task type; extract the historical node data of the corresponding task type, calculate the multi-dimensional parameter difference between the resource allocation index of the historical node data and the resource allocation index of the current node, assign sensitivity weights to the resource allocation index based on the multi-dimensional parameter difference, and output the weighted resource allocation index; Perform interactive correlation analysis on weighted resource allocation indicators and output indicator fluctuation correlation characteristics; construct task type anomaly judgment conditions by combining indicator fluctuation correlation characteristics and task types; establish a mapping relationship between task type anomaly judgment conditions and weighted resource allocation indicators, and construct anomaly correlation data sets based on the mapping relationship; identify anomaly subsets with overlapping tasks in the anomaly correlation data sets, and generate alarm intensity sequences based on the trigger response sequence of relevant indicators in the anomaly subsets; and construct adaptive alarm rules by combining relevant indicators, task type anomaly judgment conditions, and alarm intensity sequences.
4. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 3 is characterized in that: The method of constructing the task type abnormality determination condition includes: Task types include cloud gaming, mobile office, and cloud mobile phone; When the task type is cloud gaming, the abnormality criteria include the short-term GPU utilization change, the number of consecutive frame rate jitters, and the image rendering delay. When the task type is mobile office, the task type abnormality determination criteria include CPU mutation change rate, memory allocation, and work task switching frequency; When the task type is cloud phone, the task type abnormality judgment conditions include system access delay, network delay jump frequency and data interruption number.
5. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 4 is characterized in that: The method of performing abnormal response identification includes: Using adaptive alarm rules as a basis for judgment, identify abnormal responses in node synchronization operation data and output preliminary screening of abnormal data; identify the operation behavior corresponding to the preliminary screening abnormal data, and filter the abnormal processing subset based on the operation behavior; set an abnormal risk level for each abnormal processing subset, and determine the relevant abnormal indicators with abnormal risk levels higher than the preset risk level threshold as the preferred abnormal indicators; Construct a dynamic sliding window to traverse all exception processing subsets, obtain all preferred exception indicators within the dynamic sliding window at the same time, and determine the preferred exception indicator as the highest processing priority; For non-preferred abnormal indicators, the distribution density and duration of non-preferred abnormal indicators are obtained using a dynamic sliding window at the same time; the load status characteristics of the corresponding ARM computing power node are collected in real time, and the suboptimal abnormality score is calculated based on the distribution density, duration and load status characteristics. The non-preferred abnormal indicators are sorted based on the suboptimal abnormality score to obtain the suboptimal abnormality queue; the preferred abnormal indicators and the suboptimal abnormality queue are integrated to obtain the abnormal response data queue.
6. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 5, characterized in that: The communication status detection method includes: Obtain the communication path set of the ARM computing nodes corresponding to the abnormal response data queue; extract the node response data of each communication path in the communication path set, identify the semantic content of the node response data, and construct a node semantic vector matrix based on the semantic content; perform structural integrity analysis on the node semantic vector matrix, and filter the initial abnormal response data packet based on the structural integrity analysis results; Identify the request instructions corresponding to the response data of the remaining nodes, and query the original included fields based on the request instructions; calculate the mismatch rate by comparing the number of existing fields with the number of originally included fields, and perform a second-class screening of the response data of the remaining nodes based on the mismatch rate and a preset threshold to obtain the second-class abnormal response data packet; The resource occupancy rate of the response data of the remaining nodes is extracted to construct an abnormal field occupancy vector; the standard deviation of all elements in the abnormal field occupancy vector is calculated, and the skewness analysis of the abnormal field occupancy vector is performed based on the standard deviation to output the skewness coefficient; if the skewness coefficient does not meet the preset standard coefficient interval, the response data of the remaining nodes are screened into three categories to obtain three categories of abnormal response data packets; the abnormal response data packets screened each time are marked with a communication abnormality status to obtain communication status mark data.
7. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 6, characterized in that: The method of marking the abnormal communication state includes: If the structural integrity analysis result shows that the structure is incomplete, the initial abnormal response data packet at this time will be judged as a structural destruction abnormal state; if the mismatch rate is higher than the preset threshold, the second-class abnormal response data packet at this time will be judged as a semantic content abnormal state; if the skewness coefficient does not meet the preset standard coefficient range, the third-class abnormal response data packet at this time will be judged as a resource write abnormal state.
8. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 7, characterized in that: The generation node executes the response strategy in the following ways: Extract the communication status tag of the communication status mark data, match the communication status tag and the relevant abnormal indicators of the corresponding data with the preset abnormality resolution strategy database, and generate appropriate strategy instructions; Extract the relevant abnormal indicator fluctuation value of any node within the continuous preset monitoring period, and construct a state fluctuation matrix based on the relevant abnormal indicator fluctuation value; set a reasonable fluctuation range, and count the number of times all the relevant abnormal indicator fluctuation values of each node in the state fluctuation matrix exceed the reasonable fluctuation range; if the number is greater than the preset frequency threshold, directly execute the abnormal handling action; otherwise, temporarily mark the corresponding node, and record the trend change of the relevant abnormal indicators of the temporarily marked node in the subsequent continuous period to construct a trend offset function; if the accumulation of the trend offset function value in the continuous period does not exceed the preset accumulation threshold and the change direction is the correct direction, then remove the temporary mark and abandon the execution of the abnormal handling action; if the trend offset function value is continuously greater than the preset function threshold in the continuous period and the change direction is an upward trend, then directly execute the abnormal handling action; integrate appropriate strategy instructions and the encoded abnormal handling action execution process to construct a node operation response strategy.
9. The intelligent operation and maintenance monitoring method of an ARM computing power node according to claim 8, characterized in that: The method of constructing the state fluctuation matrix includes: All relevant abnormal indicators belonging to the same node are normalized to obtain normalized abnormal indicators; the normalized abnormal indicators are spliced into row vectors, and the normalized abnormal indicator row vectors corresponding to all nodes are superimposed to obtain the state fluctuation matrix.
10. An intelligent operation and maintenance monitoring system for an ARM computing power node, which is used to implement an intelligent operation and maintenance monitoring method for an ARM computing power node according to any one of claims 1 to 9, characterized in that: include: The data acquisition module is used to collect the operating status data of the ARM computing power node and perform data cleaning to obtain the node operating status data; The node synchronization module is used to perform node time synchronization processing on the node operation status data and generate node synchronization operation data; The alarm optimization module is used to optimize the alarm mechanism based on the node synchronization operation data and obtain adaptive alarm rules; The anomaly recognition module is used to perform abnormal response recognition on the node synchronization operation data in combination with the adaptive alarm rules and output the abnormal response data queue; The communication detection module is used to detect the communication status of the abnormal response data queue and obtain communication status mark data; The strategy generation module is used to generate a node operation response strategy based on the communication status mark data and send the node operation response strategy to the preset ARM computing power node control terminal; each module is connected by wired and / or wireless means.
Citation Information
Patent Citations
GPU computing power cluster monitoring method and system for large-scale multiple data centers
CN117453493A
Intelligent management and control method of IB computing power network
CN118740582A
Intelligent task alarm rule self-learning method and system based on support priority
CN119441832A
Network data risk assessment system for computer
CN119449432A
Intelligent Internet of Things public security management and control system and method based on multi-source data fusion
CN120263824A