Abnormal node judgment method and related equipment
By inputting the monitoring data and environmental noise of the target node in the large language model into the data restoration model, the predicted monitoring data is evaluated to determine the abnormal state, and the problem of difficulty in accurately determining the abnormal node in the prior art is solved, and the accuracy and robustness of the model are improved.
Patent Information
- Application Number
- CN202411998553.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to accurately determine abnormal nodes in large language models, resulting in instability and waste of resources during model training.
By obtaining the monitoring data of the target node and inputting it with environmental noise into the pre-trained data restoration model, the prediction monitoring data is output, and the evaluation results are determined whether the target node is in an abnormal state.
The accuracy of node state judgment results in the model is improved, the reliability and robustness of the model are enhanced, and the periodic characteristics of computing and network IO can be more efficiently utilized.
Smart Images

Figure CN119938444A_ABST
Abstract
Description
Background Art
[0002] Large Language Model (LLM) has become a revolutionary technology in artificial intelligence, and the latest development of LLM has demonstrated its great potential. However, LLM training also requires large-scale computing resources, high-performance networks, and longer training time spans. This large-scale model training scenario brings more complex dependencies between software and hardware and more instability, requiring deeper observability and more robust anomaly prediction models to prevent problems before they occur, issue warnings in a timely manner, or introduce automatic correction solutions.
[0003] In the related technologies, commonly used status detection methods such as heartbeat packets, log monitoring and other detection solutions need to face the balance problem between detection real-time and resource consumption. For example, heartbeat packets based on fixed actual intervals or log monitoring solutions based on fixed size segmentation and asynchronous upload do not fully consider the characteristics of computing periodicity and network IO periodicity during model training, and do not make more refined use of idle network resources during the computing cycle, and do not obtain deeper observability. At the same time, there are a large number of symmetrical nodes in large-scale clusters that execute periodic training in parallel, and the real-time comparison of their execution process and status also has strong potential for feature discovery.
[0004] In addition, the mainstream prediction models in related technologies generally use historical window data and rely on a single signal output in the last step, which may result in an unrobust model. Therefore, although some related technologies use conditional diffusion models to predict abnormal nodes, the uncertainty of time series in large-scale model training scenarios is higher, and the data performance is more heterogeneous and complex. It is difficult and costly to select appropriate and normal data for learning. Moreover, LLM training uses a large-scale cluster environment and a large number of symmetric nodes. The symmetry in each training cycle cannot be captured by the training and prediction processes of the diffusion model prediction.
[0005] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0006] The present disclosure provides a method, device, equipment, medium and computing program for determining abnormal nodes, which at least to a certain extent overcome the problem in the related art that abnormal nodes in the model cannot be accurately determined.
[0007] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0008] According to one aspect of the present disclosure, a method for determining an abnormal node is provided, comprising: obtaining monitoring data of a target node, wherein the target monitoring data is noisy data monitored and collected in real time by a monitoring point pre-set on the target node; inputting the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and outputting predicted monitoring data; evaluating the predicted monitoring data, and determining whether the target node is in an abnormal state based on the evaluation result.
[0009] In some embodiments, before inputting the target monitoring data and the collected environmental noise into a pre-trained data restoration model and outputting predicted monitoring data, the method also includes: adding pre-configured noise to the actual historical monitoring data to obtain noisy historical monitoring data; inputting the noisy historical monitoring data into a pre-built neural network model, outputting predicted noise and predicted historical monitoring data; based on a pre-defined loss function, calculating the loss value between the predicted noise and the preset noise; and updating the neural network model based on the loss value until the neural network model converges to obtain a trained data restoration model.
[0010] In some embodiments, before adding pre-configured noise to the actual historical monitoring data to obtain noisy historical monitoring data, the method also includes: obtaining a pre-constructed random mask vector, wherein the random mask vector is used to randomly mask the data during the model training process; processing the actual historical monitoring data based on the random mask vector to obtain the masked actual historical monitoring data.
[0011] In some embodiments, the loss function is expressed by the following formula:
[0012]
[0013] Among them, LOSS represents the loss function, min is used to calculate the minimum value, E represents the expected value calculation, ||·|| 2 represents the Euclidean norm square, ε represents the actual added noise, represents the prediction noise output by the data restoration model, k represents the kth computing node, and j represents the jth time step.
[0014] In some embodiments, the predictive monitoring data is evaluated and whether the target node is in an abnormal state is determined based on the evaluation result, including: evaluating the predictive monitoring data based on a preset evaluation method to obtain an outlier value of the target node; determining whether the outlier value of the target node is greater than a preset threshold; if so, determining that the target node is in an abnormal state.
[0015] In some embodiments, the preset evaluation method is an intermediate-step weighted Z-score algorithm.
[0016] According to another aspect of the present disclosure, a device for determining an abnormal node is also provided, including: a target monitoring data acquisition module, used to acquire monitoring data of the target node, wherein the target monitoring data is noisy data monitored and collected in real time by a monitoring point pre-set on the target node; a predicted monitoring data output module, used to input the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and output predicted monitoring data; a target node state determination module, used to evaluate the predicted monitoring data, and determine whether the target node is in an abnormal state based on the evaluation result.
[0017] According to another aspect of the present disclosure, an electronic device is also provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned abnormal node determination methods by executing the executable instructions.
[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for determining an abnormal node described in any one of the above is implemented.
[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned abnormal node determination methods.
[0020] The abnormal node determination method, device, equipment, medium and computing program provided in the embodiments of the present disclosure, after acquiring the monitoring data of the target node, input the target price data and the collected environmental noise into the pre-trained data restoration model, output the predicted monitoring data, and then evaluate the predicted monitoring data, and determine whether the target node is in an abnormal state according to the evaluation result. The embodiments of the present disclosure can improve the accuracy of the node state determination result in the model.
[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0023] Figure 1 A flow chart of a method for determining an abnormal node in an embodiment of the present disclosure is shown;
[0024] Figure 2 A schematic diagram of a monitoring point adding process in an embodiment of the present disclosure is shown;
[0025] Figure 3 A schematic diagram of a modeling process of an unconditional diffusion model in an embodiment of the present disclosure is shown;
[0026] Figure 4 A schematic diagram of an abnormal node determination process in an embodiment of the present disclosure is shown;
[0027] Figure 5 A schematic diagram of a device for determining an abnormal node in an embodiment of the present disclosure is shown;
[0028] Figure 6 A structural block diagram of an electronic device is shown. DETAILED DESCRIPTION
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0030] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0031] For ease of understanding, before introducing the embodiments of the present disclosure, several terms involved in the embodiments of the present disclosure are first explained as follows:
[0032] LLM: Large Language Model, refers to a deep learning model trained with a large amount of text data that can generate natural language text or understand the meaning of language text.
[0033] MFU: Model FLOPs Utilization, model computing power utilization, refers to the ratio of the matrix computing power consumed by the model's forward and reverse calculations to the node computing power.
[0034] The specific implementation of the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0035] Figure 1 A flow chart of a method for determining an abnormal node in an embodiment of the present disclosure is shown as follows: Figure 1 As shown, the method comprises the following steps:
[0036] S102, acquiring monitoring data of a target node, wherein the target monitoring data is noisy data collected and monitored in real time by a monitoring point pre-set on the target node.
[0037] In one embodiment of the present disclosure, monitoring points can be more concentratedly set at the beginning of the IO intensive process during the local training step of the target node and the reduction operation part in the gradient synchronization to reasonably utilize network resources. Specifically, the local training step may include: data segmentation tokenization and loading to the image processor GPU, forward propagation, loss calculation and back propagation, and gradient synchronization may include: gradient aggregation update. For example, monitoring points can be set in the preprocessing stage before the step is executed, the calculation stage of forward propagation, and the back propagation determined to be a computationally intensive process.
[0038] S104, input the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and output predicted monitoring data.
[0039] In one embodiment of the present disclosure, the data restoration model can be an unconditional diffusion model, which does not require additional conditional information to guide the generation process, that is, it directly generates data samples that meet the target distribution from pure noise. Based on the real-time data sequence and real-time environmental noise collected by the monitoring point, the data restoration model can output predicted monitoring data, which is the noise-free monitoring data predicted after the data restoration model is denoised.
[0040] S106, evaluating the predicted monitoring data, and determining whether the target node is in an abnormal state according to the evaluation result.
[0041] In one embodiment of the present disclosure, the predictive monitoring data can be evaluated by a pre-set evaluation method, and then it can be determined whether the target node is in an abnormal state based on the evaluation result. The periodic training characteristics of symmetric nodes can be used to achieve more effective real-time state comparison and anomaly detection, thereby significantly improving the reliability and accuracy of the model.
[0042] As can be seen from the above, after acquiring the monitoring data of the target node, the disclosed embodiment inputs the target node data and the collected environmental noise into the pre-trained data restoration model, outputs the predicted monitoring data, and then evaluates the predicted monitoring data, and determines whether the target node is in an abnormal state according to the evaluation result. The disclosed embodiment can improve the accuracy of the node state determination result in the model.
[0043] In one embodiment of the present disclosure, Figure 2 A schematic diagram of a monitoring point adding process in an embodiment of the present disclosure is shown. Figure 2 As shown in the figure, the large language model LLM is preliminarily tested to ensure that the model can run normally and achieve the expected basic performance; the more fine-grained model computing power utilization MFU and IO utilization in each training process are obtained to determine whether the condition of MFU>IO utilization is met. If so, the training process can be determined to be a computationally intensive process, and a monitoring point is inserted at the beginning of this training process. MFU refers to the ratio of the actual number of floating-point operations per second FLOPs in a forward and backward calculation of the model to the maximum FLOPs designed by the computing power card. The forward and backward calculation consumption can be calculated based on the model formula combined with the actual hyperparameter quantity.
[0044] In one embodiment of the present disclosure, the fine-grained MFU is used to calculate the subdivided processes such as data tokenization, loading to the GPU, forward propagation, backward propagation, and gradient synchronization. The MFU of the subdivided process is calculated through the actual model formula, hyperparameters, parallelization, and optimization strategies. For large models, the forward propagation FLOPs Forward =2ΦN tokens and backpropagation FLOPs Backward =4ΦN tokens Estimation is performed, where Φ is a parameter and N tokens is the number of tokens in the word segmentation. The IO utilization of hard disk, network, etc. is obtained through hardware monitoring of the segmentation process. Specifically, MFU is the ratio of the number of floating-point operations per iteration of the model (model FLOPs per iteration) to the product of the GPU single card computing power, the number of cards, and the iteration time. Model FLOPs per iteration is the total FLOPs per iteration. total Specifically, FLOP total =FLOPsForward +FLOPs Backward .
[0045] In one embodiment of the present disclosure, based on the characteristics of large-scale large-model training scenarios, during the training process, from the perspective of a single computing node, the timing has significant computing-IO periodicity characteristics. That is, after the initialization of each node is completed, the training process can be divided into local training steps (data tokenization and loading to the GPU, forward propagation, loss calculation, back propagation) and gradient synchronization (gradient aggregation update) processes. Among them, the local training step is a computationally intensive process, and the gradient synchronization is an IO-intensive process, which occupies more network resources. Therefore, in the embodiment of the present disclosure, more monitoring points are concentrated in the initial stage of the IO-intensive process in the local training step and the reduction operation part in the gradient synchronization, so as to reasonably utilize network resources.
[0046] In one embodiment of the present disclosure, considering that each process in the model has obvious computation-intensive and IO-intensive periodicity, heartbeat packets can be added to the computation-intensive process to monitor the training process more carefully. At the same time, reducing the heartbeat packets in the IO-intensive process can also avoid the occupation of the network, bus, etc., so as to balance the resource consumption allocation of monitoring and business. That is, monitoring points are added to the processes that can be determined to be computationally intensive in the local training code of each computing node. These monitoring points are located at the beginning stages of key steps such as data loading, forward propagation, loss calculation, and back propagation, as well as the reduction stage of gradient synchronization. Each step will be assigned a unique number when it is executed, which is used to identify the progress of each stage in the subsequent process. At each monitoring point, the computing node k will add a heartbeat packet at the beginning stage according to the current timestamp t. Sending mechanism, where r represents the round identifier of the current training, b represents the data batch identifier currently being processed, mb represents the micro-batch identifier, and s represents the current training step number.
[0047] In one embodiment of the present disclosure, the heartbeat packet sending mechanism may include a heartbeat packet sending function and an asynchronous disk sending function, wherein the heartbeat packet sending function is to implement a special heartbeat packet sending function on each node, which is responsible for periodically generating and sending heartbeat packets; the asynchronous disk sending function is to asynchronously write the information of the heartbeat packet into persistent storage, and then another process or thread is responsible for reading this information and sending it to the control node.
[0048] It should be noted that the embodiments of the present disclosure may adopt any of the above-mentioned heartbeat packet sending mechanisms according to actual conditions, and the present disclosure does not make any specific limitation on this.
[0049] In one embodiment of the present disclosure, before the above S103, the method also includes: adding pre-configured noise to the actual historical monitoring data to obtain noisy historical monitoring data; inputting the noisy historical monitoring data into a pre-built neural network model, outputting predicted noise and predicted historical monitoring data; based on a pre-defined loss function, calculating the loss value between the predicted noise and the preset noise; updating the neural network model based on the loss value until the neural network model converges to obtain a trained data restoration model.
[0050] In one embodiment of the present disclosure, when a large model is trained, it often has the characteristics of periodic calculation, but due to factors such as actual data distribution, network, and error failure, it may cause fluctuations in each step within the cycle. These fluctuations are regarded as noise, and the denoising process of the diffusion model can be used to monitor the noise to generate raw data. Figure 3 A schematic diagram of a modeling process of an unconditional diffusion model in an embodiment of the present disclosure is shown. Figure 3 As shown in the figure, firstly, the diffusion model is built, the model architecture and parameters are defined, and the monitoring data obtained from the monitoring points are obtained. The data are randomly denoised and trained using unsupervised learning methods to enhance the robustness and generalization ability of the model. Finally, a trained unconditional diffusion model is obtained, which is the above-mentioned data restoration model.
[0051] In one embodiment of the present disclosure, the time sequence of node k in a training cycle can be set as in, It can represent the monitoring data obtained from any monitoring point in a cycle, and N represents the number of monitoring points in a cycle.
[0052] In one embodiment of the present disclosure, a heartbeat packet may be added at the start node of a computationally intensive process. Among them, k represents the node identifier, e represents the round identifier of the current training, b represents the data batch identifier currently processed, mb represents the micro-batch identifier, s represents the current training step number, and t n represents the timestamp, i represents the number index of monitoring points, n represents the time point index of the timestamp, and M represents the total number of different features measured at each time point. Among them, t can be regarded as a variable, and the rest can be regarded as label quantities. The model as a whole can be regarded as a diffusion model of a single variable.
[0053] In one embodiment of the present disclosure, T k After adding noise according to the preset number of noise adding steps S, the noisy monitoring data is obtained. After denoising according to the preset denoising step number J, the denoised predicted monitoring data is obtained
[0054] It should be noted that the above-mentioned preset number of denoising steps and the preset number of denoising steps may be the same or different, and may be set according to actual conditions, and the embodiments of the present disclosure do not make specific limitations on this.
[0055] In one embodiment of the present disclosure, before adding pre-configured noise to the actual historical monitoring data to obtain the noisy historical monitoring data, the method also includes: obtaining a pre-constructed random mask vector, wherein the random mask vector is used to randomly mask the data during the model training process; processing the actual historical monitoring data based on the random mask vector to obtain the masked actual historical monitoring data.
[0056] In one embodiment of the present disclosure, the loss function is expressed by the following formula:
[0057]
[0058] Among them, LOSS represents the loss function, min is used to calculate the minimum value, E represents the expected value calculation, ||·|| 2 represents the Euclidean norm square, ε represents the actual added noise, represents the prediction noise output by the data restoration model, k represents the kth computing node, and j represents the jth time step.
[0059] In one embodiment of the present disclosure, an unconditional diffusion model can be used for training to collect historical monitoring data. Construct a random mask vector M = {m 1 ,m 2 ,...,m N}, M∈[0,1], the random mask vector has the same length as the historical monitoring data, and the random mask vector and the preset noise ε~N(0,I) sampled from the standard normal distribution are used to generate the random mask vector. Masking is performed to obtain masked historical monitoring data, wherein, since the vector adopts a random masking process, the noise of the masked part is known, and the noise of the unmasked part is unknown. In the disclosed embodiment, by constructing a random mask vector for random masking, the situation of partial data missing can be effectively simulated during the training process, thereby improving its applicability in the real world.
[0060] In one embodiment of the present disclosure, take ε~N(0,I) for Perform gradual noise addition. Since the calculation processes in training are independent of each other, the Markov independence assumption can be introduced. The specific calculation process can be expressed by the following formula:
[0061]
[0062] in, Indicates from the state To status The conditional probability distribution of From the initial state To status The joint probability distribution of represents the product of the conditional probabilities of all time steps, and Both represent scaling factors, α s It represents a preset value that gradually increases with the number of noise adding steps, S represents the preset number of noise adding steps, and s represents the index of the number of steps.
[0063] In one embodiment of the present disclosure, a neural network is trained by the model to predict the real noise at each time step. Given that the data is a time series of a single variable, a transformer Transformer architecture is used here, and the input is a noisy sequence. Output the prediction for each time step And predict noise The objective function is used to minimize the difference between the model output and the real noise. The loss function can be calculated by the above formula (1).
[0064] In one embodiment of the present disclosure, the above S106 includes: evaluating the predicted monitoring data based on a preset evaluation method to obtain the outlier value of the target node; judging whether the outlier value of the target node is greater than a preset threshold; if so, determining that the target node is in an abnormal state.
[0065] In one embodiment of the present disclosure, the preset evaluation method is an intermediate-step weighted Z-score algorithm.
[0066] In one embodiment of the present disclosure, the current mainstream span-leaf nodes, multi-track networks and other networking will group computing resources such as hosts or graphics cards on the hosts to achieve parallel training within and between groups, and the switches and nodes across them are symmetrical. The configuration during training (for example, in the optimization process such as model parallelism, the configuration of weight allocation and the number of graphics cards and network cards) makes the groups symmetrical at different levels. That is, symmetric nodes should be in the same operation stage at the same time, and denoising and abnormal node determination can be performed on this. Figure 4 A schematic diagram of an abnormal node determination process in an embodiment of the present disclosure is shown. Figure 4 As shown, the prediction results obtained by the diffusion model And calculate the mean square error of the prediction result. Since the preset noise added is ε~N(0,I), after multiplying two different isotropic standard Gaussian distributions, the structure still conforms to the Gaussian distribution. The intermediate step prediction noise should converge to the final prediction error, and the mean variance of the intermediate step output containing noise should also converge to the final result. When a node in the cluster is abnormal, its convergence result will be different from the convergence results of other nodes. Considering whether a node with a different convergence result is normal compared to a symmetric node, this problem can be abstracted into a sample population composed of the sample that is different from the symmetric node. The abnormal node can be determined by calculating the outlier value using the intermediate step weighted Z-score algorithm. If the outlier value obtained is greater than the preset threshold, it can be determined that the node is abnormal, thereby improving the accuracy of the prediction result.
[0067] It should be noted that the preset threshold value can be set to 3 according to conventional experience, and can also be adjusted according to actual conditions. The embodiment of the present disclosure does not specifically limit the value of the preset threshold value.
[0068] In one embodiment of the present disclosure, the mean square error of the prediction result can be calculated by the following formula:
[0069]
[0070] in, represents the mean square error of the jth time step at the kth node, represents the predicted time series at the jth time step at the kth node, represents the predicted time series of the j-1th time step at the kth node, ||·|| 2 Represents the squared Euclidean norm, which quantifies the distance between two vectors.
[0071] In one embodiment of the present disclosure, the outlier value may be calculated by the following formula:
[0072]
[0073] Among them, Z k represents an outlier, is the mean of the intermediate step outputs of each symmetric node at the jth time step, is the standard deviation of the intermediate step output of each symmetric node in the jth time step, J represents the total number of time steps, represents the mean square error of the jth time step on the kth node, K represents the total number of nodes, Represents the weighting factor.
[0074] In one embodiment of the present disclosure, by performing detailed state detection in idle network resources during the computing cycle, the periodic characteristics of computing and network IO are fully utilized, the observability of the training process is enhanced, and resources are used in a more balanced and sufficient manner; at the same time, the diffusion process is used to improve the robustness of the model. In large-scale clusters, the periodic training characteristics of symmetric nodes are utilized, and the intermediate-step weighted Z-score algorithm is used to identify outliers, thereby achieving more effective real-time state comparison and anomaly detection, thereby significantly improving the reliability and accuracy of the model, and having great potential for feature discovery.
[0075] Based on the same inventive concept, the present disclosure also provides an abnormal node determination device, as described in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0076] Figure 5 A schematic diagram of a device for determining abnormal nodes in an embodiment of the present disclosure is shown. Figure 5 As shown, the device includes: a target monitoring data acquisition module 501, a prediction monitoring data output module 502 and a target node state determination module 503.
[0077] Among them, the target monitoring data acquisition module 501 is used to obtain the monitoring data of the target node, wherein the target monitoring data is the noisy data monitored and collected in real time by the monitoring point pre-set on the target node; the predicted monitoring data output module 502 is used to input the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and output the predicted monitoring data; the target node state determination module 503 is used to evaluate the predicted monitoring data and determine whether the target node is in an abnormal state based on the evaluation results.
[0078] As can be seen from the above, after acquiring the monitoring data of the target node, the disclosed embodiment inputs the target node data and the collected environmental noise into the pre-trained data restoration model, outputs the predicted monitoring data, and then evaluates the predicted monitoring data, and determines whether the target node is in an abnormal state according to the evaluation result. The disclosed embodiment can improve the accuracy of the node state determination result in the model.
[0079] In one embodiment of the present disclosure, a preliminary test is performed on the large language model LLM to ensure that the model can run normally and achieve the expected basic performance; the more fine-grained model computing power utilization MFU and IO utilization in each training process are obtained to determine whether the condition of MFU>IO utilization is met. If so, the training process can be determined to be a computationally intensive process, and a monitoring point is inserted at the beginning of this training process. MFU refers to the ratio of the actual number of floating-point operations per second FLOPS in a forward and reverse calculation of the model to the maximum FLOPS designed by the computing power card, where the forward and reverse calculation consumption can be calculated based on the model formula combined with the actual hyperparameter quantity.
[0080] In one embodiment of the present disclosure, considering that each process in the model has obvious computation-intensive and IO-intensive periodicity, heartbeat packets can be added to the computation-intensive process to monitor the training process more carefully. At the same time, reducing the heartbeat packets in the IO-intensive process can also avoid the occupation of the network, bus, etc., so as to balance the resource consumption allocation of monitoring and business. That is, monitoring points are added to the processes that can be determined to be computationally intensive in the local training code of each computing node. These monitoring points are located at the beginning stages of key steps such as data loading, forward propagation, loss calculation, and back propagation, as well as the reduction stage of gradient synchronization. Each step will be assigned a unique number when it is executed, which is used to identify the progress of each stage in the subsequent process. At each monitoring point, the computing node k will add a heartbeat packet at the beginning stage according to the current timestamp t. Sending mechanism, where r represents the round identifier of the current training, b represents the data batch identifier currently being processed, mb represents the micro-batch identifier, and s represents the current training step number.
[0081] In one embodiment of the present disclosure, the heartbeat packet sending mechanism may include a heartbeat packet sending function and an asynchronous disk sending function, wherein the heartbeat packet sending function is to implement a special heartbeat packet sending function on each node, which is responsible for periodically generating and sending heartbeat packets; the asynchronous disk sending function is to asynchronously write the information of the heartbeat packet into persistent storage, and then another process or thread is responsible for reading this information and sending it to the control node.
[0082] It should be noted that the embodiments of the present disclosure may adopt any of the above-mentioned heartbeat packet sending mechanisms according to actual conditions, and the present disclosure does not make any specific limitation on this.
[0083] In one embodiment of the present disclosure, the device also includes: a model building module 504, which is used to add pre-configured noise to actual historical monitoring data to obtain noisy historical monitoring data; input the noisy historical monitoring data into a pre-built neural network model, and output predicted noise and predicted historical monitoring data; based on a pre-defined loss function, calculate the loss value between the predicted noise and the preset noise; based on the loss value, update the neural network model until the neural network model converges to obtain a trained data restoration model.
[0084] In one embodiment of the present disclosure, the above-mentioned model building module 504 is also used to obtain a pre-built random mask vector, wherein the random mask vector is used to randomly mask the data during the model training process; the actual historical monitoring data is processed based on the random mask vector to obtain the masked actual historical monitoring data.
[0085] In one embodiment of the present disclosure, the loss function is expressed by the above formula (1).
[0086] In one embodiment of the present disclosure, since the large model often has the characteristics of periodic calculation during training, it may cause fluctuations in each step within the cycle due to factors such as actual data distribution, network, and error failures. These fluctuations are regarded as noise, and the denoising process of the diffusion model can be used to monitor the noise to generate raw data.
[0087] In one embodiment of the present disclosure, an unconditional diffusion model can be used for training to collect historical monitoring data. Construct a random mask vector M = {m 1 ,m 2 ,...,m N}, M∈[0,1], the random mask vector has the same length as the historical monitoring data, and the random mask vector and the preset noise ε~N(0,I) sampled from the standard normal distribution are used to generate the random mask vector. Masking is performed to obtain masked historical monitoring data, wherein, since the vector adopts a random masking process, the noise of the masked part is known, and the noise of the unmasked part is unknown. In the disclosed embodiment, by constructing a random mask vector for random masking, the situation of partial data missing can be effectively simulated during the training process, thereby improving its applicability in the real world.
[0088] In one embodiment of the present disclosure, take ε~N(0,I) for The noise is added step by step. Since the calculation processes in training are independent of each other, the Markov independence assumption can be introduced. The specific calculation process can be expressed by the above formula (2).
[0089] In one embodiment of the present disclosure, the above-mentioned target node state determination module 503 is also used to evaluate the predicted monitoring data based on a preset evaluation method to obtain the outlier value of the target node; determine whether the outlier value of the target node is greater than a preset threshold; if so, determine that the target node is in an abnormal state.
[0090] In one embodiment of the present disclosure, the preset evaluation method is an intermediate-step weighted Z-score algorithm.
[0091] In one embodiment of the present disclosure, the current mainstream span-leaf nodes, multi-track networks and other networking will group computing resources such as hosts or graphics cards on the hosts to achieve parallel training within and between groups, and the switches and nodes across them are symmetrical. The configuration during training (for example, in the optimization process such as model parallelism, the configuration of weight allocation and the number of graphics cards and network cards) makes the groups symmetrical at different levels. That is, symmetric nodes should be in the same operation stage at the same time, and denoising and abnormal node determination can be performed on this.
[0092] In one embodiment of the present disclosure, the mean square error of the prediction result can be calculated by the above formula (3), and the outlier value can be calculated by the above formula (4). If the obtained outlier value is greater than a preset threshold, it can be determined that the node is abnormal, thereby improving the accuracy of the prediction result.
[0093] It should be noted that the preset threshold value can be set to 3 according to conventional experience, and can also be adjusted according to actual conditions. The embodiment of the present disclosure does not specifically limit the value of the preset threshold value.
[0094] It will be appreciated by those skilled in the art that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0095] Refer to the following Figure 6 The electronic device 600 according to this embodiment of the present disclosure is described. Figure 6 The electronic device 600 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0096] like Figure 6 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include but are not limited to: at least one processing unit 610, at least one storage unit 620, and a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610).
[0097] The storage unit stores a program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps described in the above "Exemplary Method" section of this specification according to various exemplary embodiments of the present disclosure. For example, the processing unit 610 can execute the following steps of the above method embodiment: obtaining monitoring data of the target node, wherein the target monitoring data is the noisy data collected and monitored in real time by the monitoring point pre-set on the target node; inputting the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and outputting predicted monitoring data; evaluating the predicted monitoring data, and determining whether the target node is in an abnormal state based on the evaluation result.
[0098] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0099] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0100] Bus 630 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0101] The electronic device 600 may also communicate with one or more external devices 640 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 650. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 660. As shown, the network adapter 660 communicates with other modules of the electronic device 600 via a bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0102] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0103] Based on the same inventive concept, a computer-readable storage medium is also provided in an embodiment of the present disclosure, on which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned abnormal node determination methods is implemented. Since the principle of solving the problem in the computer-readable storage medium embodiment is similar to that in the above-mentioned method embodiment, the implementation of the computer-readable storage medium embodiment can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0104] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0105] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0106] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0107] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0108] Based on the same inventive concept, a computer program product is also provided in an embodiment of the present disclosure, including a computer program product, including: a computer program or an instruction, wherein when the computer program or the instruction is executed by a processor, the abnormal node determination method of any one of the above method embodiments is implemented. Since the principle of solving the problem in the computer program product embodiment is similar to that in the above method embodiment, the implementation of the computer program product embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0109] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0110] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0111] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0112] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A method for determining abnormal nodes, characterized in that: include: Acquire monitoring data of the target node, wherein the target monitoring data is noisy data collected and monitored in real time by a monitoring point pre-set on the target node; Inputting the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and outputting predicted monitoring data; The predicted monitoring data is evaluated, and whether the target node is in an abnormal state is determined according to the evaluation result.
2. The abnormal node determination method according to claim 1, characterized in that: Before inputting the target monitoring data and the collected environmental noise into a pre-trained data restoration model and outputting the predicted monitoring data, the method further includes: Adding pre-configured noise to actual historical monitoring data to obtain noisy historical monitoring data; Inputting the noisy historical monitoring data into a pre-built neural network model, and outputting predicted noise and predicted historical monitoring data; Based on a predefined loss function, calculating a loss value between the predicted noise and the preset noise; The neural network model is updated based on the loss value until the neural network model converges to obtain a trained data restoration model.
3. The abnormal node determination method according to claim 2, characterized in that: Before adding pre-configured noise to the actual historical monitoring data to obtain noisy historical monitoring data, the method further includes: Obtaining a pre-constructed random mask vector, wherein the random mask vector is used to randomly mask data during model training; The actual historical monitoring data is processed based on the random mask vector to obtain masked actual historical monitoring data.
4. The abnormal node determination method according to claim 2, characterized in that: The loss function is expressed by the following formula: Among them, LOSS represents the loss function, min is used to calculate the minimum value, E represents the expected value calculation, ||·|| 2 represents the Euclidean norm square, ε represents the actual added noise, represents the prediction noise output by the data restoration model, k represents the kth computing node, and j represents the jth time step.
5. The abnormal node determination method according to claim 1, characterized in that: The evaluating the predicted monitoring data and determining whether the target node is in an abnormal state according to the evaluation result includes: Evaluate the predicted monitoring data based on a preset evaluation method to obtain an outlier value of the target node; Determine whether the outlier value of the target node is greater than a preset threshold; If so, it is determined that the target node is in an abnormal state.
6. The abnormal node determination method according to claim 5, characterized in that: The preset evaluation method is the intermediate step weighted Z-score algorithm.
7. A device for determining abnormal nodes, characterized in that: include: A target monitoring data acquisition module is used to acquire monitoring data of a target node, wherein the target monitoring data is noisy data collected and monitored in real time by a monitoring point pre-set on the target node; A predicted monitoring data output module, used to input the target monitoring data and the collected environmental noise into a pre-trained data restoration model, and output predicted monitoring data; The target node state determination module is used to evaluate the predicted monitoring data and determine whether the target node is in an abnormal state according to the evaluation result.
8. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; The processor is configured to execute the abnormal node determination method according to any one of claims 1 to 6 by executing the executable instructions.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the abnormal node determination method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, the abnormal node determination method according to any one of claims 1 to 6 is implemented.