A method for monitoring the abnormal operation of a new type of server
By analyzing the difficulty and degree of isolation of server running log data points, adjusting the number of isolated trees in the isolated forest algorithm, the underfit or overfitting problems of the isolated forest algorithm in server abnormal detection are solved, and the accuracy and precision of the detection are improved.
Patent Information
- Application Number
- CN202510615696.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The isolated forest algorithm is prone to underfitting or overfitting when setting the number of isolated trees, resulting in low accuracy of server abnormal detection, and the existing technology cannot effectively solve this problem.
By acquiring the running log data points of the new server, analyzing and distinguishing between difficulty and isolation, building an initial isolated forest, adjusting the number of isolated trees using local detection stability and overall detection stability, and setting the appropriate number of isolated trees to improve detection accuracy.
It improves the accuracy of the isolated forest algorithm in server abnormal detection, reduces noise interference, enhances the precision and accuracy of the detection effect, and ensures the accuracy of abnormal recognition.
Smart Images

Figure CN120123202B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing. More specifically, the present invention relates to a method for monitoring the abnormal operation of a new type of server. Background Art
[0002] With the rapid development of information technology, servers are widely used in modern enterprises, cloud computing platforms, and various data centers. The stable operation of servers is crucial for ensuring business continuity, data security, and user experience. However, during operation, servers are affected by various factors and occasionally exhibit abnormal operation phenomena. To maintain the stable operation of servers, abnormal monitoring is required.
[0003] As a commonly used anomaly detection algorithm, the Isolation Forest algorithm has good anomaly detection effects. However, the detection accuracy of this algorithm is easily affected by relevant parameters. For example, if the number of isolation trees is set too small, underfitting will occur, and some patterns in the detection model will not be fitted; if the number of isolation trees is set too large, overfitting will occur, and some interference factors in the detection model will be fitted, such as noise. Therefore, how to set an appropriate number of isolation trees becomes the research focus of the present invention.
[0004] The patent application document with the publication number CN117806912A discloses a method and system for monitoring server anomalies. The method in this patent application document realizes the anomaly detection of the server by comparing the predicted temperature with a preset threshold. Since the method in this patent application document does not involve the Isolation Forest algorithm, the technical problems of this solution cannot be solved using the method in this patent application document. Summary of the Invention
[0005] To solve the problem of how to set an appropriate number of isolation trees, the present invention proposes a method for monitoring the abnormal operation of a new type of server, which includes the following steps:
[0006] Obtain the operation log data points of the new type of server;
[0007] Obtain the discrimination difficulty , , obtain the isolation degree of each operation log data point in each dimension, cluster the isolation degrees of all operation log data points in any dimension into two categories, and obtain the membership degrees of each isolation degree to each category, respectively represent the membership degrees of the th operation log data point in the th dimension to one category and another category, represents the th operation log data point in the th dimension, Denotes the mean of the isolation degree of the j-th running log data point in all dimensions. Denotes the number of dimensions.
[0008] Construct an initial isolation forest for all running log data points using the isolation forest algorithm.
[0009] Obtain the local detection stability of each running log data point. The local detection stability characterizes the stability of the detection results of each running log data point under different numbers of isolation trees. Set weights based on the discrimination difficulty, perform weighted summation on the local detection stability to obtain the overall detection stability, and adjust the number of isolation trees in the initial isolation forest according to the overall detection stability to achieve abnormal operation monitoring.
[0010] The present invention adjusts the number of isolation trees by analyzing the detection effect of the initial isolation forest, so as to improve the accuracy of the isolation forest detection by setting a relatively accurate number of isolation trees; further, when analyzing the detection effect of the initial isolation forest, analyze each running log data point separately to improve the fineness of the detection effect analysis, and then improve the accuracy of the detection effect analysis; further, when analyzing the detection effect of the initial isolation forest, introduce the discrimination difficulty to set the evaluation weight of the detection effect of each running log data point, so as to effectively prevent the interference of factors such as noise on the detection effect evaluation and improve the accuracy of the evaluation; further, when analyzing the discrimination difficulty, introduce the variance of the isolation degree of different dimensions to accurately reflect the interference of the abnormal differences of different dimensions on abnormal recognition, and more accurately evaluate the discrimination difficulty of the isolation forest algorithm for each running log data point; further, when analyzing the discrimination difficulty, introduce the membership degree difference of each running log data point to different categories to reflect the degree to which each running log data point belongs to abnormal and normal, and then provide a data basis for accurately evaluating the discrimination difficulty.
[0011] Preferably, the constructing an initial isolation forest for all running log data points using the isolation forest algorithm includes:
[0012] Obtain the pre-obtained initial number of isolation trees.
[0013] Process all running log data points using the isolation forest algorithm based on the initial number of isolation trees to obtain the initial isolation forest.
[0014] Preferably, the obtaining the pre-obtained initial number of isolation trees includes:
[0015] Multiply the normalized value of the cumulative sum of the discrimination difficulties of all running log data points by the preset number of isolation trees to obtain the initial number of isolation trees.
[0016] The present invention sets the initial number of isolation trees by combining the discrimination difficulty, thereby considering the discrimination difficulty of the analyzed data when setting the isolation trees, so as to set more appropriate isolation trees and improve the accuracy of anomaly detection.
[0017] Preferably, the obtaining of the isolation degree of each operation log data point in each dimension includes:
[0018] For any dimension, obtain the distance between every two adjacent operation log data points in this dimension, and record the cumulative sum of the distances between all two adjacent operation log data points in this dimension as the overall isolation distance; obtain the accumulated sum of the distances between any operation log data point and its two adjacent operation log data points in this dimension as the local isolation distance of this operation log data point in this dimension; divide the local isolation distance of this operation log data point in this dimension by the overall isolation distance to obtain the isolation degree of this operation log data point in this dimension.
[0019] Preferably, the obtaining of the membership degree of each isolation degree to each category includes:
[0020] Take the normalized value of the reciprocal of the distance between any isolation degree and the cluster center of any category as the center binding degree;
[0021] Take the normalized value of the reciprocal of the product of the distance between this isolation degree and the category boundary of this category and the attribute relation flag value as the boundary binding degree;
[0022] Take the product of the center binding degree and the boundary binding degree as the membership degree of this isolation degree to this category.
[0023] When analyzing the membership degree, the present invention not only considers the distance from the cluster center, but also considers the distance from the boundary, so as to more comprehensively evaluate the degree to which the isolation degree belongs to each category.
[0024] Preferably, the obtaining of the local detection stability of each operation log data point includes:
[0025] Set a preset variable , randomly combine every T isolation trees in the initial isolation forest to obtain several isolation tree combinations, perform anomaly detection on any operation log data point based on each isolation tree combination to obtain the anomaly score under this isolation tree combination, and take the mean value of the anomaly scores of this operation log data point under all isolation tree combinations as the comprehensive anomaly score; take each integer in the interval [K - A, K] for T, where K represents the number of isolation trees included in the initial isolation forest, and A represents a preset parameter; take the difference between the comprehensive anomaly score at any value and the comprehensive anomaly score at the previous value as the anomaly detection change amount at this value, and take the reciprocal of the mean value of the anomaly detection change amounts of this operation log data point at all values as the local detection stability of this operation log data point.
[0026] The present invention accurately reflects the convergence of the detection results of the initial isolation forest by introducing the differences in the detection effects of each operation log data point under different combinations of isolation trees, and further accurately reflects the appropriate number of isolation trees in the initial isolation forest, providing a basis for setting an appropriate number of isolation trees.
[0027] Preferably, adjusting the number of isolation trees in the initial isolation forest according to the overall detection stability includes:
[0028] Taking the ratio of a preset reference value to the overall detection stability as an adjustment coefficient;
[0029] Taking the integer value of the product of the number of isolation trees in the initial isolation forest and the adjustment coefficient as the adjusted number of isolation trees.
[0030] The present invention adjusts the number of isolation trees according to the overall detection stability, thereby obtaining a more accurate number of isolation trees, and further improving the accuracy of anomaly detection.
[0031] Preferably, realizing operation anomaly monitoring includes:
[0032] Inputting the operation log data points of the newly collected new type of server into the isolation forest with the adjusted number of isolation trees to obtain anomaly detection results.
[0033] Preferably, after obtaining the anomaly detection results, it further includes:
[0034] If the anomaly detection result is that there is an anomaly, a warning reminder is issued;
[0035] If the anomaly detection result is that there is no anomaly, no warning reminder is issued.
[0036] Preferably, obtaining the operation log data points of the new type of server includes:
[0037] Obtaining each log of the new type of server and converting the data of each log;
[0038] Denoting the data points formed by the converted data of all dimensions in each log as operation log data points.
[0039] The present invention has the following beneficial effects:
[0040] The present invention adjusts the number of isolation trees by analyzing the detection effect of the initial isolation forest, thereby improving the accuracy of the isolation forest detection by setting a relatively accurate number of isolation trees;
[0041] Furthermore, when analyzing the detection effect of the initial isolation forest, by analyzing each operation log data point separately, the fineness of the detection effect analysis is improved, and further the accuracy of the detection effect analysis is improved;
[0042] Further, when analyzing the detection effect of the initial isolation forest, by introducing the discrimination difficulty to set the evaluation weight of the detection effect of each operation log data point, the interference of factors such as noise to the detection effect evaluation can be effectively prevented, and the accuracy of the evaluation can be improved;
[0043] Further, when analyzing the discrimination difficulty, by introducing the variance of the isolation degree in different dimensions to accurately reflect the interference of the abnormal differences in different dimensions to the abnormal recognition, the discrimination difficulty of the isolation forest algorithm for each operation log data point can be evaluated more accurately;
[0044] Further, when analyzing the discrimination difficulty, by introducing the membership degree difference of each operation log data point to different categories to reflect the degree to which each operation log data point belongs to abnormal and normal, and further providing a data basis for accurately evaluating the discrimination difficulty. Description of the Drawings
[0045] Figure 1 is a flowchart of the steps of a method for monitoring operation anomalies of a new type of server according to an embodiment of the present invention. Detailed Embodiment
[0046] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.
[0047] Next, the detailed embodiment of the present invention will be described in conjunction with the drawings.
[0048] Please refer to Figure 1 , which shows a flowchart of the steps of a method for monitoring operation anomalies of a new type of server provided by an embodiment of the present invention. The method includes the following steps:
[0049] S1: Obtain the operation log data points of the new type of server.
[0050] Specifically, obtain the operation logs of the new type of server, and convert the words, texts, and characters in each operation log into data. In this embodiment, the Word2vec algorithm is used for data conversion, and other embodiments can adopt other methods, which are not specifically limited in this embodiment.
[0051] The data points formed by the converted data in all dimensions of each operation log are recorded as operation log data points.
[0052] S2: Obtain the discrimination difficulty.
[0053] It should be noted that some running log data points are noise data points, and the distinguishability between these data points and normal data points is relatively small. When the number of isolation trees in the isolation forest algorithm is set too large, these data points will be misjudged as abnormal data points. And although some running log data points are abnormal data points, the distinguishability between these data points and normal data points is also relatively small. If the number of isolation trees in the isolation forest algorithm is set too small, these data points will be misjudged as normal data points, thus failing to detect these abnormal data points. Therefore, in order to set an appropriate number of isolation trees, it is necessary to analyze the detection effects of running log data points with different levels of distinguishability difficulty. First, analyze the distinguishability difficulty of each type of running log data.
[0054] Preferably, as an example, obtaining the distinguishability difficulty includes:
[0055]
[0056] Among them, obtain the isolation degree of each running log data point in each dimension, cluster the isolation degrees of all running log data points in any dimension into two categories, and obtain the membership degrees of each isolation degree to each category. Denote the th running log data point's membership degree of the isolation degree in the th dimension to a category. Denote the th running log data point's membership degree of the isolation degree in the th dimension to another category. Denote the th running log data point's isolation degree in the th dimension. Denote the mean value of the isolation degrees of the jth running log data point in all dimensions. Denote the number of dimensions. Denote the distinguishability difficulty of the jth running log data point.
[0057] It can be understood that reflects the difference in the membership degrees of the isolation degrees of each running log data point to different categories. The larger this value is, the more the running log data point tends to a certain detection result. Therefore, the distinguishability difficulty of this running log data point is relatively small. reflects the difference in the isolation degrees of each running log data point in different dimensions. The larger this value is, the greater the difference in the isolation degrees of the running log data point in different dimensions. Therefore, the possibility that the obtained anomaly detection results using different dimensions are different is relatively large. Therefore, the distinguishability difficulty of this running log data point is relatively large.
[0058] The above embodiments relate to the degree of isolation and the membership degrees of various degrees of isolation to various categories. Next, the method for determining the degree of isolation and the membership degrees of the degree of isolation to various categories will be described.
[0059] First, the method for obtaining the degree of isolation will be introduced.
[0060] Preferably, as an example, obtaining the degree of isolation of each running log data point in each dimension includes:
[0061] For any dimension, obtain the distance between every two adjacent running log data points in this dimension, and record the cumulative sum of the distances between all two adjacent running log data points in this dimension as the overall isolation distance; obtain the sum of the distances between any running log data point and its two adjacent running log data points in this dimension and record it as the local isolation distance of this running log data point in this dimension; divide the local isolation distance of this running log data point in this dimension by the overall isolation distance to obtain the degree of isolation of this running log data point in this dimension.
[0062] It should be noted that the implementation process of the isolation forest algorithm includes: obtaining the value range of the data of all data points in one dimension, randomly extracting a point within the value range as the splitting point, and using the splitting point to divide the data points into two sets until independent data points are split out. Since abnormal data points deviate from normal data points, the abnormal data points will be split out quickly. Through the above calculation process, it can be found that the degree of isolation reflects the probability of being extracted on both sides of each running log data point in one dimension. The larger this value is, the greater the probability that this running log data point is extracted, and thus the smaller the difficulty of splitting out this running log data point. Therefore, the degree of isolation calculated in this way can provide a data basis for accurately calculating the discrimination difficulty.
[0063] Next, the method for obtaining the membership degrees of the degree of isolation to various categories will be introduced.
[0064] Preferably, as an example, obtaining the membership degrees of each degree of isolation to various categories includes:
[0065] Take the normalized value of the reciprocal of the distance between any degree of isolation and the clustering center of any category as the center binding degree;
[0066] Take the normalized value of the reciprocal of the product of the distance between this degree of isolation and the category boundary of this category and the attribute relationship flag value as the boundary binding degree;
[0067] Take the product of the center binding degree and the boundary binding degree as the membership degree of this degree of isolation to this category. In this embodiment, the maximum-minimum normalization method is used for normalization processing. Other embodiments can use other normalization methods, and this embodiment does not make specific limitations.
[0068] It should be noted that when performing membership analysis, not only the distance between the data point and the center point is considered, but also the distance between the data point and the boundary is taken into account. Therefore, it can analyze the membership relationship between the data and the corresponding category more comprehensively and accurately.
[0069] It should be added that the method for setting the attribute relationship flag value includes:
[0070] If the degree of isolation belongs to a certain category, set the attribute relationship flag value between the degree of isolation and this category to -1. If the degree of isolation does not belong to a certain category, set the attribute relationship flag value between this degree of isolation and this category to 1.
[0071] S3: Use the isolation forest algorithm to construct an initial isolation forest for all running log data points.
[0072] It should be noted that in order to judge the detection effect of running log data points with different degrees of discrimination difficulty, an anomaly detection model needs to be constructed first. Since the specific number of isolation trees cannot be determined currently, a rough number of isolation trees needs to be determined first, and an anomaly detection model is constructed based on the roughly determined number of isolation trees.
[0073] Optionally, as an example, using the isolation forest algorithm to construct an initial isolation forest for all running log data points includes:
[0074] A preset number of isolation trees, take the preset number of isolation trees as the initial number of isolation trees, and based on the initial number of isolation trees, use the random forest algorithm to analyze all running log data points to obtain the initial isolation forest.
[0075] It should be noted that the preset number of isolation trees is generally a value set according to experience. However, due to the different data types for anomaly analysis, using fixed empirical values to process diverse data cannot achieve good results.
[0076] Preferably, as an example, using the isolation forest algorithm to construct an initial isolation forest for all running log data points includes:
[0077] Multiply the normalized value of the sum of the discrimination difficulties of all running log data points by the preset number of isolation trees to obtain the initial number of isolation trees;
[0078] Based on the initial number of isolation trees, use the isolation forest algorithm to process all running log data points to obtain the initial isolation forest.
[0079] It should be noted that the discrimination difficulty reflects the difficulty of anomaly detection of running log data points. The larger this value is, the more isolation trees need to be set for fitting. Therefore, the preset number of isolation trees can be adjusted according to the discrimination difficulty.
[0080] S4: Obtain the local detection stability of each running log data point. The local detection stability characterizes the stability of the detection results under different numbers of isolated trees for each running log data point. Set weights based on the discrimination difficulty, perform weighted summation on the local detection stability to obtain the overall detection stability, and adjust the number of isolated trees in the initial isolation forest according to the overall detection stability to achieve running anomaly monitoring.
[0081] S40: Obtain the local detection stability of each running log data point.
[0082] It should be noted that the above is only the roughly set number of isolated trees. To obtain a more accurate number of isolated trees, it needs to be adjusted according to the actual anomaly detection effect. First, obtain the detection effect of the initial isolation forest constructed based on the roughly set number of isolated trees on each type of running log data point. In this embodiment, the detection effect is reflected by the local detection stability.
[0083] Preferably, as an example, obtaining the local detection stability of each running log data point includes:
[0084] Set a preset variable T, randomly combine every T isolated trees in the initial isolation forest to obtain several isolated tree combinations. Based on each isolated tree combination, perform anomaly detection on any running log data point to obtain the anomaly score of the running log data point under this isolated tree combination. Take the mean of the anomaly scores of the running log data point under all isolated tree combinations as the comprehensive anomaly score of the running log data point.
[0085] Let T take each integer in the interval [K - A, K], where K represents the number of isolated trees in the initial isolation forest, and A represents a preset parameter. In this embodiment, the preset parameter is taken as 3 for description. Other embodiments can take other values, and this embodiment does not make specific restrictions. Obtain the comprehensive anomaly score of the running log data point under each value. Take the difference between the comprehensive anomaly score under any value and the comprehensive anomaly score under the previous value as the anomaly detection change amount of the running log data point under this value. Take the reciprocal of the mean of the anomaly detection change amounts of the running log data point under all values as the local detection stability of the running log data point;
[0086] It should be noted that the smaller the local detection stability, the less the initial isolation forest has converged, and thus the greater the possibility of underfitting in the initial isolation forest; the larger this value, the more likely it is that stable detection results can be obtained by only using part of the isolated trees in the initial isolation forest, and thus the greater the possibility of overfitting in the initial isolation forest.
[0087] It should be noted that performing anomaly detection on each running log data point based on the isolated tree combination is a prior art and will not be elaborated here.
[0088] S41: Set weights based on the discrimination difficulty, and perform weighted summation on the local detection stability to obtain the overall detection stability.
[0089] Preferably, as an example, setting weights based on the discrimination difficulty and performing weighted summation on the local detection stability to obtain the overall detection stability includes:
[0090] Taking the natural constant as the base and the negative value of the product of the discrimination difficulty of each operation log data point and the preset ratio as the exponent for calculation to obtain the weight of each operation log data point. Taking the weight of each operation log data point as the weight, performing weighted summation on the local detection stability of all operation log data points and then performing normalization processing to obtain the overall detection stability. In this embodiment, the preset ratio is taken as 0.05 for description, and other values can be taken in other embodiments, and this embodiment does not make specific limitations.
[0091] It should be noted that when the discrimination difficulty is too large, it indicates that the operation log data point is more likely to be interfered by noise, and the data interfered by noise is not suitable for evaluating the detection effect of the initial isolation forest. Therefore, the weight of the data with large discrimination difficulty needs to be adjusted downwards.
[0092] It should be further noted that the overall detection stability reflects the stability of the initial isolation forest anomaly detection. The larger this value is, the more stable the initial isolation forest is, and the greater the probability of overfitting of this initial isolation forest. The smaller this value is, the more unstable the initial isolation forest is, and the greater the probability of underfitting of this initial isolation forest. Therefore, the fitting situation of the initial isolation forest can be judged according to the overall detection stability.
[0093] S42: Adjust the number of isolation trees in the initial isolation forest according to the overall detection stability.
[0094] It should be noted that in order to obtain a more accurate number of isolation trees, the number of isolation trees needs to be adjusted according to the detection effect.
[0095] Preferably, as an example, adjusting the number of isolation trees in the initial isolation forest according to the overall detection stability includes:
[0096] Taking the ratio of the preset reference value to the overall detection stability as the adjustment coefficient;
[0097] Taking the integer value of the product of the number of isolation trees in the initial isolation forest and the adjustment coefficient as the adjusted number of isolation trees.
[0098] If the adjusted number of isolation trees is less than the number of isolation trees in the initial isolation forest, randomly select C isolation trees in the initial isolation forest for removal, and use the isolation forest composed of the remaining isolation trees as the isolation forest after adjusting the number of isolation trees.
[0099] If the number of isolated trees after adjustment is greater than the number of isolated trees in the initial isolated forest, then construct another C isolated trees and supplement the newly constructed isolated trees into the initial isolated forest to obtain the isolated forest with the adjusted number of isolated trees. C represents the absolute value of the difference between the number of isolated trees after adjustment and the number of isolated trees in the initial isolated forest.
[0100] It should be noted that if the overall detection stability is too large, it indicates that the number of isolated trees in the initial isolated forest is too large and there is a greater possibility of overfitting. Therefore, the number of isolated trees can be reduced; if the overall detection stability is too small, it indicates that the number of isolated trees in the initial isolated forest is too small and there is a greater possibility of underfitting. Therefore, the number of isolated trees can be increased.
[0101] S43: To achieve operation anomaly monitoring.
[0102] Preferably, as an example, to achieve operation anomaly monitoring, it includes:
[0103] Input the operation log data points of the newly collected new-type server into the isolated forest with the adjusted number of isolated trees to obtain the anomaly detection result.
[0104] If the anomaly detection result indicates the existence of an anomaly, then issue a warning reminder;
[0105] If the anomaly detection result indicates the non-existence of an anomaly, then do not issue a warning reminder.
[0106] So far, this embodiment is completed.
[0107] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for monitoring the abnormal operation of a new type of server, characterized in that, Including: Obtaining the operation log data points of a new type of server; Obtain the discrimination difficulty , , obtain the isolation degree of each operation log data point in each dimension, cluster the isolation degrees of all operation log data points in any dimension into two categories, and obtain the membership degrees of each isolation degree to each category respectively represent the membership degrees of the -th operation log data point in the -th dimension to one category and another category represents the -th operation log data point in the -th dimension represents the mean value of the isolation degrees of the j-th operation log data point in all dimensions represents the number of dimensions Obtaining the isolation degree of each operation log data point in each dimension, including: for any dimension, obtaining the distance between every two adjacent operation log data points in this dimension, and recording the cumulative sum of the distances between all two adjacent operation log data points in this dimension as the overall isolation distance; obtaining the cumulative sum of the distances between any operation log data point and its two adjacent operation log data points in this dimension as the local isolation distance of this operation log data point in this dimension; dividing the local isolation distance of this operation log data point in this dimension by the overall isolation distance to obtain the isolation degree of this operation log data point in this dimension; Using the Isolation Forest algorithm to construct an initial isolation forest for all operation log data points; Obtaining the local detection stability of each operation log data point, where the local detection stability characterizes the stability of the detection results of each operation log data point under different numbers of isolation trees, setting weights based on the discrimination difficulty, performing a weighted sum on the local detection stability to obtain the overall detection stability, and adjusting the number of isolation trees in the initial isolation forest according to the overall detection stability to achieve operation anomaly monitoring.
2. The abnormal operation monitoring method of a new type of server according to claim 1, wherein The using the Isolation Forest algorithm to construct an initial isolation forest for all operation log data points includes: Obtaining the pre-obtained initial number of isolation trees; Based on the initial number of isolation trees, using the Isolation Forest algorithm to process all operation log data points to obtain an initial isolation forest.
3. The abnormal operation monitoring method for a new type of server according to claim 2, wherein The obtaining the pre-obtained initial number of isolation trees includes: Multiplying the normalized value of the cumulative sum of the discrimination difficulties of all operation log data points by a preset number of isolation trees to obtain the initial number of isolation trees.
4. A method for monitoring abnormal operation of a new type of server according to claim 1, characterized in that, The obtaining the membership degree of each isolation degree to each category includes: Taking the normalized value of the reciprocal of the distance between any isolation degree and the cluster center of any category as the center binding degree; Taking the normalized value of the reciprocal of the product of the distance between this isolation degree and the category boundary of this category and the attribute relationship flag value as the boundary binding degree; Taking the product of the center binding degree and the boundary binding degree as the membership degree of this isolation degree to this category.
5. A method for monitoring the abnormal operation of a new type of server according to claim 1, characterized in that, The obtaining the local detection stability of each operation log data point includes: Set preset variables , randomly combine every T isolated trees in the initial isolation forest to obtain a number of isolated tree combinations, perform anomaly detection on any running log data point based on each isolated tree combination to obtain the anomaly score under this isolated tree combination, and take the mean of the anomaly scores of this running log data point under all isolated tree combinations as the comprehensive anomaly score; Take each integer between the interval [K - A, K], where K represents the number of isolated trees in the initial isolation forest, and A represents a preset parameter; take the difference between the comprehensive anomaly score at any value and the comprehensive anomaly score at the previous value as the anomaly detection change amount at this value, and take the reciprocal of the mean of the anomaly detection change amounts of this running log data point at all values as the local detection stability of this running log data point.
6. A method for monitoring the abnormal operation of a new type of server according to claim 1, characterized in that, The adjusting the number of isolation trees in the initial isolation forest according to the overall detection stability includes: Taking the ratio of a preset reference value to the overall detection stability as the adjustment coefficient; Taking the integer value of the product of the number of isolation trees in the initial isolation forest and the adjustment coefficient as the adjusted number of isolation trees.
7. A method for monitoring abnormal operation of a new type of server according to claim 1, characterized in that, The achieving operation anomaly monitoring includes: Inputting the operation log data points of the newly collected new type of server into the isolation forest with the adjusted number of isolation trees to obtain an anomaly detection result.
8. A method for monitoring the abnormal operation of a new type of server according to claim 7, characterized in that, After obtaining the anomaly detection result, it further includes: If the anomaly detection result is that there is an anomaly, issuing a warning reminder; If the anomaly detection result is that there is no anomaly, not issuing a warning reminder.
9. The operation anomaly monitoring method for a new type of server according to claim 1, characterized in that, The obtaining the operation log data points of the new type of server includes: Obtaining each log of the new type of server and performing data conversion on each log; Recording the data points formed by the converted data of all dimensions in each log as operation log data points.
Citation Information
Patent Citations
Server abnormity monitoring method and system
CN117806912A
Server log anomaly detection method and system based on isolated forest algorithm
CN110958222A