5G-R core network virtual machine fault prediction method and system based on random forest algorithm

By applying the fault prediction method of random forest algorithm in 5G-R core network virtual machines, the potential failure of virtual machines is extracted and predicted, and the problem that traditional monitoring cannot be warning in advance is solved, and the reliability and service continuity of virtual machines are improved.

CN120029717APending Publication Date: 2025-05-23CHINA RAILWAY DESIGN GRP CO LTD +1

Patent Information

Application Number
CN202510059571.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional virtual machine monitoring can only alarm after a failure occurs, and cannot provide early warning of failure, which affects the reliability and service continuity of the virtual machine.

Method used

The fault prediction method based on the random forest algorithm is adopted, and the monitoring data of the 5G-R core network virtual machine is collected and preprocessed, characteristic data related to the virtual machine status is extracted, and a random forest model is built for training and tuning. Finally, the virtual machine status is predicted within each time period, and whether there is a potential fault and alarm is issued.

Benefits of technology

It realizes early warning of virtual machine failures, avoids interruption of virtual machine services, and improves the reliability and service continuity of virtual machines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029717A_ABST
    Figure CN120029717A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of big data and machine learning, and discloses a 5G-R core network virtual machine fault prediction method and system based on a random forest algorithm, and the method comprises the steps: carrying out the feature extraction of preprocessed monitoring data based on mutual information and chi-square test, and obtaining the feature data strongly related to the state of a 5G-R core network virtual machine; constructing a random forest model, and training a random forest to obtain an optimal random forest model; and based on the optimal random forest model and the acquired monitoring indexes, predicting the state of the virtual machine in each time period, judging whether a potential fault exists or not, and if so, giving an alarm. According to the method, on the basis of various index data collected by the virtual machine prometheus, the random forest algorithm is adopted, the virtual machine fault prediction model is established, the fault probability of the virtual machine is predicted, migration and backup of key services are completed in time, interruption of the virtual machine services is avoided, and the reliability of the virtual machine is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of big data and machine learning, and relates to a 5G-R core network virtual machine fault prediction method and system based on a random forest algorithm. Background Art

[0002] With the rapid development of cloud technology, the application of cloud technology by public network operators is becoming more and more extensive. The application of cloud technology not only improves the service quality and efficiency of the operator, but also makes it easier to cope with changes in data traffic and the diversification of customer needs. In order to improve the service capabilities of railway operations, the 5G-R core network adopts a cloud deployment solution. As an important part of the railway network, the 5G-R core network has a more important 24-hour service capability. The bottom layer of the cloud server relies on virtual machine support. The reliability and performance of the virtual machine directly affect the use of upper-layer cloud services. Traditional virtual machine monitoring can only achieve "fault repair", that is, only when the virtual machine fails, will an alarm be processed, and early warning of the failure cannot be achieved. Summary of the invention

[0003] The purpose of the present invention is to solve the problem that traditional virtual machine monitoring in the prior art can only achieve "fault repair", that is, alarm processing is only performed when a virtual machine fails, and early warning of the failure cannot be achieved, and a 5G-R core network virtual machine fault prediction method and system based on a random forest algorithm is provided.

[0004] In order to achieve the above object, the present invention adopts the following technical solutions:

[0005] The 5G-R core network virtual machine fault prediction method based on the random forest algorithm includes:

[0006] Collect monitoring data on the 5G-R core network virtual machine and pre-process the collected monitoring data;

[0007] Based on mutual information and chi-square test, feature extraction is performed on the pre-processed monitoring data to obtain feature data that is strongly correlated with the status of the 5G-R core network virtual machine;

[0008] Constructing a random forest model, dividing the acquired feature data into a training set and a test set;

[0009] The random forest model is trained based on the training set, and the performance of the model is evaluated through the test set. The hyperparameters of the random forest model are tuned based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, and then the optimal random forest model is obtained.

[0010] Based on the optimal random forest model and the collected monitoring indicators, the virtual machine status is predicted in each time period to determine whether there is a potential failure. If so, an alarm is issued.

[0011] A further improvement of the present invention is:

[0012] Furthermore, monitoring data on the 5G-R core network virtual machine is collected, specifically: the Prometheus service is deployed on the 5G-R core network virtual machine, and monitoring indicators are regularly collected from the 5G-R core network virtual machine through the exporter module of Prometheus; the monitoring indicators include: timestamp, IP address, host name, running time, memory, number of CPU cores, 5-minute load, CPU utilization, memory utilization, disk partition utilization, disk read and write rate, network download rate, network upload rate and virtual machine status.

[0013] Further, the collected monitoring data is preprocessed, specifically: cleaning, filtering and feature extraction of the collected data, including timestamp transformation, data standardization, missing value processing and data conversion;

[0014] The timestamp transformation is to convert the timestamp into a standard time format yyyy-MM-dd HH:mm:ss, and extract the year, month, day, hour, minute, and second derived attributes;

[0015] The data standardization is to keep the dimensions of each attribute of running time, memory, disk read and write rate, network download rate, and network upload rate uniform and perform data conversion;

[0016] The unit of the running time is unified as seconds, the unit of the memory is unified as bytes, the unit of the disk read and write rate is unified as bytes per second, and the unit of the network download rate and the network upload rate is unified as kilobytes per second;

[0017] The missing value processing is for sample data with missing 5-minute load attributes, and the data is marked as NAN; the data is converted to quantify the virtual machine status attributes, 0 is marked as normal, and 1 is marked as abnormal.

[0018] Furthermore, feature data strongly related to the state of the 5G-R core network virtual machine is obtained, specifically: 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate are selected as feature variables, and the virtual machine state is used as the target variable to calculate the mutual information and chi-square test;

[0019] According to the mutual information calculation results, the mutual information calculation results are sorted, and the front feature vector is selected as the feature variable with strong correlation with the target variable virtual machine state;

[0020] Based on the chi-square test calculation results, the relationship between the chi-square test calculation results and the preset thresholds is judged in turn to obtain the characteristic variables with strong correlation with the target variable virtual machine state;

[0021] The final characteristic variables are obtained by comprehensively considering the mutual information calculation results and the chi-square test calculation results.

[0022] Furthermore, the mutual information and chi-square test are calculated as follows:

[0023] The mutual information formula is as follows:

[0024]

[0025] Among them, P(x,y) is the joint probability of variables X and Y taking the values ​​of x and y at the same time; P(x) is the marginal probability of variable X; P(y) is the marginal probability of variable Y;

[0026] The chi-square test formula is as follows:

[0027]

[0028] Where O is the actual observed frequency; E is the expected frequency; x 2 is the chi-square statistic, which is used to measure the deviation of the actual value from the expected value.

[0029] Furthermore, a random forest model is constructed, specifically:

[0030] Based on the Bootstrap sampling method, the sample data collected from Prometheus is sampled with replacement to generate multiple sub-sample sets;

[0031] A decision tree is constructed for each sub-sample set. At each splitting node, several features in the sub-sample set are randomly selected for splitting to obtain a decision tree.

[0032] Integrate the prediction results of multiple decision trees to achieve multi-tree voting to generate the final prediction; in classification tasks, the model generates prediction results through majority voting:

[0033] f(x) = majority_vote(h 1 (x),h 2 (x),...,h m (x))

[0034] Among them, h i (x) represents the prediction result of the i-th decision tree for input x, and m is the number of decision trees.

[0035] Furthermore, the random forest model is trained based on the training set, and the performance of the model is evaluated through the test set, specifically:

[0036] The cross entropy loss function is used to measure the difference between the model prediction value and the true label. The formula is as follows:

[0037]

[0038] Where N is the total number of samples; y i is the actual label of sample i, which is 0 or 1; z i is the probability that sample i is predicted to be a fault, and the output value is between [0, 1];

[0039] The training process is as follows:

[0040] Each decision tree in the random forest randomly selects a subset from the original feature set for splitting. j The splitting of the decision tree is optimized by calculating the Gini index or information gain of the feature; for each sample x in the training set i , each decision tree T j Output a predicted label h j (x i ); calculate the prediction error of each decision tree through the cross entropy loss function; vote on the prediction results of all decision trees, and the final prediction result is the category with the most votes; adjust the weight and splitting method of the decision tree by minimizing the loss function or cross entropy, and evaluate the performance of the model through the test set. If the performance of the random forest model is not good, adjust the parameters of the random forest to optimize the random forest model.

[0041] Furthermore, it is determined whether there is a potential fault. Specifically, the current monitoring data of the system is collected, and 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate are selected as feature variables. The input features are passed through the splitting rules of each decision tree in turn to determine whether each decision tree has a potential fault, and the prediction results of each tree are counted. If the percentage of decision trees with potential faults in the total decision trees exceeds the set threshold, the final prediction label is obtained.

[0042] The 5G-R core network virtual machine fault prediction system based on the random forest algorithm includes:

[0043] A preprocessing module, wherein the preprocessing module collects monitoring data on the 5G-R core network virtual machine and preprocesses the collected monitoring data;

[0044] A feature extraction module, wherein the feature extraction module extracts features from the preprocessed monitoring data based on mutual information and chi-square test to obtain feature data that is strongly correlated with the state of the 5G-R core network virtual machine;

[0045] A construction module, wherein the construction module constructs a random forest model and divides the acquired feature data to obtain a training set and a test set;

[0046] A tuning module, wherein the tuning module trains the random forest model based on the training set and evaluates the performance of the model through the test set; the hyperparameters of the random forest model are tuned based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, thereby obtaining the best random forest model;

[0047] The prediction module predicts the state of the virtual machine in each time period based on the optimal random forest model and the collected monitoring indicators, determines whether there is a potential failure, and if so, issues an alarm.

[0048] Compared with the prior art, the present invention has the following beneficial effects:

[0049] The present invention establishes a 5G-R core network virtual machine failure prediction model through data collection, data preprocessing, data feature selection and privilege enhancement, model construction and evaluation, model application and feedback process. Based on the various indicator data collected by virtual machine prometheus, a random forest algorithm is used to establish a virtual machine failure prediction model to predict the probability of virtual machine failure, complete the migration and backup of key services in time, avoid the interruption of virtual machine services, and improve the reliability of virtual machines. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0051] Figure 1 It is a flow chart of a 5G-R core network virtual machine fault prediction method based on a random forest algorithm of the present invention;

[0052] Figure 2 It is a structural schematic diagram of a 5G-R core network virtual machine fault prediction system based on a random forest algorithm of the present invention;

[0053] Figure 3 This is a schematic diagram of the framework of the 5G-R core network virtual machine fault prediction method based on the random forest algorithm of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0055] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0056] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0057] In the description of the embodiments of the present invention, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. indicate an orientation or positional relationship based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of the invention is usually placed when in use, it is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. In addition, the terms "first", "second", etc. are only used to distinguish the description, and cannot be understood as indicating or implying relative importance.

[0058] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", which does not mean that the structure must be completely horizontal, but can be slightly tilted.

[0059] In the description of the embodiments of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0060] The present invention is further described in detail below in conjunction with the accompanying drawings:

[0061] See also Figure 1 The present invention discloses a 5G-R core network virtual machine fault prediction method based on a random forest algorithm, comprising:

[0062] S101, collecting monitoring data on the 5G-R core network virtual machine and preprocessing the collected monitoring data;

[0063] Collect monitoring data on the 5G-R core network virtual machine, specifically: deploy the Prometheus service on the 5G-R core network virtual machine, and regularly collect monitoring indicators from the 5G-R core network virtual machine through the Prometheus exporter module; the monitoring indicators include: timestamp, IP address, host name, running time, memory, number of CPU cores, 5-minute load, CPU utilization, memory utilization, disk partition utilization, disk read and write rate, network download rate, network upload rate and virtual machine status.

[0064] The collected monitoring data is preprocessed, specifically: the collected data is cleaned, filtered and feature extracted, including timestamp transformation, data standardization, missing value processing and data conversion.

[0065] The timestamp transformation is: converting the timestamp into a standard time format yyyy-MM-dd HH:mm:ss, and extracting the year, month, day, hour, minute, and second derived attributes;

[0066] The data standardization is to keep the dimensions of each attribute of running time, memory, disk read and write rate, network download rate, and network upload rate uniform and perform data conversion;

[0067] The unit of the running time is unified as seconds, the unit of the memory is unified as bytes, the unit of the disk read and write rate is unified as bytes per second, and the unit of the network download rate and the network upload rate is unified as kilobytes per second;

[0068] The missing value processing is for sample data with missing 5-minute load attributes, and the data is marked as NAN; the data is converted to quantify the virtual machine status attributes, 0 is marked as normal, and 1 is marked as abnormal.

[0069] S102, extracting features from the preprocessed monitoring data based on mutual information and chi-square test to obtain feature data that is strongly correlated with the state of the 5G-R core network virtual machine;

[0070] The 5-minute load, CPU usage, memory usage, disk partition usage, disk read / write rate, network download rate, and network upload rate were selected as feature variables, and the virtual machine status was selected as the target variable to calculate the mutual information and chi-square test;

[0071] According to the mutual information calculation results, the mutual information calculation results are sorted, and the front feature vector is selected as the feature variable with strong correlation with the target variable virtual machine state;

[0072] Based on the chi-square test calculation results, the relationship between the chi-square test calculation results and the preset thresholds is judged in turn to obtain the characteristic variables with strong correlation with the target variable virtual machine state;

[0073] The final characteristic variables are obtained by comprehensively considering the mutual information calculation results and the chi-square test calculation results.

[0074] The mutual information and chi-square test are calculated as follows:

[0075]

[0076] Among them, P(x,y) is the joint probability of variables X and Y taking the values ​​of x and y at the same time; P(x) is the marginal probability of variable X; P(y) is the marginal probability of variable Y;

[0077] The chi-square test formula is as follows:

[0078]

[0079] Where O is the actual observed frequency; E is the expected frequency; x 2 is the chi-square statistic, which is used to measure the deviation of the actual value from the expected value.

[0080] S103, constructing a random forest model, dividing the acquired feature data to obtain a training set and a test set;

[0081] Construct a random forest model, specifically:

[0082] Based on the Bootstrap sampling method, the sample data collected from Prometheus is sampled with replacement to generate multiple sub-sample sets;

[0083] A decision tree is constructed for each sub-sample set. At each splitting node, several features in the sub-sample set are randomly selected for splitting to obtain a decision tree.

[0084] Integrate the prediction results of multiple decision trees to achieve multi-tree voting to generate the final prediction; in classification tasks, the model generates prediction results through majority voting:

[0085] f(x) = majority_vote(h 1 (x),h 2 (x),...,h m (x)

[0086] Among them, h i(x) represents the prediction result of the i-th decision tree for input c, and m is the number of decision trees.

[0087] S104, training the random forest model based on the training set, and evaluating the performance of the model through the test set; tuning the hyperparameters of the random forest model based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, and then obtaining the optimal random forest model;

[0088] The cross entropy loss function is used to measure the difference between the model prediction value and the true label. The formula is as follows:

[0089]

[0090] Where N is the total number of samples; y i is the actual label of sample i, which is 0 or 1; z i is the probability that sample i is predicted to be a fault, and the output value is between [0, 1];

[0091] The training process is as follows:

[0092] Each decision tree in the random forest randomly selects a subset from the original feature set for splitting. j The splitting of the decision tree is optimized by calculating the Gini index or information gain of the feature; for each sample x in the training set i , each decision tree T j Output a predicted label h j (x i ); calculate the prediction error of each decision tree through the cross entropy loss function; vote on the prediction results of all decision trees, and the final prediction result is the category with the most votes; adjust the weight and splitting method of the decision tree by minimizing the loss function or cross entropy, and evaluate the performance of the model through the test set. If the performance of the random forest model is not good, adjust the parameters of the random forest to optimize the random forest model.

[0093] S105, based on the optimal random forest model and the collected monitoring indicators, the virtual machine status is predicted in each time period to determine whether there is a potential failure. If so, an alarm is issued.

[0094] Collect the current monitoring data of the system, select 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate as feature variables, input features through the splitting rule of each decision tree in turn, judge whether each decision tree has potential faults, and count the prediction results of each tree. If the percentage of decision trees with potential faults in the total decision trees exceeds the set threshold, the final prediction label is obtained.

[0095] See also Figure 2 The present invention discloses a 5G-R core network virtual machine fault prediction system based on a random forest algorithm, comprising:

[0096] A preprocessing module, wherein the preprocessing module collects monitoring data on the 5G-R core network virtual machine and preprocesses the collected monitoring data;

[0097] A feature extraction module, wherein the feature extraction module extracts features from the preprocessed monitoring data based on mutual information and chi-square test to obtain feature data that is strongly correlated with the state of the 5G-R core network virtual machine;

[0098] A construction module, wherein the construction module constructs a random forest model and divides the acquired feature data to obtain a training set and a test set;

[0099] A tuning module, wherein the tuning module trains the random forest model based on the training set and evaluates the performance of the model through the test set; the hyperparameters of the random forest model are tuned based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, thereby obtaining the best random forest model;

[0100] The prediction module predicts the state of the virtual machine in each time period based on the optimal random forest model and the collected monitoring indicators, determines whether there is a potential failure, and if so, issues an alarm.

[0101] Example:

[0102] See also Figure 3 The present invention proposes a 5G-R core network virtual machine fault prediction method based on a random forest algorithm, using key operating indicators of the 5G-R core network virtual machine (such as timestamp, IP address, host name, running time, memory, number of CPU cores, 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, download rate, upload rate and status), and constructing a fault prediction model through the random forest algorithm to monitor and predict the operating status of the core network virtual machine in real time.

[0103] Data collection: The 5G-R core network virtual machine deploys the Prometheus service. Through the exporter module of Prometheus, relevant monitoring indicators are regularly collected from the 5G-R core network virtual machine, including timestamp, IP address, host name, uptime, memory, number of CPU cores, 5-minute load, CPU usage (CPU used (%)), memory usage (Memory used (%)), disk partition usage (Partition used (%)), disk read and write rate (Disk Read and Disk Write), network download rate (Download), network upload rate (Upload) and virtual machine state (State). In the 5G-R core network environment, these indicators can reflect the resource utilization and operating status of the virtual machine. The following is the sample data collected:

[0104]

[0105] Data preprocessing: The data volume and monitoring frequency of 5G-R core network virtual machines are high, so it is necessary to clean, filter and extract features from the data collected by Prometheus. The data preprocessing stage can be used to ensure data quality and consistency, and provide reliable basic data for subsequent feature extraction and model training. According to the characteristics of the collected data, the main methods used include timestamp transformation, data standardization, missing value processing, and data conversion.

[0106] Data preprocessing: The data volume and monitoring frequency of 5G-R core network virtual machines are high, so it is necessary to clean, filter and extract features from the data collected by Prometheus. The data preprocessing stage can be used to ensure data quality and consistency, and provide reliable basic data for subsequent feature extraction and model training. According to the characteristics of the collected data, the main methods used include timestamp transformation, data standardization, missing value processing, and data conversion.

[0107] Timestamp conversion is to convert the timestamp (Timestamp) 1730697000 into the standard time format yyyy-MM-ddHH:mm:ss, and extract the year, month, day, hour, minute, and second derived attributes. The sample data is as follows: Based on the online timestamp conversion tool wiicha.com, the timestamp is converted.

[0108] {Timestamp=1730697000,

[0109] DateYmdHms=2024-11-04 13:10:00,

[0110] DateYmd=2024-11-04,

[0111] DateHms=13:10:00

[0112] }

[0113] Unified dimensions: Keep the dimensions of the attributes of uptime, memory, disk read and write rates (Disk Read and DiskWrite), network download rate (Download), and network upload rate (Upload) unified and perform data conversion. Uptime is unified into seconds (s), memory is unified into bytes (B), disk read and write rates (Disk Read and Disk Write) are unified into bytes per second (B / s), and network download rate (Download) and network upload rate (Upload) are unified into kilobytes per second (kbps).

[0114] Missing value processing: For sample data with missing 5-minute load (5m Load) attributes, the data is marked as NAN.

[0115] Data conversion: Quantify the virtual machine state attributes, with 0 marking normal and 1 marking exception.

[0116] Data feature selection and extraction: In terms of data feature selection, a combination of mutual information and chi-square test is used. Mutual information can capture the nonlinear relationship between features and target variables, while the chi-square test focuses on the independence between categorical variables and is suitable for processing categorical features.

[0117] Based on the operation and maintenance experience of 5G-R railway, we selected 5-minute load (5m Load), CPU usage (CPU used(%)), memory usage (Memory used(%)), disk partition usage (Partition used(%)), disk read and write rates (Disk Read and Disk Write), network download rate (Download), and network upload rate (Upload) as feature variables, and virtual machine state (State) as the target variable, and calculated the mutual information and chi-square test.

[0118] Mutual information is a quantitative method to measure the statistical dependence between two random variables. The mutual information formula is as follows:

[0119]

[0120] Among them, P(x,y) is the joint probability of variables X and Y taking the values ​​of x and y at the same time; P(x) is the marginal probability of variable X; P(y) is the marginal probability of variable Y;

[0121] Taking 1100 sample data collected by Prometheus as an example, the steps to calculate the mutual information value between CPU usage and fault occurrence are as follows:

[0122] The continuous CPU usage characteristic values ​​are discretized and divided into the following intervals: [0-30], (30-60], (60-100].

[0123] The joint probability between statistical discrete variables is shown in the following table:

[0124]

[0125]

[0126] The mutual information value of each combination item is calculated according to the mutual information calculation method.

[0127] Combination Item Joint probability P(x,y) Marginal probability P(x) Marginal probability P(y) Mutual Information Value [0-30], the fault is 0.1764 0.3409 0.8245 -0.0106 [0-30], Fault No 0.1636 0.3409 0.1755 0.159 (30-60], the fault is 0.3482 0.3564 0.8245 0.0646 (30-60], fault 0.0082 0.3564 0.1755 -0.0112 (60-100], the fault is 0.3 0.3036 0.8245 0.0565 (60-100], fault 0.0036 0.3036 0.1755 -0.0088

[0128] The final mutual information value accumulated for each combination item is 0.2495.

[0129] The chi-square test is used to detect the independence between categorical variables. The chi-square test formula is as follows:

[0130]

[0131] Where O is the actual observed frequency; E is the expected frequency; x 2 is the chi-square statistic, which is used to measure the deviation of the actual value from the expected value.

[0132] Taking the above sample data as an example, the chi-square check value between CPU usage and fault occurrence is calculated. The mutual information value of each combination item is as follows:

[0133] Observation frequency (O) Expected frequency (E) <![CDATA[(O―E) 2 / E]]> [0-30], the fault is 294 0.378 [0-30], Fault No 180 0.680 (30-60], the fault is 223 6.17 (30-60], fault 69 11.29 (60-100], the fault is 200 15.25 (60-100], fault 34 28.23

[0134] The final chi-square test value accumulated for each combination item is 61.02.

[0135] The following is the calculation result based on 1100 sample data collected by Prometheus:

[0136] Feature variables Mutual Information Chi-square value p-value 5 minutes load 0.1398 11.09 0.281 CPU usage (%) 0.2495 61.02 0.001 Memory usage (%) 0.2862 63.03 0.001 Disk partition usage (%) 0.1087 12.34 0.284 Disk read rate 0.2465 62.50 0.003 Disk write rate 0.2550 61.23 0.002 Network download speed 0.0065 41.10 0.290 Network upload rate 0.0079 39.50 0.450

[0137] The mutual information results show that the CPU usage (%), memory usage (%), disk read / write rate and the target variable virtual machine status have the strongest relationship. The chi-square test results show that the p-values ​​of CPU usage (%) and memory usage (%) are less than 0.05, indicating that they have a significant relationship with the virtual machine status. Although the p-value of disk read / write rate is slightly higher, it is still close to the significance level. The features finally selected are CPU usage (%), memory usage (%) and disk read / write rate, which are the features most related to the target variable virtual machine status.

[0138] Model construction: The random forest algorithm is suitable for processing complex, high-dimensional data. It achieves classification by integrating multiple decision trees, and can capture multiple characteristic patterns in the 5G-R core network virtual machine data to improve the accuracy of fault prediction. The following are the steps to build a fault prediction model based on the random forest algorithm:

[0139] Generate sub-sample sets: Based on the Bootstrap sampling method, the sample data collected from Prometheus is sampled with replacement to generate multiple sub-sample sets. The size of these sub-sample sets is generally the same as the original data set, but because it is sampled with replacement, each sample set may contain repeated data samples.

[0140] Construct a decision tree: construct a decision tree for each sub-sample set; at each splitting node, randomly select several features in the sub-sample set for splitting to obtain a decision tree; this method enables each decision tree to be generated under different feature combinations, thereby enhancing the robustness of the model.

[0141] Decision tree ensemble: Integrate the prediction results of multiple decision trees and implement multi-tree voting to generate the final prediction. In classification tasks, the model generates prediction results through majority voting:

[0142] f(x) = majority_vote(h 1 (x),h 2 (x),...,h m (x))

[0143] Among them, h i (x) represents the prediction result of the i-th decision tree for input x, and m is the number of decision trees.

[0144] The random forest algorithm can process high-dimensional data and has good generalization ability and interpretability, so the random forest algorithm is selected as the prediction model for construction.

[0145] Optimization and hyperparameter tuning: The model's hyperparameters are tuned through cross-validation and grid search to find the best combination of parameters such as the number of trees (n_estimators) and the maximum number of features (max_features). This can significantly improve the performance of the model, especially the sensitivity in detecting small probability failure events.

[0146] Model evaluation: After model training is completed, the model performance is evaluated using the following evaluation indicators to ensure that the model can provide high-quality predictions in actual deployment: The model evaluation method uses accuracy, precision, recall, and F1Score scores as evaluation indicators.

[0147] TP: True Positive — The number of samples predicted to be positive and actually positive.

[0148] TN: True Negative – The number of samples predicted to be negative and actually negative.

[0149] FP: False Positive – The number of samples predicted to be positive but actually negative.

[0150] FN: False Negative – The number of samples predicted to be negative but actually positive.

[0151] Accuracy: reflects the accuracy of the model's prediction of the overall sample.

[0152]

[0153] Precision refers to the proportion of actual failures in cases where the model predicts failures, and is used to measure the accuracy of the prediction.

[0154]

[0155] Recall rate: It indicates the proportion of faults predicted by the model when the fault actually occurs, which can reflect the model's ability to capture fault events.

[0156]

[0157] F1 score: The harmonic mean of precision and recall, used as a comprehensive evaluation indicator:

[0158]

[0159] In the evaluation, special attention is paid to the recall and F1 score because it is more important for a fault prediction system to capture as many fault events as possible rather than false positives.

[0160] Based on the sample data, the optimal parameters are obtained through the grid search method, and the model is retrained and evaluated using the optimal parameters to improve the performance of the model. The optimal grid parameters and model evaluation results are as follows:

[0161]

[0162] Model application and fault prediction: This invention is mainly used for virtual machine fault prediction in 5G-R core network, helping operation and maintenance personnel to take necessary measures before faults occur through real-time monitoring and early warning.

[0163] Determine whether there is a potential fault, specifically: collect the current monitoring data of the system, select 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate as feature variables, input features through the splitting rules of each decision tree in turn, determine whether each decision tree has a potential fault, and count the prediction results of each tree. If the percentage of decision trees with potential faults in the total decision trees exceeds the set threshold, the final prediction label is obtained.

[0164] For example, the first tree splitting rule is: CPU usage > 80% → memory usage > 90% → fault prediction is 1, then it is considered that there is a potential fault;

[0165] The second tree splitting rule is: disk read / write rate > 100MB / s → disk partition usage > 70% → fault prediction is 0, then it is considered that there is no potential fault;

[0166] The prediction results of each tree are counted, and if the percentage of decision trees with potential faults exceeds the set threshold, the final prediction label is obtained. For example: if 60% of the trees predict fault 1, the final prediction is likely to be a fault.

[0167] In the 5G-R core network application scenario, this method can achieve the following functions:

[0168] Real-time prediction: The model predicts the state of the virtual machine in each time period based on the monitoring indicators collected by Prometheus. Once the model detects signs of potential failure, it can trigger an early warning signal.

[0169] Early warning mechanism: The system sends early warning information to the operation and maintenance team through the alarm interface to prompt potential virtual machine failures so that the operation and maintenance personnel can handle them in advance. The early warning information includes key information such as the virtual machine's IP, host name, failure characteristics, and timestamp to help quickly locate the problem.

[0170] Intelligent operation and maintenance: This fault prediction method can be integrated into the operation and maintenance system of the 5G-R core network to achieve visual monitoring of the operating status of virtual machines, facilitating the analysis and optimization of system resource allocation. For example, when the load is too high or resource usage is close to the threshold, the operation and maintenance system can automatically trigger resource reallocation or migration to ensure the continuous stability of network services.

[0171] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A 5G-R core network virtual machine fault prediction method based on a random forest algorithm, characterized in that: include: Collect monitoring data on the 5G-R core network virtual machine and pre-process the collected monitoring data; Based on mutual information and chi-square test, feature extraction is performed on the pre-processed monitoring data to obtain feature data that is strongly correlated with the status of the 5G-R core network virtual machine; Constructing a random forest model, dividing the acquired feature data into a training set and a test set; The random forest model is trained based on the training set, and the performance of the model is evaluated through the test set. The hyperparameters of the random forest model are tuned based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, and then the optimal random forest model is obtained. Based on the optimal random forest model and the collected monitoring indicators, the virtual machine status is predicted in each time period to determine whether there is a potential failure. If so, an alarm is issued.

2. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 1 is characterized in that: The collecting of monitoring data on the 5G-R core network virtual machine specifically comprises: deploying the Prometheus service on the 5G-R core network virtual machine, and regularly collecting monitoring indicators from the 5G-R core network virtual machine through the exporter module of Prometheus; the monitoring indicators include: timestamp, IP address, host name, running time, memory, number of CPU cores, 5-minute load, CPU utilization, memory utilization, disk partition utilization, disk read and write rate, network download rate, network upload rate and virtual machine status.

3. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 2 is characterized in that: The preprocessing of the collected monitoring data is specifically: cleaning, filtering and feature extraction of the collected data, including timestamp transformation, data standardization, missing value processing and data conversion; The timestamp transformation is to convert the timestamp into a standard time format yyyy-MM-dd HH:mm:ss, and extract the year, month, day, hour, minute, and second derived attributes; The data standardization is to keep the dimensions of each attribute of running time, memory, disk read and write rate, network download rate, and network upload rate uniform and perform data conversion; The unit of the running time is unified as seconds, the unit of the memory is unified as bytes, the unit of the disk read and write rate is unified as bytes per second, and the unit of the network download rate and the network upload rate is unified as kilobytes per second; The missing value processing is for sample data with missing 5-minute load attributes, and the data is marked as NAN; the data is converted to quantify the virtual machine status attributes, 0 is marked as normal, and 1 is marked as abnormal.

4. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 3 is characterized in that: The obtaining of feature data strongly related to the state of the 5G-R core network virtual machine is specifically as follows: selecting 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate as feature variables, and the virtual machine state as the target variable, and calculating mutual information and chi-square test; According to the mutual information calculation results, the mutual information calculation results are sorted, and the front feature vector is selected as the feature variable with strong correlation with the target variable virtual machine state; Based on the chi-square test calculation results, the relationship between the chi-square test calculation results and the preset thresholds is judged in turn to obtain the characteristic variables with strong correlation with the target variable virtual machine state; The final characteristic variables are obtained by comprehensively considering the mutual information calculation results and the chi-square test calculation results.

5. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 4 is characterized in that: The calculation of mutual information and chi-square test is specifically as follows: The mutual information formula is as follows: Among them, P(x,y) is the joint probability of variables X and Y taking the values ​​of x and y at the same time; P(x) is the marginal probability of variable X; P(y) is the marginal probability of variable Y; The chi-square test formula is as follows: Where O is the actual observed frequency; E is the expected frequency; x 2 is the chi-square statistic, which is used to measure the deviation of the actual value from the expected value.

6. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 5 is characterized in that: The random forest model is constructed as follows: Based on the Bootstrap sampling method, the sample data collected from Prometheus is sampled with replacement to generate multiple sub-sample sets; A decision tree is constructed for each sub-sample set. At each splitting node, several features in the sub-sample set are randomly selected for splitting to obtain a decision tree. Integrate the prediction results of multiple decision trees to achieve multi-tree voting to generate the final prediction; in classification tasks, the model generates prediction results through majority voting: f(x)=majority_vote(h1(x),h2(x),...,h m (x)) Among them, h i (x) represents the prediction result of the i-th decision tree for input x, and m is the number of decision trees.

7. The 5G-R core network virtual machine fault prediction method based on random forest according to claim 6 is characterized in that: The random forest model is trained based on the training set, and the performance of the model is evaluated through the test set, specifically: The cross entropy loss function is used to measure the difference between the model prediction value and the true label. The formula is as follows: Where N is the total number of samples; y i is the actual label of sample i, which is 0 or 1; z i is the probability that sample i is predicted to be a fault, and the output value is between [0, 1]; The training process is as follows: Each decision tree in the random forest randomly selects a subset from the original feature set for splitting. j The splitting of the decision tree is optimized by calculating the Gini index or information gain of the feature; for each sample x in the training set i , each decision tree T j Output a predicted label h j (x i ); calculate the prediction error of each decision tree through the cross entropy loss function; vote on the prediction results of all decision trees, and the final prediction result is the category with the most votes; adjust the weight and splitting method of the decision tree by minimizing the loss function or cross entropy, and evaluate the performance of the model through the test set. If the performance of the random forest model is not good, adjust the parameters of the random forest to optimize the random forest model.

8. The 5G-R core network virtual machine fault prediction method based on random forest algorithm according to claim 7 is characterized in that: The determination of whether there is a potential fault is specifically as follows: the current monitoring data of the system is collected, 5-minute load, CPU usage, memory usage, disk partition usage, disk read and write rate, network download rate and network upload rate are selected as feature variables, the input features are sequentially passed through the splitting rule of each decision tree, and it is determined whether each decision tree has a potential fault, and the prediction results of each tree are counted. If the percentage of decision trees with potential faults in the total decision trees exceeds the set threshold, the final prediction label is obtained.

9. A 5G-R core network virtual machine fault prediction system based on a random forest algorithm, characterized in that: include: A preprocessing module, wherein the preprocessing module collects monitoring data on the 5G-R core network virtual machine and preprocesses the collected monitoring data; A feature extraction module, wherein the feature extraction module extracts features from the preprocessed monitoring data based on mutual information and chi-square test to obtain feature data that is strongly correlated with the state of the 5G-R core network virtual machine; A construction module, wherein the construction module constructs a random forest model and divides the acquired feature data to obtain a training set and a test set; A tuning module, wherein the tuning module trains the random forest model based on the training set and evaluates the performance of the model through the test set; the hyperparameters of the random forest model are tuned based on cross-validation and grid search to obtain the best hyperparameters of the random forest model, thereby obtaining the best random forest model; The prediction module predicts the state of the virtual machine in each time period based on the optimal random forest model and the collected monitoring indicators, determines whether there is a potential failure, and if so, issues an alarm.

Citation Information

Patent Citations

  • Turnout fault diagnosis method based on random forest

    CN111046931A

  • Multi-factor cloud service storage device error prediction

    CN112771504A

  • Transmission monitoring system for relay protection overhaul test of intelligent substation

    CN118171195A

  • Cloud computing virtual machine communication optimization method

    CN118784485A

  • Tunnel unfavorable geology identification method and system based on Bayesian optimization random forest

    CN119004191A

Cited By

  • Ground wire state monitoring and early warning system and method based on random forest algorithm

    CN121461601A