Model Training Method, Apparatus, and System
By conducting incremental training on local point analysis equipment and offline training, using local point feature data to customize machine learning models, the model adaptation problem is solved, and efficient model adaptation and the effect of reducing training costs is achieved.
Patent Information
- Application Number
- CN201910878280.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-09-17
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2039-09-17
AI Technical Summary
In the prior art, when the machine learning model is directly deployed on the local point analysis device after cloud training, it cannot effectively adapt to the needs of the local point equipment, and the training cost is high, the model is low in versatility, and generalization cannot be achieved.
The feature data obtained by the local point analysis device is used for incremental training, combined with offline training, the model is customized to meet the needs of the local point equipment, and the model structure is optimized through node splitting to reduce computing resource occupation and storage space.
It realizes high adaptability and flexibility of machine learning models on local point devices, reduces training costs, and improves the generalization ability and prediction efficiency of the model.
Smart Images

Figure CN112529204B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Artificial Intelligence (AI), and particularly to a model training method, apparatus, and system. Background Art
[0002] Machine learning refers to enabling a machine to train a machine learning model based on training samples, so that the machine learning model has the ability to predict the category of samples other than the training samples.
[0003] Currently, a data analysis system includes multiple analysis devices for data analysis. The multiple analysis devices may include cloud analysis devices and local analysis devices. The deployment method of a machine learning model in this system includes: performing offline training of the model by a cloud analysis device, and then directly deploying the offline-trained model on the local analysis device.
[0004] However, the trained model may not be able to effectively adapt to the requirements of the local analysis device. Summary of the Invention
[0005] Embodiments of this application provide a model training method, apparatus, and system. The technical solutions are as follows:
[0006] In a first aspect, a model training method is provided, which is applied to a local analysis device and includes:
[0007] Receiving a machine learning model sent by a first analysis device. Optionally, the first analysis device is a cloud analysis device;
[0008] Performing incremental training on the machine learning model based on a first training sample set, where the feature data in the first training sample set is the feature data of the local network corresponding to the local analysis device.
[0009] On the one hand, the feature data in the first training sample set is the feature data obtained from the local network corresponding to the local analysis device, which is more suitable for the application scenario of the local analysis device. Using the first training sample set including the feature data obtained from the corresponding local network by the local analysis device for model training can make the trained machine learning model more suitable for the requirements of the local analysis device itself, realize the customization of the model, and improve the application flexibility of the model. On the other hand, by combining offline training and incremental training to train the machine learning model, when the category or pattern of the feature data obtained by the local analysis device changes, incremental training of the machine learning model can be performed to realize the flexible adjustment of the machine learning model, so as to ensure that the trained machine learning model meets the requirements of the local analysis device. Therefore, the model training method provided by the embodiments of this application can effectively adapt to the requirements of the local analysis device compared with the related technologies.
[0010] Optionally, after receiving the machine learning model sent by the first analysis device, the method further includes:
[0011] Using the machine learning model to predict classification results;
[0012] Optionally, the method further includes: sending prediction information to the evaluation device, where the prediction information includes the predicted classification results, for the evaluation device to evaluate whether the machine learning model has deteriorated based on the prediction information. In one example, the local point analysis device can send prediction information including the predicted classification results to the evaluation device after each prediction of classification results using the machine learning model; in another example, the local point analysis device can also periodically send prediction information including the classification results obtained in the current period to the evaluation device; in yet another example, the local point analysis device can also send prediction information including the obtained classification results to the evaluation device after the number of obtained classification results reaches a threshold; in still another example, the local point analysis device can also send prediction information including the currently obtained classification results to the evaluation device within a set time period, so as to avoid interfering with user services.
[0013] The incrementally training the machine learning model based on the first training sample set includes:
[0014] After receiving the training instruction sent by the evaluation device, incrementally training the machine learning model based on the first training sample set, where the training instruction is used to indicate training the machine learning model.
[0015] Optionally, the machine learning model is used to predict classification results for prediction data composed of one or more key performance indicator (KPI) feature data; the KPI feature data is feature data of a KPI time series or KPI data;
[0016] The prediction information further includes the KPI category corresponding to the KPI feature data in the prediction data, the identifier of the device to which the prediction data belongs, and the acquisition time of the KPI data corresponding to the prediction data.
[0017] Optionally, the method further includes:
[0018] When the performance of the incrementally trained machine learning model does not meet the performance compliance condition, sending a retraining request to the first analysis device, where the retraining request is used to request the first analysis device to retrain the machine learning model.
[0019] Optionally, the machine learning model is a tree model. The incremental training of the machine learning model based on the first training sample set includes:
[0020] For any training sample in the first training sample set, start traversing from the root node of the machine learning model and perform the following traversal process:
[0021] When the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, add an associated second node. The first node is any non-leaf node in the machine learning model, and the second node is the parent node or child node of the first node;
[0022] When the current splitting cost of the first node is not less than the historical splitting cost of the first node, traverse the nodes in the subtree of the first node and determine the traversed node as the new first node, and perform the traversal process again until the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, or the target depth is traversed;
[0023] Wherein, the current splitting cost of the first node is the cost of splitting the node based on the first training sample. The first training sample is any training sample in the first training sample set. The first training sample includes feature data in one or more feature dimensions, and the feature data is numerical data. The historical splitting cost of the first node is the cost of splitting the node based on the historical training sample set of the first node. The historical training sample set of the first node is the set of samples divided to the first node in the historical training sample set of the machine learning model.
[0024] Optionally, the current splitting cost of the first node is negatively correlated with the size of the first numerical distribution range. The first numerical distribution range is a distribution range determined based on the feature values in the first training sample and the second numerical distribution range; the second numerical distribution range is the distribution range of the feature values in the historical training sample set of the first node, and the historical splitting cost of the first node is negatively correlated with the size of the second numerical distribution range.
[0025] Optionally, the current splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the first numerical distribution range, and the historical splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the second numerical distribution range.
[0026] During the incremental training process, node splitting is performed based on the numerical distribution range of the training sample set, without the need to access a large number of historical training samples. Therefore, the memory and computing resources occupied are effectively reduced, and the training cost is lowered. Moreover, by carrying the relevant information of each node through the foregoing node information, the lightweight of the machine learning model can be achieved, which is more convenient for the deployment of the machine learning model and realizes the effective generalization of the model.
[0027] Optionally, adding the associated second node includes:
[0028] Determine the span range of the feature values of the first numerical distribution range on each feature dimension;
[0029] Add the second node based on the first splitting point on the first splitting dimension, where the numerical range in the first numerical distribution range that is not greater than the value of the first splitting point on the first splitting dimension is divided into the left child node of the second node, and the numerical range in the first numerical distribution range that is greater than the value of the first splitting point on the first splitting dimension is divided into the right child node of the second node. The first splitting dimension is the splitting dimension determined among the various feature dimensions based on the span range of the feature values on the various feature dimensions, and the first splitting point is the numerical point determined for splitting on the first splitting dimension of the first numerical distribution range;
[0030] When the first splitting dimension is different from the second splitting dimension, the second node is the parent node or child node of the first node, the second splitting dimension is the historical splitting dimension of the first node in the machine learning model, and the second splitting point is the historical splitting point of the first node in the machine learning model;
[0031] When the first splitting dimension is the same as the second splitting dimension and the first splitting point is located on the right side of the second splitting point, the second node is the parent node of the first node, and the first node is the left child node of the second node;
[0032] When the first splitting dimension is the same as the second splitting dimension and the first splitting point is located on the left side of the second splitting point, the second node is the left child node of the first node.
[0033] Optionally, the first splitting dimension is a feature dimension randomly selected from the various feature dimensions of the first numerical distribution range, or the first splitting dimension is the feature dimension with the largest span among the various feature dimensions of the first numerical distribution range;
[0034] And / or, the first splitting point is a numerical point randomly selected on the first splitting dimension of the first numerical distribution range.
[0035] Optionally, adding the associated second node includes:
[0036] When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is greater than the first sample number threshold, add the second node;
[0037] The method further includes:
[0038] When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is not greater than the first sample number threshold, stop the incremental training of the machine learning model.
[0039] Optionally, the method further includes:
[0040] Merge the first non-leaf node and the second non-leaf node in the machine learning model, and merge the first leaf node and the second leaf node to obtain a streamlined machine learning model, which is used to predict the classification result;
[0041] Alternatively, receive the streamlined machine learning model sent by the first analysis device, where the streamlined machine learning model is obtained by the first analysis device merging the first non-leaf node and the second non-leaf node, and merging the first leaf node and the second leaf node in the machine learning model;
[0042] Wherein, the first leaf node is the child node of the first non-leaf node, the second leaf node is the child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets assigned on the same feature dimension are adjacent.
[0043] The streamlined machine learning model has a simpler structure, reduces the branch levels of the tree, and prevents the tree from being too deep. Although the model architecture changes, it does not affect its prediction result, can save storage space, and improve prediction efficiency. And overfitting of the model can be prevented through the streamlining process. Further, if the streamlined model is only used for sample analysis, historical split information may not be recorded in its node information, which can further reduce the size of the model itself and improve the model prediction efficiency.
[0044] Optionally, each node in the machine learning model stores corresponding node information. The node information of any node in the machine learning model includes label distribution information, which is used to reflect the proportion of labels of different categories of samples in the total number of labels in the historical training sample set divided into the corresponding node. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, which is the information used for splitting the corresponding node.
[0045] Optionally, the historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical splitting cost;
[0046] The label distribution information includes: the number of labels of the same category of samples in the historical training sample set divided into the corresponding node and the total number of labels; or, the proportion of labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels.
[0047] Optionally, the first training sample set includes samples that meet the low discrimination condition and are screened from the samples obtained by the game point analysis device. The low discrimination condition includes at least one of the following:
[0048] The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, where 1 < n < m, and m is the total number of probabilities obtained by predicting the sample using the machine learning model;
[0049] Or, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than the second difference threshold;
[0050] Or, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the third difference threshold;
[0051] Or, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the fourth difference threshold;
[0052] Or, the probability distribution entropy E of multiple classification results obtained by predicting the sample using the machine learning model is greater than the specified distribution entropy threshold, and the E satisfies:
[0053]
[0054] where, xi represents the i-th classification result, P(x i ) represents the probability of the i-th classification result of the predicted sample, b is the specified base, and 0 ≤ P(x i ) ≤ 1.
[0055] In a second aspect, a model training method is provided, which is applied to the first analysis device. For example, the first analysis device can be a cloud analysis device. The method includes:
[0056] Performing offline training based on a historical training sample set to obtain a machine learning model;
[0057] Sending the machine learning model to multiple local analysis devices for the local analysis devices to perform incremental training on the machine learning model based on a first training sample set. The feature data in the training sample set used by any local analysis device to train the machine learning model is the feature data of the local network corresponding to the any local analysis device.
[0058] In the embodiments of the present application, the first analysis device can distribute the trained machine learning model to each local analysis device for incremental training by each local analysis device, ensuring the performance of the machine learning models on each local analysis device. In this way, the first analysis device does not need to train a corresponding machine learning model for each local analysis device, effectively reducing the overall training duration of the first analysis device. Moreover, the model obtained by offline training can be used as the basis for incremental training on each local analysis device, improving the generality of the model obtained by offline training, thereby realizing model generalization and reducing the overall training cost of the first analysis device.
[0059] Optionally, the historical training sample set is a set of training samples sent by multiple local analysis devices.
[0060] Optionally, after sending the machine learning model to the local analysis device, the method further includes:
[0061] Receiving a retraining request sent by the local analysis device, and retraining the machine learning model based on the training sample set sent by the local analysis device that sends the retraining request;
[0062] Or, receiving a retraining request sent by the local analysis device, and retraining the machine learning model based on the training sample set sent by the local analysis device that sends the retraining request and the training sample sets sent by other local analysis devices;
[0063] Or, receiving the training sample sets sent by at least two local analysis devices, and retraining the machine learning model based on the received training sample sets.
[0064] Optionally, the machine learning model is a tree model, and the machine learning model is obtained by performing offline training based on a historical training sample set, including:
[0065] Obtain a historical training sample set with determined labels, where the training samples in the historical training sample set include feature data of one or more feature dimensions, and the feature data is numerical data;
[0066] Create a root node;
[0067] Use the root node as the third node, and perform the offline training process until the splitting cut-off condition is reached;
[0068] Determine a classification result for each leaf node to obtain the machine learning model;
[0069] Wherein, the offline training process includes:
[0070] Perform splitting of the third node to obtain a left child node and a right child node of the third node;
[0071] Use the left child node as the updated third node, divide the historical training sample set into the left sample set of the left child node as the updated historical training sample set, and perform the offline training process again;
[0072] Use the right child node as the updated third node, divide the historical training sample set into the right sample set of the right child node as the updated historical training sample set, and perform the offline training process again.
[0073] In an optional manner, the performing splitting of the third node to obtain a left child node and a right child node of the third node includes:
[0074] Perform splitting of the third node based on the numerical distribution range of the historical training sample set to obtain a left child node and a right child node of the third node, where the numerical distribution range of the historical training sample set is the distribution range of the feature values in the historical training sample set.
[0075] Optionally, the performing splitting of the third node based on the numerical distribution range of the historical training sample set to obtain a left child node and a right child node of the third node includes:
[0076] Determine a third splitting dimension among the feature dimensions of the historical training sample set;
[0077] Determine a third splitting point on the third splitting dimension of the historical training sample set;
[0078] Divide the numerical range in the third numerical distribution range where the value in the third splitting dimension is not greater than the value of the third splitting point to the left child node, and divide the numerical range in the third numerical distribution range where the value in the third splitting dimension is greater than the value of the third splitting point to the right child node. The third numerical distribution range is the distribution range of the feature values in the historical training sample set of the third node.
[0079] During the offline training process, node splitting is performed based on the numerical distribution range of the training sample set, without accessing a large number of historical training samples. Therefore, the occupation of memory and computing resources is effectively reduced, and the training cost is lowered. And by carrying the relevant information of each node through the foregoing node information, the lightweight of the machine learning model can be realized, which is more convenient for the deployment of the machine learning model and achieves the effective generalization of the model.
[0080] In another alternative way, since the first analysis device has obtained the foregoing historical training sample set for training, it is also possible to directly use the samples for node splitting. The foregoing splitting of the third node to obtain the left and right child nodes of the third node can be replaced by: dividing the samples in the historical training sample set where the feature value in the third splitting dimension is not greater than the value of the third splitting point to the left child node, and dividing the samples in the historical training sample set where the feature value in the third splitting dimension is greater than the value of the third splitting point to the right child node.
[0081] Optionally, the splitting termination condition includes at least one of the following:
[0082] The current splitting cost of the third node is greater than the splitting cost threshold, so as to avoid excessive splitting of the tree and reduce the operation overhead;
[0083] Or, the number of samples in the historical training sample set is less than the second sample number threshold. This situation indicates that the data volume of the historical training sample set is already small and is no longer sufficient to support effective node splitting. At this time, stop the offline training process to reduce the operation overhead;
[0084] Or, the number of splitting times corresponding to the third node is greater than the splitting times threshold. This situation indicates that the current machine learning model has reached the upper limit of the splitting times. At this time, stop the offline training process to reduce the operation overhead;
[0085] Or, the depth of the third node in the machine learning model is greater than the depth threshold, so as to realize the control of the depth of the machine learning model;
[0086] Alternatively, the proportion of the number of the label with the largest proportion among the labels corresponding to the historical training sample set in the total number of labels corresponding to the historical training sample set is greater than a specified proportion threshold. This situation indicates that the number of the label with the largest proportion has reached the classification condition, and an accurate classification result can be determined based on this. At this time, the offline training process is stopped, which can reduce unnecessary splitting and reduce the computational overhead.
[0087] Optionally, the current splitting cost of the third node is negatively correlated with the size of the distribution range of the feature values in the historical training sample set.
[0088] Optionally, the current splitting cost of the third node is the reciprocal of the sum of the spans of the feature values of the historical training sample set in each feature dimension.
[0089] Optionally, the method further includes:
[0090] merging the first non-leaf node and the second non-leaf node in the machine learning model, and merging the first leaf node and the second leaf node to obtain a refined machine learning model for predicting the classification result, where the first leaf node is the child node of the first non-leaf node, the second leaf node is the child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets allocated on the same feature dimension are adjacent;
[0091] sending the refined machine learning model to the local point analysis device for the local point analysis device to predict the classification result based on the refined machine learning model.
[0092] Optionally, each node in the machine learning model stores node information correspondingly, and the node information of any node in the machine learning model includes label distribution information, which is used to reflect the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels, where the total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, which is the information used for splitting the corresponding node.
[0093] Optionally, the historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical splitting cost;
[0094] The label distribution information includes: the number of labels of the same category in the samples in the historical training sample set divided into the corresponding node and the total number of the labels; or, the proportion of the labels of different categories in the samples in the historical training sample set divided into the corresponding node in the total number of the labels.
[0095] Optionally, the first training sample set includes samples that meet the low discrimination condition selected from the samples obtained by the in-game point analysis device, and the low discrimination condition includes at least one of the following:
[0096] The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than the first difference threshold, and the target probability set includes the probabilities of the top n classification results arranged in descending order of probability, 1 < n < m, where m is the total number of probabilities obtained by predicting the sample using the machine learning model;
[0097] Or, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than the second difference threshold;
[0098] Or, the absolute value of the difference between the highest probability and the lowest probability in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the third difference threshold;
[0099] Or, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the fourth difference threshold;
[0100] Or, the probability distribution entropy E of multiple classification results obtained by predicting the sample using the machine learning model is greater than the specified distribution entropy threshold, and the E satisfies:
[0101]
[0102] where x i represents the i-th classification result, P(x i ) represents the probability of the i-th classification result of the predicted sample, b is the specified base number, and 0 ≤ P(x i ) ≤ 1.
[0103] In a third aspect, a model training device is provided. The device includes: a plurality of functional modules: the plurality of functional modules interact with each other to implement the methods in the first aspect and its various embodiments above. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be arbitrarily combined or divided based on specific implementations.
[0104] Fourthly, a model training device is provided. The device includes a plurality of functional modules which interact with each other to implement the methods in the second aspect and its various embodiments. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be arbitrarily combined or divided based on specific implementations.
[0105] Fifthly, a model training device is provided, including a processor and a memory.
[0106] The memory is used to store a computer program which includes program instructions.
[0107] The processor is used to call the computer program to implement the model training method as described in any one of the first aspect; or to implement the model training method as described in any one of the second aspect.
[0108] Sixthly, a computer storage medium is provided. Instructions are stored on the computer storage medium. When the instructions are executed by a processor, the model training method as described in any one of the first aspect is implemented; or the model training method as described in any one of the second aspect is implemented.
[0109] Seventhly, a chip is provided. The chip includes programmable logic circuits and / or program instructions. When the chip runs, the model training method as described in any one of the first aspect is implemented; or the model training method as described in any one of the second aspect is implemented.
[0110] Eighthly, a computer program product is provided. Instructions are stored in the computer program product. When the instructions run on a computer, the computer is caused to execute the model training method as described in any one of the first aspect; or the computer is caused to execute the model training method as described in any one of the second aspect.
[0111] The beneficial effects brought by the technical solutions provided in the embodiments of the present application are:
[0112] The model training method provided by the embodiments of the present application is such that the local point analysis device receives the machine learning model sent by the first analysis device, and can perform incremental training on the machine learning model based on the first training sample set obtained from the local point network corresponding to the local point analysis device. On the one hand, the feature data in the first training sample set is the feature data obtained from the local point network corresponding to the local point analysis device, which is more adapted to the application scenario of the local point analysis device. Using the first training sample set including the feature data obtained by the local point analysis device from the corresponding local point network for model training can make the trained machine learning model more adapted to the own needs of the local point analysis device, realize the customization of the model, and improve the application flexibility of the model. On the other hand, by combining offline training and incremental training to train the machine learning model, when the category or pattern of the feature data obtained by the local point analysis device changes, incremental training of the machine learning model can be carried out to realize the flexible adjustment of the machine learning model, so as to ensure that the trained machine learning model meets the needs of the local point analysis device. Therefore, the model training method provided by the embodiments of the present application can effectively adapt to the needs of the local point analysis device compared with the related technologies.
[0113] Further, the first analysis device can distribute the trained machine learning model to each local point analysis device for incremental training to ensure the performance of the machine learning model on each local point analysis device. In this way, the first analysis device does not need to train the corresponding machine learning model for each local point analysis device, effectively reducing the overall training duration of the first analysis device. Moreover, the model obtained by offline training can be used as the basis for incremental training on each local point analysis device, improving the generality of the model obtained by offline training, thus realizing model generalization and reducing the overall training cost of the first analysis device.
[0114] The structure of the simplified machine learning model is simpler, reducing the number of branches of the tree and preventing the tree from being too deep. Although the model architecture changes, it does not affect its prediction results, can save storage space, and improve prediction efficiency. And through the simplification process, overfitting of the model can be prevented. Further, if the simplified model is only used for sample analysis, the historical split information can not be recorded in its node information, which can further reduce the size of the model itself and improve the model prediction efficiency.
[0115] In the embodiments of the present application, during incremental training or offline training, node splitting is performed based on the numerical distribution range of the training sample set, without a large number of accesses to historical training samples. Therefore, the occupation of memory and computing resources is effectively reduced, and the training cost is reduced. And by carrying the relevant information of each node through the aforementioned node information, the lightweight of the machine learning model can be realized, which is more convenient for the deployment of the machine learning model and realizes the effective generalization of the model. Description of the Drawings
[0116] Figure 1 It is a schematic diagram of an application scenario involved in the model training method provided by an embodiment of the present application;
[0117] Figure 2 It is another schematic diagram of an application scenario involved in the model training method provided by an embodiment of the present application;
[0118] Figure 3 It is yet another schematic diagram of an application scenario involved in the model training method provided by an embodiment of the present application;
[0119] Figure 4 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0120] Figure 5 It is a flowchart of a method for controlling a key point analysis device to perform incremental training on a machine learning model based on an evaluation result of a classification result provided by an embodiment of the present application;
[0121] Figure 6 It is a schematic diagram of a tree structure provided by an embodiment of the present application;
[0122] Figure 7 It is a schematic diagram of the splitting principle of a tree model provided by an embodiment of the present application;
[0123] Figure 8 It is another schematic diagram of the splitting principle of a tree model provided by an embodiment of the present application;
[0124] Figure 9 It is yet another schematic diagram of the splitting principle of a tree model provided by an embodiment of the present application;
[0125] Figure 10 It is still another schematic diagram of the splitting principle of a tree model provided by an embodiment of the present application;
[0126] Figure 11 It is a schematic diagram of the splitting principle of a tree model provided by another embodiment of the present application;
[0127] Figure 12 It is another schematic diagram of the splitting principle of a tree model provided by another embodiment of the present application;
[0128] Figure 13 It is yet another schematic diagram of the splitting principle of a tree model provided by another embodiment of the present application;
[0129] Figure 14 It is still another schematic diagram of the splitting principle of a tree model provided by another embodiment of the present application;
[0130] Figure 15It is a schematic diagram of the splitting principle of a tree model provided by another embodiment of the present application;
[0131] Figure 16 It is a schematic diagram of the splitting principle of another tree model provided by another embodiment of the present application;
[0132] Figure 17 It is a schematic diagram of the splitting principle of yet another tree model provided by another embodiment of the present application;
[0133] Figure 18 It is a schematic diagram of the incremental training effect of a traditional machine learning model;
[0134] Figure 19 It is a schematic diagram of the incremental training effect of the machine learning model provided by the embodiment of the present application;
[0135] Figure 20 It is a schematic diagram of the structure of a model training device provided by the embodiment of the present application;
[0136] Figure 21 It is a schematic diagram of the structure of another model training device provided by the embodiment of the present application;
[0137] Figure 22 It is a schematic diagram of the structure of yet another model training device provided by the embodiment of the present application;
[0138] Figure 23 It is a schematic diagram of the structure of still another model training device provided by the embodiment of the present application;
[0139] Figure 24 It is a schematic diagram of the structure of a model training device provided by another embodiment of the present application;
[0140] Figure 25 It is a schematic diagram of the structure of another model training device provided by another embodiment of the present application;
[0141] Figure 26 It is a schematic diagram of the structure of yet another model training device provided by another embodiment of the present application;
[0142] Figure 27 It is a block diagram of an analysis device provided by the embodiment of the present application. Detailed Description of the Embodiment
[0143] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0144] For the convenience of readers' understanding, the machine learning algorithms involved in the model training method provided by the embodiments of the present application are briefly introduced.
[0145] Machine learning algorithms, as an important branch in the field of AI, have been widely applied in many fields. From the perspective of learning methods, machine learning algorithms can be classified into several categories: supervised learning algorithms, unsupervised learning algorithms, semi-supervised learning algorithms, and reinforcement learning algorithms. Supervised learning algorithms refer to the ability to learn an algorithm or establish a pattern based on training data and use this algorithm or pattern to infer new instances. Training data, also known as training samples, consists of input data and expected outputs. The model of a machine learning algorithm, also known as a machine learning model, has an expected output called a label, which can be a predicted classification result (referred to as a classification label). The difference between unsupervised learning algorithms and supervised learning algorithms is that the training samples of unsupervised learning algorithms do not have given labels, and the machine learning algorithm model analyzes the training samples to obtain certain results. Semi-supervised learning algorithms have some training samples with labels and some without labels, and the unlabeled data is much more than the labeled data. Reinforcement learning algorithms continuously try in the environment to achieve the maximum expected benefit, and through the rewards or punishments given by the environment, generate choices that can obtain the maximum benefit.
[0146] It should be noted that each training sample includes one-dimensional or multi-dimensional feature data, that is, feature data including one or more features. By way of example, in the scenario of predicting the classification result of key performance indicator (KPI) data, the feature data can specifically be KPI feature data. KPI feature data refers to the feature data generated based on KPI data. The KPI feature data can be the feature data of the KPI time series, that is, the data obtained by extracting the features of the KPI time series; the KPI feature data can also directly be KPI data. Among them, KPI can specifically be network KPI, and network KPI can include various categories of KPI such as central processing unit (CPU) utilization rate, optical power, network traffic, packet loss rate, delay, and / or user access number. When the KPI feature data is the feature data of the KPI time series, the KPI feature data can specifically be the feature data extracted from the time series of the KPI data of any of the foregoing KPI categories. For example, a training sample includes network KPI feature data with 2 features, namely the maximum value and the weighted average value of the corresponding network KPI time series. When the KPI feature data is KPI data, the KPI feature data can specifically be the KPI data of any of the foregoing KPI categories. For example, a training sample includes network KPI feature data with 3 features, namely CPU utilization rate, packet loss rate, and delay. Further, in the scenario of applying supervised learning algorithms or semi-supervised learning algorithms, the training sample can also include a label. For example, in the scenario of predicting the classification result of KPI data as described above, assuming that the classification result is used to indicate whether the data sequence is abnormal, a training sample also includes a label: "abnormal" or "normal".
[0147] It should be noted that the foregoing time series is a special data sequence, which is a set of data arranged in time sequence. The time sequence is usually the order of data generation, and the data in the time series is also called data points. Usually, the time interval between each data point in a time series is a constant value, so the time series can be analyzed and processed as discrete time data.
[0148] The training methods of current machine learning algorithms are divided into offline learning (also known as offline training) and online learning.
[0149] In the offline learning (also known as offline training) method, it is necessary to batch input the samples in the training sample set into the machine learning model for model training, and the amount of data required for training is large. Offline learning is usually used to train large or complex models, so the training process is often time-consuming and the amount of data processed is large.
[0150] In the online learning (also known as online training) method, it is necessary to use samples in the training sample set in small batches or one by one for model training, and the amount of data required for training is small. Online learning is often applied to scenarios with high requirements for immediacy. The incremental learning (also known as incremental training) method is a special online learning method, which not only requires the model to have the ability to immediately learn new patterns, but also requires the model to have the ability to resist forgetting, that is, the model should be able to remember the patterns learned in history and learn new patterns at the same time.
[0151] In the practical tasks of machine learning, it is necessary to select representative samples to form a sample set to build a machine learning model. Usually, in the labeled sample data, samples with strong correlation with the category are selected as the sample set. Among them, the label is used to identify the sample data, such as identifying the category of the sample data. In the embodiments of the present application, the data used for machine learning model training are all sample data. In the following text, the training data are called training samples, the training sample set is called the training sample set, and in some content, the sample data are simply called samples.
[0152] Figure 1 is a schematic diagram of an application scenario involved in the model training method provided by the embodiments of the present application. As Figure 1 shown, this application scenario includes multiple analysis devices, and the multiple analysis devices include analysis device 101 and multiple analysis devices 102. Each analysis device is used to perform a series of data analysis processes such as data mining and / or data modeling. Figure 1 The numbers of analysis device 101 and analysis devices 102 in are only for illustration and do not limit the application scenario involved in the model training method provided by the embodiments of the present application.
[0153] Among them, the analysis device 101 can specifically be a cloud analysis device (also known as a cloud analysis platform), which can be a computer, a server, a server cluster composed of several servers, or a cloud computing service center, and is deployed at the backend of the service network. The analysis device 102 can specifically be a local point analysis device (also known as a local point analysis platform), which can be a server, a server cluster composed of several servers, or a cloud computing service center. In this application scenario, the model training system involved in the model training method includes multiple local point networks, and the local point network can be a core network or an edge network. The users of each local point network can be operators or enterprise customers. Multiple analysis devices 102 and multiple local point networks can correspond one by one. Each analysis device 102 is used to provide data analysis services for the corresponding local point network. Each analysis device 102 can be located within the corresponding local point analysis network or outside the corresponding local point analysis network. Each analysis device 102 is connected to the analysis device 101 through a wired network or a wireless network. The communication network involved in the embodiments of the present application is a second-generation (2G) communication network, a third-generation (3G) communication network, a Long Term Evolution (LTE) communication network, or a fifth-generation (5G) communication network, etc.
[0154] In addition to performing data analysis, the analysis device 101 is also used to manage some or all of the services of the analysis device 102, collect a training sample set, provide data analysis services for the analysis device 102, etc. The analysis device 101 can train a machine learning model based on the collected training sample set (this process is the aforementioned offline learning method), and then deploy the machine learning model in each local point analysis device, and the local point analysis device performs incremental training (this process is the aforementioned online learning method). Based on different training samples, different machine learning models can be trained, and different machine learning models can implement different classification functions. For example, functions such as anomaly detection, prediction, network security protection, application identification, or user experience evaluation (i.e., evaluating the user's experience) can be implemented.
[0155] Furthermore, as Figure 2 shown, Figure 2 is another schematic diagram of an application scenario involved in the model training method provided by the embodiments of the present application. In Figure 1Based on this, the application scenario further includes network device 103. Each analysis device 102 can manage network devices 103 in a network (also known as a local network). There is a wired or wireless connection between the analysis device 102 and the network devices 103 it manages. The network device 103 can be a router, a switch, a base station, or the like. There is a wired or wireless connection between the network device 103 and the analysis device 102. The network device 103 is used to upload the collected data to the analysis device 102, such as various KPI time series. The analysis device 102 is used to extract and use the data from the network device 103, such as determining the tags of the obtained time series. Optionally, the data uploaded by the network device 103 to the analysis device 102 may further include various types of log data and device status data, etc.
[0156] Such as Figure 3 shown, Figure 3 is a schematic diagram of another application scenario involved in the model training method provided by an embodiment of the present application. Based on Figure 1 or Figure 2 this, the application scenario further includes an evaluation device 104 ( Figure 3 The drawing is based on Figure 2 as the basis of the application scenario, but this is not limited thereto). There is a wired or wireless connection between the evaluation device 104 and the analysis device 102. The evaluation device 104 is used to evaluate the classification result of the analysis device 102 using the machine learning model for data classification, and control the local analysis device to perform incremental training of the machine learning model based on the evaluation result.
[0157] Based on Figures 1 to 3 the scenario shown, the application scenario may further include a storage device, which is used to store the data provided by the network device 103 or the analysis device 102. The storage device can be a distributed storage device, and the analysis device 102 or the analysis device 101 can read and write the data stored in the storage device. In this way, when there is a large amount of data in the application scenario, the storage device is used for data storage, which can reduce the load of the analysis device (such as the analysis device 102 or the analysis device 101) and improve the data analysis efficiency of the analysis device. The storage device can be used to store the data with determined tags, and these data with determined tags can be used as samples for model training. It should be noted that when the amount of data in the application scenario is small, the storage device may not be set up.
[0158] Optionally, the application scenario further includes a management device, such as a network management device (also known as a network management platform) or a third-party management device, which is used to provide configuration feedback and sample annotation feedback, and is usually managed by operation and maintenance personnel. For example, the management device can be a computer, a server, a server cluster composed of several servers, or a cloud computing service center, which can be an operations support system (OSS) or other network devices connected to the analysis device. Optionally, the foregoing analysis device can perform feature data selection and model update for each machine learning model, and feedback the selected feature data and the model update result to the management device, and the management device decides whether to retrain the model.
[0159] Furthermore, the model training method provided by the embodiments of the present application can be used in an anomaly detection scenario. Anomaly detection refers to the detection of patterns that do not meet expectations. The data sources for anomaly detection include applications, processes, operating systems, devices, or networks. For example, the object of the anomaly detection can be the foregoing KPI data sequence. When the model training method provided by the embodiments of the present application is applied in an anomaly detection scenario, the analysis device 102 can be a network analyzer, the machine learning model maintained by the analysis device 102 is an anomaly detection model, and the determined label is an anomaly detection label, and the anomaly detection label includes two classification labels, namely: "normal" and "abnormal".
[0160] In the anomaly detection scenario, the foregoing machine learning model can be a model based on algorithms of statistics and data distribution (such as the N-Sigma algorithm), a model based on distance / density algorithms (such as the Local Outlier Factor algorithm), a tree model (such as Isolation forest (Iforest)), or a model based on prediction algorithms (such as the Autoregressive Integrated Moving Average model (ARIMA)), etc.
[0161] In the related art, in a data analysis system, an off-line training of a model is performed by a cloud analysis device, and then the off-line trained model is directly deployed on a local analysis device. However, the trained model may not be effectively adapted to the requirements of the local analysis device, such as the requirements for prediction performance (such as accuracy or recall rate). On the one hand, the training samples in the historical training sample set used by the cloud analysis device are usually pre-configured fixed training samples, which may not conform to the requirements of the local analysis device; on the other hand, even if the machine learning model obtained by training conforms to the requirements of the local analysis device when it is first deployed on the local analysis device, over time, due to changes in the categories or patterns of the feature data obtained by the local analysis device, the machine learning model obtained by training no longer conforms to the requirements of the local analysis device.
[0162] Also in the related art, the trained machine learning model can only be targeted at a single local analysis device. When the cloud analysis device serves multiple local analysis devices, corresponding machine learning models need to be trained separately for each local analysis device. The trained models have low generality, cannot achieve model generalization, and have high training costs.
[0163] The embodiment of the present application provides a model training method. In the following embodiments, it is assumed that the foregoing analysis device 101 is the first analysis device, and the analysis device 102 is the local analysis device. The local analysis device receives the machine learning model sent by the first analysis device and can perform incremental training on the machine learning model based on the first training sample set obtained from the local network corresponding to the local analysis device. On the one hand, the feature data in the first training sample set is the feature data obtained from the local network corresponding to the local analysis device, which is more adapted to the application scenario of the local analysis device. Using the first training sample set including the feature data obtained by the local analysis device from the corresponding local network for model training can make the trained machine learning model more adapted to the requirements of the local analysis device itself (that is, the requirements of the local network corresponding to the local analysis device), realize the customization of the model, and improve the application flexibility of the model; on the other hand, by combining off-line training and incremental training to train the machine learning model, incremental training of the machine learning model can be performed when the categories or patterns of the feature data obtained by the local analysis device change, realizing flexible adjustment of the machine learning model, so as to ensure that the trained machine learning model meets the requirements of the local analysis device. Therefore, the model training method provided by the embodiment of the present application can be effectively adapted to the requirements of the local analysis device compared with the related art.
[0164] Further, the first analysis device can distribute the trained machine learning model to each local analysis device for incremental training to ensure the performance of the machine learning model on each local analysis device. In this way, the first analysis device does not need to train a corresponding machine learning model for each local analysis device, effectively reducing the overall training duration of the first analysis device. Moreover, the model obtained through offline training can serve as the basis for incremental training on each local analysis device, improving the generality of the model obtained through offline training, thereby realizing model generalization and reducing the overall training cost of the first analysis device.
[0165] An embodiment of this application provides a model training method, which can be applied to Figures 1 to 3 any of the illustrated application scenarios. The machine learning model can be used to predict classification results. For example, it can be a binary classification model. For the sake of distinction, in the subsequent embodiments of this application, the classification results determined by the artificial or label migration method are referred to as labels, and the results predicted by the machine learning model itself are referred to as classification results. Substantially the same, both are used to identify the category of the corresponding sample. The application scenarios of this model training method usually include multiple local analysis devices. In the embodiments of this application, a model training method is described by taking one local analysis device as an example, and the actions of other local analysis devices can refer to the actions of this local analysis device. As Figure 4 shown, the method includes:
[0166] Step 401, the first analysis device performs offline training based on the historical training sample set to obtain a machine learning model.
[0167] The first analysis device can continuously collect training samples to obtain a training sample set, and perform offline training based on the collected training sample set (which can be called the historical training sample set) to obtain a machine learning model. By way of example, the historical training sample set can be a set of training samples sent by multiple local analysis devices. The machine learning model trained based on this can adapt to the needs of the multiple local analysis devices, and the generality of the trained model is relatively high, thus ensuring model generalization.
[0168] Please refer to the foregoing Figure 2 and Figure 3 , the training samples can be obtained by the local analysis device from the data collected and uploaded by the network device and transmitted to the first analysis device by the local analysis device. The training samples can also be obtained by the first analysis device in other ways, such as obtaining from the data stored in the storage device. The embodiments of this application do not make any limitations in this regard.
[0169] Among them, the training samples can have various forms. Correspondingly, the first analysis device can obtain training samples in various ways. The embodiments of this application take the following two optional ways as examples for illustration:
[0170] In the first optional mode, the training samples obtained by the foregoing first analysis device may include data determined based on a time series. For example, it includes data determined based on a KPI time series. Usually, each training sample in the historical training sample set corresponds to a time series, and each training sample may include feature data obtained by extracting one or more features from the corresponding time series. The number of features corresponding to each training sample is the same as the number of feature data of the training sample (that is, the features and the feature data are in one-to-one correspondence). Among them, the features in the training sample refer to the features possessed by the corresponding time series, which may include data features and / or extraction features.
[0171] Among them, the data feature is the self-feature of the data in the time series. For example, the data feature includes the data arrangement period, the data change trend, or the data fluctuation, etc. Correspondingly, the feature data of the data feature includes: the data of the data arrangement period, the data change trend data, or the data fluctuation data, etc. The data arrangement period refers to the period involved in the data arrangement in the time series if the data in the time series is arranged periodically. For example, the data of the data arrangement period includes the period duration (that is, the time interval between the initiations of two periods) and / or the number of periods; the data change trend data is used to reflect the change trend of the data arrangement in the time series (that is, the data change trend). For example, the data change trend data includes: continuous growth, continuous decline, rising first and then falling, falling first and then rising, or conforming to a normal distribution, etc.; the data fluctuation data is used to reflect the fluctuation state of the data in the time series (that is, the data fluctuation). For example, the data fluctuation data includes the function representing the fluctuation curve of the time series, or the specified value of the time series, such as the maximum value, the minimum value, or the average value.
[0172] Feature extraction is the process of extracting features from the data in the time series. For example, feature extraction includes statistical features, fitting features, or frequency domain features, etc. Correspondingly, the feature data for feature extraction includes statistical feature data, fitting feature data, or frequency domain feature data, etc. Statistical features refer to the statistical features of the time series. Statistical features are divided into quantitative features and attribute features. Among them, quantitative features are further divided into measurement features and counting features. Quantitative features can be directly represented by numerical values. For example, the consumption values of various resources such as CPU, memory, and IO resources are measurement features; while the number of times of anomalies and the number of normally working devices are counting features; attribute features cannot be directly represented by numerical values, such as whether the device has an anomaly or whether the device has a downtime, etc. The features in statistical features are the indicators to be considered during statistics. For example, the statistical feature data includes moving average, weighted average, etc.; Fitting features are the features during time series fitting, and the fitting feature data is used to reflect the features used for fitting the time series. For example, the fitting feature data includes the algorithms used for fitting, such as ARIMA; Frequency domain features are the features of the time series in the frequency domain, and the frequency domain features are used to reflect the features of the time series in the frequency domain. For example, the frequency domain feature data includes: data on the law followed by the distribution of the time series in the frequency domain, such as the proportion of high-frequency components in the time series. Optionally, the frequency domain feature data can be obtained by performing wavelet decomposition on the time series.
[0173] Assume that the feature data in the training sample is obtained from the first time series. Then the data acquisition process may include: determining the target features to be extracted, extracting the feature data of the determined target features in the first time series, and obtaining a training sample composed of the data of the obtained target features. Exemplarily, the target features to be extracted are determined based on the application scenarios involved in the model training method. In an alternative exemplary embodiment, the target features are pre-configured features, such as features configured by the user. In another alternative exemplary embodiment, the target features are one or more of the specified features. For example, the specified features are the aforementioned statistical features.
[0174] It should be noted that the user can preset specified features in advance. However, for the first time series, it may not have all the specified features. The first analysis device can screen the features belonging to the specified features in the first time series as target features. For example, the target features include statistical features: time series decomposition - seasonal component (time seriesdecompose_seasonal, Tsd_seasonal), moving average, weighted average, time series classification, maximum value, minimum value, quantile, variance, standard deviation, year-on-year (yoy, referring to comparison with the same historical period), daily volatility, bucket entropy, sample entropy, moving average, exponential moving average, Gaussian distribution feature, or T-distribution feature, etc. One or more of them. Correspondingly, the target feature data includes the data of one or more of these statistical features;
[0175] And / or, the target features include fitting features: autoregressive fitting error, Gaussian process regression fitting error, or neural network fitting error, etc. One or more of them. Correspondingly, the target feature data includes the data of one or more of these fitting features;
[0176] And / or, the target features include frequency domain features: the proportion of high-frequency components in the time series; Correspondingly, the target feature data includes the data of the proportion of high-frequency components in the time series, and this data can be obtained by performing wavelet decomposition on the time series.
[0177] Table 1 is a schematic description of a sample in the historical training sample set. In Table 1, each training sample in the historical training sample set includes the feature data of one or more features of the KPI time series, and each training sample corresponds to a KPI time series. In Table 1, the training sample with the identification (ID) of KPI_1 includes the feature data of 4 features, and the feature data of these 4 features are respectively: moving average (Moving_average), weighted average (Weighted_mv), time series decomposition - seasonal component (time series decompose_seasonal, Tsd_seasonal), and year-on-year (yoy). The KPI time series corresponding to this training sample is (x1, x2, ……, xn) (this time series is usually obtained by sampling the data of a KPI category), and the corresponding label is "abnormal".
[0178] Table 1
[0179]
[0180] In the second optional manner, the training samples obtained by the foregoing first analysis device may include data with certain characteristics itself, which is the data obtained itself. For example, the training samples include KPI data. As mentioned above, assuming that the KPI is a network KPI, each sample may include network KPI data of one or more network KPI categories, that is, the corresponding feature of the sample is the KPI category.
[0181] Table 2 is a schematic illustration of a sample in the historical training sample set. In Table 2, each training sample in the historical training sample set includes network KPI data of one or more features. In Table 2, each training sample corresponds to multiple network KPI data obtained at the same acquisition moment. In Table 2, the training sample with the identity identification (ID) of KPI_2 includes feature data of 4 features, and the feature data of the 4 features are respectively: network traffic, CPU utilization rate, packet loss rate, and delay, and the corresponding label is "normal".
[0182] Table 2
[0183]
[0184] The feature data corresponding to each feature in Table 1 and Table 2 above is usually numerical data, that is, each feature has a feature value. For the convenience of description, the feature value is not shown in Table 1 and Table 2. Assuming that the historical training sample set stores the feature data in a fixed format, and its corresponding features may be pre-set features, the feature data of the historical training sample set can be stored in the format of Table 1 or Table 2. In the actual implementation of the embodiments of the present application, the samples in the historical training sample set may also have other forms, and the embodiments of the present application do not limit this.
[0185] It should be noted that before offline training, the first analysis device may preprocess the samples in the collected training sample set, and then perform the foregoing offline training based on the preprocessed training sample set. This preprocessing process is used to process the collected samples into samples that meet the preset conditions, and this preprocessing process may include one or more of sample deduplication, data cleaning, and data completion.
[0186] The offline training process described in step 401 is also called the model learning process, which is the learning process for a machine learning model to perform its relevant classification functions. In an alternative approach, this offline training process is a process of training an initial learning model to obtain a machine learning model; in another alternative approach, this offline training process is a process of establishing a machine learning model, that is, the machine learning model after offline training is the initial learning model, and the embodiments of the present application do not make any limitations in this regard. After completing the offline training, the first analysis device can also perform a model evaluation process on the trained machine learning model to evaluate whether the machine learning model meets the performance standard conditions. When the machine learning model meets the performance standard conditions, the following step 402 is executed. When the machine learning model does not meet the performance standard conditions, at least one retraining of the machine learning model can be performed until the machine learning model meets the performance standard conditions, and then the following step 402 is executed.
[0187] In one example, the first analysis device can set a first performance standard threshold based on user requirements, and compare the parameter value of the positive performance parameter of the trained machine learning model with the first performance standard threshold. When the positive performance parameter value is greater than the first performance standard threshold, it is determined that the machine learning model meets the performance standard conditions; when the positive performance parameter value is not greater than the first performance standard threshold, it is determined that the machine learning model does not meet the performance standard conditions. This positive performance parameter is positively correlated with the quality of the performance of the machine learning model, that is, the larger the parameter value of this positive performance parameter, the better the performance of the machine learning model. For example, the positive performance parameter is an index characterizing the model performance such as accuracy, recall rate, precision rate, or f-score (f-measure), and for another example, the first performance standard threshold is 90%. The accuracy = the number of correct predictions / the total number of predictions.
[0188] In another example, the first analysis device can set a first performance degradation threshold based on user requirements, and compare the parameter value of the negative performance parameter of the trained machine learning model with the first performance degradation threshold. When the negative performance parameter value is greater than the first performance degradation threshold, it is determined that the machine learning model does not meet the performance standard conditions; when the negative performance parameter value is greater than the first performance degradation threshold, it is determined that the machine learning model meets the performance standard conditions. This negative performance parameter is negatively correlated with the quality of the performance of the machine learning model, that is, the larger the parameter value of this negative performance parameter, the worse the performance of the machine learning model. For example, the negative performance parameter is the classification result error rate (also called the misjudgment rate), and the first performance degradation threshold is 20%. The misjudgment rate = the number of incorrect predictions / the total number of predictions.
[0189] For example, a specified number of test samples are input into a machine learning model to obtain a specified number of classification results. Based on these specified number of classification results, the accuracy rate or misjudgment rate is statistically calculated. In the foregoing formula, the total number of prediction times is the aforementioned specified number. Whether the predicted classification result is correct or incorrect can be determined by the operation and maintenance personnel based on expert experience.
[0190] For example, if the specified number is 100 times, and among them, the prediction is incorrect 20 times, then the misjudgment rate is 20 / 100 = 20%. If the first performance degradation threshold is 10%, it is determined that the machine learning model does not meet the performance compliance condition.
[0191] The foregoing retraining process can be an offline training process or an online training process (such as an incremental training process). The training samples used in this retraining process can be the same as or different from the training samples used in the previous training process. The embodiments of the present application do not make any limitations in this regard.
[0192] Step 402: The first analysis device sends the machine learning model to multiple local analysis devices.
[0193] The first analysis device can provide the machine learning model to each local analysis device in different ways. The embodiments of the present application are described by the following two examples: In an optional example, after receiving the model acquisition request sent by the local analysis device, the first analysis device sends the machine learning model to the local analysis device, and this model acquisition request is used to request the first analysis device to obtain the machine learning model; in another optional example, after training the machine learning model, the first analysis device actively pushes the machine learning model to the local analysis device.
[0194] Exemplarily, the first analysis device may include a model deployment module, and this model deployment module has established communication connections with each local analysis device, and the machine learning model can be deployed to each local analysis device through this model deployment module.
[0195] Step 403: The local analysis device uses the machine learning model to predict the classification result.
[0196] As described above, different machine learning models can respectively implement different functions. These functions are all realized by predicting the classification result. The classification results corresponding to different functions are different. After receiving the machine learning model sent by the first analysis device, the local analysis device can use the machine learning model to predict the classification result.
[0197] Exemplarily, if it is necessary to predict the classification result of the online data of the local analysis device, the data that needs to be predicted for the classification result may include the KPI of the CPU and / or the KPI of the memory.
[0198] Suppose it is necessary to perform anomaly detection on the online data of the game point analysis device, that is, to predict whether the classified result indicates that the data is abnormal. Then the game point analysis device can periodically execute the anomaly detection process. After performing anomaly detection on the online data, the anomaly detection results output by the machine learning model are shown in Table 3 and Table 4. Table 3 and Table 4 record the anomaly detection results of the data to be detected obtained at different acquisition times (also known as data generation times), where the different acquisition times include T1 to TN (N is an integer greater than 1), and the anomaly detection results indicate whether the corresponding data to be detected is abnormal. Among them, the data to be detected in Table 3 and Table 4 both include one-dimensional feature data. Table 3 records the anomaly detection results of the data to be detected of the KPI with the feature category of CPU; Table 4 records the anomaly detection results of the data to be detected of the KPI with the feature category of memory. Suppose 0 represents normal and 1 represents abnormal. The time interval between every two acquisition times from T1 to TN is a preset time period. Taking the acquisition time T1 as an example, at this time, the KPI of CPU in Table 3 is 0, and the KPI of memory in Table 4 is 1, indicating that the KPI of CPU collected at the acquisition time T1 is normal, and the KPI of memory collected at the acquisition time T1 is abnormal.
[0199] Table 3
[0200] Collection time KPI of CPU T1 0 T2 0 ... ... TN 0
[0201] Table 4
[0202] Collection time KPI of memory T1 1 T2 0 ... ... TN 1
[0203] Step 404: The game point analysis device performs incremental training on the machine learning model based on the first training sample set.
[0204] There are various situations for the game point analysis device to obtain the first training sample set. For example, the game point analysis device can periodically obtain the first training sample set; for another example, after the game point analysis device receives a sample set acquisition instruction sent by the operation and maintenance personnel of the game point analysis device, or receives a sample set acquisition instruction sent by the first analysis device or the aforementioned management device, the game point analysis device obtains the first training sample set, and this sample set acquisition instruction is used to indicate obtaining the first training sample set; for yet another example, when the machine learning model deteriorates on the game point analysis device, the game point analysis device obtains the first training sample set.
[0205] Under normal circumstances, the game point analysis device only performs incremental training on the machine learning model when the machine learning model deteriorates, which can reduce the training duration and avoid affecting the user's business. The triggering mechanism for this incremental training (that is, the detection mechanism for detecting whether the model deteriorates) can include the following two situations:
[0206] In the first case, as Figure 3As shown, the application scenario of this model training method also includes an evaluation device, which can control the local point analysis device to perform incremental training on the machine learning model based on the evaluation result of the classification result. For example Figure 5 , this process includes:
[0207] Step 4041, the local point analysis device sends the prediction information to the evaluation device.
[0208] In one example, the local point analysis device can send the prediction information to the evaluation device after each prediction of the classification result using the machine learning model. The prediction information includes the predicted classification result. In another example, the local point analysis device can also periodically send the prediction information to the evaluation device, and the prediction information includes the classification results obtained in the current period. In yet another example, the local point analysis device can also send the prediction information to the evaluation device after the number of obtained classification results reaches a threshold, and the prediction information includes the obtained classification results. In still another example, the local point analysis device can also send the prediction information to the evaluation device within a set time period, and the prediction information includes the currently obtained classification results. For example, this time period can be a user-set time period, or a period when the user's business occurrence frequency is lower than a specified frequency threshold, such as 0:00 - 5:00. This can avoid interfering with the user's business.
[0209] It should be noted that in different application scenarios, the foregoing prediction information can also carry other information to facilitate the evaluation device to effectively evaluate each classification result and ensure the accuracy of the evaluation.
[0210] For example, in the scenario of predicting the classification result of the KPI data sequence, the machine learning model is used to predict the classification result of the data to be predicted composed of one or more KPI feature data. Among them, the KPI feature data is the feature data of the KPI time series or the KPI data. Correspondingly, the prediction information also includes: the identifier of the device to which the data to be predicted belongs (that is, the device that generates the KPI data corresponding to the data to be predicted, such as a network device), the KPI category corresponding to the data to be predicted, and the collection time of the KPI data corresponding to the data to be predicted. Based on this information, it is possible to determine the device, KPI category, and collection time of the KPI data corresponding to each classification result, so as to accurately determine whether the KPI data collected at different collection times is abnormal.
[0211] Among them, when the KPI feature data is the feature data of the KPI time series, the KPI data corresponding to the data to be predicted is the data in the KPI time series. Then, the KPI category corresponding to the data to be predicted is the category of the KPI time series, and the acquisition moment is the acquisition moment of any data in the KPI time series, or can also be the acquisition moment of the data at a specified position, such as the acquisition moment of the last data. For example, assume that the KPI category of the KPI time series is packet loss rate, and the time series: (x1, x2, ……, xn) represents that the packet loss rates collected in one acquisition period are x1, x2, ……, xn respectively. Assume that the structure of the data to be predicted is similar to Table 1, and the data is (1, 2, 3, 4), representing that the moving average is 1, the weighted average is 2, the periodic component of time series decomposition is 3, and the periodic yoy is 4. Assume that the acquisition moment of the KPI data corresponding to the data to be predicted is the acquisition moment of the last data in the KPI time series. Then, the KPI category corresponding to the data to be predicted is packet loss rate, and the acquisition moment of the KPI data corresponding to the data to be predicted is the acquisition moment of xn.
[0212] When the KPI feature data is KPI data, the KPI category corresponding to the data to be predicted is the KPI category of the KPI data, the KPI data corresponding to the data to be predicted is the KPI data itself, and the acquisition moment is the acquisition moment of the KPI data. For example, assume that the structure of the data to be predicted is similar to Table 2, and the data to be predicted is (100, 20%, 3, 4), representing that the network traffic is 100, the CPU utilization rate is 20%, the packet loss rate is 3, and the latency is 4. Then, the KPI categories corresponding to the data to be predicted are network traffic, CPU utilization rate, packet loss rate, and latency, and the acquisition moment of the KPI data corresponding to the data to be predicted is the acquisition moment of (100, 20%, 3, 4), and this acquisition moment is usually the same acquisition moment.
[0213] After receiving the prediction information, the evaluation device can present at least the classification result and the data to be predicted in the prediction information, or can also present all the content in the prediction information for the operation and maintenance personnel to mark whether the classification result is correct or wrong according to expert experience.
[0214] Step 4042: The evaluation device evaluates whether the machine learning model deteriorates based on the prediction information.
[0215] Exemplarily, the evaluation device can evaluate whether the machine learning model deteriorates based on the prediction information when the received classification results reach a specified quantity threshold; or can also periodically evaluate whether the machine learning model deteriorates based on the prediction information. Correspondingly, the evaluation period can be one week or one month, etc. In the embodiments of the present application, this evaluation process is similar to the model evaluation process in the foregoing step 401 in principle.
[0216] In one example, the evaluation device may set a second performance compliance threshold based on user requirements, and compare the parameter value of the positive performance parameter of the trained machine learning model with the second performance compliance threshold. When the positive performance parameter value is greater than the second performance compliance threshold, it is determined that the machine learning model has not deteriorated; when the positive performance parameter value is not greater than the second performance compliance threshold, it is determined that the machine learning model has deteriorated. The positive performance parameter is positively correlated with the quality of the performance of the machine learning model, that is, the larger the parameter value of the positive performance parameter, the better the performance of the machine learning model. For example, the positive performance parameter is an index characterizing the model performance such as accuracy, recall rate, precision rate, or f-score, and the second performance compliance threshold is 90%. The second performance compliance threshold and the aforementioned first performance compliance threshold may be the same or different. The calculation method of the accuracy rate may refer to the calculation method of the accuracy rate provided in the model evaluation process in the aforementioned 401.
[0217] In another example, the evaluation device may set a second performance deterioration threshold based on user requirements, and compare the parameter value of the negative performance parameter of the trained machine learning model with the second performance deterioration threshold. When the negative performance parameter value is greater than the second performance deterioration threshold, it is determined that the machine learning model has deteriorated; when the negative performance parameter value is greater than the second performance deterioration threshold, it is determined that the machine learning model has not deteriorated. The negative performance parameter is negatively correlated with the quality of the performance of the machine learning model, that is, the larger the parameter value of the negative performance parameter, the worse the performance of the machine learning model. For example, the negative performance parameter is the classification result error rate (also called misjudgment rate), and the second performance deterioration threshold is 20%. The second performance deterioration threshold and the aforementioned first performance deterioration threshold may be the same or different. The calculation method of the misjudgment rate may refer to the calculation method of the misjudgment rate provided in the model evaluation process in the aforementioned 401.
[0218] For example, the evaluation device obtains multiple classification results. The accuracy rate or misjudgment rate is statistically calculated based on the obtained multiple classification results. The total number of prediction times in the accuracy rate or misjudgment rate is the number of the obtained classification results. As mentioned above, the correctness or error of the predicted classification results can be determined by the operation and maintenance personnel of the evaluation device.
[0219] In the anomaly detection scenario, the false positive rate can also be obtained in other ways. In this scenario, the local analysis device is also communicatively connected to the management device. When the classification result output by the machine learning model of the local analysis device is "anomaly", an alarm message will be sent to the management device. This alarm message is used to indicate that the sample data is abnormal and carries the sample data with the classification result of "anomaly". The management device will identify the sample data and the classification result in the alarm message. If the classification result is incorrect, the classification result will be updated (i.e., the classification result: "anomaly" is updated to "normal"), indicating that this alarm message is a false alarm message. The number of false alarm messages is the number of errors in the predicted classification result. The local analysis device or the management device can feedback the number of false alarm messages in each evaluation period to the evaluation device, or report the false alarm messages to the evaluation device, and the evaluation device will count the number of false alarm messages. Then, based on the number of classification results obtained during the evaluation period counted by the evaluation device, the false positive rate is calculated using the aforementioned false positive rate calculation formula.
[0220] Step 4043: After determining that the machine learning model has deteriorated, the evaluation device sends a training instruction to the local analysis device.
[0221] This training instruction is used to indicate training the machine learning model.
[0222] Optionally, when the evaluation device determines that the machine learning model has not deteriorated, no action is taken.
[0223] Step 4044: After receiving the training instruction sent by the evaluation device, the local analysis device incrementally trains the machine learning model based on the first training sample set.
[0224] In the second case, the local analysis device itself can evaluate whether the machine learning model has deteriorated based on the prediction information, and after the machine learning model has deteriorated, incrementally train the machine learning model based on the first training sample set. This evaluation process can refer to the aforementioned step 4042.
[0225] It is worth noting that the local analysis device can also use other triggering mechanisms for incremental training. For example, when at least one of the following triggering conditions is met, incremental training is performed: reaching the incremental training cycle, or receiving a training instruction sent by the operation and maintenance personnel of the local analysis device, or receiving a training instruction sent by the first analysis device. This training instruction is used to indicate performing incremental training.
[0226] In an embodiment of the present application, the first training sample set may include sample data that is directly extracted from the data obtained by the local point analysis device based on set rules and whose labels are determined. For example, the first training sample set may include time series data or time series feature data obtained by the local point analysis device from a network device. Moreover, the labels of the first training sample set may be presented to the operation and maintenance personnel by the local point analysis device, or the aforementioned management device, or the aforementioned first analysis device, and the operation and maintenance personnel perform label annotation according to expert experience.
[0227] Referring to the foregoing step 401, there can be various forms of training samples. Correspondingly, the local point analysis device can obtain training samples in various ways. In an embodiment of the present application, the following two optional ways are taken as examples for illustration:
[0228] In the first optional way, the training samples in the first training sample set obtained by the local point analysis device may include data determined based on a time series. For example, it includes data determined based on a KPI time series. Referring to the structure of the foregoing historical training sample set, usually, each training sample in the first training sample set corresponds to a time series, and each training sample may include feature data obtained by extracting one or more features from the corresponding time series. The features corresponding to each training sample are the same as the number of the feature data of the training sample (that is, the features and the feature data are in one-to-one correspondence). Among them, the features in the training sample refer to the features possessed by the corresponding time series, which may include data features and / or extraction features.
[0229] In an optional example, the local point analysis device may receive the time series sent by the network device (i.e., the network device it manages) connected to the local point analysis device in the corresponding local point network; in another optional example, the local point analysis device has an input / output (I / O) interface and receives the time series in the corresponding local point network through this I / O interface; in still another optional example, the local point analysis device may read the time series from its corresponding storage device, and this storage device is used to store the time series pre-obtained by the local point analysis device in the corresponding local point network.
[0230] Assume that the feature data in the training sample is obtained from a second time series. Then this data acquisition process may refer to the foregoing process of obtaining the training samples in the historical training sample set from the first time series. For example, determine the target features to be extracted, extract the feature data of the determined target features in the second time series, and obtain the first training sample composed of the data of the obtained target features. This is not elaborated in the embodiments of the present application.
[0231] In the second optional method, the training samples obtained by the local point analysis device may include data with certain characteristics, which are the data itself obtained. For example, the training samples include KPI data. As described above, assuming that the KPI is a network KPI, each sample may include network KPI data of one or more network KPI categories, that is, the feature corresponding to the sample is the KPI category.
[0232] The process of the local point analysis device obtaining training samples may refer to the process of the first analysis device obtaining training samples in the foregoing step 401, and the structure of the training samples in the obtained first training sample set may also refer to the structure of the training samples in the foregoing historical training sample set. This application embodiment will not elaborate further.
[0233] Generally, when performing incremental training on a machine learning model, the sample data with a collection time closer to the current time has a greater impact on the machine learning model. Then, if the quality of the sample data in the first training sample set used for incremental training of the machine learning model is poor, the finally trained machine learning model may overwrite the previously trained machine learning model with better performance, resulting in a deviation in the performance of the machine learning model.
[0234] Therefore, the local point analysis device can perform certain screening on the sample data obtained by itself to select sample data with better quality as training samples, and provide these training samples to the operation and maintenance personnel for label annotation to obtain sample data with labels. Thereby improving the performance of the trained machine learning model. This application refers to this screening function as the active learning function.
[0235] The machine learning model predicts the classification result based on probability theory, that is, predicts the probabilities of multiple classification results existing, and takes the classification result with the highest probability as the final classification result. For example, a machine learning model based on the binary classification principle selects the classification result with a larger probability (such as 0 or 1) as the final classification result output. Among them, binary classification means that the classification result of the machine learning model has two types.
[0236] Taking the anomaly detection of the CPU's KPI for online data (i.e., the type of sample data input into the machine learning model is the CPU's KPI) as an example, as shown in Table 5, Table 5 records the probabilities of different classification results of the CPU's KPI obtained at different collection times. The different collection times include T1 to TN (N is a positive integer greater than 1), and the different classification results include two results: "normal" and "abnormal". 0_prob represents the probability of being predicted as normal, and 1_prob represents the probability of being predicted as abnormal. For the collection time T1, 1_prob is 0.51 and 0_prob is 0.49. Since 1_prob is greater than 0_prob, the machine learning model determines that the final classification result of the CPU's KPI collected at this collection time T1 is 1, that is, the CPU's KPI collected at time T1 is abnormal.
[0237] Table 5
[0238] Collection time 1_prob 0_prob Classification result T1 0.49 0.51 0 T2 0.9 0.1 1 ... ... ... ... TN 0.51 0.49 1
[0239] However, when the probabilities of multiple predicted classification results are relatively close, although the final classification result determined by the machine learning model is the classification result with the highest probability, the gap with the classification result with the second-highest probability is very small, resulting in poor reliability of the final classification result determined by the machine learning model. In practical applications of the embodiments of the present application, when the probabilities of multiple predicted classification results are significantly different from each other, the reliability of the final classification result determined by the machine learning model is higher.
[0240] Still taking Table 5 as an example, the difference between the probability that the machine learning model predicts the CPU's KPI obtained at the collection time T1 as 0 and the probability as 1 is only 0.02, and the two are very close, indicating that the prediction result of the machine learning model for the sample data obtained at the collection time T1 is unreliable. The difference between the probability that the machine learning model predicts the CPU's KPI obtained at the collection time T2 as 0 and the probability as 1 is relatively large, indicating that the prediction result of the machine learning model for the sample data obtained at the collection time T2 is reliable.
[0241] As can be seen from the above, when the probabilities of multiple classification results predicted by the machine learning model for a sample are significantly different from each other, that is, the discrimination degree of the probabilities is very large, the machine learning model can already determine the accurate classification result, and such a sample does not need to be trained anymore; when the discrimination degree is small, the machine learning model cannot determine the accurate classification result. Such a sample can determine its label in a manual or label transfer manner, so as to give the accurate classification result (which can also be considered as the ideal classification result). Using the sample with the determined label as the training sample for training can improve the reliability of the classification result of the machine learning model for such a sample.
[0242] Exemplarily, embodiments of the present application use low discrimination conditions to screen the first training sample set, that is, the first training sample set includes samples that meet the low discrimination conditions screened from the samples obtained by the local point analysis device. The low discrimination conditions include at least one of the following:
[0243] Condition 1: The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using a machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, where 1 < n < m, and m is the total number of probabilities obtained by the machine learning model predicting the sample. Under this condition, samples with insufficient discrimination among n classification results can be screened out.
[0244] Condition 2: The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using a machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, and n is an integer greater than 1. Under this condition, samples with insufficient discrimination of classification results can be screened out.
[0245] Condition 3: The absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using a machine learning model is less than the third difference threshold. Under this condition, samples with insufficient discrimination of multiple classification results can be screened out.
[0246] Condition 4: The absolute value of the difference between any two probabilities among the probabilities of multiple classification results obtained by predicting the sample using a machine learning model is less than the fourth difference threshold. Under this condition, samples with insufficient discrimination of multiple classification results can be screened out.
[0247] Condition 5: The probability distribution entropy E of the probabilities of multiple classification results obtained by predicting the sample using a machine learning model is greater than the specified distribution entropy threshold, and the E satisfies:
[0248]
[0249] where x i represents the i-th classification result, P(x i ) represents the probability of the i-th classification result of the predicted sample, b is the specified base, such as 2 or the constant e, 0 ≤ P(x i ) ≤ 1, and ∑ represents summation.
[0250] Suppose the first sample is any sample to be predicted obtained by the game point analysis device. For the aforementioned condition 1, a machine learning model can be first used to predict the first sample to obtain the probabilities of multiple classification results. The value range of this probability is from 0 to 1. Sort the probabilities of the multiple classification results in descending order of probability. Screen the probabilities of the first n classification results from the sorted probabilities to obtain a target probability set. Calculate the absolute value of the difference between every two probabilities in the target probability set, and compare the calculated absolute value of the difference with the first difference threshold. When the absolute value of the difference between any two probabilities is less than the first difference threshold, the first sample is determined to be a sample that meets the low discrimination condition.
[0251] For example, n = 2, the first difference threshold is 0.3. Using a machine learning model to predict the sample X, the probabilities of 3 classification results (i.e., m = 3) are 0.32, 0.33, and 0.35 respectively. Then the target probability set includes: 0.33 and 0.35. The absolute value of the difference between the two is less than the first difference threshold, so the sample X is a sample that meets the low discrimination condition.
[0252] For the aforementioned condition 2, a machine learning model can be first used to predict the first sample to obtain the probabilities of multiple classification results. The value range of this probability is from 0 to 1. Then calculate the absolute value of the difference between every two probabilities, and compare the calculated absolute value of the difference with the second difference threshold. When the absolute value of the difference between any two probabilities is less than the second difference threshold, the first sample is determined to be a sample that meets the low discrimination condition.
[0253] For example, in a binary classification scenario, if the absolute value of the difference between the first probability and the second probability corresponding to the first sample is less than the second difference threshold, then the first sample meets the low discrimination condition. The first probability is the probability that the classification result of the first sample is the first classification result predicted by the first tree model, and the second probability is the probability that the classification result of the first sample is the second classification result predicted by the first tree model. Continuing to refer to Table 5, suppose the first sample is the KPI of the CPU obtained at the acquisition time TN, the second difference threshold is 0.1, and the machine learning model predicts that the probabilities of the KPI of the CPU obtained at the acquisition time TN being 0 and 1 are 0.51 and 0.49 respectively. That is, the first probability and the second probability are one of 0.51 and 0.49 respectively, and they only differ by 0.02. The absolute value of the difference between the two is less than 0.1, so it can be determined that the first sample meets the low discrimination condition.
[0254] For the aforementioned condition 3, a machine learning model can be first used to predict the first sample to obtain the probabilities of multiple classification results. The highest probability and the lowest probability can be selected from the probabilities of the multiple classification results, the absolute value of the difference between the two probabilities can be calculated, and the calculated absolute value of the difference can be compared with the third difference threshold. When the absolute value of the difference is less than the third difference threshold, the first sample is determined to be a sample that meets the low discrimination condition.
[0255] For example, the third difference threshold is 0.2. Using a machine learning model to predict the probabilities of 3 classification results of sample Y, which are 0.33, 0.33, and 0.34 respectively. Then the maximum probability and the minimum probability are 0.34 and 0.33 respectively. The absolute value of the difference between the two is less than the third difference threshold, so sample Y is a sample that meets the low discrimination condition.
[0256] For the aforementioned condition 4, a machine learning model can be first used to predict the first sample to obtain the probabilities of multiple classification results. The absolute value of the difference between each two of the probabilities of the multiple classification results can be calculated, and the calculated absolute value of the difference can be compared with the fourth difference threshold. When the absolute value of the difference between any two probabilities is less than the fourth difference threshold, the first sample is determined to be a sample that meets the low discrimination condition.
[0257] For example, the fourth difference threshold is 0.2. Using a machine learning model to predict the probabilities of 3 classification results of sample Z, which are 0.33, 0.33, and 0.34 respectively. The absolute value of the difference between any two probabilities is less than 0.2, so sample Z is a sample that meets the low discrimination condition.
[0258] For the aforementioned condition 5, the probability distribution is a description of a random variable. Different random variables have the same or different probability distributions. The probability distribution entropy is a description of different probability distributions. In the embodiments of the present application, the probability distribution entropy is positively correlated with the uncertainty of the probability. The greater the probability distribution entropy, the greater the uncertainty of the probability. For example, for a binary classification machine learning model, the probabilities of the two classification results of predicting a sample are both 50%. The probability distribution entropy takes the maximum value, but finally it is impossible to select an actual probability-reliable classification result as the final classification result.
[0259] It can be seen from this that when the probability distribution entropy E reaches a certain level, such as a specified distribution entropy threshold, effective probability discrimination cannot be achieved. Therefore, this formula can be used to effectively screen out probabilities with low discrimination.
[0260] When the aforementioned machine learning model is a binary classification model, the aforementioned formula one can be:
[0261] E = -P(x1)log b P(x1) + P(x2)log b P(x2). (Formula two);
[0262] Among them, x1 represents the first classification result, and x2 represents the second classification result. For example, in the anomaly detection scenario, x1 represents the classification result "normal", and x2 represents the classification result "abnormal". The meanings of other parameters can be referred to the aforementioned formula one.
[0263] Step 405: When the performance of the machine learning model after incremental training does not meet the performance compliance condition, the local point analysis device triggers the first analysis device to retrain the machine learning model.
[0264] The machine learning model after incremental training may have poor performance due to poor quality of training samples or other reasons. In this case, the first analysis device still needs to retrain the machine learning model. Generally, the first analysis device is an analysis device that supports offline training. The data volume of the training sample set it obtains is much larger than that of the first training sample set of the local point analysis device. The training duration that the first analysis device can perform is also much longer than the allowed training duration of the local point analysis device. The computing performance of the first analysis device is also greater than that of the local point analysis device. Therefore, when the performance of the machine learning model after incremental training does not meet the performance compliance condition, retraining the machine learning model by the first analysis device can obtain a machine learning model with better performance.
[0265] The action of evaluating whether the performance of the machine learning model after incremental training meets the performance compliance condition can be performed by the first analysis device. This process can refer to the process of evaluating whether the machine learning model meets the performance compliance condition in the aforementioned step 401. Among them, the machine learning model meeting the performance compliance condition means that the machine learning model has not deteriorated, and the machine learning model not meeting the performance compliance condition means whether the machine learning model has deteriorated. The action of evaluating whether the performance of the machine learning model after incremental training meets the performance compliance condition can also be performed by the evaluation device or the local point analysis device. This process can refer to the process of detecting whether the machine learning model has deteriorated in the aforementioned step 404. When the action of evaluating whether the performance of the machine learning model after incremental training meets the performance compliance condition is performed by other devices (such as the first analysis device or the evaluation device) other than the local point analysis device, after the other device completes the evaluation action, it needs to send the evaluation result to the local point analysis device for the local point analysis device to determine whether the performance of the machine learning model after incremental training meets the performance compliance condition. This application example will not be elaborated further.
[0266] Exemplarily, the process in which the local point analysis device triggers the first analysis device to retrain the machine learning model may include: the local point analysis device sends a retraining request to the first analysis device, and this retraining request is used to request the first analysis device to retrain the machine learning model; after receiving this retraining request, the first analysis device retrains the machine learning model based on the retraining request. In this case, the local point analysis device may also send a training sample set obtained from the corresponding local point network for the first analysis device to retrain the machine learning model based on the training sample set. This training sample set may be carried in the foregoing retraining request or may be sent to the first analysis device through independent information, and the embodiments of this application do not make any limitations in this regard. Correspondingly, the retraining process of the first analysis device may include the following two optional methods:
[0267] In the first optional method, after receiving the retraining request sent by the local point analysis device, the first analysis device may retrain the machine learning model based on the training sample set sent by the local point analysis device that sent the retraining request.
[0268] Among them, the training sample set sent by the local point analysis device is a training sample set obtained from the local point network corresponding to the local point analysis device, and it at least includes the foregoing first training sample set. In this way, retraining using a training sample set including the feature data obtained by the local point analysis device can make the trained machine learning model more adaptable to the needs of the local point analysis device itself, achieve model customization, and improve the application flexibility of the model.
[0269] In the second optional method, the first analysis device receives the retraining request sent by the local point analysis device and retrains the machine learning model based on the training sample set sent by the local point analysis device that sent the retraining request and the training sample sets sent by other local point analysis devices.
[0270] Among them, the training sample set used for retraining not only includes the training sample set obtained by the foregoing local point analysis device from the corresponding local point network, but also includes the training sample sets obtained by other local point analysis devices in their respective corresponding local point networks. Therefore, the sample sources of the training sample set used for retraining are more extensive, the data types are more diverse, the machine learning model obtained by retraining is more adaptable to the needs of multiple local point analysis devices, the generality of the model obtained by offline training is improved, thereby realizing model generalization, and the overall training cost of the first analysis device is reduced.
[0271] It should be noted that, in addition to the foregoing two optional methods, the local point analysis device may also not send a retraining request and only send the training sample set obtained from the corresponding local point network. Correspondingly, the first analysis device may also use the following third optional method for retraining:
[0272] In the third optional manner, the first analysis device receives the training sample sets sent by at least two local analysis devices, and retrains the machine learning model based on the received training sample sets.
[0273] Exemplarily, the first analysis device may retrain the machine learning model after receiving the training sample sets sent by a specified number of local analysis devices (for example, all local analysis devices that have established a communication connection with the first analysis device), or when the training cycle is reached, or when a sufficient number of training samples are obtained (that is, the number of obtained training samples is greater than the training data volume threshold).
[0274] In the foregoing three optional manners, there may be other timings for the local analysis device to send the training sample set (such as the first training sample set) obtained from the corresponding local network to the first analysis device. For example, the local analysis device may periodically upload the obtained training sample set; or, the local analysis device may upload the training sample set after receiving a sample set upload instruction sent by an operation and maintenance personnel or after receiving a sample set upload instruction sent by the first analysis device, and the sample set upload instruction is used to indicate uploading the obtained training sample set to the first analysis device. The first analysis device may retrain the machine learning model based on the collected training sample sets. This retraining process may be an offline training process or an incremental training process. The training samples used in this retraining process may be the same as or different from the training samples used in the previous training process, and the embodiments of the present application do not limit this.
[0275] It should be noted that in the embodiments of the present application, the foregoing steps 401 and 404 may be executed periodically, that is, this application scenario supports periodic offline training or incremental training. Among them, after the machine learning model of the first analysis device is evaluated and determined to meet the performance compliance condition, it may be sent to at least one local analysis device in the manner of step 402. For example, it is only sent to the local analysis device that sent the foregoing retraining request to the first analysis device; or, it is sent to the local analysis device that provides the training sample set for retraining; or, it is sent to all or specified local analysis devices that have established a communication connection with the first analysis device, and so on. For the local analysis device that receives the retrained machine learning model, if the machine learning model trained by itself by this local analysis device also meets the performance compliance condition, the local analysis device may screen the target machine learning model from the obtained machine learning models to screen a better machine learning model (such as the machine learning model with the highest performance index) for predicting the classification result. Usually, the local analysis device selects the latest machine learning model as the target machine learning model to adapt to the classification requirements of the current application scenario.
[0276] As described above, the machine learning model can be of multiple types. Among them, the tree model is a relatively common machine learning model. The tree model includes multiple associated nodes. For the convenience of readers' understanding, the embodiments of the present application briefly introduce the tree model. In the tree model, each node includes a node element and several branches pointing to subtrees; the subtree on the left side of a node is called the left subtree of the node, and the subtree on the right side of the node is called the right subtree; the root of the subtree of a node is called the child node of the node, also known as the child node; if a node is the child node of another node, then the other node is the parent node of this node, also known as the parent node; the depth or level of a certain node refers to the number of edges of the longest simple path from the root node to this node. For example, the depth (also known as the height or level) of the root node is 1, and the depth of the child node of the root is 2, and so on; leaf node: also called the terminal node, is a node with a degree of 0; the degree of a node refers to the number of subtrees of the node; non-leaf nodes are nodes other than leaf nodes, including the root node and the nodes between the root node and the leaf nodes. A binary tree is a tree structure in which each node has at most two subtrees, and it is a relatively common tree structure. The machine learning model in the embodiments of the present application can be a binary tree model. For example, the isolation forest model.
[0277] As Figure 6 shown, Figure 6 FIG. 1 is a schematic tree structure provided by an embodiment of the present application. The tree structure includes nodes P1 to P5, where P1 is the root node, P3, P4, and P5 are leaf nodes, P1 and P2 are non-leaf nodes, and the depth of the tree is 2. The machine learning model is formed by two node splits of nodes P1 and P3. Node split means that the training sample set corresponding to a node is divided into at most two subsets at a split point in a certain split dimension, which can be regarded as the node splitting out at most two child nodes, and each child node corresponds to a subset. That is, the way of dividing the training sample set corresponding to a node into child nodes is called splitting. In the actual application of the embodiments of the present application, there are multiple ways to represent the tree structure, Figure 6 FIG. 1 is only a schematic tree structure, and it can also have Figure 8 or Figure 9 other representation methods, etc. The present application does not limit the representation method of the tree structure.
[0278] Currently, when training a machine learning model using a training sample set, it is necessary to traverse the feature data of all feature dimensions in the training sample set, and then determine the split parameters of the non-leaf nodes in the machine learning model, such as the split dimension and the split point, and train the machine learning model based on the determined split parameters of the non-leaf nodes.
[0279] Since it is necessary to traverse the feature data in the training samples when training a machine learning model, and usually the amount of data in the training samples is very large, the training efficiency of the machine learning model is thus relatively low.
[0280] In the embodiments of the present application, when the aforementioned machine learning model is a tree model, on the basis of supporting incremental training of the machine learning model, the training efficiency of the machine learning model can also be improved. In the subsequent embodiments, taking the machine learning model as a tree model as an example, the foregoing steps will be explained. During the offline training or incremental training process of the tree model, the splitting of the tree model is involved, and its main principle is to cut (split) the space corresponding to one or more samples (also called the sample space). As mentioned above, each training sample includes the feature data of one or more features. In the tree model, the feature (i.e., the feature category) corresponding to a training sample is the dimension where the feature data of the training sample is located in the corresponding space. Therefore, in the tree model, considering the space concept, the feature corresponding to a training sample is also called the feature dimension. In the embodiments of the present application, a training sample includes one-dimensional or multi-dimensional feature data, which means that the training sample includes the feature data of one or more feature dimensions. For example, a training sample includes the feature data of two-dimensional features, which is also called that the training sample includes the feature data of two feature dimensions, and the space corresponding to the training sample is a two-dimensional space (i.e., a plane). Another example is referring to Table 1. In Table 1, a training sample includes the feature data of 4 feature dimensions, and the space corresponding to the training sample is a 4-dimensional space.
[0281] As Figure 7 shown, Figure 7 FIG. 10 is a schematic diagram of the splitting principle of a tree model provided by an embodiment of the present application. The tree model is split based on the Mondrian process, and it uses a random hyperplane to cut the data space. Each cut can generate two subspaces, and then continue to use a random hyperplane to cut each subspace, and so on until there is only one sample point in each subspace. In this way, clusters with a very high density can be cut many times before stopping, and points with a very low density are easily stopped in a subspace early, which finally corresponds to a leaf node of the tree. Figure 7 Suppose Figure 6 the samples in the training sample set corresponding to the machine learning model in include two-dimensional feature data, that is, the feature data of two feature dimensions, and the feature dimensions are the feature dimension x and the feature dimension y respectively. The training sample set includes the samples (a1, b1), (a1, b2), (a2, b1). The splitting dimension of the first node splitting is the feature dimension x, and the splitting point is a3. The sample space where the training sample set is located is cut into 2 subspaces, corresponding to Figure 6That is, the left subtree and the right subtree of the P1 node; the splitting dimension of the second node splitting is the feature dimension y, and the splitting point is b3, corresponding to Figure 6 That is, the left subtree and the right subtree of the P2 node. It can be seen from this that the sample set: {(a1, b1), (a1, b2), (a2, b1)} is respectively divided into three subspaces. Among them, when the foregoing feature data is the feature data of the time series, the foregoing feature dimension x and feature dimension y can be any two feature dimensions in the foregoing data features and / or extracted features (specific features thereof refer to the foregoing embodiments). For example, if the feature dimension x is the cycle duration in the data arrangement cycle and the feature dimension y is the moving average value in the statistical features, the feature data (a1, b1) refers to that the cycle duration is a1 and the moving average value is b1. When the foregoing feature data is data with certain features itself, such as network KPI data, the foregoing feature dimension x and feature dimension y can be any two KPI categories in the foregoing KPI categories (specific features thereof refer to the foregoing embodiments). For example, if the feature dimension x is network traffic and the feature dimension y is CPU utilization rate, the feature data (a1, b1) refers to that the network traffic is a1 and the CPU utilization rate is b1.
[0282] It should be noted that since the feature data is numerical data, each feature data has a corresponding value. In the embodiments of the present application, the value of the feature data is referred to as a feature value in the following text.
[0283] In the embodiments of the present application, each node in the machine learning model can correspondingly store node information, so that when the machine learning model is retrained later, node splitting can be performed based on the node information, and the classification result of the leaf node can be determined. By way of example, the node information of any node in the machine learning model includes label distribution information, and the label distribution information is used to reflect the proportion of the labels of different classes of samples in the historical training sample set divided into the corresponding node in the total number of labels, and the total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, and the historical splitting information is the information used for splitting the corresponding node. It should be noted that since the leaf node in the machine learning model is the node that has not been split currently and has no subtree, its historical splitting information is empty. During the retraining process, if it is split, the leaf node becomes a non-leaf node, and historical splitting information needs to be added to it.
[0284] Exemplarily, the foregoing historical split information includes: one or more of the position information of the corresponding node in the machine learning model, the split dimension, the split point, the numerical distribution range of the historical training sample set partitioned to the corresponding node, and the historical split cost. Among them, the position information of the corresponding node in the machine learning model is used to uniquely locate the node in the machine learning model. For example, this information includes: the layer number of the node, the identifier of the node, and / or the branch relationship of the node. The identifier of the node is used to uniquely identify the node in the machine learning model, which can be assigned to the node when the node is generated, and the identifier can be composed of numbers and / or characters. When any node has a parent node, the branch relationship of the any node includes the identifier of the parent node of the any node and the relationship description with the parent node. For example Figure 6 the branch relationship of node P2 in Figure 6 includes: (node P1: parent node); when any node has child nodes, the branch relationship of the any node includes the identifiers of the child nodes of the any node and the relationship description with the child nodes. For example Figure 6 the branch relationship of node P2 in Figure 6 also includes: (node P4: left child node, node P5: right child node). The split dimension of any node is the feature dimension for splitting in the historical sample data set partitioned to the any node, and the split point is the numerical point for splitting. In the example Figure 6 in Figure 6 , the split dimension of node P1 is x and the split point is a3; the split dimension of node P2 is y and the split point is b3. Each non-leaf node has only one split dimension and one split point. The numerical distribution range of the historical training sample set partitioned to the corresponding node is the distribution range of the feature values in the historical training sample set corresponding to the node. For example Figure 1 in Figure 1 , the numerical distribution range of node P3 is [a3, a2], which can also be expressed as a3 - a2. The historical split cost is the split cost determined by the corresponding node based on the numerical distribution range of the historical training sample set. For a specific explanation, please refer to the following text.
[0285] The foregoing label distribution information includes: the number of labels of the same category in the historical training sample set divided into the corresponding node and the total number of the foregoing labels; or, the proportion of the labels of different categories in the historical training sample set divided into the corresponding node in the total number of labels. Wherein, the total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into any one of the nodes, and the ratio of the number of labels of the same category in the historical training sample set to the total number of labels is the proportion of the labels of this category in the total number of labels. For example, in an anomaly detection scenario, there are 10 labels for the samples divided into node P1. Among them, there are 2 "normal" labels 0 and 8 anomaly labels "1". Then the label distribution information includes: label 0: 2; label 1: 8, total number of labels: 10 (based on this, the proportion of the labels of different categories in the historical training sample set divided into the corresponding node in the total number of labels can be determined). Or, the label distribution information includes: label 0: 20%; label 1: 80%. It should be noted that here is only a schematic introduction to the representation method of the label distribution information. In actual implementation, there may be other ways of representing the label distribution information, which is not limited in the embodiments of the present application.
[0286] Further, the node information of the leaf node may further include a classification result, that is, the label finally determined for the leaf node. For example Figure 6 the node information of nodes P4 and P5 in
[0287] By storing the node information for each node, complete training information can be provided for subsequent model training, reducing the complexity of obtaining relevant information during model training and improving the model training efficiency.
[0288] Especially when the node information includes the numerical distribution range of the historical training sample set, during the subsequent retraining process, effective model retraining can be performed only based on the numerical distribution range of the historical training sample set, without obtaining the actual numerical values of the feature data in the historical training sample set, effectively reducing the training complexity.
[0289] Assume that the machine learning model is a binary tree model, and the offline training process is the process of establishing the machine learning model. Then the training process of the machine learning model in the foregoing step 401 includes:
[0290] Step A1, obtain the historical training sample set with determined labels.
[0291] In one case, the labels of the samples in the historical training sample set may not be labeled when acquired by the first analysis device. In this case, the first analysis device can present the samples to the operation and maintenance personnel for label annotation. In another case, the labels of the samples in the historical training sample set have been labeled when acquired by the first analysis device. For example, for the samples acquired by the first analysis device from the aforementioned storage device, the first analysis device can directly use this training sample set for model training.
[0292] Step A2: Create a root node.
[0293] Step A3: Take the root node as the third node and execute the offline training process until the splitting cut-off condition is reached. The offline training process includes:
[0294] Step A31: Split the third node to obtain the left child node and the right child node of the third node.
[0295] Step A32: Take the left child node as the updated third node, divide the historical training sample set into the left sample set of the left child node as the updated historical training sample set, and execute the offline training process again.
[0296] Step A33: Take the right child node as the updated third node, divide the historical training sample set into the right sample set of the right child node as the updated historical training sample set, and execute the offline training process again.
[0297] Step A4: Determine the classification result for each leaf node to obtain a machine learning model.
[0298] In the embodiments of the present application, when the third node in the foregoing steps reaches the splitting cut-off condition, it has no child nodes and thus can be used as a leaf node. When a node is a leaf node, the classification result of the leaf node can be determined based on the number of labels of the same category in the historical training sample set divided into this node and the total number of labels in the historical training sample set divided into this node; or based on the proportion of the labels of different categories in the total number of labels in the historical training sample set. The determination method of this classification result is still based on the aforementioned probability theory principle, that is, the label with the highest proportion or the largest number is used as the final classification result. Among them, the proportion of any label in the total number of labels is the ratio of the total number of this any label to the total number of labels. For example, if the number of "abnormal" labels in a leaf node is 7 and the number of "normal" labels is 3, then the proportion of the "abnormal" label in the total number of labels is 70%, and the proportion of the "normal" label in the total number of labels is 30%. The final classification result is "abnormal". Save this classification result in the node information of the leaf node.
[0299] In the traditional iForest model, the classification result is calculated based on the average height of leaf nodes on each tree. In the embodiments of the present application, the classification result of the leaf node is to use the label with the highest proportion or the largest number as the final classification result, which is relatively accurate and has a small computational cost.
[0300] In the foregoing step A31, the third node can be split based on the numerical distribution range of the historical training sample set to obtain the left child node and the right child node of the third node. The numerical distribution range of the historical training sample set reflects the density of the samples in the historical training sample set. When the sample distribution is relatively dispersed, the numerical distribution range is larger; when the sample distribution is relatively concentrated, the numerical distribution range is smaller.
[0301] In the embodiments of the present application, the samples in the historical training sample set may include at least one-dimensional feature data, and the numerical distribution range of the historical training sample set is the distribution range of the feature values in the historical training sample set, that is, the numerical distribution range of the historical training sample set can be characterized by the minimum value and the maximum value of the feature values on each feature dimension. In the embodiments of the present application, the feature data included in the samples in the historical training sample set is numerical data, such as decimal values, binary values, or vectors. For example, the samples in the historical training sample set include one-dimensional feature data, and the feature values included in the historical training sample set are: 1, 3,..., 7, 10, where the minimum value is 1 and the maximum value is 10. Then the numerical distribution range of the historical training sample set is [1, 10], which can also be expressed as 1 - 10.
[0302] The feature data included in the samples in the historical training sample set may originally be numerical data, or may be obtained by converting non-numerical data through a specified algorithm. For example, feature dimensions such as data change trend, data fluctuation, statistical features, or fitting features, which cannot be initially represented numerically, can be converted into numerical data through a specified algorithm. For example, for the feature data: high, it can be converted into numerical data: 2; for the feature data: medium, it can be converted into numerical data: 1; for the feature data: low, it can be converted into numerical data: 0. Using the historical training sample set including numerical values for node splitting can simplify the computational complexity and improve the operation efficiency.
[0303] Exemplarily, the process of splitting the third node based on the numerical distribution range of the historical training sample set to obtain the left child node and the right child node of the third node may include:
[0304] Step A311: Determine the third splitting dimension among the feature dimensions of the historical training sample set.
[0305] In the first alternative implementation, the third splitting dimension is a randomly selected feature dimension among the feature dimensions of the historical training sample set.
[0306] In the second alternative implementation, the third splitting dimension is the feature dimension with the largest span among the feature dimensions of the historical training sample set. For example, the span of the feature values on each feature dimension is the difference between the maximum and minimum values of the feature values on that feature dimension.
[0307] For example, the feature dimensions of the historical training sample set can be sorted in descending order of span first, and then the feature dimension corresponding to the span with the topmost sorting is selected as the third splitting dimension.
[0308] For a feature dimension with a large span, the probability of its splitting is relatively high. Splitting a node on this feature dimension can accelerate the model convergence speed and avoid splitting of ineffective feature dimensions. Therefore, by selecting the feature dimension with the largest span as the third splitting dimension, the probability of effective splitting of the machine learning model can be increased, and the overhead of node splitting can be saved.
[0309] As Figure 8 shown, it is assumed in the embodiment of the present application that the samples in the historical training sample set include two-dimensional feature data, which corresponds to a two-dimensional space, and there are two feature dimensions, x1 and x2. The span ranges of the feature values on each feature dimension are [x1_min, x1_max] and [x2_min, x2_max] respectively. The corresponding spans are x1_max – x1_min and x2_max – x2_min. Comparing these two spans, assuming x1_max – x1_min > x2_max – x2_min, then the feature dimension x1 is selected as the third splitting dimension.
[0310] Based on the same principle as the aforementioned second alternative implementation, the third splitting dimension is the feature dimension with the largest span ratio among the feature dimensions of the historical training sample set. The span ratio d of the feature values on any feature dimension satisfies the ratio formula: d = h / z, where h is the span of the historical training sample set on this any feature dimension, and z is the sum of the spans of the feature values on each feature dimension.
[0311] Taking Figure 8 as an example, then z = (x1_max – x1_min) + (x2_max – x2_min), the span ratio dx1 of the feature dimension x1 is (x1_max – x1_min) / z, and the span ratio dx2 of the feature dimension x2 is (x2_max – x2_min) / z; assuming dx1 > dx2, then the feature dimension x2 is selected as the third splitting dimension.
[0312] Step A312: Determine a third splitting point on the third splitting dimension of the historical training sample set.
[0313] Exemplarily, the third splitting point is a numerical point randomly selected on the third splitting dimension of the historical training sample set. This can achieve an equiprobable split on the third splitting dimension.
[0314] In the actual implementation of the embodiments of the present application, other methods may also be used to select the third splitting node, and the embodiments of the present application do not limit this.
[0315] Step A313: Split the third node based on the third splitting point of the third splitting dimension, where the numerical range in the third numerical distribution range whose value on the third splitting dimension is not greater than the value of the third splitting point is divided into the left child node, and the numerical range in the third numerical distribution range whose value on the third splitting dimension is greater than the value of the third splitting point is divided into the right child node.
[0316] The third numerical distribution range is the distribution range of the feature values in the historical training sample set, which is composed of the span ranges of the feature values of the historical training set on each feature dimension. In this way, when performing node splitting, only the minimum and maximum values of the numerical values on each feature dimension need to be obtained, the amount of data obtained is small, the calculation is simple, and the training efficiency of the model is relatively high.
[0317] Still taking... Figure 8 as an example, assuming that the selected third splitting point is x1_value ∈ [x1_min, x1_max], then the numerical range on the feature dimension x1 that is not greater than x1_value, that is, [x1_min, x1_value], is divided into the left child node P2 of the third node P1, and the numerical range on the feature dimension x1 that is greater than x1_value, that is, [x1_value, x1_max], is divided into the right child node P3 of the third node P1.
[0318] Since the splitting of the foregoing nodes is only based on the numerical distribution range of the feature data rather than the feature data itself, when performing node splitting, only the minimum and maximum values of the numerical values on each feature dimension need to be obtained, the amount of data obtained is small, the calculation is simple, and the training efficiency of the model is relatively high.
[0319] It should be noted that since the first analysis device has obtained the foregoing historical training sample set for training, the sample can also be directly used for node splitting. The method of using the numerical distribution range for node division in step A313 can be replaced by: dividing the samples in the historical training sample set whose feature values on the third splitting dimension are not greater than the value of the third splitting point into the left child node, and dividing the samples in the historical training sample set whose feature values on the third splitting dimension are greater than the value of the third splitting point into the right child node.
[0320] If no restrictions are added to the splitting of the machine learning model, it may cause the depth of the machine learning model to increase without limit, and it may keep iterating until each leaf node has sample points with the same label or only contains one sample point, and then the splitting of the machine learning model stops. By setting a splitting cut-off condition in the embodiments of the present application, the depth of the tree can be controlled to avoid over-splitting of the tree.
[0321] Optionally, the splitting cut-off condition includes at least one of the following:
[0322] Condition 1: The current splitting cost of the third node is greater than the splitting cost threshold.
[0323] Condition 2: The number of samples in the historical training sample set is less than the second sample number threshold.
[0324] Condition 3: The splitting times corresponding to the third node are greater than the splitting times threshold.
[0325] Condition 4: The depth of the third node in the machine learning model is greater than the depth threshold.
[0326] Condition 5: The proportion of the number of the label with the largest proportion in the labels corresponding to the historical training sample set in the total number of labels corresponding to the historical training sample set is greater than the specified proportion threshold.
[0327] For the aforementioned Condition 1, the embodiments of the present application propose a concept of splitting cost. During the offline training process of the machine learning model, the current splitting cost of any node is negatively correlated with the size of the numerical distribution range of the training sample set of any node, and the training sample set of any node is the set of samples obtained by dividing the training sample set for training the machine learning model to any node. Exemplarily, the current splitting cost of any node is the reciprocal of the sum of the spans of the feature values of the samples in the training sample set of any node in each feature dimension. Then, for the third node, the current splitting cost of the third node is negatively correlated with the size of the distribution range of the feature values in the historical training sample set, that is, the larger the numerical distribution range, the smaller the splitting cost. The current splitting cost of the third node is the reciprocal of the sum of the spans of the feature values of the historical training sample set in each feature dimension. Exemplarily, the splitting cost threshold can be positive infinity.
[0328] The current splitting cost of the third node satisfies the cost calculation formula:
[0329] Where max j - min j represents the maximum value minus the minimum value of the feature values in the j-th feature dimension in the numerical distribution range of the historical training sample set, that is, the span of the feature values in this feature dimension, and N is the total number of feature dimensions.
[0330] Then asFigure 8 , the splitting cost of node P1 is 1 / z, where z = (x1_max – x1_min) + (x2_max – x2_min). As Figure 9 , during the incremental training process, the splitting cost can be calculated each time a node is split and compared with the splitting cost threshold. Figure 9 It is assumed that the initial value of the splitting cost is 0 and the splitting cost threshold is positive infinity. The calculated splitting costs along the depth direction of the tree are 0, COST1 (the first node split), COST2 (the second node split), and COST3 (the third node split), etc. Among them, the number of splits is positively correlated with the splitting cost. The more splits, the higher the splitting cost.
[0331] By adopting this condition 1, when the splitting cost reaches a certain level, no more node splitting is performed, which can avoid excessive splitting of the tree and reduce the computational overhead.
[0332] For the aforementioned condition 2, when the number of samples in the historical training sample set of the third node is less than the second sample number threshold, it indicates that the amount of data in the historical training sample set is already small and is no longer sufficient to support effective node splitting. At this time, the offline training process is stopped, which can reduce the computational overhead. For example, the second sample threshold can be 2 or 3.
[0333] For the aforementioned condition 3, the splitting times corresponding to the third node are the total number of splits from the first split of the root node to the current split of this third node. When the splitting times corresponding to the third node are greater than the splitting times threshold, it means that the current machine learning model has reached the upper limit of the splitting times. At this time, the offline training process is stopped, which can reduce the computational overhead.
[0334] For the aforementioned condition 4, when the depth of the third node in the machine learning model is greater than the depth threshold, the offline training process is stopped, which can achieve the control of the depth of the machine learning model.
[0335] For condition 5, when the proportion of the number of the label with the largest proportion in the labels corresponding to the historical training sample set in the total number of labels is greater than the specified proportion threshold, it indicates that the number of the label with the largest proportion has reached the classification condition, and the accurate classification result can be determined based on this. At this time, the offline training process is stopped, which can reduce unnecessary splits and reduce the computational overhead.
[0336] Optionally, in the foregoing step 404, the process of the local point analysis device incrementally training the machine learning model based on the first training sample set is actually a process of sequentially inputting multiple training samples in the first training sample set into the machine learning model (that is, inputting one training sample at a time), and performing multiple training processes. Each training process is the same, and each training process is actually a traversal process of nodes. For any training sample in the first training sample set, this traversal process is executed. In the embodiment of the present application, taking the foregoing first training sample as an example, assuming that the first training sample is any training sample in the first training sample set, and it includes feature data of one or more feature dimensions. Referring to the structure of the samples in the foregoing historical training sample set, the feature data in the first training sample set is numerical data, which may originally be numerical data or may be obtained by converting non-numerical data through a specified algorithm. Assuming that the first node is any non-leaf node in the machine learning model, starting from the root node of the machine learning model and traversing, the traversal process in step 404 is described as follows:
[0337] Step B1, when the current splitting cost of the first node being traversed is less than the historical splitting cost of the first node, add an associated second node, where the second node is the parent node or the child node of the first node.
[0338] Among them, in the incremental training process, the current splitting cost of the first node is the cost of splitting the node by the first node based on the first training sample (that is, the cost of adding an associated second node to the first node. In this case, the node splitting of the first node refers to adding a new branch to the first node), and the historical splitting cost of the first node is the cost of splitting the node by the first node based on the historical training sample set of the first node. The historical training sample set of the first node is the set of samples divided to the first node in the historical training sample set of the machine learning model. Then referring to the foregoing 401, if the current incremental training is the first incremental training after receiving the machine learning model, and the first node is any of the foregoing third nodes, the historical training sample set of the first node is the historical training sample set corresponding to the third node.
[0339] In the embodiment of the present application, the current splitting cost of the first node and the historical splitting cost of the first node can be directly compared to add an associated second node when the current splitting cost of the first node is less than the historical splitting cost of the first node; further, the difference obtained by subtracting the historical splitting cost of the first node from the current splitting cost of the first node can be obtained first, and it can be determined whether the absolute value of the difference is greater than a specified difference threshold, so as to ensure that node splitting is performed only when the current splitting cost of the first node is much less than the historical splitting cost of the first node, which can save training costs and improve training efficiency.
[0340] Among them, the current splitting cost of the first node is negatively correlated with the size of the first numerical distribution range, and the first numerical distribution range is a distribution range determined based on the feature values in the first training samples and the second numerical distribution range. The second numerical distribution range is the distribution range of the feature values in the historical training sample set of the first node. Optionally, the first numerical distribution range is a distribution range determined based on the union of the first training samples and the second numerical distribution range. For example, the samples in the historical training sample set of the first node include feature data of two feature dimensions. The span range of the feature values in feature dimension x is [1, 10], and the span range of the feature values in feature dimension y is [5, 10]. Then the second numerical distribution range includes the span range of the feature values on feature dimension x: [1, 10] and the span range of the feature values on feature dimension y: [5, 10]; if the feature value of the first training sample on feature dimension x is 9 and the feature value on feature dimension y is 13, then the union of the first training sample and the second numerical distribution range is obtained for different feature dimensions respectively. Then the span range of the feature values of the first numerical distribution range on each feature dimension includes: [1, 10] on feature dimension x and [5, 13] on feature dimension y.
[0341] Exemplarily, the current splitting cost of the first node is the reciprocal of the sum of the spans of the feature values of the first numerical distribution range on each feature dimension. The calculation method of the current splitting cost of the first node can refer to the calculation method of the current splitting cost of the foregoing third node. For example, it is calculated using the foregoing cost calculation formula (i.e., formula three), except that the corresponding numerical distribution range in the formula is replaced by the first numerical distribution range from the foregoing historical training sample set. This embodiment of the present application will not elaborate further.
[0342] Exemplarily, the historical splitting cost of the first node is the reciprocal of the sum of the spans of the feature values of the samples in the historical training sample set of the first node on each feature dimension. The calculation method of the historical splitting cost of the first node can refer to the calculation method of the current splitting cost of the foregoing third node. For example, it is calculated using the foregoing cost calculation formula, except that the corresponding numerical distribution range in the formula is replaced by the numerical distribution range of the historical training sample set of the first node. This embodiment of the present application will not elaborate further.
[0343] Among them, the process of adding an associated second node may include:
[0344] Step B11: Determine the span range of the feature values of the first numerical distribution range on each feature dimension.
[0345] Step B12: Add a second node based on the first split point on the first split dimension. Among them, the numerical range in the first numerical distribution range whose value on the first split dimension is not greater than the value of the first split point is divided into the left child node of the second node, and the numerical range in the first numerical distribution range whose value on the first split dimension is greater than the value of the first split point is divided into the right child node of the second node.
[0346] Take, for Figure 10 example. Assume that the split point of the selected second node P4 is y1_value ∈ [y1_min, y1_max]. Then, the numerical range on the feature dimension y1 that is less than or equal to y1_value in the first numerical distribution range, that is, [y1_min, y1_value], is divided into the left child node P5 of the second node P4, and the numerical range on the feature dimension y1 that is greater than y1_value in the first numerical distribution range, that is, [y1_value, y1_max], is divided into the right child node P6 of the second node P4.
[0347] The process of splitting the second node can refer to the process of splitting the third node in the foregoing step A313, and this embodiment of the present application will not elaborate on this.
[0348] Among them, the foregoing first split dimension is the split dimension determined among each feature dimension based on the span range of the feature values on each feature dimension, and the first split point is the numerical point for splitting determined on the first split dimension of the first numerical distribution range.
[0349] Exemplarily, in the first optional manner, the first split dimension is a randomly selected feature dimension among each feature dimension of the first numerical distribution range. In the second optional manner, the first split dimension is the feature dimension with the largest span among each feature dimension of the first numerical distribution range. The corresponding principle can refer to the foregoing step A311, and this embodiment of the present application will not elaborate further.
[0350] Optionally, the first split point is a randomly selected numerical point on the first split dimension of the first numerical distribution range. This can achieve an equal-probability split on the first split dimension.
[0351] In the actual implementation of this embodiment of the present application, other methods can also be used to select the first split node, and this embodiment of the present application does not limit this.
[0352] In the first case, when the first split dimension is different from the second split dimension, the second node is the parent node or child node of the first node, that is, the second node is located above or below the first node.
[0353] If the second splitting dimension is the historical splitting dimension of the first node in the machine learning model, and the second splitting point is the historical splitting point of the first node in the machine learning model, then referring to the foregoing steps A311 and A312, when the first node is any of the foregoing third nodes, the second splitting dimension is the foregoing third splitting dimension, and the second splitting point is the foregoing third splitting point.
[0354] In the second case, when the first splitting dimension is the same as the second splitting dimension, and the first splitting point is to the right of the second splitting point, the second node is the parent node of the first node, and the first node is the left child node of the second node.
[0355] In the third case, when the first splitting dimension is the same as the second splitting dimension, and the first splitting point is to the left of the second splitting point, the second node is the left child node of the first node.
[0356] As before, the node information of each non-leaf node may include a splitting dimension and a splitting point. For the convenience of description, the subsequent embodiments of this application represent that the splitting dimension is u and the splitting point is v in the format of "u>v".
[0357] For the first case described above, as Figure 9 and Figure 11 shown, Figure 11 is a machine learning model before adding the second node, including nodes Q1 and Q3, Figure 9 is Figure 11 the machine learning model after adding the second node to the shown machine learning model. Assume that node Q1 is the first node, its second splitting dimension is x2, and its second splitting point is 0.2; Q2 is the second node, its first splitting dimension is x1, and its first splitting point is 0.7. Since the splitting dimensions of the first node and the second node are different, as Figure 9 shown, the newly added second node becomes the parent node of the first node.
[0358] As Figure 9 and Figure 12 shown, Figure 12 is another machine learning model before adding the second node, including nodes Q1 and Q2, Figure 9 is Figure 12 the machine learning model after adding the second node to the shown machine learning model. Assume that node Q1 is the first node, its second splitting dimension is x2, and its second splitting point is 0.2; Q3 is the second node, its first splitting dimension is x1, and its first splitting point is 0.4. Since the splitting dimensions of the first node and the second node are different, the newly added second node becomes the child node of the first node.
[0359] For the second case described above, as Figure 13 and Figure 14 shown, Figure 13A machine learning model before adding a second node, including nodes Q4 and Q6, Figure 14 is Figure 13 the machine learning model after adding a second node to the machine learning model shown. Assume that node Q4 is the first node, its second splitting dimension is x1, and the second splitting point is 0.2; Q5 is the second node, its first splitting dimension is x1, and the first splitting point is 0.7. Since the splitting dimensions of the first node and the second node are the same, and the first splitting point is on the right side of the second splitting point, the newly added second node serves as the parent node of the first node, and the first node is the left child node of the second node.
[0360] For the aforementioned third case, as Figure 13 and Figure 15 shown, Figure 13 A machine learning model before adding a second node, including nodes Q4 and Q6, Figure 15 is Figure 13 the machine learning model after adding a second node to the machine learning model shown. Assume that node Q4 is the first node, its second splitting dimension is x1, and the second splitting point is 0.2; Q7 is the second node, its first splitting dimension is x1, and the first splitting point is 0.1. Since the splitting dimensions of the first node and the second node are the same, and the first splitting point is on the left side of the second splitting point, the newly added second node serves as the left child node of the first node.
[0361] It should be noted that after the aforementioned second node is added, the child nodes that are not the child nodes of the first node are leaf nodes, and the classification results of these leaf nodes need to be determined. That is, when the second node is the parent node of the first node, the other child node of the second node is a leaf node; when the first node is the parent node of the second node, both child nodes of the second node are leaf nodes.
[0362] The method for determining the classification results of leaf nodes during the incremental training process can refer to the method for determining the classification results of leaf nodes during the aforementioned offline training process. Based on the number of labels of the same category of samples in the historical training sample set and the total number of labels, the classification results of the leaf nodes are determined. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the leaf nodes; or, based on the proportion of the labels of different categories of samples in the historical training sample set in the total number of labels, the classification results of the leaf nodes are determined. This application example will not elaborate on this.
[0363] As described above, node information is stored corresponding to each node in the machine learning model. In the incremental training process, historical splitting information obtained from the node information of the first node, such as the splitting dimension, splitting point, and numerical distribution range of the historical training sample set, is used to determine whether to add a second node to the first node to achieve fast incremental training; when determining that a node is a leaf node, the classification result of the leaf node can be quickly determined based on the label distribution information in the node information.
[0364] Furthermore, in the offline training process, after determining each node, corresponding node information can be stored for each node, or after the entire machine learning training is completed, corresponding node information can be stored for each node for subsequent retraining. In the incremental training process, after adding the second node, it is necessary to save the corresponding node information for the second node. Since the purpose of adding the second node is to separate samples of different categories, the branch where the newly added second node is located is a branch that does not exist in the branches of the original machine learning model and belongs to a newly added branch, so it does not affect the original branch distribution. Therefore, for nodes having a connection relationship with the second node, such as the position information in the node information corresponding to the parent node or child node is updated correspondingly, and other information in the node information remains unchanged. In this way, the incremental training of the machine learning model can be carried out while minimizing the impact on other nodes.
[0365] It is worth noting that before adding the associated second node in step B1, it is also possible to detect whether the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is greater than the first sample number threshold. When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is greater than the first sample number threshold, add the second node. In each incremental training process, the number of the first training samples is 1.
[0366] When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is not greater than the first sample number threshold, stop the incremental training of the machine learning model. That is, do not execute the above step of adding the associated second node. In this way, the second node is added and node splitting is performed only when the number of samples is relatively large, so as to avoid invalid node splitting and reduce the overhead of computing resources. And since splitting nodes with too few samples may lead to a decline in the prediction performance of the machine learning model, by setting this first sample number threshold, the accuracy of the model can be guaranteed.
[0367] Step B2: When the current splitting cost of the first node is not less than the historical splitting cost of the first node, traverse each node in the subtree of the first node, and determine the traversed node as the new first node, and execute the traversal process again until the current splitting cost of the traversed first node is less than the historical splitting cost of the first node, or the target depth is traversed. At this time, stop the traversal process. It should be noted that when the current splitting cost of the first node is not less than the historical splitting cost of the first node, and the subtree of the first node is traversed, that is, when the first node is a leaf node, the above traversal process is also stopped.
[0368] For the updated first node, when the current splitting cost of the first node is less than the historical splitting cost of the first node, add a second node associated with the first node. The process of adding the second node can refer to the aforementioned Step B1, and this application embodiment will not elaborate on it. Stopping the traversal process when reaching the target depth can avoid over-splitting of the tree model and prevent the tree from having too many levels.
[0369] It is worth noting that in the incremental training process, the historical training sample set of the machine learning model refers to the training sample set before the current training process, which is relative to the first training sample input currently. For example, if this incremental training process is the first incremental training process after Step 401, then the historical training sample set described in this incremental training process is the same as the historical training sample set in the aforementioned Step 401; if this incremental training process is the wth (w is an integer greater than 1) incremental training process after Step 401, then the historical training sample set described in this incremental training process is the set of the historical training sample set in the aforementioned Step 401 and the training samples input in the previous w - 1 incremental training processes.
[0370] The embodiment of this application provides a training method. Through the aforementioned incremental training method, online incremental training of the machine learning model can be achieved. And since each node stores node information correspondingly, incremental training can be carried out without obtaining a large number of samples, thus realizing a lightweight machine learning model.
[0371] It is worth noting that the analysis device maintaining the machine learning model can streamline the machine learning model when the model streamlining condition is met, making the structure of the streamlined machine learning model simpler and having higher operation efficiency when performing predictions. In the embodiment of this application, the principle of model streamlining is actually the principle of finding connected domains, that is, merging the divided spaces in the machine learning model that can belong to the same connected domain. This model streamlining process includes:
[0372] Merge the first non-leaf node and the second non-leaf node in the machine learning model, and merge the first leaf node and the second leaf node to obtain a refined machine learning model for predicting classification results. Herein, the first leaf node is the child node of the first non-leaf node, the second leaf node is the child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets allocated on the same feature dimension are adjacent.
[0373] For example Figure 16 , Figure 16 Assume that the samples in the training sample set corresponding to the machine learning model include feature data of two feature dimensions, namely feature dimension x and feature dimension y. The training sample set includes: samples M(a1, b1), N(a1, b2), Q(a2, b1), U(a4, b4). The splitting dimension of the first node splitting is feature dimension x, and the splitting point is a3, which cuts the sample space where the two-dimensional samples are located into 2 subspaces, corresponding to Figure 16 which are the left subtree and the right subtree of node Q8; the splitting dimension of the second node splitting is feature dimension y, and the splitting point is b3, corresponding to Figure 16 which are the left subtree and the right subtree of node Q9; the splitting dimension of the third node splitting is feature dimension y, and the splitting point is b4, corresponding to Figure 16 which are the left subtree and the right subtree of node Q10. It can be seen therefrom that the spaces where samples M(a1, b1), N(a1, b2), Q(a2, b1), U(a4, b4) are located are divided into 4 subspaces, namely subspaces 1-4. From Figure 16 the space splitting schematic diagram on the right, it can be seen that the classification results of the leaf nodes Q91 and Q101 corresponding to subspaces 3 and 4 are both c, and the span ranges of their feature values are adjacent, so the two can be merged to form a connected domain, and the merged subspace does not affect the actual classification result of the machine learning model. From Figure 16 the machine learning model on the left, it can be seen that the leaf nodes Q91 of the non-leaf node Q9 and the leaf nodes Q101 of the non-leaf node Q10 have the same label, both c, and the span ranges of the feature values on the y-axis are respectively adjacent [b4, b3] and [b3, b2]. Therefore, the non-leaf node Q9 and the non-leaf node Q10 are respectively the aforementioned first non-leaf node and the second non-leaf node, and their leaf nodes are respectively the aforementioned first leaf node and the second leaf node. For example Figure 17As shown, the final subspace 3 and subspace 4 are merged to form a new subspace 3, the non-leaf node Q9 and the non-leaf node Q10 are merged to form a new non-leaf node Q12, and the leaf node Q91 and the leaf node Q101 are merged to form a new leaf node Q121. The corresponding node information is also merged. Among them, the merging of node information is actually taking the union of the corresponding parameters (i.e., parameters of the same type) in the node information. For example, the span ranges [b4, b3] and [b3, b2] on the y-axis mentioned above are merged into [b4, b2].
[0374] The refined machine learning model structure is simpler, reducing the number of tree branch layers and preventing the tree from being too deep. Although the model architecture changes, it does not affect its prediction results, can save storage space, and improve prediction efficiency. And through the refinement process, overfitting of the model can be prevented. Optionally, this refinement process can be executed periodically, and this refinement process needs to be executed in the order from the bottom layer to the top layer (also known as from large depth to small depth) of the machine learning model.
[0375] This model refinement process can be executed by the first analysis device after the aforementioned step 401 or 405. The refined machine learning model can be sent to the local analysis device for the local analysis device to perform sample analysis based on this machine learning model, that is, to predict the classification result. Since the size of the refined model itself (i.e., the memory size occupied by the model itself) becomes smaller, when using this model to predict the classification result, the prediction speed is faster than that of the unrefined model, the prediction efficiency is higher, and the transmission overhead of this model is also reduced accordingly. Further, if the refined model is only used for sample analysis, historical split information can not be recorded in its node information, which can further reduce the size of the model itself and improve the model prediction efficiency.
[0376] It should be noted that in the foregoing Figures 8 to 17 , when the foregoing feature data is feature data of a time series, any feature dimension is any one of the foregoing data features and / or extracted features (the specific features refer to the foregoing embodiments). When the foregoing feature data is data with certain features itself, such as network KPI data, any feature dimension is any one of the foregoing KPI categories. For example, when the delay data is used as the feature data, the feature dimension of this feature data is delay. Another example is that when the packet loss rate data is used as the feature data, the feature dimension of this feature data is the packet loss rate. The specific definition can refer to the foregoing embodiments and the explanations in Figure 7 . This application embodiment will not elaborate on this.
[0377] Since the model needs to use a machine learning model with a complete structure during training, rather than a streamlined machine learning model. It can be seen that the machine learning model sent to the local point analysis device in the aforementioned step 402 is the unstreamlined machine learning model directly obtained in step 401, so as to support the local point analysis device to perform incremental training on this machine learning model. In another alternative, the machine learning model sent to the local point analysis device in the aforementioned step 402 can also be a streamlined machine learning model, but this machine learning model needs to additionally carry unmerged node information, so that the local point analysis device can restore the unstreamlined machine learning model based on the streamlined machine learning model and the unmerged node information to perform incremental training on this machine learning model.
[0378] The model streamlining process can also be executed by the local point analysis device after the aforementioned step 404. The streamlined machine learning model can be used for sample analysis, that is, the prediction of classification results. When incremental training is performed again later, the model used is the unstreamlined machine learning model.
[0379] It should be noted that the foregoing embodiments of the present application take the local point analysis device directly performing incremental training on the machine learning model based on the first training sample set obtained from the local point analysis device as an example for illustration. In the actual implementation of the embodiments of the present application, the foregoing local point analysis device can also indirectly perform incremental training on the machine learning model based on the first training sample set obtained from the local point analysis device. In one implementation, the local point analysis device can send the current machine learning model and the first training sample set to the first analysis device, and the first analysis device performs incremental training on the machine learning model based on the first training sample set and sends the trained machine learning model to the local point analysis device. This incremental training process can refer to the aforementioned step 404, and the embodiments of the present application will not be elaborated; in another implementation, the local point analysis device can send the first training sample set to the first analysis device, and the first analysis device integrates the first training sample set and the historical training samples used to train this machine learning model to obtain a new historical training sample, and based on this historical training sample, performs offline training on the initial machine learning model, and its training result is the same as the result of performing incremental training based on the first training sample set. This offline training process can refer to the aforementioned step 401, and the embodiments of the present application will not be elaborated.
[0380] In the traditional model training method, once the machine learning model is deployed on the local analysis device after offline training, incremental training cannot be performed. However, in the model training method provided by the embodiments of the present application, the machine learning model supports incremental training and can have good self-adaptability to new training samples. Especially for the anomaly detection scenario, it can have good self-adaptability to the emergence of new anomaly patterns and samples with new labels, and the trained model can accurately detect different anomaly patterns. Thus, the generalization of the model is realized, the prediction performance is guaranteed, and the user experience is effectively improved.
[0381] Furthermore, if the traditional model training method is applied to the application scenario provided by the embodiments of the present application, a large number of samples need to be collected on the local analysis device where the machine learning model is deployed, and batch training of the samples is performed. Since the training of the machine learning model needs to be able to access historical training samples, a large number of historical training samples also need to be stored, resulting in the consumption of a large amount of memory and computing resources, and the training cost is relatively high.
[0382] In the embodiments of the present application, during the incremental training or offline training process, node splitting is performed based on the numerical distribution range of the training sample set, and there is no need to access a large number of historical training samples. Therefore, the occupation of memory and computing resources is effectively reduced, and the training cost is reduced. And through the foregoing node information, relevant information of each node can be carried, which can realize the lightweight of the machine learning model, make it more convenient for the deployment of the machine learning model, and realize the effective generalization of the model.
[0383] As Figure 18 and Figure 19 shown, Figure 18 is a schematic diagram of the incremental training effect of a traditional machine learning model, Figure 19 and Figure 18 is a schematic diagram of the incremental training effect of the machine learning model provided by the embodiments of the present application. The horizontal axis represents the percentage of the input training samples in the total amount of training samples, and the vertical axis represents the performance index reflecting the model performance. The larger the value of this index, the better the performance of the model. Figure 19 In
[0384] The sequence of steps of the model training method provided by the embodiments of the present application can be appropriately adjusted, and the steps can also be increased or decreased accordingly according to the situation. Any person skilled in the art in the technical field disclosed in the present application can easily think of the changed methods, which should be covered within the protection scope of the present application, so they will not be elaborated here.
[0385] The embodiments of the present application provide a model training device 50, as Figure 20 shown, which is applied to a local point analysis device and includes:
[0386] A receiving module 501, configured to receive a machine learning model sent by a first analysis device;
[0387] An incremental training module 502, configured to perform incremental training on the machine learning model based on a first training sample set, where the feature data in the first training sample set is the feature data of the local point network corresponding to the local point analysis device.
[0388] The embodiments of the present application provide a model training device. The receiving module receives a machine learning model sent by a first analysis device, and the incremental training module performs incremental training on the machine learning model based on a first training sample set obtained from the local point network corresponding to the local point analysis device. On the one hand, the feature data in the first training sample set is the feature data obtained from the local point network corresponding to the local point analysis device, which is more suitable for the application scenario of the local point analysis device. Using the first training sample set including the feature data obtained from the corresponding local point network of the local point analysis device for model training can make the trained machine learning model more suitable for the own needs of the local point analysis device, realize the customization of the model, and improve the application flexibility of the model. On the other hand, by combining offline training and incremental training to train the machine learning model, when the category or mode of the feature data obtained by the local point analysis device changes, incremental training of the machine learning model can be performed to realize the flexible adjustment of the machine learning model, so as to ensure that the trained machine learning model meets the needs of the local point analysis device. Therefore, the model training device provided by the embodiments of the present application can effectively adapt to the needs of the local point analysis device compared with the related technologies.
[0389] Optionally, as Figure 21 shown, the device 50 further includes:
[0390] A prediction module 503, configured to perform prediction of classification results using the machine learning model after receiving the machine learning model sent by the first analysis device;
[0391] The first sending module 504 is configured to send prediction information to the evaluation device, where the prediction information includes the predicted classification result, so that the evaluation device can evaluate whether the machine learning model deteriorates based on the prediction information.
[0392] The incremental training module 502 is configured to:
[0393] After receiving the training instruction sent by the evaluation device, perform incremental training on the machine learning model based on the first training sample set, where the training instruction is used to indicate training of the machine learning model.
[0394] Optionally, the machine learning model is used to predict the classification result of the data to be predicted composed of one or more key performance indicator (KPI) feature data; the KPI feature data is the feature data of the KPI time series or the KPI data.
[0395] The prediction information further includes the KPI category corresponding to the KPI feature data in the data to be predicted, the identifier of the device to which the data to be predicted belongs, and the acquisition time of the KPI data corresponding to the data to be predicted.
[0396] Optionally, as Figure 22 shown, the apparatus 50 further includes:
[0397] The second sending module 505 is configured to send a retraining request to the first analysis device when the performance of the machine learning model after incremental training does not meet the performance compliance condition, where the retraining request is used to request the first analysis device to retrain the machine learning model.
[0398] Optionally, the machine learning model is a tree model, and the incremental training module 502 is configured to:
[0399] For any training sample in the first training sample set, start traversing from the root node of the machine learning model and perform the following traversal process:
[0400] When the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, add an associated second node, where the first node is any non-leaf node in the machine learning model, and the second node is the parent node or child node of the first node.
[0401] When the current splitting cost of the first node is not less than the historical splitting cost of the first node, traverse the nodes in the subtree of the first node and determine the traversed node as the new first node, and perform the traversal process again until the current splitting cost of the first node traversed is less than the historical splitting cost of the first node or the target depth is traversed.
[0402] Among them, the current splitting cost of the first node is the cost for the first node to perform node splitting based on the first training sample, where the first training sample is any training sample in the first training sample set, the first training sample includes feature data in one or more feature dimensions, the feature data is numerical data, the historical splitting cost of the first node is the cost for the first node to perform node splitting based on the historical training sample set of the first node, and the historical training sample set of the first node is the set of samples divided into the first node in the historical training sample set of the machine learning model.
[0403] Optionally, the current splitting cost of the first node is negatively correlated with the size of the first numerical distribution range, and the first numerical distribution range is a distribution range determined based on the feature values in the first training sample and the second numerical distribution range; the second numerical distribution range is the distribution range of the feature values in the historical training sample set of the first node, and the historical splitting cost of the first node is negatively correlated with the size of the second numerical distribution range.
[0404] Optionally, the current splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the first numerical distribution range, and the historical splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the second numerical distribution range.
[0405] Optionally, the incremental training module 502 is configured to:
[0406] Determine the span range of the feature values on each feature dimension in the first numerical distribution range;
[0407] Add the second node based on the first splitting point on the first splitting dimension, where the numerical range in the first numerical distribution range whose value on the first splitting dimension is not greater than the value of the first splitting point is divided into the left child node of the second node, and the numerical range in the first numerical distribution range whose value on the first splitting dimension is greater than the value of the first splitting point is divided into the right child node of the second node. The first splitting dimension is the splitting dimension determined among the feature dimensions based on the span range of the feature values on each feature dimension, and the first splitting point is the numerical point determined on the first splitting dimension of the first numerical distribution range for splitting;
[0408] When the first splitting dimension is different from the second splitting dimension, the second node is the parent node or child node of the first node, the second splitting dimension is the historical splitting dimension of the first node in the machine learning model, and the second splitting point is the historical splitting point of the first node in the machine learning model;
[0409] When the first splitting dimension is the same as the second splitting dimension, and the first splitting point is on the right side of the second splitting point, the second node is the parent node of the first node, and the first node is the left child node of the second node;
[0410] When the first splitting dimension is the same as the second splitting dimension, and the first splitting point is on the left side of the second splitting point, the second node is the left child node of the first node.
[0411] Optionally, the first splitting dimension is a feature dimension randomly selected from the feature dimensions within the first numerical distribution range, or the first splitting dimension is the feature dimension with the largest span among the feature dimensions within the first numerical distribution range;
[0412] And / or, the first splitting point is a numerical point randomly selected on the first splitting dimension within the first numerical distribution range.
[0413] Optionally, the incremental training module 502 is configured to:
[0414] When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is greater than the first sample number threshold, add the second node;
[0415] The apparatus further includes:
[0416] A stop module, configured to stop the incremental training of the machine learning model when the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is not greater than the first sample number threshold.
[0417] In an alternative implementation, as Figure 23 shown, the apparatus 50 further includes:
[0418] A merging module 506, configured to merge the first non-leaf node and the second non-leaf node in the machine learning model, and merge the first leaf node and the second leaf node, to obtain a refined machine learning model, and the refined machine learning model is used for predicting classification results.
[0419] In another alternative implementation, the receiving module 501 is further configured to receive the refined machine learning model sent by the first analysis device, where the refined machine learning model is obtained by the first analysis device merging the first non-leaf node and the second non-leaf node in the machine learning model, and merging the first leaf node and the second leaf node;
[0420] Wherein, the first leaf node is a child node of the first non-leaf node, the second leaf node is a child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets assigned on the same feature dimension are adjacent.
[0421] Optionally, each node in the machine learning model stores node information correspondingly. The node information of any node in the machine learning model includes label distribution information, and the label distribution information is used to reflect the proportion of the labels of different categories of samples in the total number of labels in the historical training sample set partitioned to the corresponding node. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set partitioned to the any node. The node information of any non-leaf node further includes historical splitting information, and the historical splitting information is the information used for splitting by the corresponding node.
[0422] Optionally, the historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set partitioned to the corresponding node, and the historical splitting cost;
[0423] The label distribution information includes: the number of labels of the same category of samples in the historical training sample set partitioned to the corresponding node and the total number of labels; or, the proportion of the labels of different categories of samples in the historical training sample set partitioned to the corresponding node in the total number of labels.
[0424] Optionally, the first training sample set includes samples that meet the low discrimination condition and are screened from the samples obtained by the game point analysis device. The low discrimination condition includes at least one of the following:
[0425] The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, 1 < n < m, and m is the total number of probabilities obtained by predicting the sample using the machine learning model;
[0426] Or, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than the second difference threshold;
[0427] Or, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the third difference threshold;
[0428] Or, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the fourth difference threshold;
[0429] Alternatively, the probability distribution entropy E of multiple classification results of a sample predicted by the machine learning model is greater than a specified distribution entropy threshold, and the E satisfies:
[0430]
[0431] where x i represents the i-th classification result, P(x i ) represents the probability of predicting the i-th classification result of the sample, b is a specified base number, and 0 ≤ P(x i ) ≤ 1.
[0432] An embodiment of the present application provides a model training device 60, as Figure 24 shown, applied to a first analysis device, including:
[0433] An offline training module 601, configured to perform offline training based on a historical training sample set to obtain a machine learning model;
[0434] A sending module 602, configured to send the machine learning model to multiple local analysis devices for the local analysis devices to perform incremental training on the machine learning model based on a first training sample set, and the feature data in the training sample set used by any local analysis device to train the machine learning model is the feature data of the local network corresponding to the any local analysis device.
[0435] The sending module can distribute the machine learning model trained by the offline training module to each local analysis device for incremental training by each local analysis device, ensuring the performance of the machine learning models on each local analysis device. In this way, the first analysis device does not need to train a corresponding machine learning model for each local analysis device, effectively reducing the overall training duration of the first analysis device, and the model obtained by offline training can be used as the basis for incremental training on each local analysis device, improving the generality of the model obtained by offline training, thereby realizing model generalization and reducing the overall training cost of the first analysis device.
[0436] Moreover, the local point analysis device receives the machine learning model sent by the first analysis device, and can perform incremental training on the machine learning model based on the first training sample set obtained from the local point network corresponding to the local point analysis device. On the one hand, the feature data in the first training sample set is the feature data obtained from the local point network corresponding to the local point analysis device, which is more suitable for the application scenario of the local point analysis device. Using the first training sample set including the feature data obtained by the local point analysis device from the corresponding local point network for model training can make the trained machine learning model more suitable for the needs of the local point analysis device itself, realize the customization of the model, and improve the application flexibility of the model. On the other hand, by combining offline training and incremental training to train the machine learning model, incremental training of the machine learning model can be carried out when the category or mode of the feature data obtained by the local point analysis device changes, realizing the flexible adjustment of the machine learning model, so as to ensure that the trained machine learning model meets the needs of the local point analysis device. Therefore, the model training method provided in the embodiments of the present application can effectively adapt to the needs of the local point analysis device compared with the related art.
[0437] Optionally, the historical training sample set is a set of training samples sent by multiple local point analysis devices.
[0438] As Figure 25 shown, the device 60 further includes:
[0439] A receiving module 603, configured to:
[0440] After sending the machine learning model to the local point analysis device, receive the retraining request sent by the local point analysis device, and retrain the machine learning model based on the training sample set sent by the local point analysis device that sends the retraining request;
[0441] Or, receive the retraining request sent by the local point analysis device, and retrain the machine learning model based on the training sample set sent by the local point analysis device that sends the retraining request and the training sample sets sent by other local point analysis devices;
[0442] Or, receive the training sample sets sent by at least two local point analysis devices, and retrain the machine learning model based on the received training sample sets.
[0443] Optionally, the machine learning model is a tree model, and the offline training module is configured to:
[0444] Obtain a historical training sample set with determined labels, where the training samples in the historical training sample set include feature data of one or more feature dimensions, and the feature data is numerical data;
[0445] Create a root node;
[0446] Use the root node as the third node and perform the offline training process until the splitting cutoff condition is reached;
[0447] Determine the classification results for each leaf node to obtain the machine learning model;
[0448] Among them, the offline training process includes:
[0449] Split the third node to obtain the left child node and the right child node of the third node;
[0450] Use the left child node as the updated third node, divide the historical training sample set into the left sample set of the left child node as the updated historical training sample set, and perform the offline training process again;
[0451] Use the right child node as the updated third node, divide the historical training sample set into the right sample set of the right child node as the updated historical training sample set, and perform the offline training process again.
[0452] Optionally, the offline training module 601 is used for:
[0453] Based on the numerical distribution range of the historical training sample set, split the third node to obtain the left child node and the right child node of the third node. The numerical distribution range of the historical training sample set is the distribution range of the feature values in the historical training sample set.
[0454] Optionally, the offline training module 601 is used for:
[0455] Determine the third splitting dimension among the feature dimensions of the historical training sample set;
[0456] Determine the third splitting point on the third splitting dimension of the historical training sample set;
[0457] Divide the numerical range in the third numerical distribution range that is not greater than the value of the third splitting point on the third splitting dimension into the left child node, and divide the numerical range in the third numerical distribution range that is greater than the value of the third splitting point on the third splitting dimension into the right child node. The third numerical distribution range is the distribution range of the feature values in the historical training sample set of the third node.
[0458] Optionally, the splitting cutoff condition includes at least one of the following:
[0459] The current splitting cost of the third node is greater than the splitting cost threshold;
[0460] Alternatively, the number of samples in the historical training sample set is less than a second sample number threshold;
[0461] Alternatively, the number of splitting times corresponding to the third node is greater than a splitting times threshold;
[0462] Alternatively, the depth of the third node in the machine learning model is greater than a depth threshold;
[0463] Alternatively, the proportion of the number of the label with the largest proportion among the labels corresponding to the historical training sample set in the total number of labels corresponding to the historical training sample set is greater than a specified proportion threshold.
[0464] Optionally, the current splitting cost of the third node is negatively correlated with the size of the distribution range of the feature values in the historical training sample set.
[0465] Optionally, the current splitting cost of the third node is the reciprocal of the sum of the spans of the feature values of the historical training sample set in each feature dimension.
[0466] Optionally, as Figure 26 shown, the apparatus 60 further includes:
[0467] A merging module 604, configured to merge a first non-leaf node and a second non-leaf node in the machine learning model, and merge a first leaf node and a second leaf node, to obtain a refined machine learning model, where the refined machine learning model is used to predict a classification result, where the first leaf node is a child node of the first non-leaf node, the second leaf node is a child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets allocated in the same feature dimension are adjacent;
[0468] The sending module 602 is further configured to send the refined machine learning model to the local point analysis device for the local point analysis device to predict a classification result based on the refined machine learning model.
[0469] Optionally, each node in the machine learning model stores node information correspondingly, and the node information of any node in the machine learning model includes label distribution information, where the label distribution information is used to reflect the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels, and the total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node, and the node information of any non-leaf node further includes historical splitting information, where the historical splitting information is the information used for splitting by the corresponding node.
[0470] Optionally, the historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set partitioned to the corresponding node, and the historical splitting cost;
[0471] The label distribution information includes: the number of labels of the same category of samples in the historical training sample set partitioned to the corresponding node and the total number of the labels; or, the proportion of the labels of different categories of samples in the historical training sample set partitioned to the corresponding node in the total number of the labels.
[0472] Optionally, the first training sample set includes samples that meet the low discrimination condition and are screened from the samples obtained by the in-game point analysis device, and the low discrimination condition includes at least one of the following:
[0473] The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than a first difference threshold, where the target probability set includes the probabilities of the top n classification results arranged in descending order of probability, 1 < n < m, and m is the total number of probabilities obtained by predicting the sample using the machine learning model;
[0474] Or, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than a second difference threshold;
[0475] Or, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than a third difference threshold;
[0476] Or, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than a fourth difference threshold;
[0477] Or, the probability distribution entropy E of the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is greater than a specified distribution entropy threshold, and the E satisfies:
[0478]
[0479] where, x i represents the i-th classification result, P(x i ) represents the probability of the i-th classification result of the predicted sample, b is a specified base number, and 0 ≤ P(x i ) ≤ 1.
[0480] Figure 27 is a block diagram of a model training device provided by an embodiment of the present application. The model training device may be the aforementioned analysis device, such as an in-game point analysis device, or the aforementioned first analysis device. As Figure 27As shown in the figure, the analysis device 70 includes: a processor 701 and a memory 702.
[0481] The memory 701 is used to store a computer program, and the computer program includes program instructions.
[0482] The processor 702 is used to call the computer program to implement the model training method provided by the embodiments of the present application.
[0483] Optionally, the network device 70 further includes a communication bus 703 and a communication interface 704.
[0484] Among them, the processor 701 includes one or more processing cores. The processor 701 executes various functional applications and data processing by running the computer program.
[0485] The memory 702 can be used to store a computer program. Optionally, the memory can store an operating system and application program units required for at least one function. The operating system can be an operating system such as Real Time eXecutive (RTX), LINUX, UNIX, WINDOWS, or OS X.
[0486] There can be multiple communication interfaces 704. The communication interface 704 is used to communicate with other storage devices or network devices. For example, in the embodiments of the present application, the communication interface 704 can be used to receive sample data sent by network devices in a communication network.
[0487] The memory 702 and the communication interface 704 are respectively connected to the processor 701 through the communication bus 703.
[0488] The embodiments of the present application provide a computer storage medium, on which instructions are stored. When the instructions are executed by a processor, the model training method provided by the embodiments of the present application is implemented.
[0489] The embodiments of the present application provide a model training system, including: a first analysis device and multiple local analysis devices;
[0490] The first analysis device includes the model training device described in any of the foregoing embodiments; the local analysis device includes the model training device described in any of the foregoing embodiments. For example, the deployment of each device in the model training system can refer to the deployment of each device in the application scenario shown in the foregoing Figures 1 to 3 As shown, for example, the model training system further includes: one or more of a network device, an evaluation device, a storage device, and a management device. The introduction of related devices refers to the foregoing Figures 1 to 3 , and the embodiments of the present application will not elaborate on this.
[0491] In the embodiments of the present application, A referring to B means that A can be the same as B or can be simply modified based on B.
[0492] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product, and the computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server, data center, etc. that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid-state drive), etc.
[0493] The foregoing are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A model training method, characterized in that, Applied to a local point analysis device, where the local point analysis device is any one of a plurality of local point analysis devices, including: Receiving a machine learning model sent by a first analysis device, where the machine learning model is obtained by the first analysis device through offline training based on a historical training sample set, and the historical training sample set is a set of training samples sent by the plurality of local point analysis devices; Based on a first training sample set, performing incremental training on the machine learning model, where the feature data in the first training sample set is the feature data of the local point network corresponding to the local point analysis device.
2. The method according to claim 1, wherein After receiving the machine learning model sent by the first analysis device, the method further includes: Using the machine learning model to predict a classification result; Sending prediction information to an evaluation device, where the prediction information includes the predicted classification result, for the evaluation device to evaluate whether the machine learning model has deteriorated based on the prediction information; The performing incremental training on the machine learning model based on the first training sample set includes: After receiving a training instruction sent by the evaluation device, performing incremental training on the machine learning model based on the first training sample set, where the training instruction is used to indicate training the machine learning model.
3. The method according to claim 2, characterized in that, The machine learning model is used to predict a classification result for data to be predicted composed of one or more key performance indicator (KPI) feature data; the KPI feature data is the feature data of a KPI time series or KPI data; The prediction information further includes the KPI category corresponding to the KPI feature data in the data to be predicted, the identifier of the device to which the data to be predicted belongs, and the acquisition time of the KPI data corresponding to the data to be predicted.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: When the performance of the machine learning model after incremental training does not meet the performance compliance condition, sending a retraining request to the first analysis device, where the retraining request is used to request the first analysis device to retrain the machine learning model.
5. The method according to any one of claims 1 to 3, characterized in that, The machine learning model is a tree model, and the performing incremental training on the machine learning model based on the first training sample set includes: For any training sample in the first training sample set, starting from the root node of the machine learning model, performing the following traversal process: When the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, adding an associated second node, where the first node is any non-leaf node in the machine learning model, and the second node is the parent node or child node of the first node; When the current splitting cost of the first node is not less than the historical splitting cost of the first node, traversing the nodes in the subtree of the first node and determining the traversed node as the new first node, and performing the traversal process again until the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, or the target depth is traversed. Among them, the current splitting cost of the first node is the cost for the first node to perform node splitting based on the first training sample. The first training sample is any training sample in the first training sample set. The first training sample includes feature data of one or more feature dimensions, and the feature data is numerical data. The historical splitting cost of the first node is the cost for the first node to perform node splitting based on the historical training sample set of the first node. The historical training sample set of the first node is the set of samples in the historical training sample set of the machine learning model that are partitioned to the first node.
6. The method according to claim 5, wherein The current splitting cost of the first node is negatively correlated with the size of the first numerical distribution range. The first numerical distribution range is a distribution range determined based on the feature values in the first training sample and the second numerical distribution range. The second numerical distribution range is the distribution range of the feature values in the historical training sample set of the first node. The historical splitting cost of the first node is negatively correlated with the size of the second numerical distribution range.
7. The method according to claim 6, wherein The current splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the first numerical distribution range. The historical splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the second numerical distribution range.
8. The method according to claim 6, characterized in that, Adding the associated second node includes: Determining the span range of the feature values on each feature dimension in the first numerical distribution range; Adding the second node based on the first splitting point on the first splitting dimension. Among them, the numerical range in the first numerical distribution range whose value on the first splitting dimension is not greater than the value of the first splitting point is partitioned to the left child node of the second node, and the numerical range in the first numerical distribution range whose value on the first splitting dimension is greater than the value of the first splitting point is partitioned to the right child node of the second node. The first splitting dimension is the splitting dimension determined among the various feature dimensions based on the span range of the feature values on each feature dimension. The first splitting point is the numerical point determined on the first splitting dimension of the first numerical distribution range for splitting. When the first splitting dimension is different from the second splitting dimension, the second node is the parent node or child node of the first node. The second splitting dimension is the historical splitting dimension of the first node in the machine learning model. When the first splitting dimension is the same as the second splitting dimension, and the first splitting point is on the right side of the second splitting point, the second node is the parent node of the first node, and the first node is the left child node of the second node. The second splitting point is the historical splitting point of the first node in the machine learning model. When the first splitting dimension is the same as the second splitting dimension, and the first splitting point is on the left side of the second splitting point, the second node is the left child node of the first node.
9. The method according to claim 6, wherein The first splitting dimension is a feature dimension randomly selected from among the feature dimensions within the first numerical distribution range, or the first splitting dimension is the feature dimension with the largest span among the feature dimensions within the first numerical distribution range; And / or, the first splitting point is a numerical point randomly selected on the first splitting dimension within the first numerical distribution range.
10. The method according to claim 5, wherein Adding the associated second node includes: When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is greater than the first sample number threshold, adding the second node; The method further includes: When the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is not greater than the first sample number threshold, stopping the incremental training of the machine learning model.
11. The method according to claim 5, characterized in that, The method further includes: Merging the first non-leaf node and the second non-leaf node in the machine learning model, and merging the first leaf node and the second leaf node to obtain a refined machine learning model, where the refined machine learning model is used to predict classification results; Alternatively, receiving the refined machine learning model sent by the first analysis device, where the refined machine learning model is obtained by the first analysis device merging the first non-leaf node and the second non-leaf node in the machine learning model, and merging the first leaf node and the second leaf node; Wherein, the first leaf node is a child node of the first non-leaf node, the second leaf node is a child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets allocated on the same feature dimension are adjacent.
12. The method according to claim 5, wherein, Each node in the machine learning model stores node information correspondingly. The node information of any node in the machine learning model includes label distribution information, where the label distribution information is used to reflect the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, where the historical splitting information is the information used by the corresponding node for splitting.
13. The method according to claim 12, characterized in that, The historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical splitting cost; The label distribution information includes: the number of labels of the same category of samples in the historical training sample set divided into the corresponding node and the total number of labels; or, the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels.
14. The method according to any one of claims 1 to 3, characterized in that, The first training sample set includes samples that meet the low discrimination condition selected from the samples obtained by the local point analysis device. The low discrimination condition includes at least one of the following: The absolute value of the difference between any two probabilities in the set of target probabilities obtained by predicting a sample using the machine learning model is less than a first difference threshold, where the set of target probabilities includes the probabilities of the top n classification results arranged in descending order of probability, 1 < n < m, and m is the total number of probabilities obtained by the machine learning model predicting the sample; Alternatively, the absolute value of the difference between any two probabilities among the probabilities obtained by predicting a sample using the machine learning model is less than a second difference threshold; Alternatively, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting a sample using the machine learning model is less than a third difference threshold; Alternatively, the absolute value of the difference between any two probabilities among the probabilities of multiple classification results obtained by predicting a sample using the machine learning model is less than a fourth difference threshold; Alternatively, the probability distribution entropy E of the multiple classification results obtained by predicting a sample using the machine learning model is greater than a specified distribution entropy threshold, where the E satisfies: Among them, x i represents the i-th classification result, and P(x i ) represents the probability of the i-th classification result of the predicted sample. b is the specified base, and 0 ≤ P(x i ) ≤ 1.
15. A model training method, characterized in that, Applied to a first analysis device, it includes: Performing offline training based on a historical training sample set to obtain a machine learning model, where the historical training sample set is a set of training samples sent by multiple local analysis devices; Sending the machine learning model to the multiple local analysis devices for any one of the multiple local analysis devices to perform incremental training on the machine learning model based on a first training sample set, and the feature data in the training sample set used by any one of the local analysis devices to train the machine learning model is the feature data of the local network corresponding to any one of the local analysis devices.
16. The method according to claim 15, wherein After sending the machine learning model to the local analysis device, the method further includes: Receiving a retraining request sent by the local analysis device and retraining the machine learning model based on the training sample set sent by the local analysis device that sent the retraining request; Alternatively, receiving a retraining request sent by the local analysis device and retraining the machine learning model based on the training sample set sent by the local analysis device that sent the retraining request and the training sample sets sent by other local analysis devices; Alternatively, receiving the training sample sets sent by at least two of the local analysis devices and retraining the machine learning model based on the received training sample sets.
17. The method according to claim 15 or 16, characterized in that The machine learning model is a tree model, and the performing offline training based on a historical training sample set to obtain a machine learning model includes: Obtaining a historical training sample set with determined labels, where the training samples in the historical training sample set include feature data in one or more feature dimensions, and the feature data is numerical data; Creating a root node; Taking the root node as a third node and performing an offline training process until a splitting cutoff condition is reached; Determining classification results for each leaf node to obtain the machine learning model; Wherein, the offline training process includes: Splitting the third node to obtain a left child node and a right child node of the third node; Take the left child node as the updated third node, divide the historical training sample set into the left sample set of the left child node as the updated historical training sample set, and execute the offline training process again; Take the right child node as the updated third node, divide the historical training sample set into the right sample set of the right child node as the updated historical training sample set, and execute the offline training process again.
18. The method according to claim 17, wherein The splitting of the third node to obtain the left child node and the right child node of the third node includes: Based on the numerical distribution range of the historical training sample set, split the third node to obtain the left child node and the right child node of the third node, where the numerical distribution range of the historical training sample set is the distribution range of the feature values in the historical training sample set.
19. The method according to claim 18, wherein The splitting of the third node based on the numerical distribution range of the historical training sample set to obtain the left child node and the right child node of the third node includes: Determine the third splitting dimension in each feature dimension of the historical training sample set; Determine the third splitting point on the third splitting dimension of the historical training sample set; Divide the numerical range in the third numerical distribution range whose value on the third splitting dimension is not greater than the value of the third splitting point into the left child node, and divide the numerical range in the third numerical distribution range whose value on the third splitting dimension is greater than the value of the third splitting point into the right child node, where the third numerical distribution range is the distribution range of the feature values in the historical training sample set of the third node.
20. The method according to claim 17, characterized in that, The splitting termination condition includes at least one of the following: The current splitting cost of the third node is greater than the splitting cost threshold; Or, the number of samples in the historical training sample set is less than the second sample number threshold; Or, the number of splitting times corresponding to the third node is greater than the splitting times threshold; Or, the depth of the third node in the machine learning model is greater than the depth threshold; Or, the proportion of the number of the label with the largest proportion in the labels corresponding to the historical training sample set in the total number of labels corresponding to the historical training sample set is greater than the specified proportion threshold.
21. The method according to claim 20, characterized in that, The current splitting cost of the third node is negatively correlated with the size of the distribution range of the feature values in the historical training sample set.
22. The method according to claim 21, wherein The current splitting cost of the third node is the reciprocal of the sum of the spans of the feature values of the historical training sample set in each feature dimension.
23. The method according to claim 17, characterized in that, The method further includes: Merge the first non-leaf node and the second non-leaf node in the machine learning model, and merge the first leaf node and the second leaf node to obtain a refined machine learning model, where the refined machine learning model is used to predict the classification result. Among them, the first leaf node is the child node of the first non-leaf node, the second leaf node is the child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets assigned on the same feature dimension are adjacent; Send the refined machine learning model to the local point analysis device for the local point analysis device to predict the classification result based on the refined machine learning model.
24. The method according to claim 17, wherein Each node in the machine learning model stores corresponding node information. The node information of any node in the machine learning model includes label distribution information, which is used to reflect the proportion of the labels of different categories of samples in the total number of labels in the historical training sample set divided into the corresponding node. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, which is the information used for splitting the corresponding node.
25. The method according to claim 24, wherein The historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical splitting cost; The label distribution information includes: the number of labels of the same category of samples in the historical training sample set divided into the corresponding node and the total number of labels; or, the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels.
26. The method according to claim 15 or 16, characterized in that, The first training sample set includes samples that meet the low discrimination condition selected from the samples obtained by the local point analysis device. The low discrimination condition includes at least one of the following: The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, where 1 < n < m, and m is the total number of probabilities obtained by the machine learning model for predicting the sample; Or, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than the second difference threshold; Or, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the third difference threshold; Or, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the fourth difference threshold; Or, the probability distribution entropy E of the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is greater than the specified distribution entropy threshold, and the E satisfies: where x i represents the i-th classification result, and P(x i ) represents the probability of the i-th classification result of the predicted sample. b is the specified base, and 0 ≤ P(x i ) ≤ 1.
27. A model training device, characterized in that Applied to the local point analysis device, the local point analysis device is any one of multiple local point analysis devices, including: A receiving module, configured to receive the machine learning model sent by the first analysis device. The machine learning model is obtained by the first analysis device through offline training based on the historical training sample set, and the historical training sample set is a set of training samples sent by the multiple local point analysis devices; An incremental training module, configured to perform incremental training on the machine learning model based on the first training sample set. The feature data in the first training sample set is the feature data of the local point network corresponding to the local point analysis device.
28. The device according to claim 27, characterized in that, The device further includes: A prediction module, configured to, after receiving the machine learning model sent by the first analysis device, use the machine learning model to predict classification results; A first sending module, configured to send prediction information to an evaluation device, where the prediction information includes the predicted classification results, so that the evaluation device can evaluate whether the machine learning model has deteriorated based on the prediction information; The incremental training module is configured to: After receiving the training instruction sent by the evaluation device, perform incremental training on the machine learning model based on the first training sample set, where the training instruction is used to indicate training of the machine learning model.
29. The device according to claim 28, characterized in that, The machine learning model is used to predict classification results for data to be predicted composed of one or more key performance indicator (KPI) feature data; the KPI feature data is feature data of a KPI time series or KPI data; The prediction information further includes the KPI category corresponding to the KPI feature data in the data to be predicted, the identifier of the device to which the data to be predicted belongs, and the acquisition time of the KPI data corresponding to the data to be predicted.
30. The device according to any one of claims 27 to 29, characterized in that, The apparatus further includes: A second sending module, configured to send a retraining request to the first analysis device when the performance of the incrementally trained machine learning model does not meet the performance compliance condition, where the retraining request is used to request the first analysis device to retrain the machine learning model.
31. The device according to any one of claims 27 to 29, characterized in that The machine learning model is a tree model, and the incremental training module is configured to: For any training sample in the first training sample set, start traversing from the root node of the machine learning model and perform the following traversal process: When the current splitting cost of the first node traversed is less than the historical splitting cost of the first node, add an associated second node, where the first node is any non-leaf node in the machine learning model, and the second node is the parent node or child node of the first node; When the current splitting cost of the first node is not less than the historical splitting cost of the first node, traverse the nodes in the subtree of the first node and determine the traversed node as the new first node, and perform the traversal process again until the current splitting cost of the first node traversed is less than the historical splitting cost of the first node or the target depth is traversed; Wherein, the current splitting cost of the first node is the cost of splitting the first node based on the first training sample, the first training sample is any training sample in the first training sample set, the first training sample includes feature data of one or more feature dimensions, the feature data is numerical data, the historical splitting cost of the first node is the cost of splitting the first node based on the historical training sample set of the first node, and the historical training sample set of the first node is the set of samples divided into the first node in the historical training sample set of the machine learning model.
32. The device according to claim 31, characterized in that, The current splitting cost of the first node is negatively correlated with the size of the first numerical distribution range, which is a distribution range determined based on the feature values in the first training sample and the second numerical distribution range; the second numerical distribution range is the distribution range of the feature values in the historical training sample set of the first node, and the historical splitting cost of the first node is negatively correlated with the size of the second numerical distribution range.
33. The apparatus according to claim 32, wherein The current splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the first numerical distribution range, and the historical splitting cost of the first node is the reciprocal of the sum of the spans of the feature values on each feature dimension in the second numerical distribution range.
34. The device according to claim 32, characterized in that, The incremental training module is used to Determine the span range of the feature values on each feature dimension in the first numerical distribution range; Add the second node based on the first splitting point on the first splitting dimension, where the numerical range in the first numerical distribution range that is not greater than the value of the first splitting point on the first splitting dimension is divided into the left child node of the second node, and the numerical range in the first numerical distribution range that is greater than the value of the first splitting point on the first splitting dimension is divided into the right child node of the second node. The first splitting dimension is the splitting dimension determined from among the feature dimensions based on the span range of the feature values on each feature dimension, and the first splitting point is the numerical point determined for splitting on the first splitting dimension of the first numerical distribution range; When the first splitting dimension is different from the second splitting dimension, the second node is the parent node or child node of the first node, and the second splitting dimension is the historical splitting dimension of the first node in the machine learning model; When the first splitting dimension is the same as the second splitting dimension and the first splitting point is to the right of the second splitting point, the second node is the parent node of the first node, and the first node is the left child node of the second node. The second splitting point is the historical splitting point of the first node in the machine learning model; When the first splitting dimension is the same as the second splitting dimension and the first splitting point is to the left of the second splitting point, the second node is the left child node of the first node.
35. The device according to claim 32, wherein The first splitting dimension is a randomly selected feature dimension from among the feature dimensions in the first numerical distribution range, or the first splitting dimension is the feature dimension with the largest span among the feature dimensions in the first numerical distribution range; And / or, the first splitting point is a randomly selected numerical point on the first splitting dimension of the first numerical distribution range.
36. The device according to claim 31, wherein The incremental training module is used to: When the sum of the number of samples in the historical training sample set of the first node and the number of the first training sample is greater than the first sample number threshold, add the second node; The device further includes: A stop module, configured to stop the incremental training of the machine learning model when the sum of the number of samples in the historical training sample set of the first node and the number of the first training samples is not greater than the first sample number threshold.
37. The device according to claim 31, characterized in that, The apparatus further includes: A merging module, configured to merge a first non-leaf node and a second non-leaf node in the machine learning model, and merge a first leaf node and a second leaf node, to obtain a refined machine learning model, where the refined machine learning model is used to predict classification results; Alternatively, the receiving module is further configured to receive the refined machine learning model sent by the first analysis device, where the refined machine learning model is obtained by the first analysis device merging a first non-leaf node and a second non-leaf node in the machine learning model, and merging a first leaf node and a second leaf node; Wherein, the first leaf node is a child node of the first non-leaf node, the second leaf node is a child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets allocated in the same feature dimension are adjacent.
38. The device according to claim 31, characterized in that, Each node in the machine learning model stores node information correspondingly, and the node information of any node in the machine learning model includes label distribution information, where the label distribution information is used to reflect the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels, and the total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, where the historical splitting information is the information used for splitting by the corresponding node.
39. The device according to claim 38, characterized in that, The historical splitting information includes: the position information of the corresponding node in the machine learning model, the splitting dimension, the splitting point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical splitting cost; The label distribution information includes: the number of labels of the same category of samples in the historical training sample set divided into the corresponding node and the total number of labels; or, the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels.
40. The device according to any one of claims 27 to 29, characterized in that, The first training sample set includes samples that meet the low discrimination condition selected from the samples obtained by the local point analysis device, and the low discrimination condition includes at least one of the following: The absolute value of the difference between any two probabilities in the target probability set obtained by predicting a sample using the machine learning model is less than the first difference threshold, where the target probability set includes the probabilities of the top n classification results arranged in descending order of probability, 1 < n < m, and m is the total number of probabilities obtained by the machine learning model predicting a sample; Alternatively, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting a sample using the machine learning model is less than the second difference threshold; Alternatively, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting a sample using the machine learning model is less than the third difference threshold; Alternatively, the absolute value of the difference between any two probabilities among the probabilities of multiple classification results predicted by the machine learning model for the sample is less than a fourth difference threshold; Alternatively, the probability distribution entropy E of the multiple classification results predicted by the machine learning model for the sample is greater than a specified distribution entropy threshold, and the E satisfies: Among them, x i represents the i-th classification result, and P(x i ) represents the probability of the i-th classification result of the predicted sample. b is the specified base, and 0 ≤ P(x i ) ≤ 1.
41. A model training device, characterized in that, Applied to a first analysis device, it includes: An offline training module, configured to perform offline training based on a historical training sample set to obtain a machine learning model, where the historical training sample set is a set of training samples sent by multiple local analysis devices; A sending module, configured to send the machine learning model to multiple local analysis devices, so that any one of the multiple local analysis devices performs incremental training on the machine learning model based on a first training sample set, and the feature data in the training sample set used by any one of the local analysis devices to train the machine learning model is the feature data of the local network corresponding to the local analysis device.
42. The device according to claim 41, characterized in that, The device further includes: A receiving module, configured to: After sending the machine learning model to the local analysis device, receive a retraining request sent by the local analysis device, and perform retraining on the machine learning model based on the training sample set sent by the local analysis device that sends the retraining request; Alternatively, receive a retraining request sent by the local analysis device, and perform retraining on the machine learning model based on the training sample set sent by the local analysis device that sends the retraining request and the training sample sets sent by other local analysis devices; Alternatively, receive the training sample sets sent by at least two of the local analysis devices, and retrain the machine learning model based on the received training sample sets.
43. The device according to claim 41 or 42, characterized in that, The machine learning model is a tree model, and the offline training module is configured to: Obtain a historical training sample set with determined labels, where the training samples in the historical training sample set include feature data of one or more feature dimensions, and the feature data is numerical data; Create a root node; Use the root node as the third node, and execute the offline training process until a splitting cut-off condition is reached; Determine classification results for each leaf node to obtain the machine learning model; Wherein, the offline training process includes: Perform splitting of the third node to obtain a left child node and a right child node of the third node; Use the left child node as the updated third node, divide the historical training sample set into the left sample set of the left child node as the updated historical training sample set, and execute the offline training process again; Use the right child node as the updated third node, divide the historical training sample set into the right sample set of the right child node as the updated historical training sample set, and execute the offline training process again.
44. The apparatus according to claim 43, wherein The offline training module is configured to: Perform splitting of the third node based on the numerical distribution range of the historical training sample set to obtain a left child node and a right child node of the third node, where the numerical distribution range of the historical training sample set is the distribution range of the feature values in the historical training sample set.
45. The device according to claim 44, characterized in that, The offline training module is used for: Determine a third splitting dimension among the feature dimensions of the historical training sample set; Determine a third splitting point on the third splitting dimension of the historical training sample set; Divide the numerical range in the third numerical distribution range whose value on the third splitting dimension is not greater than the value of the third splitting point into the left child node, and divide the numerical range in the third numerical distribution range whose value on the third splitting dimension is greater than the value of the third splitting point into the right child node. The third numerical distribution range is the distribution range of the feature values in the historical training sample set of the third node.
46. The apparatus according to claim 44, wherein, The splitting termination condition includes at least one of the following: The current splitting cost of the third node is greater than the splitting cost threshold; Or, the number of samples in the historical training sample set is less than the second sample number threshold; Or, the splitting times corresponding to the third node are greater than the splitting times threshold; Or, the depth of the third node in the machine learning model is greater than the depth threshold; Or, the proportion of the number of the label with the largest proportion in the labels corresponding to the historical training sample set in the total number of labels corresponding to the historical training sample set is greater than the specified proportion threshold.
47. The device according to claim 46, characterized in that, The current splitting cost of the third node is negatively correlated with the size of the distribution range of the feature values in the historical training sample set.
48. The device according to claim 47, characterized in that, The current splitting cost of the third node is the reciprocal of the sum of the spans of the feature values of the historical training sample set on each feature dimension.
49. The device according to claim 44, characterized in that, The device further includes: A merging module, configured to merge a first non-leaf node and a second non-leaf node in the machine learning model, and merge a first leaf node and a second leaf node to obtain a refined machine learning model for predicting classification results. The first leaf node is a child node of the first non-leaf node, the second leaf node is a child node of the second non-leaf node, the first leaf node and the second leaf node include the same classification result, and the span ranges of the feature values of the historical training sample sets assigned on the same feature dimension are adjacent; The sending module is further configured to send the refined machine learning model to the local point analysis device for the local point analysis device to predict classification results based on the refined machine learning model.
50. The device according to claim 44, characterized in that, Each node in the machine learning model stores node information correspondingly. The node information of any node in the machine learning model includes label distribution information, which is used to reflect the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of labels. The total number of labels is the total number of labels corresponding to the samples in the historical training sample set divided into the any node. The node information of any non-leaf node further includes historical splitting information, which is the information used for splitting by the corresponding node.
51. The device according to claim 50, characterized in that, The historical split information includes: the location information of the corresponding node in the machine learning model, the split dimension, the split point, the numerical distribution range of the historical training sample set divided into the corresponding node, and the historical split cost; The label distribution information includes: the number of labels of the same category of samples in the historical training sample set divided into the corresponding node and the total number of the labels; or, the proportion of the labels of different categories of samples in the historical training sample set divided into the corresponding node in the total number of the labels.
52. The device according to claim 41 or 42, characterized in that, The first training sample set includes samples that meet the low discrimination condition and are screened from the samples obtained by the in-game point analysis device. The low discrimination condition includes at least one of the following: The absolute value of the difference between any two probabilities in the target probability set obtained by predicting the sample using the machine learning model is less than the first difference threshold. The target probability set includes the probabilities of the top n classification results arranged in descending order of probability, where 1 < n < m, and m is the total number of probabilities obtained by predicting the sample using the machine learning model; Alternatively, the absolute value of the difference between any two probabilities in the probabilities obtained by predicting the sample using the machine learning model is less than the second difference threshold; Alternatively, the absolute value of the difference between the highest probability and the lowest probability among the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the third difference threshold; Alternatively, the absolute value of the difference between any two probabilities in the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is less than the fourth difference threshold; Alternatively, the probability distribution entropy E of the probabilities of multiple classification results obtained by predicting the sample using the machine learning model is greater than the specified distribution entropy threshold, and the E satisfies: where x i represents the i-th classification result, and P(x i ) represents the probability of the i-th classification result of the predicted sample. b is the specified base, and 0 ≤ P(x i ) ≤ 1.
53. A model training device, characterized in that, Including: A processor and a memory; The memory is used to store a computer program, and the computer program includes program instructions; The processor is used to call the computer program to implement the model training method according to any one of claims 1 to 14; or, to implement the model training method according to any one of claims 15 to 26.
54. A computer storage medium, characterized in that, Instructions are stored on the computer storage medium, and when the instructions are executed by the processor, the model training method according to any one of claims 1 to 14 is implemented; or, the model training method according to any one of claims 15 to 26 is implemented.
55. A model training system, characterized in that, Including: A first analysis device and multiple in-game point analysis devices; The first analysis device includes the model training device according to any one of claims 27 to 40; The in-game point analysis device includes the model training device according to any one of claims 41 to 52.