Prediction method and device based on stacking and machine learning model
By statistically processing the prediction results of the basic model and combining the logistic regression calculation of the metamodel, the problem of limited improvement of prediction accuracy in the existing technology is solved, and higher prediction accuracy is achieved.
Patent Information
- Application Number
- CN202411942373.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing prediction technologies based on stacking and machine learning models have limited effectiveness in improving prediction accuracy, due to the accuracy of the basic model and the limitations of logistic regression algorithms.
By obtaining business data and multiple base models, input business data to each base model separately to obtain prediction results, perform statistical processing to generate statistical results, and input prediction results and statistical results into the meta-model to generate final prediction results.
By increasing the data dimensions and processing capabilities of metamodel input, the accuracy of predictions of machine learning models and stacking technology is improved.
Smart Images

Figure CN120012036A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of machine learning technology, and in particular to a prediction method based on stacking and machine learning models, a prediction device based on stacking and machine learning models, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the advent of the era of artificial intelligence (AI), the development of many businesses is inseparable from machine learning (ML) models. By training ML models and using the trained models to predict and analyze business, accurate prediction and analysis of business can be achieved.
[0003] In order to further improve the accuracy of ML model prediction, stacking technology is usually used. Specifically, multiple trained ML models are prepared in advance as base models, and a logistic regression algorithm is prepared as a meta-model. When a prediction is required based on a certain business data, the business data is used as input and input into each base model. Each base model processes the business data and outputs the prediction results. The prediction results output by each base model are then used as input to the meta-model and input into the meta-model together. The meta-model uses a logistic regression algorithm to process each prediction result and obtain and output the final prediction result.
[0004] However, stacking technology is limited by the accuracy of each ML model and the limitations of the logistic regression algorithm, and has limited effect on improving prediction accuracy. It can be seen that how to achieve a significant improvement in prediction accuracy is an urgent problem that needs to be solved. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a prediction method based on stacking and machine learning models, a prediction device based on stacking and machine learning models, an electronic device, a computer-readable storage medium and a computer program product to improve the accuracy of predictions based on machine learning models and stacking technology.
[0006] In order to solve the above technical problems, the embodiments of the present application provide the following technical solutions:
[0007] The first aspect of the present application provides a prediction method based on stacking and machine learning models, the method comprising: obtaining business data and multiple base models, the business data being the data to be predicted, and the base models being used to make predictions based on the business data; inputting the business data into multiple base models respectively to obtain multiple prediction results corresponding to the outputs of the multiple base models; performing statistical processing on the multiple prediction results to obtain statistical results; inputting the multiple prediction results and the statistical results into a meta-model to obtain a final prediction result output by the meta-model, the meta-model being used to make predictions based on the outputs of the multiple base models, the meta-model being trained based on the initial meta-model using training data, the training data comprising the training outputs of the multiple base models and the statistical results of the training outputs.
[0008] Compared with the prior art, the prediction method based on stacking and machine learning models provided in the first aspect of the present application, after multiple base models output prediction results, the prediction results output by each base model are statistically analyzed, and then the statistical results and the prediction results of each base model are input into the meta-model together, so that the meta-model can further analyze and process the prediction results of multiple base models in combination with the statistical results. Because the meta-model was previously only able to perform logistic regression calculations based on the prediction results of multiple base models, the input data dimension and processing power are limited. This time, while inputting the prediction results of multiple base models into the meta-model, the statistical results of multiple prediction results are also input, so that the data dimension and processing power processed by the meta-model can be improved, thereby improving the accuracy of predictions of machine learning models and stacking technology.
[0009] In some modified implementations of the first aspect of the present application, multiple prediction results are statistically processed to obtain statistical results, including: based on each prediction result among the multiple prediction results, respectively determining the physical measurement value of the base model corresponding to each prediction result to obtain multiple physical measurement values, and performing preferential processing among the multiple physical measurement values to obtain an optimal measurement value; performing nonlinear processing on at least two prediction results among the multiple prediction results to obtain a nonlinear result; and determining the optimal measurement value and the nonlinear result as statistical results.
[0010] When performing specific statistics, each prediction result can be transformed into a metric and the best one can be selected, and nonlinear statistics can be performed on all or two of the multiple prediction results. This can increase the dimension of the statistical results, enrich the data input into the meta-model, and thus improve the accuracy of the prediction.
[0011] In some modified implementations of the first aspect of the present application, the physical measurement value includes a confidence level, and the nonlinear processing includes at least one of the following: calculating the mode, calculating the variance, calculating the mean, calculating the difference between two values, calculating the sum of two values, and the optimal model.
[0012] Parameters such as confidence and mode can simply and accurately count multiple prediction results, thereby improving the efficiency and accuracy of the final prediction.
[0013] In some modified implementations of the first aspect of the present application, the metamodel includes multiple preset algorithms; multiple prediction results and statistical results are input into the metamodel, including: multiple prediction results, statistical results and specified algorithm names are input into the metamodel, so that the metamodel searches for a target algorithm corresponding to the specified algorithm name in multiple preset algorithms, and then processes the multiple prediction results and statistical results based on the target algorithm.
[0014] Algorithms can be specified in the metamodel, which improves the flexibility of metamodel construction. In addition, by specifying the algorithm name, the algorithm can be specified in the metamodel, which improves the convenience of algorithm specification in the metamodel.
[0015] In some modified implementations of the first aspect of the present application, before the business data is input into multiple base models respectively, the method also includes: determining the prediction accuracy parameter of each base model, and determining the prediction similarity between each base model; based on the prediction accuracy parameter and the prediction similarity, deleting one or more base models from the multiple base models to obtain a preset number of target base models, so that the business data can be input into the target base models respectively, wherein the base models with lower prediction accuracy parameters and higher prediction similarity will be deleted first, and the priority of the prediction accuracy parameter is higher than the priority of the prediction similarity.
[0016] Before using multiple base models for prediction, multiple base models are screened. Among multiple base models, base models with low prediction accuracy and similar to other base models are deleted. In this way, it is possible to avoid the base model with low prediction accuracy from causing adverse interference to the prediction results of similar base models with high prediction accuracy, thereby improving the accuracy of the final prediction.
[0017] In some modified implementations of the first aspect of the present application, based on the prediction accuracy parameter and the similarity, one or more base models among the multiple base models are deleted to obtain a preset number of target base models, including: taking the base model with the largest prediction accuracy parameter among the multiple base models as the current base model, and taking the base models among the multiple base models other than the current base model as other base models; among the other base models, deleting the base model with the highest prediction similarity to the current base model to obtain multiple remaining base models including the current base model; if the number of the multiple remaining base models is greater than the preset number, continuing to take the base model with the largest prediction accuracy parameter among the multiple remaining base models as the new current base model, and taking the base models among the multiple remaining base models other than the new current base model as new other base models, and among the new other base models, deleting the base model with the highest prediction similarity to the new current base model to obtain new remaining base models including the new current base model, until the number of new remaining base models is equal to the preset number.
[0018] In the screening of multiple base models, the current base model is selected in turn according to the prediction accuracy parameter sorting, and the base model most similar to the current base model is deleted from other base models until the number of remaining base models meets the requirement. In this way, the base model with low prediction accuracy and similar to other base models can be deleted quickly and accurately, improving the efficiency of base model optimization.
[0019] In some modified implementations of the first aspect of the present application, each base model can output multiple prediction results based on data corresponding to multiple known real results, and the method also includes: if the prediction similarity between the two base models is less than or equal to the similarity threshold, then determine the data quantity representation, category quantity representation and / or data generation distance time representation of the prediction difference between the multiple prediction results corresponding to the two base models respectively; based on the prediction accuracy parameter and the data quantity representation, category quantity representation and / or data generation distance time representation, delete one or more base models from the multiple base models to obtain a preset number of target base models, so that the business data can be input into the target base model respectively, wherein the base model with a lower prediction accuracy parameter, a smaller data quantity representation, a smaller category quantity representation, and a smaller data generation distance time representation will be deleted first, and the priority of the prediction accuracy parameter is higher than the priority of the data quantity representation, the category quantity representation and the data generation distance time representation.
[0020] When the base models are not similar to each other, we can select and delete the base models with a small number of predicted difference data, a small number of categories, and a short data generation distance. This can ensure that the differences between the remaining base models are maximized, thereby improving the accuracy of subsequent predictions.
[0021] The second aspect of the present application provides a prediction device based on stacking and machine learning models, and the device includes: an acquisition module, used to acquire business data and multiple base models, the business data is the data to be predicted, and the base model is used to make predictions based on the business data; a first processing module, used to input the business data into multiple base models respectively, and obtain multiple prediction results corresponding to the outputs of the multiple base models; a second processing module, used to perform statistical processing on the multiple prediction results to obtain statistical results; a prediction module, used to input the multiple prediction results and the statistical results into a meta-model, and obtain the final prediction result output by the meta-model, the meta-model is used to make predictions based on the outputs of multiple base models, and the meta-model is trained based on the initial meta-model using training data, and the training data includes the training outputs of multiple base models and the statistical results of the training outputs.
[0022] A third aspect of the present application provides an electronic device, which includes a processor, a memory and a bus. The processor and the memory communicate with each other through the bus, and the processor is used to call program instructions in the memory to execute the method in the first aspect.
[0023] A fourth aspect of the present application provides a computer-readable storage medium, which includes a stored program, and when the program is executed, controls a device where the computer-readable storage medium is located to execute the method in the first aspect.
[0024] A fifth aspect of the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a device, the method in the first aspect is implemented.
[0025] The prediction device based on stacking and machine learning models provided in the second aspect of this application, the electronic device provided in the third aspect, the computer-readable storage medium provided in the fourth aspect, and the computer program product provided in the fifth aspect have the same or similar beneficial effects as the prediction method based on stacking and machine learning models provided in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become easy to understand. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0027] Figure 1 A schematic diagram of a scenario architecture of a prediction method based on stacking and a machine learning model in an embodiment of the present application;
[0028] Figure 2The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 1 ;
[0029] Figure 3 Schematic diagram of the overall architecture of the prediction method based on stacking and machine learning model in the embodiment of the present application;
[0030] Figure 4 The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 2 ;
[0031] Figure 5 The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 3 ;
[0032] Figure 6 The structure of the prediction device based on stacking and machine learning model in the embodiment of the present application is shown in FIG. Figure 1 ;
[0033] Figure 7 The structure of the prediction device based on stacking and machine learning model in the embodiment of the present application is shown in FIG. Figure 2 ;
[0034] Figure 8 Schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0035] The exemplary embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0036] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in this application should have the common meanings understood by technicians in the field to which this application belongs.
[0037] Currently, the prediction accuracy using stacking technology cannot be significantly improved due to the limitations of the prediction accuracy of each base model and the limitations of the logistic regression algorithm in the meta-model.
[0038] In view of this, the embodiments of the present application provide a prediction method based on stacking and machine learning models, a prediction device based on stacking and machine learning models, an electronic device, a computer-readable storage medium, and a computer program product. After each base model outputs a prediction result, multiple prediction results are first statistically processed, and then the statistical results and each prediction result are input into the meta-model for prediction. In this way, the dimension of the input data in the meta-model can be increased, and the input of statistical results into the meta-model can increase the diversity of the meta-model's data processing methods. Overall, the accuracy of the prediction can be significantly improved.
[0039] First, the application scenario of the prediction method based on stacking and machine learning model provided in the embodiment of the present application is described.
[0040] Figure 1 This is a schematic diagram of the scenario architecture of the prediction method based on stacking and machine learning model in the embodiment of the present application, see Figure 1 As shown, the architecture may include: multiple base models and a meta-model.
[0041] The base model can make predictions based on the input and output the prediction results.
[0042] The meta-model can make decisions based on the outputs of multiple base models and output the decision results.
[0043] In practical applications, both the base model and the meta-model are ML models, specifically Auto ML models. The specific types of the base model and the meta-model are not limited here.
[0044] When business data needs to be predicted, the business data is input into each base model. Each base model predicts the business data and outputs the prediction results. The prediction results of multiple base models are then counted to obtain statistical results. The prediction results and statistical results of each base model are then input into the meta-model so that the meta-model makes decisions based on multiple prediction results and statistical results. Finally, the output of the meta-model is the final prediction result of the business data.
[0045] It should be noted here that the various models, data and their processing processes involved in the embodiments of this application have obtained relevant authorization in advance and are legal and compliant.
[0046] Next, the prediction method based on stacking and machine learning model provided in the embodiment of the present application is described in detail.
[0047] Figure 2 The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 1 , see Figure 2 As shown, the method may include:
[0048] S21: Acquire business data and multiple base models.
[0049] Business data is the data to be predicted. In other words, it is necessary to make inferences and judgments on the content in the business data. For example, the business data is a picture, and the face in the picture needs to be recognized.
[0050] The base model is used to make predictions based on business data. Since the base model is generally an ML model, the base model here is the trained model.
[0051] The base model can be provided at the same time as the business data provided by the user, or the default configuration can be used. When the base model uses the default configuration, the default base model is generally an untrained model. When the user provides business data, he or she also needs to provide training data with known prediction results of the same type as the business data. First, the model with the default configuration is trained with the training data, and after obtaining the base model, the subsequent steps are carried out.
[0052] S22: Inputting the business data into a plurality of base models respectively, and obtaining a plurality of prediction results corresponding to the outputs of the plurality of base models.
[0053] Since the base model is obtained by training with the same training data as the business data, when the business data is input into the base model, the base model can make corresponding predictions on the business data and output the prediction results.
[0054] Different base models have different network architectures or specific algorithms, and the prediction results output for the same business data may also be different. Input the business data into different base models, and different base models will output prediction results. The number of prediction results can be obtained as many base models as there are.
[0055] It should be noted that the base models used in the embodiments of the present application are all known base models, and the specific content and processing methods of the base models will not be described in detail here.
[0056] S23: Perform statistical processing on the multiple prediction results to obtain statistical results.
[0057] The statistics here may refer to processing procedures that are not involved in the algorithm in the meta-model. For example: the meta-model involves a logistic regression algorithm, but the logistic regression algorithm does not involve nonlinear calculations. Therefore, the statistics in this step may refer to various nonlinear calculations. Another example: the meta-model involves a logistic regression algorithm, and the logistic regression algorithm is a regression calculation based on multiple results. Therefore, the statistics in this step may also refer to various non-regression calculations. Of course, as long as the calculation is different from the logistic regression calculation but is related to the accuracy of the prediction results, it can be used as a statistical method in this step, and they will not be enumerated here one by one.
[0058] S24: Input the multiple prediction results and statistical results into the meta-model to obtain the final prediction result output by the meta-model.
[0059] The metamodel is used to make predictions based on the outputs of multiple base models. By inputting the prediction results of multiple base models and the statistical results of multiple prediction results into the metamodel, the metamodel can make predictions based on multiple preset results and statistical results. Since the metamodel takes more statistical results into account when making predictions, it not only increases the dimension of the metamodel input data, but also enriches the calculation method of the metamodel in disguise, thereby improving the accuracy of the final prediction results output by the metamodel.
[0060] Since the previous meta-model only used simple algorithms such as logistic regression and only processed the prediction results output by multiple base models, statistical results were added to the input of this meta-model, the dimension of data input increased, and the meta-model needed to be pre-trained based on the original data and the newly added data before it could be put into use.
[0061] During specific training, the initial meta-model is first constructed, and then the training data is input into multiple base models respectively, and then the prediction results output by the multiple base models are counted to obtain statistical results. Then, the statistical results, multiple prediction results, and the known true results of the training data are input into the initial meta-model for training to obtain a meta-model that can be used. During formal prediction, the output of the business data in multiple base models and the statistical results of each output are input into the meta-model. The output of the meta-model is the final prediction result of the business data.
[0062] From the above content, it can be seen that the prediction method based on stacking and machine learning models provided in the embodiment of the present application, after multiple base models output prediction results, the prediction results output by each base model are counted, and then the statistical results and the prediction results of each base model are input into the meta-model together, so that the meta-model is further analyzed and processed in combination with the statistical results based on the prediction results of multiple base models. Because the meta-model was previously only able to perform logistic regression calculations based on the prediction results of multiple base models, the input data dimension and processing power are limited. This time, while inputting the prediction results of multiple base models into the meta-model, the statistical results of multiple prediction results are also input, so that the data dimension and processing power processed by the meta-model can be improved, thereby improving the accuracy of the prediction of the machine learning model and stacking technology.
[0063] Furthermore, as a Figure 2 As a refinement and extension of the method shown, the embodiment of the present application also provides a prediction method based on stacking and machine learning models.
[0064] Figure 3 This is a schematic diagram of the overall architecture of the prediction method based on stacking and machine learning model in the embodiment of the present application, see Figure 3 As shown, the architecture may include three parts: input, optimization and output.
[0065] At the input, the user is required to provide multiple base models, data sets (training set, validation set and test set) and evaluation indicators.
[0066] Multiple base models can use the historical training models (trained model history) of any Auto ML platform, or the historical training models (trained model history) that the developer has tried.
[0067] In the data set, the training set and validation set are used to train the meta-model. The test set is the actual data set to be predicted.
[0068] Evaluation indicators, that is, the evaluation function used by the meta-model, for example: logloss function.
[0069] In the optimization part, it mainly includes: model screening module, feature construction module and Auto ML meta-model optimization module.
[0070] The model screening module is used to screen out a certain number of models with the best performance and large differences among the multiple base models provided. The number can be determined according to the actual situation, for example: 5. In practical applications, model screening can be performed through the loss function.
[0071] The feature construction module can further process multiple prediction results based on the prediction results of multiple base models after screening to construct more new feature values.
[0072] The metamodel optimization module of Auto ML mainly includes two functions. One is to optimize the metamodel based on the training set and the validation set, and the other is to use the optimized metamodel for prediction based on the test set.
[0073] The training set and validation set are input into multiple base models after screening, and multiple prediction results are obtained accordingly. Multiple prediction results are input into the feature construction module to obtain statistical results. Multiple prediction results based on the training set and validation set, as well as statistical results and known real results in the training set and validation set are input into the meta-model optimization module of Auto ML to complete the optimization of the meta-model.
[0074] The test set is input into multiple base models after screening, and multiple prediction results are obtained accordingly. Multiple prediction results are input into the feature construction module to obtain statistical results. Multiple prediction results based on the test set and statistical results are input into the optimized meta-model to obtain the final prediction result.
[0075] At the output, not only the final prediction results can be output, but also a set of models and functions (i.e., the Pipline of the optimal stacking model) including multiple base models after screening, feature construction methods, and parameters in the meta-model after optimization can be output, so that users can use the Pipline set by themselves to obtain the final prediction results of the test set.
[0076] Figure 4 The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 2 , see Figure 4 As shown, the method may include:
[0077] S41: Acquire business data and multiple base models.
[0078] When a user needs to make a prediction (eg, classification) on the content in business data, the user may input the business data and a plurality of base models for prediction.
[0079] The business data here can include the actual data to be predicted, as well as the data used for model training with known prediction results. In other words, the business data can include training sets, validation sets, and test sets. The training set and validation set are the data used for model training with known prediction results. The test set is the data to be predicted.
[0080] Since different base models have different prediction accuracy for business data, and the base model with more inaccurate prediction among two base models with similar prediction results will affect the accuracy of the final prediction, for example: the business data is an image, and it is necessary to identify the face of a in the image. Base model 1 identifies face 1, and base model 2 identifies face 2. In fact, the face of a is face 1. Base model 2 can also identify faces, and it identifies face 2 that is more similar to face a. If base model 2 is still used to participate in the final prediction, it is easy for face 2 to reduce the probability of using face 1 as the final prediction result, thereby reducing the accuracy of the prediction. At this time, it is necessary to delete the base models with low prediction accuracy and similar to other base models among multiple base models.
[0081] S42: Determine the prediction accuracy parameter of each base model, and determine the prediction similarity between any two base models.
[0082] For determining the prediction accuracy parameter, the data in the training set and the validation set can be input into the base model, and the base model can output the prediction result based on the data, and the similarity between the prediction result and the known true result of the data in the training set and the validation set is calculated, and then the prediction accuracy parameter of the base model is determined based on the similarity calculation result. The higher the calculated similarity, the greater the prediction accuracy parameter of the base model.
[0083] In order to improve the accuracy of determining the prediction accuracy parameters of the base model, the above similarity calculation can be performed using multiple data in the training set and the validation set, and the similarity calculation results can be averaged to determine the prediction accuracy parameters of the corresponding base model.
[0084] Of course, in order to improve the efficiency of determining the prediction accuracy parameters of the base model, a small amount of data in the training set and the validation set can be used to perform the above similarity calculations respectively, and the similarity calculation results can be averaged to determine the prediction accuracy parameters of the corresponding base model. Even more, a specified data in the training set and the validation set can be used to perform the above similarity calculations, so as to determine the prediction accuracy parameters of the corresponding base model according to the similarity calculation results.
[0085] The determination of prediction similarity needs to be done between two base models. First, one or more data can be selected from the training set, validation set, and test set, and then the selected data is input into each base model to obtain the prediction results of each base model. Then, the similarity between the corresponding prediction results of the two base models is calculated. In this way, the prediction similarity between the two base models is obtained.
[0086] If one data is selected from the training set, validation set, and test set, the prediction similarity between the two base models is the similarity between the prediction results of the base models based on the data. If multiple data are selected from the training set, validation set, and test set, the prediction similarity between the two base models is the average of the similarities between the prediction results of one base model based on each data and the prediction results of another base model based on each data.
[0087] As for the amount of data selected from the training set, validation set, and test set, similar to the determination of the prediction accuracy parameter, it needs to be determined comprehensively based on the requirements of accuracy and efficiency.
[0088] S43: Based on the prediction accuracy parameter and the prediction similarity, one or more base models from the multiple base models are deleted to obtain a preset number of target base models.
[0089] Among them, the lower the prediction accuracy parameter, the higher the prediction similarity of the base model, the higher the priority of being deleted. The priority of the prediction accuracy parameter is higher than the priority of the prediction similarity. That is to say, among two similar base models with relatively high prediction accuracy parameters and one base model with low prediction accuracy parameter, the base model with low prediction accuracy parameter among the two similar base models is deleted. In this way, the base models with large prediction differences can be retained to the maximum extent, and the adverse interference of similar base models can be reduced.
[0090] When deleting, you can sort the prediction similarity from high to low, and delete the base models with low prediction accuracy parameters in the two similar base models in turn. You can also sort the prediction accuracy parameters from high to low, and delete the base models that are most similar to the current base model in turn. Every time you delete a base model, you can confirm whether the remaining base models meet the quantity requirements. If they do, you can stop deleting. If they do not, continue to delete to achieve precise control of the number of base model screening. You can also determine whether the remaining base models meet the quantity requirements after deleting multiple base models to improve the efficiency of model deletion. In order to retain the base models with large prediction differences to the maximum extent and reduce the adverse interference of similar base models, you can use the prediction accuracy parameter sorting to delete them.
[0091] Specifically, the above step S43 may include:
[0092] Step A1: taking a base model with the largest prediction accuracy parameter among multiple base models as the current base model, and taking base models other than the current base model among the multiple base models as other base models.
[0093] Step A2: Among other base models, delete the base model with the highest prediction similarity to the current base model to obtain a plurality of remaining base models including the current base model.
[0094] Step A3: If the number of the multiple remaining base models is greater than the preset number, then continue to use the base model with the largest prediction accuracy parameter among the multiple remaining base models as the new current base model, and use the base models among the multiple remaining base models except the new current base model as new other base models, and delete the base model with the highest prediction similarity with the new current base model from the new other base models to obtain new remaining base models including the new current base model, until the number of new remaining base models is equal to the preset number.
[0095] For example, assuming that the number of multiple base models is i, and the number of base models after screening is 5. These i base models all have corresponding prediction accuracy parameters. Sort these i prediction accuracy parameters in order from large to small. For the first-ranked base model, it has corresponding prediction similarities with the remaining i-1 base models. Among these i-1 prediction similarities, the corresponding base model with the highest prediction similarity is selected and deleted. At this time, i-1 base models remain. If the number of i is 10, there are still 9 base models left at this time, and the standard of remaining 5 base models has not been met, and the base models continue to be deleted. Next, for the second-ranked base model, it has corresponding prediction similarities with the remaining i-2 base models. Among these i-2 prediction similarities, the corresponding base model with the highest prediction similarity is selected and deleted. At this time, i-2 base models remain. Repeat this process until the number of remaining i base models is 5. The number of base models after screening, that is, the preset number, can be determined according to actual conditions and is not specifically limited here.
[0096] If the predictions between two models are not similar, that is, the prediction similarity between two base models is less than or equal to the similarity threshold, the base models will be deleted in the above way, which will result in the deletion of base models with large differences. This will reduce the accuracy of subsequent meta-model comprehensive decision-making. Therefore, other methods need to be used to delete base models.
[0097] Specifically, before the above step A2, the method may further include:
[0098] Step A01: If the prediction similarity between the two base models is less than or equal to the similarity threshold, then determine the data quantity representation, category quantity representation and / or data generation distance time representation of the prediction difference between the multiple prediction results corresponding to the two base models respectively.
[0099] If the prediction similarities between any two base models are less than or equal to the similarity threshold, it means that the predictions of these models are not similar. The base models can be deleted no longer by sorting according to the prediction accuracy parameter, but by sorting according to the prediction similarity. In the order from large to small, among the prediction differences between the current base model and each other base model, the corresponding other base models with the smallest prediction difference can be deleted.
[0100] If, in the process of deleting base models according to the prediction accuracy parameter, for the current base model, if it is found that the prediction similarities with other multiple base models are relatively low, and the largest prediction similarity is less than the similarity threshold, then the base model with the smallest prediction difference with other base models is deleted from the other multiple base models. Therefore, it is necessary to further determine the prediction difference value representation between the multiple prediction results corresponding to each of the two base models, that is, the data quantity representation, category quantity representation and / or data generation distance time representation of the prediction difference.
[0101] The so-called number of prediction difference data may refer to the number of prediction results of a base model that are different from the prediction results of another base model. The number of prediction difference data may be the number of prediction difference data or the proportion of the number of prediction difference data in the number of prediction results.
[0102] The number of categories of prediction differences may refer to the number of categories in which multiple prediction results of a base model differ from multiple prediction results corresponding to another base model. The number of categories of prediction differences may be represented by the number of categories of prediction differences or the proportion of the number of categories of prediction differences in the total number of categories of multiple prediction results.
[0103] The so-called prediction difference data generation distance time may refer to the maximum time difference in the generation time of samples (i.e., data input to the base model) corresponding to the differences between the multiple prediction results of one base model and the multiple prediction results corresponding to another base model. The prediction difference data generation distance time representation may be the prediction difference data generation distance time, or the proportion of the prediction difference data generation distance time in the maximum generation distance time between all data.
[0104] For the convenience of calculation, the data generation distance may be a specific time. When each data in the business data is sorted according to the generation time, the data index may be used instead of the data generation time.
[0105] The data quantity representation, category quantity representation and data generation distance and duration representation of the predicted difference can be used one by one or in combination, and are not limited here.
[0106] Step A02: based on the prediction accuracy parameter and the data quantity representation, the category quantity representation and / or the data generation distance duration representation, one or more base models among the multiple base models are deleted to obtain a preset number of target base models.
[0107] Among them, the lower the prediction accuracy parameter, the smaller the data quantity representation, the smaller the category number representation, and the smaller the data generation distance time representation of the base model, the more priority it will be deleted. The priority of the prediction accuracy parameter is higher than the priority of the data quantity representation, the category number representation, and the data generation distance time representation.
[0108] That is to say, multiple base models are sorted according to the prediction accuracy parameter. For the current base model, if the prediction similarity with other base models is less than the similarity threshold, then among the other base models, the base model with the smallest number of data, the smallest number of categories, or the shortest data generation distance difference from the current base model is selected for deletion.
[0109] For example, assume that the following Table 1 shows the results of classification prediction of some data using base model 1 and base model 2 respectively.
[0110] Table 1 Classification results of some data using base model 1 and base model 2
[0111]
[0112]
[0113] In Table 1, data indexes 0-19 represent 20 data respectively. Base model 1 and base model 2 make classification predictions for each data respectively. 0, 1, 2, 3 indicate the 4 different classifications predicted.
[0114] If base model 1 is used as the current base model, and the prediction similarity between base model 1 and base model 2 is not high, and the prediction similarity between base model 1 and base model 3 is not high either, then base models are no longer selected for deletion from base model 2 and base model 3 based on prediction similarity. Instead, base models are selected for deletion from base model 2 and base model 3 based on the number of data, number of categories, and / or data generation distance of the prediction difference.
[0115] Compared with base model 1, base model 2 has 4 data differences in category judgment. Compared with base model 1, base model 3 has 5 data differences in category judgment. According to the number of data with predicted differences, base model 2 with the least number of predicted difference data is selected for deletion.
[0116] Compared with base model 1, base model 2 has 4 data with different category judgments, namely, category 1 and 3 are judged as category 2, and category 2 and 1 are judged as category 0, and the number of predicted difference categories is 2. Compared with base model 1, base model 3 has 5 data with different category judgments, namely, category 1 is judged as category 2, category 3 is judged as category 1, category 2 is judged as category 0, and category 1 is judged as category 0, and the number of predicted difference categories is 3. According to the number of predicted difference categories, base model 2 with a small number of predicted difference categories is selected for deletion.
[0117] Compared with base model 1, base model 2 has 4 data with different classifications. The largest difference in generation time between these 4 data is 10 days. Compared with base model 1, base model 3 has 5 data with different classifications. The largest difference in generation time between these 5 data is 2 days. According to the generation time of the predicted difference data, base model 3 with the shortest generation time of the predicted difference data is selected and deleted.
[0118] If the number of data, number of categories, or data generation distance length of the predicted difference is used, one of the two base models can be directly selected for deletion. However, if the number of data, number of categories, and data generation distance length of the predicted difference are used, if the base models selected for the number of data, number of categories, and data generation distance length of the predicted difference are different, it is difficult to select one base model for deletion. In this case, the number of data, number of categories, and data generation distance length of the predicted difference can be expressed through a function, and then the base model to be deleted can be determined based on the calculation result of the function.
[0119] The specific functions are as follows:
[0120] F(x) = accuracy + difference sample distribution coefficient × difference category coefficient formula (1)
[0121] Among them, F(x) is used to represent the similarity.
[0122] Accuracy, used to characterize the number of data points that predict differences. Accuracy = number of different predictions / total number of predictions.
[0123] The difference sample distribution coefficient is used to characterize the distance between the data generation of the prediction difference. The difference sample distribution coefficient = the index of the samples with inconsistent prediction results minus the total number of records.
[0124] The difference category coefficient is used to characterize the number of categories with different predictions. Difference category coefficient = total number of predicted different categories / total number of predicted categories.
[0125] The smaller the F(x), the higher the probability that the corresponding base model will be deleted. For the current base model, each other base model has an F(x) with the current base model. From these F(x), select the base model corresponding to the F(x) with the smallest value and delete it, and the most similar base model corresponding to the current base model is deleted. After multiple deletions, the number of remaining base models reaches the preset number, and then stop deleting to obtain the target base model.
[0126] S44: Inputting the business data into the target base models respectively, and obtaining multiple prediction results corresponding to the outputs of the multiple base models.
[0127] Generally speaking, there are multiple target base models. The business data is input into each base model respectively, and each base model has a corresponding output prediction result. At this time, multiple prediction results of the business data under different base models are obtained.
[0128] S45: Based on each prediction result among the multiple prediction results, respectively determine the physical metric value of the base model corresponding to each prediction result to obtain multiple physical metric values, and perform optimal processing among the multiple physical metric values to obtain an optimal metric value.
[0129] Through the prediction results, the prediction of the base model can also be characterized in reverse. That is, the prediction results are converted into physical measurement values of the base model, and then the best is selected from various physical measurement values. The specific content of the physical measurement value can be determined according to the actual prediction requirements. The physical measurement value can include but is not limited to confidence, accuracy, etc.
[0130] The confidence of the base model can be determined by using the characteristics of the base model. That is, the model outputs the confidence while outputting the prediction result. Therefore, the confidence of the corresponding base model can be found from the data packet of the prediction result, and the maximum confidence found can be determined as one of the statistical results.
[0131] The specific value of the confidence can directly reflect the "unconfidence" of the base model for this data. The specific value of the maximum confidence can reflect the least "confidence" of all base models for this data. In multi-classification scenarios, for difficult samples, the maximum confidence is often only higher than 0.5 and cannot reach above 0.98.
[0132] S46: Perform nonlinear processing on at least two prediction results among the multiple prediction results to obtain nonlinear results.
[0133] Any processing method other than linear processing can be nonlinear processing.
[0134] In nonlinear processing, all of the multiple prediction results may be processed, or some of the prediction results may be processed first and then the processed prediction results may be integrated.
[0135] The nonlinear processing may include at least one of the following: calculating a mode, calculating a variance, calculating a mean, calculating a pairwise difference, calculating a pairwise sum, and an optimal model.
[0136] The mode here may refer to the prediction result with the largest number of identical prediction results among multiple prediction results, which can reflect the maximum possibility of judgment of multiple "differentiated" base models.
[0137] Variance, here it can refer to the variance of multiple prediction results. The larger the variance, the more difficult it is to judge the business data.
[0138] The average value here can refer to the average result of multiple prediction results, and can also reflect the maximum possibility of judgments of multiple "differentiated" base models.
[0139] The difference between the two models can refer to the difference between the prediction results of the two base models.
[0140] The pairwise sum, here can refer to the sum of the prediction results of the two base models.
[0141] The optimal model is the pseudo-optimal model index. If it is in the training phase, the pseudo-optimal model index is the index corresponding to the model with the highest confidence among the correct models. If it is in the testing phase, the base model with the highest confidence among the base models with the majority prediction results is used.
[0142] It should be noted that the above steps S45 and S46 can be executed simultaneously to improve prediction efficiency. Steps S45 and S46 can also be executed asynchronously to avoid mutual influence between the steps and improve prediction accuracy.
[0143] S47: Determine the physical measurement value and nonlinear result of each base model as a statistical result.
[0144] The above statistical results can more fully characterize the relationship between each model and the relationship between each prediction result in addition to the prediction results of each base model, so as to improve the accuracy of the prediction.
[0145] It should be noted that the business data processed in the above step S44 may include a training set, a validation set, and a test set. The results corresponding to the data in the training set and the validation set are known, while the results corresponding to the data in the test set are unknown, and are also the data that actually needs to be predicted this time. In other words, in the above step S44, each data in the training set, the validation set, and the test set needs to be predicted using the target base model, and each data corresponds to multiple prediction results equal to the number of the target base models.
[0146] Correspondingly, in the above step S47, each data in the training set, the validation set and the test set corresponds to a statistical result.
[0147] Compared with the previous meta-model that only made decisions based on multiple prediction results, the meta-model has added statistical results of multiple prediction results in the decision-making process this time, and the data dimension processed by the meta-model has increased. In order to enable the meta-model to make accurate decisions, before inputting multiple prediction results and statistical results corresponding to the test set into the meta-model for prediction, the meta-model is first trained using multiple prediction results and statistical results of each data in the training set and validation set as well as the known results of each data.
[0148] The metamodel may include multiple preset algorithms. Users can select these preset algorithms according to actual needs. Alternatively, if the user does not specify, the metamodel directly selects the preset algorithms according to the default rules.
[0149] In practical applications, the preset algorithm may be a multi-classification algorithm such as a logistic regression algorithm and a Bayesian classification algorithm.
[0150] S48: Input the multiple prediction results, statistical results and the specified algorithm name into the meta-model, so that the meta-model searches for the target algorithm corresponding to the specified algorithm name in the multiple preset algorithms, and then processes the multiple prediction results and statistical results based on the target algorithm.
[0151] After receiving multiple prediction results, statistical results and specified algorithm names, the meta-model first searches for the corresponding algorithm, i.e., the target algorithm, among multiple preset algorithms according to the specified algorithm name. If no unique algorithm is accurately found according to the specified algorithm name, the meta-model searches for the most similar algorithm among multiple preset algorithms according to the specified algorithm name to ensure smooth prediction.
[0152] After the target algorithm is determined, the meta-model is trained using multiple prediction results, statistical results, and known results corresponding to the data in the training set and validation set of the business data, that is, the parameters in the target algorithm are optimized to obtain the trained meta-model. The data in the test set of the business data is then input into the trained meta-model, and the prediction results output by the trained meta-model through the processing of the target algorithm after parameter optimization are more accurate final prediction results.
[0153] Finally, the prediction method based on stacking and machine learning model provided in the embodiment of the present application is explained again with a specific example.
[0154] Figure 5 The process diagram of the prediction method based on stacking and machine learning model in the embodiment of the present application is as follows: Figure 3 , see Figure 5 As shown in the figure, when the user needs to predict certain data, the data is used as a test set, and the data with known results of the same type as the data is used as a training set and a validation set, and then the training set, validation set and test set are input into the system as business data. At the same time, multiple base models are also input into the system. Multiple base models may include: Trained model of SVM, MLP, LightBGM, ..., KNN, LSTM, etc.
[0155] The business data and these base models first enter the model screening module of the system. The model screening module calculates the prediction accuracy parameters and prediction similarity of these base models based on the business data, and selects the TOP5 best models from these base models based on the prediction accuracy parameters and prediction similarity, namely, the trained model of SVM, MLP, LightBGM, KNN and LSTM.
[0156] After the TOP5 optimal models are determined, the business data can be input into the TOP5 optimal models respectively. The prediction category output by the Trained model of SVM is 1, the prediction category output by MLP is 2, the prediction category output by LightBGM is 2, the prediction category output by KNN is 0, and the prediction category output by LSTM is 2. At this time, a 5-dimensional array is obtained, that is, [1, 2, 2, 0, 2].
[0157] Next, the prediction categories corresponding to the five optimal models are input into the feature construction module of the system. Based on the prediction categories corresponding to the five optimal models, the feature construction module generates the maximum confidence of the prediction categories of the five optimal models, the mode of the prediction categories of the five optimal models, the variance of the prediction categories of the five optimal models, the average of the prediction categories of the five optimal models, the pseudo-optimal model index, the pairwise difference of the prediction categories corresponding to the five optimal models, and the pairwise sum of the prediction categories corresponding to the five optimal models. The prediction categories corresponding to the five optimal models belong to the original features. The maximum confidence of the prediction categories of the five optimal models, the mode of the prediction categories of the five optimal models, the variance of the prediction categories of the five optimal models, the average of the prediction categories of the five optimal models, the pseudo-optimal model index, the pairwise difference of the prediction categories corresponding to the five optimal models, and the pairwise sum of the prediction categories corresponding to the five optimal models belong to the newly added features.
[0158] The original features and the newly added features are input into the system's Auto ML metamodel optimization module. The Auto ML metamodel optimization module includes a search space and a search process. The search space may include a naive Bayes search space and a logistic regression search space, etc. The search process may include an evolutionary algorithm search strategy and an early stopping strategy, etc. The naive Bayes search space involves a Bayesian kernel. The Bayesian kernel may include GaussianNB, MultinomialNB, BernouliNB, etc. The logistic regression search space involves logistic regression parameters. The logistic regression parameters may include regularization type, regularization strength, maximum number of iterations, number of iterations, classification method, etc. After optimizing the parameters involved in the AutoML metamodel using the original features and newly added features corresponding to the training set and the validation set, the optimized Auto ML metamodel is used to process the original features and newly added features corresponding to the test set. The output of the Auto ML metamodel optimization module is the final prediction result.
[0159] At this point, the prediction method based on stacking and machine learning model provided in the embodiments of the present application has been fully explained.
[0160] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application also provides a prediction device based on stacking and machine learning models.
[0161] Figure 6 The structure of the prediction device based on stacking and machine learning model in the embodiment of the present application is shown in FIG. Figure 1 , see Figure 6 As shown, the device may include: an acquisition module 61 , a first processing module 62 , a second processing module 63 and a prediction module 64 .
[0162] The acquisition module 61 is used to acquire business data and multiple base models, where the business data is data to be predicted and the base models are used to make predictions based on the business data.
[0163] The first processing module 62 is used to input the business data into multiple base models respectively to obtain multiple prediction results corresponding to the outputs of the multiple base models.
[0164] The second processing module 63 is used to perform statistical processing on the multiple prediction results to obtain statistical results.
[0165] The prediction module 64 is used to input multiple prediction results and statistical results into the meta-model to obtain the final prediction result output by the meta-model. The meta-model is used to make predictions based on the outputs of multiple base models. The meta-model is trained based on the initial meta-model using training data. The training data includes the training outputs of multiple base models and the statistical results of the training outputs.
[0166] Furthermore, as a Figure 6 As a refinement and extension of the device shown, the embodiment of the present application also provides a prediction device based on stacking and machine learning models.
[0167] Figure 7 The structure of the prediction device based on stacking and machine learning model in the embodiment of the present application is shown in FIG. Figure 2 , see Figure 7 As shown, the device may include:
[0168] The acquisition module 71 is used to acquire business data and multiple base models, where the business data is data to be predicted and the base models are used to make predictions based on the business data.
[0169] The screening module 72 includes: a determining unit 721 and a deleting unit 722 .
[0170] The determination unit 721 is used to determine the prediction accuracy parameter of each base model and determine the prediction similarity between any two base models.
[0171] The deletion unit 722 is used to delete one or more base models from the multiple base models based on the prediction accuracy parameter and the prediction similarity, and obtain a preset number of target base models so that the business data can be input into the target base models respectively, wherein the base models with lower prediction accuracy parameters and higher prediction similarity are deleted with higher priority, and the priority of the prediction accuracy parameter is higher than the priority of the prediction similarity.
[0172] The deletion unit 722 is specifically used to use the base model with the largest prediction accuracy parameter among multiple base models as the current base model, and use the base models other than the current base model among the multiple base models as other base models; among the other base models, delete the base model with the highest prediction similarity with the current base model to obtain multiple remaining base models including the current base model; if the number of the multiple remaining base models is greater than the preset number, continue to use the base model with the largest prediction accuracy parameter among the multiple remaining base models as the new current base model, and use the base model other than the new current base model among the multiple remaining base models as new other base models, and among the new other base models, delete the base model with the highest prediction similarity with the new current base model to obtain new remaining base models including the new current base model, until the number of new remaining base models is equal to the preset number.
[0173] The deletion unit 722 is also used to determine the data quantity representation, category quantity representation and / or data generation distance time representation of the prediction difference between multiple prediction results corresponding to each of the two base models if the prediction similarity between the two base models is less than or equal to the similarity threshold; based on the prediction accuracy parameter and the data quantity representation, category quantity representation and / or data generation distance time representation, delete one or more base models from the multiple base models to obtain a preset number of target base models, so that the business data can be input into the target base models respectively, wherein the base models with lower prediction accuracy parameters, smaller data quantity representation, smaller category quantity representation and smaller data generation distance time representation are deleted first, and the priority of the prediction accuracy parameter is higher than the priority of the data quantity representation, the category quantity representation and the data generation distance time representation.
[0174] The first processing module 73 is used to input the business data into multiple base models respectively to obtain multiple prediction results corresponding to the outputs of the multiple base models.
[0175] The second processing module 74 includes: a first processing unit 741 , a second processing unit 742 and a collection unit 743 .
[0176] The first processing unit 741 is used to determine the physical metric value of the base model corresponding to each prediction result based on each prediction result in the multiple prediction results, obtain multiple physical metric values, and perform optimal processing among the multiple physical metric values to obtain the optimal metric value.
[0177] The second processing unit 742 is configured to perform nonlinear processing on at least two prediction results among the multiple prediction results to obtain nonlinear results.
[0178] The aggregation unit 743 is used to determine the optimal metric value and the nonlinear result as a statistical result.
[0179] The physical measurement value includes confidence. The nonlinear processing includes at least one of the following: calculating the mode, calculating the variance, calculating the mean, calculating the difference between two pairs, calculating the sum of two pairs, and the optimal model.
[0180] When the meta-model includes multiple preset algorithms, the prediction module 75 is used to input multiple prediction results, statistical results and specified algorithm names into the meta-model, so that the meta-model searches for the target algorithm corresponding to the specified algorithm name in the multiple preset algorithms, and then processes the multiple prediction results and statistical results based on the target algorithm.
[0181] It should be noted here that the description of the above device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the device embodiment of the present application, please refer to the description of the method embodiment of the present application for understanding.
[0182] Based on the same inventive concept, an embodiment of the present application also provides an electronic device.
[0183] Figure 8 This is a schematic diagram of the structure of the electronic device in the embodiment of the present application, see Figure 8 As shown, the electronic device may include: a processor 81, a memory 82, and a bus 83. The processor 81 and the memory 82 communicate with each other through the bus 83. The processor 81 is used to call program instructions in the memory 82 to execute the methods in one or more of the above embodiments.
[0184] It should be noted that the description of the above electronic device embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the electronic device embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0185] Based on the same inventive concept, an embodiment of the present application further provides a computer-readable storage medium, which may include: a stored program that controls the device where the storage medium is located to execute the method in one or more of the above embodiments when the program is running.
[0186] It should be noted here that the description of the above computer-readable storage medium embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the computer-readable storage medium embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0187] Based on the same inventive concept, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by the device where they are located, the methods in one or more of the above embodiments are implemented.
[0188] It should be noted here that the description of the above computer program product embodiment is similar to the description of the above method embodiment, and has similar beneficial effects as the method embodiment. For technical details not disclosed in the computer program product embodiment of this application, please refer to the description of the method embodiment of this application for understanding.
[0189] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A prediction method based on stacking and machine learning model, characterized in that: The method comprises: Acquire business data and multiple base models, wherein the business data is data to be predicted, and the base models are used to make predictions based on the business data; Inputting the business data into the multiple base models respectively to obtain multiple prediction results outputted by the multiple base models; Performing statistical processing on the multiple prediction results to obtain statistical results; The multiple prediction results and the statistical results are input into a meta-model to obtain a final prediction result output by the meta-model, wherein the meta-model is used to make predictions based on the outputs of the multiple base models, and the meta-model is trained based on an initial meta-model using training data, wherein the training data includes the training outputs of the multiple base models and the statistical results of the training outputs.
2. The method according to claim 1, characterized in that The statistical processing of the plurality of prediction results to obtain statistical results includes: Based on each of the multiple prediction results, respectively determine the physical metric value of the base model corresponding to each prediction result to obtain multiple physical metric values, and perform optimal processing among the multiple physical metric values to obtain an optimal metric value; Performing nonlinear processing on at least two prediction results among the multiple prediction results to obtain nonlinear results; The optimal metric value and the nonlinear result are determined as the statistical result.
3. The method according to claim 2, characterized in that The physical measurement value includes a confidence level, and the nonlinear processing includes at least one of the following: calculating a mode, calculating a variance, calculating a mean, calculating a pairwise difference, calculating a pairwise sum, and an optimal model.
4. The method according to claim 1, characterized in that: The meta-model includes a plurality of preset algorithms; the step of inputting the plurality of prediction results and the statistical results into the meta-model includes: The multiple prediction results, the statistical results and the specified algorithm name are input into the meta-model, so that the meta-model searches for a target algorithm corresponding to the specified algorithm name among the multiple preset algorithms, and then processes the multiple prediction results and the statistical results based on the target algorithm.
5. The method according to any one of claims 1 to 4, characterized in that Before inputting the business data into the multiple base models respectively, the method further includes: Determining a prediction accuracy parameter for each base model, and determining prediction similarities between any two base models; Based on the prediction accuracy parameter and the prediction similarity, one or more base models among the multiple base models are deleted to obtain a preset number of target base models, so that the business data can be input into the target base models respectively, wherein the base models with lower prediction accuracy parameters and higher prediction similarity are deleted with higher priority, and the priority of prediction accuracy parameters is higher than that of prediction similarity.
6. The method according to claim 5, characterized in that The deleting one or more base models from the plurality of base models based on the prediction accuracy parameter and the similarity to obtain a preset number of target base models includes: Using a base model with the largest prediction accuracy parameter among the multiple base models as a current base model, and using base models among the multiple base models except the current base model as other base models; Among the other base models, delete the base model with the highest prediction similarity to the current base model to obtain a plurality of remaining base models including the current base model; If the number of the multiple remaining base models is greater than the preset number, the base model with the largest prediction accuracy parameter among the multiple remaining base models will continue to be used as the new current base model, and the base models among the multiple remaining base models except the new current base model will be used as new other base models, and among the new other base models, the base model with the highest prediction similarity with the new current base model will be deleted to obtain new remaining base models including the new current base model, until the number of the new remaining base models is equal to the preset number.
7. The method according to claim 5, characterized in that The method further comprises: If the prediction similarity between the two base models is less than or equal to the similarity threshold, then determining the data quantity representation, category quantity representation and / or data generation distance time length representation of the prediction difference between the multiple prediction results corresponding to the two base models respectively; Based on the prediction accuracy parameter and the data quantity representation, category quantity representation and / or data generation distance time representation, one or more base models among the multiple base models are deleted to obtain a preset number of target base models, so that the business data can be input into the target base models respectively, wherein the base models with lower prediction accuracy parameters, smaller data quantity representation, smaller category quantity representation and smaller data generation distance time representation are deleted first, and the priority of prediction accuracy parameters is higher than that of data quantity representation, category quantity representation and data generation distance time representation.
8. A prediction device based on stacking and machine learning model, characterized in that: The device comprises: An acquisition module, used to acquire business data and a plurality of base models, wherein the business data is data to be predicted, and the base models are used to make predictions based on the business data; A first processing module, used for inputting the business data into the multiple base models respectively to obtain multiple prediction results outputted correspondingly by the multiple base models; A second processing module, used for performing statistical processing on the multiple prediction results to obtain statistical results; A prediction module is used to input the multiple prediction results and the statistical results into a meta-model to obtain a final prediction result output by the meta-model. The meta-model is used to make predictions based on the outputs of the multiple base models. The meta-model is trained based on an initial meta-model using training data. The training data includes the training outputs of the multiple base models and the statistical results of the training outputs.
9. An electronic device, characterized in that: The electronic device includes a processor, a memory and a bus, the processor and the memory communicate with each other through the bus, and the processor is used to call program instructions in the memory to execute the method as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, and when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
11. A computer program product, characterized in that The computer program product comprises a computer program or instructions, and when the computer program or instructions are executed by a device, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Machine learning method and device fusing artificial experience and ensemble learning strategy
CN112598134A
Flow prediction method and device based on stacking algorithm, and related equipment
CN114647684A
Construction and prediction method of carbon emission prediction model based on Stacking algorithm and medium
CN115860173A
Subway facility maintenance management method and system based on Internet of Things technology
CN116205636A
Training method of error predictor and transaction data prediction method and device
CN117056694A