Device and method for recommending pipelines for crowd intelligence model

By designing a multi-module collaborative device and method, the sampling scores of new pipelines are calculated using the diversity values ​​between algorithms, correlation values ​​between features and distance values ​​between hyperparameters of the same algorithm, the problem of lack of accurate recommendation of Zhongzhi model pipelines in the existing technology is solved, and a more efficient pipeline recommendation process is achieved.

CN120069128APending Publication Date: 2025-05-30IND TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311722219.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2023-12-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The lack of methods for accurately recommending pipelines for smart models in the prior art has led to data scientists finding the best hyperparameter combination manually, which is time-consuming and inefficient.

Method used

A device and method are designed, including multiple modules: data extraction and pipeline initialization, pipeline efficiency evaluation, pipeline sampling score calculation, pipeline recommendation and multi-intelligent model recommendation. Through the coordinated work of these modules, the sampling scores of new pipelines are calculated using the inter-algorithm diversity values, inter-either correlation values ​​and the distance values ​​of the same algorithm hyperparameters, thereby recommending the best pipeline.

Benefits of technology

This achieves more accurate acquisition of sampling scores for newly entered pipelines, thereby more accurately recommending pipelines for the smart model, reducing the time and effort of data scientists to operate manually.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069128A_ABST
    Figure CN120069128A_ABST
Patent Text Reader

Abstract

The invention provides a device and a method for recommending pipelines for a crowd intelligence model. The method comprises the following steps: the data extraction and pipeline initialization module generates an initial pipeline by using an algorithm; the pipeline efficiency evaluation module obtains a prediction result corresponding to the initial pipeline by using the data set, and obtains an accuracy rate corresponding to the initial pipeline by using the data set; the pipeline sampling score calculation module obtains an inter-algorithm diversity value, an inter-feature correlation value and a same-algorithm hyper-parameter distance value by using the prediction result, the accuracy rate and the new pipeline, and obtains a sampling score corresponding to the new pipeline by using the inter-algorithm diversity value, the inter-feature correlation value and the same-algorithm hyper-parameter distance value; the pipeline recommendation module determines a recommended new pipeline according to the sampling score; and the crowd intelligence model recommendation module determines a target recommendation new pipeline by using a crowd intelligence model technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and method for a crowdsourcing model recommendation pipeline. Background Art

[0002] Artificial intelligence (AI) technologies are based on and constructed by combining models trained by machine learning, deep learning, ensemble learning, or reinforcement learning. In general artificial intelligence machine learning applications, the data science processing process is that data scientists first collect a large amount of data. After data exploration, data processing, selection of model algorithms, feature engineering, evaluation of models, and continuous adjustment of parameters, a useful model is finally trained. However, whether it is constructing training data, selecting model algorithms, or finding combinations of hyperparameters based on experience, data scientists need to manually complete the best machine learning model in a batch execution manner. After the machine learning model is completed, the same manual batch method is also used for data preprocessing, obtaining model inferences, and data postprocessing to integrate and apply, and this batch process is continuously repeated to maintain the accuracy of the inference results. To reduce the time-consuming manual search by data scientists for the best combination of hyperparameters, automated machine learning (AutoML) technology that can more efficiently search for combinations of hyperparameters has gradually received attention.

[0003] In the automated machine learning (AutoML) model technology, crowdsourcing models are often used as one of the techniques for selecting the best combination from multiple combinations of hyperparameters. However, there is currently a lack of a method that can accurately recommend pipelines (i.e., combinations of hyperparameters) for crowdsourcing models. Summary of the Invention

[0004] The present invention provides an apparatus and method for recommending a pipeline for a crowdsourcing model, which can more accurately recommend a pipeline for the crowdsourcing model.

[0005] The apparatus for a crowdsourced model recommendation pipeline according to the present invention includes a storage medium and a processor. The storage medium stores a plurality of modules, where the plurality of modules include a data extraction and pipeline initialization module, a pipeline performance evaluation module, a pipeline sampling score calculation module, a pipeline recommendation module, and a crowdsourced model recommendation module. The processor is coupled to the storage medium and accesses and executes the plurality of modules, where the data extraction and pipeline initialization module generates an initial pipeline using an algorithm; the pipeline performance evaluation module obtains a prediction result corresponding to the initial pipeline using a data set, and obtains an accuracy rate corresponding to the initial pipeline using the data set; the pipeline sampling score calculation module obtains an inter-algorithm diversity value, an inter-feature correlation value, and an inter-algorithm hyperparameter distance value using the prediction result, the accuracy rate, and a new incoming pipeline, and the pipeline sampling score calculation module obtains a sampling score corresponding to the new incoming pipeline using the inter-algorithm diversity value, the inter-feature correlation value, and the inter-algorithm hyperparameter distance value; the pipeline recommendation module determines a recommended new incoming pipeline among the new incoming pipelines using the sampling score; and the crowdsourced model recommendation module determines a target recommended new incoming pipeline among the recommended new incoming pipelines using crowdsourced model technology.

[0006] The method for a crowdsourced model recommendation pipeline according to the present invention includes the following steps: generating an initial pipeline using an algorithm by the data extraction and pipeline initialization module; obtaining a prediction result corresponding to the initial pipeline using a data set by the pipeline performance evaluation module, and obtaining an accuracy rate corresponding to the initial pipeline using the data set; obtaining an inter-algorithm diversity value, an inter-feature correlation value, and an inter-algorithm hyperparameter distance value using the prediction result, the accuracy rate, and a new incoming pipeline by the pipeline sampling score calculation module, and obtaining a sampling score corresponding to the new incoming pipeline using the inter-algorithm diversity value, the inter-feature correlation value, and the inter-algorithm hyperparameter distance value by the pipeline sampling score calculation module; determining a recommended new incoming pipeline among the new incoming pipelines using the sampling score by the pipeline recommendation module; and determining a target recommended new incoming pipeline among the recommended new incoming pipelines using crowdsourced model technology by the crowdsourced model recommendation module.

[0007] Based on the above, the apparatus and method for a crowdsourced model recommendation pipeline according to the present invention can, after generating an initial pipeline and obtaining the prediction result and accuracy rate of the initial pipeline, then use the inter-algorithm diversity value, the inter-feature correlation value, and the inter-algorithm hyperparameter distance value to obtain the sampling score of the new incoming pipeline. Then, the sampling score of the new incoming pipeline can be used to recommend a pipeline for the crowdsourced model. In this way, the apparatus and method for a crowdsourced model recommendation pipeline according to the present invention can obtain the sampling score of the new incoming pipeline more accurately, and thus can recommend a pipeline for the crowdsourced model more accurately. Description of the Drawings

[0008] Figure 1 is a schematic diagram of an apparatus for a crowdsourced model recommendation pipeline according to an embodiment of the present invention.

[0009] Figure 2 It is a flowchart of a method for a crowdsourcing model recommendation pipeline according to an embodiment of the present invention.

[0010] Figure 3 It is Figure 2 an exemplary embodiment of step S210 shown.

[0011] Figure 4 It is Figure 2 an exemplary embodiment of step S220 shown.

[0012] Figures 5A - 5C It is Figure 2 an exemplary embodiment of step S230 shown.

[0013] Figure 6 It is Figure 2 an exemplary embodiment of step S240 shown.

[0014] Figure 7 It is Figure 2 an exemplary embodiment of step S260 shown.

[0015] Description of reference numerals

[0016] 1: Device for crowdsourcing model recommendation pipeline

[0017] 20: Storage medium

[0018] 21: Data extraction and pipeline initialization module

[0019] 23: Pipeline performance evaluation module

[0020] 25: Pipeline sampling fraction calculation module

[0021] 27: Pipeline recommendation module

[0022] 29: Crowdsourcing model recommendation module

[0023] 40: Processor

[0024] 60: Transceiver

[0025] S210, S220, S230, S240, S250, S260, S231, S232, S233: Steps

[0026] P 1 、P 2 、P 3 、P 4 、P 5 、P 6 : Initial pipeline

[0027] P 7 、P 8, P 9 , P 10 : New incoming pipeline

[0028] C 1 : First algorithm

[0029] C 2 : Second algorithm

[0030] x 1 , x 2 , x 3 , x 4 : Data

[0031] f 1 , f 2 , f 3 : Feature

[0032] y: Data x 1 , x 2 , x 3 , x 4 's true label

[0033] μ: Mean value

[0034] σ: Standard deviation

[0035] K N , K 6 : Initial pipeline efficiency probability distribution matrix

[0036] k: New incoming pipeline efficiency probability distribution matrix

[0037] K ij : Hybrid Kernel

[0038] i: The i-th pipeline

[0039] j: The j-th pipeline

[0040] D ij , D 76 , D 11 , D 13 , D 16 , D 71 , D 73 , D 76 : Diversity value between algorithms

[0041] F ij , F 76 , F 11 , F 13 , F 16 , F 71 , F 73 , F 76 : Correlation value between features

[0042] H ij 、H 11 、H 13 、H 16 、H 71 、H 73 、H 76 : Distance values between hyperparameters of the same algorithm Detailed implementation manners

[0043] Figure 1 is a schematic diagram of apparatus 1 for a crowdsourcing model recommendation pipeline according to an embodiment of the present invention. Please refer to Figure 1 . Apparatus 1 may include a storage medium 20 and a processor 40. The processor 40 is coupled to the storage medium 20. In other embodiments, apparatus 1 may further include a transceiver 60 coupled to the processor 40.

[0044] The storage medium 20 is, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar components, or a combination of the above components, and is used to store multiple modules or various application programs executable by the processor 40. In this embodiment, the storage medium 20 may store a data extraction and pipeline initialization module 21, a pipeline performance evaluation module 23, a pipeline sampling fraction calculation module 25, a pipeline recommendation module 27, and a crowdsourcing model recommendation module 29. The functions of these modules will be described later.

[0045] The processor 40 is, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose micro control unit (MCU), microprocessor, digital signal processor (DSP), programmable controller, application specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components, or a combination of the above components. The processor 40 can access and execute multiple modules and various application programs stored in the storage medium 20.

[0046] The transceiver 60 transmits and receives signals in a wireless or wired manner.

[0047] Figure 2 is a flowchart of a method for a crowdsourcing model recommendation pipeline according to an embodiment of the present invention, wherein the method can be implemented by Figure 1 the device 1 shown. Please refer to Figure 1 and Figure 2 .

[0048] In step S210, the data extraction and pipeline initialization module 21 can generate an initial pipeline using an algorithm.

[0049] Figure 3 is Figure 2 an implementation example of step S210 shown. Please refer to Figure 1 , Figure 2 and Figure 3. In this embodiment, the data extraction and pipeline initialization module 21 can receive an algorithm, a hyperparameter range, an initial number, a pipeline target number, and a training time through the transceiver 60. Then, the data extraction and pipeline initialization module 21 can use the algorithm, the hyperparameter range, the initial number, the pipeline target number, and the training time to generate an initial pipeline. Further, the algorithm can include a feature selection algorithm and a model algorithm. Specifically, the feature selection algorithm can be SelectPercentile, SelectKBest, VarianceThreshold UnivariateFeature Selection, or Recursive Feature Elimination (RFE). On the other hand, the model algorithm can be Support Vector Regression (SVR), Support Vector Machine (SVM), Random Forest (RF), Decision Tree, Extra Trees, AdaBoost, Gradient Boosting, XGBoost, or K-Nearest-Neighbor (KNN).

[0050] Here, it is assumed that the initial number received by the data extraction and pipeline initialization module 21 through the transceiver 60 is 3 (i.e., the data extraction and pipeline initialization module 21 needs to generate 3 initial pipelines for each algorithm), and the training time received through the transceiver 60 is, for example, 60 minutes. As Figure 3 shown, after the data extraction and pipeline initialization module 21 also receives the algorithms (i.e., SelectPercentile, Support Vector Machine, and Random Forest) and the hyperparameter range through the transceiver 60, the data extraction and pipeline initialization module 21 can generate the initial pipeline P 1 , the initial pipeline P 2 , the initial pipeline P 3 , the initial pipeline P 4 , the initial pipeline P 5 , and the initial pipeline P 6 . Specifically, in this embodiment, the algorithm can include a first algorithm (i.e., a combination of SelectPercentile and Support Vector Machine) and a second algorithm (i.e., a combination of SelectPercentile and Random Forest), where the first algorithm is different from the second algorithm. Further, the initial pipeline can include a first initial pipeline (the initial pipeline P 1 , the initial pipeline P 2 , and the initial pipeline P 3 ) and a second initial pipeline (the initial pipeline P 4, initial pipeline P 5 and initial pipeline P 6 ). The initial hyperparameters can correspond to the initial pipeline. The initial hyperparameters can include a first initial hyperparameter and a second initial hyperparameter. The first initial pipeline can include the first initial hyperparameter corresponding to the first algorithm, and the second initial pipeline can include the second initial hyperparameter corresponding to the second algorithm. For example, as Figure 3 shown, the first initial pipeline "initial pipeline P 1 " can include the first initial hyperparameters 16384 and 3.79e-5 corresponding to the first algorithm C 1 "combination of selecting percentile and support vector machine". Specifically, 16384 can be the value of the initial hyperparameter c of the support vector machine, and 3.79e-5 can be the value of the initial hyperparameter gamma of the support vector machine. On the other hand, the second initial pipeline "initial pipeline P 4 " can include the second initial hyperparameters 16 and 11 corresponding to the second algorithm C 2 (combination of selecting percentile and random forest). Specifically, 16 and 11 can be the values of the initial hyperparameters of the random forest.

[0051] Please go back to Figure 2 . In step S220, the pipeline performance evaluation module 23 can obtain the prediction result corresponding to the initial pipeline by using the data set, and can obtain the accuracy corresponding to the initial pipeline by using the data set.

[0052] Figure 4 Yes Figure 2 is an implementation example of step S220 shown. Please also refer to Figure 1 , Figure 2 , Figure 3 and Figure 4 . In this embodiment, the data extraction and pipeline initialization module 21 can receive the data set through the transceiver 60. In one embodiment, the pipeline performance evaluation module 23 can use a kernel-based method (such as Gaussian Process) and the data set to obtain the prediction result (corresponding to the initial pipeline), and can use the kernel-based method and the data set to obtain the accuracy (corresponding to the initial pipeline). In other words, the pipeline performance evaluation module 23 can use the data set to train and test (predict) the initial pipeline P 1 , initial pipeline P 2 , initial pipeline P 3 , initial pipeline P 4 , initial pipeline P 5 and initial pipeline P 6 respectively to obtain the prediction results and accuracies of these initial pipelines. For example, as Figure 4As shown, the data set may include data x 1 , data x 2 , data x 3 , data x 4 and so on, 4 pieces of data. And the data set may include multiple features (feature f 1 , feature f 2 and feature f 3 and so on, 3 features). Further, assume that the initial pipeline P 3 's prediction result for data x 1 is "0", the initial pipeline P 3 's prediction result for data x 2 is "1", the initial pipeline P 3 's prediction result for data x 3 is "1", and the initial pipeline P 3 's prediction result for data x 4 is "0". Then, the pipeline performance evaluation module 23 can obtain the accuracy of the initial pipeline P 3 . For example, the pipeline performance evaluation module 23 can use classification evaluation metrics to obtain the accuracy of a specific initial pipeline. The classification evaluation metrics may include accuracy, F1-score, and the area under the curve (AUC, The area under the Receiver Operating Characteristic curve). Here, assume that the pipeline performance evaluation module 23 obtains the accuracy of the initial pipeline P 3 as "100%". The pipeline performance evaluation module 23 can use a similar method to obtain the prediction results and accuracies of other initial pipelines. It should be noted here that although this embodiment uses "classification" as an implementation example for illustration, the present invention is not limited thereto. In other embodiments, for the implementation example of "regression", the pipeline performance evaluation module 23 can use regression evaluation metrics to obtain the accuracy of a specific pipeline. The regression evaluation metrics may include root mean square error (RMSE, Root means square error), mean square error (MSE, Mean square error), R-square, mean absolute error (MAE, Mean absolute error), and mean absolute percentage error (MAPE, Mean absolute percentage error).

[0053] Furthermore, the prediction results may include a first prediction result and a second prediction result, and the accuracy rates may include a first accuracy rate and a second accuracy rate. The first prediction result may correspond to a first initial pipeline, and the first accuracy rate may correspond to the first initial pipeline. The second prediction result may correspond to a second initial pipeline, and the second accuracy rate may correspond to the second initial pipeline. The first initial pipeline may include a first best-accuracy initial pipeline and a first other initial pipeline, where the first accuracy rate of the first best-accuracy initial pipeline is greater than the first accuracy rate of the first other initial pipeline, and the first best-accuracy initial pipeline corresponds to the first best-accuracy initial pipeline prediction result. The second initial pipeline includes a second best-accuracy initial pipeline and a second other initial pipeline, where the second accuracy rate of the second best-accuracy initial pipeline is greater than the second accuracy rate of the second other initial pipeline, and the second best-accuracy initial pipeline corresponds to the second best-accuracy initial pipeline prediction result. For example, as Figure 4 shown, since the highest accuracy rate among the first initial pipelines (initial pipeline P 1 , initial pipeline P 2 , and initial pipeline P 3 ) is that of initial pipeline P 1 and initial pipeline P 3 , thus initial pipeline P 1 and initial pipeline P 3 are the above-mentioned first best-accuracy initial pipelines, and initial pipeline P 2 is the above-mentioned first other initial pipeline. Similarly, initial pipeline P 6 is the above-mentioned second best-accuracy initial pipeline, and initial pipeline P 4 and initial pipeline P 5 are the above-mentioned second other initial pipelines. Further, the first best-accuracy initial pipeline prediction result is the first prediction result "0, 1, 1, 0" of the first best-accuracy initial pipeline (initial pipeline P 1 and initial pipeline P 3 ). On the other hand, the second best-accuracy initial pipeline prediction result is the second prediction result "0, 0, 1, 0" of the second best-accuracy initial pipeline (initial pipeline P 6 ).

[0054] It should be noted here that although the above steps S210 and S220 take two algorithms, such as a first algorithm (i.e., a combination of selecting percentiles and a support vector machine) and a second algorithm (i.e., a combination of selecting percentiles and a random forest), as implementation examples, the number of algorithms in the present invention can be adjusted according to actual needs. In other words, the number of algorithms can be two or more than two.

[0055] Please return to Figure 2。In step S230, the pipeline sampling fraction calculation module 25 can utilize the prediction results, accuracy, and the diversity value between algorithms, the correlation value between features, and the distance value between hyperparameters of the same algorithm obtained from the newly incoming pipeline, and the pipeline sampling fraction calculation module 25 can obtain the sampling fraction corresponding to the newly incoming pipeline by using the diversity value between algorithms, the correlation value between features, and the distance value between hyperparameters of the same algorithm.

[0056] Figures 5A - 5C Yes Figure 2 An implementation example of step S230 shown in. Please also refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figures 5A - 5C .

[0057] As Figure 5A shown, the pipeline sampling fraction calculation module 25 can first predict the performance probability distribution of each pipeline. For example, the pipeline sampling fraction calculation module 25 can use a kernel-based method (such as GaussianProcess) to calculate the initial pipeline performance probability distribution matrix K N , and can calculate the performance probability distribution matrix k of a specific newly incoming pipeline. Further, when calculating the initial pipeline performance probability distribution matrix K N and the performance probability distribution matrix k of the newly incoming pipeline, the pipeline sampling fraction calculation module 25 can consider the diversity value between algorithms (D ij ), the correlation value between features (F ij ), and the distance value between hyperparameters of the same algorithm (H ij ). Then, the pipeline sampling fraction calculation module 25 can use the diversity value between algorithms, the correlation value between features, and the distance value between hyperparameters of the same algorithm to establish a hybrid kernel (K ij ).

[0058] As Figure 5B shown, the pipeline sampling fraction calculation module 25 can receive the newly incoming pipeline through the transceiver 60. In this embodiment, the newly incoming pipeline P 7 is, for example, corresponding to the combination of the first algorithm (selecting percentile and support vector machine). Then, in step S231, the pipeline sampling fraction calculation module 25 can obtain the diversity value between algorithms by using the prediction results of the first accuracy best initial pipeline and the prediction results of the second accuracy best initial pipeline. In one embodiment, the diversity value between algorithms can include cosine similarity and contingency table. As described in the above Figure 4 embodiment, the prediction result of the first accuracy best initial pipeline is the first accuracy best initial pipeline (initial pipeline P 1and the initial pipeline P 3 )'s first prediction result "0, 1, 1, 0". On the other hand, the second best accuracy initial pipeline prediction result is the second prediction result "0, 0, 1, 0" of the second best accuracy initial pipeline (initial pipeline P 6 ). In this embodiment, since the newly incoming pipeline P 7 corresponds to the first algorithm (a combination of selecting percentiles and support vector machines), the pipeline sampling score calculation module 25 can utilize the first prediction result "0, 1, 1, 0" and the second prediction result "0, 0, 1, 0" to obtain the inter-algorithm diversity value D 76 . For example, the pipeline sampling score calculation module 25 can use the cosine similarity between the first prediction result "0, 1, 1, 0" and the second prediction result "0, 0, 1, 0" as the inter-algorithm diversity value D 76 .

[0059] Please continue to refer to Figure 5B . As described above Figure 4 in the embodiment, the data set may include multiple features (i.e., feature f 1 , feature f 2 , and feature f 3 and other 3 features). The initial feature set may correspond to the initial pipeline, and the initial feature set may include at least one of the multiple features. On the other hand, the newly incoming feature set may correspond to the newly incoming pipeline, and the newly incoming feature set may include at least one of the multiple features. Then, in step S232, the pipeline sampling score calculation module 25 can utilize the initial feature set and the incoming feature set to obtain the inter-feature correlation value. In one embodiment, the inter-feature correlation value may include the absolute value of the Pearson correlation coefficient, the absolute value of the Spearman correlation coefficient, the number of feature intersections divided by the total number of features, the Euclidean Distance, and the Mahalanobis Distance. Specifically, as Figure 5B shown, assuming that the initial feature set corresponding to the initial pipeline P 6 is feature f 1 and feature f 2 , and assuming that the newly incoming feature set corresponding to the newly incoming pipeline P 7 is feature f 2 and feature f 3 . The pipeline sampling score calculation module 25 can, for example, utilize the number of intersections of the features (i.e., the above-mentioned number of feature intersections and the total number of features) to obtain the inter-feature correlation value F 76 .

[0060] Please continue to refer to Figure 5B . In step S233, the pipeline sampling fraction calculation module 25 can obtain the distance value between hyperparameters of the same algorithm by using the initial hyperparameters and the newly introduced hyperparameters. Specifically, the initial hyperparameters can correspond to the initial pipeline. On the other hand, the newly introduced hyperparameters can correspond to the newly introduced pipeline. Further, the distance value between hyperparameters of the same algorithm can include the Radial Basis Function kernel (RBF kernel), the Laplace kernel, the Matern kernel, and the Rational Quadratic Kernel. Specifically, the newly introduced pipeline P 7 in this embodiment corresponds to the first algorithm (a combination of percentile selection and support vector machine), and the initial pipeline P 6 corresponds to the second algorithm (a combination of percentile selection and random forest). In other words, the algorithm of the newly introduced pipeline P 7 is different from the algorithm of the initial pipeline P 6 , so the pipeline sampling fraction calculation module 25 can obtain the distance value H 76 between hyperparameters of the same algorithm to be 0. It should be noted that the formula in step S233 is an implementation example of the above-mentioned Radial Basis Function kernel.

[0061] After the pipeline sampling fraction calculation module 25 finishes performing the above steps S231, S232, and S233, the pipeline sampling fraction calculation module 25 can establish a Hybrid Kernel by using the diversity value between algorithms, the correlation value between features, and the distance value between hyperparameters of the same algorithm. Specifically, the pipeline sampling fraction calculation module 25 can obtain the initial pipeline performance probability distribution matrix K Figure 5B as shown in 6 and the newly introduced pipeline performance probability distribution matrix K.

[0062] Please refer to Figure 5C . After establishing the Hybrid Kernel, the pipeline sampling fraction calculation module 25 can obtain the sampling fraction corresponding to the newly introduced pipeline by using the sampling function and the Hybrid Kernel. In one embodiment, the sampling function can include Expected Improvement (EI), Upper Confidence Bound (UCB), Probability of Improvement (POI), and Entropy Search (ES). As Figure 5C shown, the pipeline sampling fraction calculation module 25 can, for example, obtain the sampling fraction corresponding to the newly introduced pipeline P 7The sampling fraction is the UCB value of 1e-5.

[0063] Please go back to Figure 2 . In step S240, the pipeline recommendation module 27 can use the sampling fraction to determine the recommended new pipelines in the new pipelines.

[0064] Figure 6 Yes Figure 2 An exemplary implementation of step S240 shown. Please also refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figures 5A - 5C and Figure 6 . In this embodiment, it is assumed that the pipeline sampling fraction calculation module 25 receives the new pipeline P 7 , new pipeline P 8 , new pipeline P 9 and new pipeline P 10 through the transceiver 60, and it is assumed that the pipeline sampling fraction calculation module 25 has calculated the sampling fractions of the new pipeline P Figure 6 shown for the new pipeline P 7 , new pipeline P 8 , new pipeline P 9 and new pipeline P 10 respectively. The pipeline recommendation module 27 can use the new pipeline with the highest sampling fraction as the recommended new pipeline. In other words, the pipeline recommendation module 27 can determine that the recommended new pipeline is the new pipeline P 9 .

[0065] Please go back to Figure 2 . The pipeline recommendation module 27 can use the preset running time and the sampling fraction to determine the recommended new pipelines in the new pipelines. Specifically, in step S250, the pipeline recommendation module 27 can determine whether steps S220 to S240 have been executed for more than the preset running time.

[0066] If the pipeline recommendation module 27 determines that steps S220 to S240 have not been executed for more than the preset running time (the judgment result of step S250 is "no"), then the device 1 of the present invention can re-execute step S220.

[0067] On the other hand, if the pipeline recommendation module 27 determines that steps S220 to S240 have been executed for more than the preset running time (the judgment result of step S250 is "yes"), then in step S260, the crowdsourcing model recommendation module 29 can use the crowdsourcing model technology to determine the target recommended new pipeline in the recommended new pipelines.

[0068] Figure 7 Yes Figure 2An exemplary embodiment of step S260 shown. Please also refer to Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figures 5A - 5C , Figure 6 and Figure 7 . The crowd-wisdom model recommendation module 29 can use a preset number and crowd-wisdom learning technology to determine the target recommended new pipelines in the recommended new pipelines. For example, assume that the preset number is 5, and assume that the pipeline pool includes the initial pipelines (initial pipeline P 1 ~initial pipeline P 6 ) and new pipelines (new pipeline P 7 ~new pipeline P 10 ) of the foregoing embodiment. The crowd-wisdom model recommendation module 29 can use crowd-wisdom model technology to select 5 target recommended new pipelines to optimize the effect of the crowd-wisdom model. For example, as Figure 7 shown, assume that the crowd-wisdom model recommendation module 29 selects the initial pipeline P 1 a total of 3 times, and the crowd-wisdom model recommendation module 29 selects the initial pipeline P 2 a total of 1 time, and the crowd-wisdom model recommendation module 29 selects the new pipeline P 7 a total of 1 time. Based on this, the weight of the initial pipeline P 1 can be 0.6, and the weight of the initial pipeline P 2 can be 0.2, and the weight of the new pipeline P 7 can be 0.2.

[0069] Tables 1 and 2 use publicly available datasets to verify the classification effect and regression effect of the present invention. Compared with the international open-source software AutoSklearn and the well-known commercial software H2O, the present invention can, in the case of significantly reducing the number of attempts, recommend pipelines with similar effects for the crowd-wisdom model.

[0070] Table 1 Classification effect: Accuracy (number of attempts)

[0071]

[0072] Table 2 Regression effect: Root mean square error RMSE (number of attempts)

[0073]

[0074] In summary, after generating an initial pipeline and obtaining the prediction results and accuracy rate of the initial pipeline, the apparatus and method for a crowdsourcing model recommendation pipeline of the present invention can then use the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm to obtain the sampling score of the new incoming pipeline. Then, the sampling score of the new incoming pipeline can be used to recommend a pipeline for the crowdsourcing model. In this way, the apparatus and method for a crowdsourcing model recommendation pipeline of the present invention can obtain the sampling score of the new incoming pipeline more accurately, thereby being able to recommend a pipeline for the crowdsourcing model more accurately.

Claims

1. An apparatus for a crowdsourcing model recommendation pipeline, characterized in that, it includes: a storage medium storing a plurality of modules, wherein the plurality of modules include a data extraction and pipeline initialization module, a pipeline performance evaluation module, a pipeline sampling score calculation module, a pipeline recommendation module, and a crowdsourcing model recommendation module; and a processor coupled to the storage medium and accessing and executing the plurality of modules, wherein the data extraction and pipeline initialization module generates at least one initial pipeline using at least one algorithm; the pipeline performance evaluation module obtains at least one prediction result corresponding to the initial pipeline using a data set, and obtains at least one accuracy rate corresponding to the initial pipeline using the data set; the pipeline sampling score calculation module obtains at least one algorithm diversity value, at least one feature correlation value, and at least one distance value between hyperparameters of the same algorithm using the prediction result, the accuracy rate, and a new incoming pipeline, and the pipeline sampling score calculation module obtains at least one sampling score corresponding to the new incoming pipeline using the algorithm diversity value, the feature correlation value, and the distance value between hyperparameters of the same algorithm; the pipeline recommendation module determines a recommended new incoming pipeline among the new incoming pipelines using the sampling score; the crowdsourcing model recommendation module determines at least one target recommended new incoming pipeline among the recommended new incoming pipelines using crowdsourcing model technology.

2. The apparatus according to claim 1, characterized in that, it further includes a transceiver coupled to the processor, wherein the data extraction and pipeline initialization module receives the algorithm, at least one hyperparameter range, an initial number, a pipeline target number, and a training time through the transceiver; the data extraction and pipeline initialization module generates the initial pipeline using the algorithm, the hyperparameter range, the initial number, the pipeline target number, and the training time.

3. The apparatus according to claim 1, characterized in that, the pipeline sampling score calculation module obtains the algorithm diversity value using the prediction result of the first accuracy rate best initial pipeline and the prediction result of the second accuracy rate best initial pipeline.

4. The apparatus according to claim 3, characterized in that, initial hyperparameters correspond to the initial pipeline, wherein the algorithm includes a first algorithm and a second algorithm, and the first algorithm is different from the second algorithm; the initial pipeline includes a first initial pipeline and a second initial pipeline; the initial hyperparameters include a first initial hyperparameter and a second initial hyperparameter; the first initial pipeline includes the first initial hyperparameter corresponding to the first algorithm, and the second initial pipeline includes the second initial hyperparameter corresponding to the second algorithm; the prediction result includes a first prediction result and a second prediction result, and the accuracy rate includes a first accuracy rate and a second accuracy rate; the first prediction result corresponds to the first initial pipeline, and the first accuracy rate corresponds to the first initial pipeline, the second prediction result corresponds to the second initial pipeline, and the second accuracy rate corresponds to the second initial pipeline; The first initial pipeline includes a first best-accuracy initial pipeline and at least one first other initial pipeline, wherein the first accuracy of the first best-accuracy initial pipeline is greater than the first accuracy of the first other initial pipelines, and wherein the first best-accuracy initial pipeline corresponds to the prediction result of the first best-accuracy initial pipeline; The second initial pipeline includes a second best-accuracy initial pipeline and at least one second other initial pipeline, wherein the second accuracy of the second best-accuracy initial pipeline is greater than the second accuracy of the second other initial pipelines, and wherein the second best-accuracy initial pipeline corresponds to the prediction result of the second best-accuracy initial pipeline.

5. The apparatus according to claim 1, wherein, it further includes a transceiver coupled to the processor, wherein the data set includes a plurality of features, wherein the initial feature set corresponds to the initial pipeline, and the initial feature set includes at least one of the plurality of features, and wherein the data extraction and pipeline initialization module receives the data set through the transceiver; the pipeline sampling score calculation module receives the new pipeline through the transceiver, wherein the new feature set corresponds to the new pipeline, and the new feature set includes at least one of the plurality of features; the pipeline sampling score calculation module obtains the inter-feature correlation value by using the initial feature set and the new feature set.

6. The apparatus according to claim 1, wherein, the initial hyperparameters correspond to the initial pipeline, and the new hyperparameters correspond to the new pipeline, and wherein the pipeline sampling score calculation module obtains the distance value between the same-algorithm hyperparameters by using the initial hyperparameters and the new hyperparameters.

7. The apparatus according to claim 1, wherein, the pipeline sampling score calculation module establishes a Hybrid Kernel by using the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between the same-algorithm hyperparameters; the pipeline sampling score calculation module obtains the sampling score corresponding to the new pipeline by using a sampling function and the Hybrid Kernel.

8. The apparatus according to claim 7, wherein, the sampling function includes Expected Improvement (EI), Upper Confidence Bound (UCB), Probability of Improvement (POI), and Entropy Search (ES).

9. The apparatus according to claim 1, wherein, the algorithms include a feature selection algorithm and a model algorithm.

10. The apparatus according to claim 1, wherein, the pipeline performance evaluation module obtains the prediction result by using a kernel-based method and the data set, and obtains the accuracy by using the kernel-based method and the data set.

11. The apparatus according to claim 1, It is characterized in that the diversity value among the algorithms includes cosine similarity and contingency table.

12. The device according to claim 1, it is characterized in that the correlation value among the features includes the absolute value of the Pearson correlation coefficient, the absolute value of the Spearman correlation coefficient, the number of feature intersections divided by the total number of features, Euclidean Distance, and Mahalanobis Distance.

13. The device according to claim 1, it is characterized in that the distance value among the hyperparameters of the same algorithm includes RBF kernel (Radial Basis Function kernel), Laplace kernel, Matern kernel, and Rational Quadratic Kernel.

14. The device according to claim 1, it is characterized in that the pipeline recommendation module determines the recommended new pipelines in the new pipelines by using the preset running time and the sampling fraction.

15. The device according to claim 1, it is characterized in that the crowdsourcing model recommendation module determines the target recommended new pipelines in the recommended new pipelines by using the preset number and the crowdsourcing learning technology.

16. A method for recommending pipelines for a crowdsourcing model, applicable to a device including a storage medium and a processor, it is characterized in that the storage medium stores a plurality of modules, where the plurality of modules include a data extraction and pipeline initialization module, a pipeline performance evaluation module, a pipeline sampling fraction calculation module, a pipeline recommendation module, and a crowdsourcing model recommendation module, and the method includes the following steps: The data extraction and pipeline initialization module generates at least one initial pipeline by using at least one algorithm; The pipeline performance evaluation module obtains at least one prediction result corresponding to the initial pipeline by using a data set, and obtains at least one accuracy corresponding to the initial pipeline by using the data set; The pipeline sampling fraction calculation module obtains at least one diversity value among the algorithms, at least one correlation value among the features, and at least one distance value among the hyperparameters of the same algorithm by using the prediction result, the accuracy, and the new pipelines, and the pipeline sampling fraction calculation module obtains at least one sampling fraction corresponding to the new pipelines by using the diversity value among the algorithms, the correlation value among the features, and the distance value among the hyperparameters of the same algorithm; The pipeline recommendation module determines the recommended new pipelines in the new pipelines by using the sampling fraction; and The intelligent model recommendation module determines at least one target recommended new pipeline in the recommended new pipelines by using the intelligent model technology.

17. The method according to claim 16, wherein, the steps of the pipeline sampling score calculation module obtaining the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm by using the prediction result, the accuracy rate, and the new pipeline include: the pipeline sampling score calculation module obtains the inter-algorithm diversity value by using the prediction result of the first accuracy-optimal initial pipeline and the prediction result of the second accuracy-optimal initial pipeline.

18. The method according to claim 16, wherein, the device further includes a transceiver, wherein the data set includes a plurality of features, the initial feature set corresponds to the initial pipeline, and the initial feature set includes at least one of the plurality of features, and the steps of the pipeline sampling score calculation module obtaining the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm by using the prediction result, the accuracy rate, and the new pipeline include: the data extraction and pipeline initialization module receives the data set through the transceiver; the pipeline sampling score calculation module receives the new pipeline through the transceiver, wherein the new feature set corresponds to the new pipeline, and the new feature set includes at least one of the plurality of features; and the pipeline sampling score calculation module obtains the inter-feature correlation value by using the initial feature set and the new feature set.

19. The method according to claim 16, wherein, the initial hyperparameters correspond to the initial pipeline, the new hyperparameters correspond to the new pipeline, and the steps of the pipeline sampling score calculation module obtaining the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm by using the prediction result, the accuracy rate, and the new pipeline include: the pipeline sampling score calculation module obtains the distance value between hyperparameters of the same algorithm by using the initial hyperparameters and the new hyperparameters.

20. The method according to claim 16, wherein, the steps of the pipeline sampling score calculation module obtaining the sampling score corresponding to the new pipeline by using the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm include: the pipeline sampling score calculation module establishes a Hybrid Kernel by using the inter-algorithm diversity value, the inter-feature correlation value, and the distance value between hyperparameters of the same algorithm; and the pipeline sampling score calculation module obtains the sampling score corresponding to the new pipeline by using the sampling function and the Hybrid Kernel.