Experimental design method automatic selection system based on machine learning

Through the automatic selection system of experimental design methods based on machine learning, the problem of experimental design method selection relying on manual experience is solved, automated and optimized experimental design is realized, and experimental efficiency and quality are improved. It is suitable for industries such as industry, medical research and development, materials science, and agricultural research.

CN120724104APending Publication Date: 2025-09-30ZHANGQI (CHENGDU) INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510739450.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The selection of existing experimental design methods mainly relies on manual experience and lacks scientific selection basis, resulting in low experimental efficiency and waste of resources. In particular, it is impossible to fully consider various factors when designing complex experimental problems. The existing experimental design method selection mechanism lacks learning ability and cannot optimize the selection strategy from historical experimental data.

Method used

An automatic selection system for experimental design methods based on machine learning is adopted. Through the feature extraction module, model training module, weight calculation module and feedback learning module, the CatBoost algorithm is used to train the model, dynamically select the best experimental design method, and continuously optimize the selection strategy through feedback learning.

Benefits of technology

It realizes the automatic selection of the most suitable experimental design method according to the experimental characteristics, reduces the time of manual judgment, avoids the waste of resources, improves the quality and efficiency of experimental design, has good scalability and adaptability, and can handle the experimental design needs of different scales and complexities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120724104A_ABST
    Figure CN120724104A_ABST
Patent Text Reader

Abstract

The invention discloses an experimental design method automatic selection system based on machine learning, and relates to the technical field of experimental design, and the system comprises a feature extraction module which is used for extracting key features from an experimental design problem; the model training module is used for selecting a model by using a CatBoost algorithm training method; the weight calculation module is used for calculating the weight of each experiment design method; and the dynamic selection module is used for selecting an optimal experiment design method according to the weight. According to the invention, through a machine learning algorithm and a dynamic weight mechanism, intelligent selection of experimental design methods is realized, the most suitable experimental design method is automatically selected, manual judgment time is reduced, a suitable method is selected according to experimental characteristics, resource waste is avoided, the system has good expandability and adaptability, and the method is suitable for popularization and application. According to the method, experimental design requirements of different scenes can be met, the system can continuously optimize selection strategies through a feedback learning mechanism, experimental design requirements of different scales and complexity can be processed, and the quality and efficiency of experimental design are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of experimental design, and in particular to an automatic selection system for experimental design methods based on machine learning. Background Art

[0002] In the field of experimental design, commonly used experimental design methods include Latin hypercube (LHS), Sobol sequence, orthogonal design, full factorial design and Bayesian optimization. Each method has its applicable scenarios and limitations. At present, the selection of experimental design methods mainly relies on manual experience and lacks scientific selection basis, which easily leads to low experimental efficiency and waste of resources. Especially when facing complex experimental design problems, manual selection methods often cannot fully consider various factors, resulting in suboptimal experimental design.

[0003] In existing technologies, the selection of experimental design methods is often based on empirical rules, lacks quantitative evaluation models, and does not fully consider computing resource limitations, resulting in significant differences in the computational complexity of different experimental design methods. The existing experimental design method selection mechanism lacks learning ability and cannot optimize the selection strategy from historical experimental data. Therefore, an automatic selection system for experimental design methods based on machine learning is urgently needed. Summary of the Invention

[0004] The purpose of the present invention is to provide an automatic selection system for experimental design methods based on machine learning in order to solve the problem of how to automatically select the most suitable experimental design method based on characteristics such as the number of experimental factors and the number of levels, and to continuously optimize the selection strategy based on feedback from experimental results.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a system for automatically selecting experimental design methods based on machine learning, comprising:

[0006] Feature extraction module: extract key features from experimental design problems;

[0007] Model training module: Use CatBoost algorithm training method to select the model;

[0008] Weight calculation module: calculate the weight of each experimental design method;

[0009] Dynamic selection module: select the best experimental design method based on weights;

[0010] Feedback learning module: updates the model based on experimental results.

[0011] Preferably, the experimental design characteristics include the number of factors (n), the number of levels (L), the number of samples (s), the adequacy of computing resources (r) and the optimization requirement (o), and the feature vector in the feature extraction module is expressed as: X = [n, L, s, r, o].

[0012] Preferably, the number of factors is the number of independent variables in the experiment, the number of levels is the number of possible values ​​for each factor, the number of samples is the maximum number of experiments allowed, the computing resource sufficiency is a value between 0 and 1, indicating the available computing resources, and the optimization requirement is a Boolean value indicating whether the objective function needs to be optimized.

[0013] Preferably, the CatBoost algorithm is a machine learning algorithm based on gradient boosting decision trees, which is suitable for processing classification problems. The CatBoost training model regards the selection of experimental design methods as a classification problem.

[0014] Preferably, the CatBoost training model training step includes:

[0015] S1. Prepare training data: collect historical experimental data, extract features and labels;

[0016] S2, model configuration: iterations: 100, learning rate: 0.1, tree depth: 6, loss function: MultiClass;

[0017] S3, model training: fit the model using training data;

[0018] S4. Model evaluation: Model performance was evaluated using cross-validation.

[0019] Preferably, the weight calculation module is the core of the system, and includes three parts: basic weight, conditional weight and complexity penalty.

[0020] Preferably, the dynamic selection module selects the best method according to the calculated weights, and the algorithm steps are as follows:

[0021] S1. Normalized weight: Calculate the total weight S = ∑mTotalWeightm, normalized weight NormalizedWeightm = TotalWeightm / S;

[0022] S2. Threshold determination: Calculate the maximum weight Max=maxmTotalWeightm and the minimum weight Min=minmTotalWeightm. If Max-Min<5%, randomly select a method; otherwise, proceed to S3.

[0023] S3, Probability selection: According to the weighted probability selection method, the probability of method m being selected is NormalizedWeight m , mathematical expression: Where P(m) represents the probability of selecting method m.

[0024] Preferably, the feedback learning module continuously optimizes the model by recording experimental results, and the steps are as follows:

[0025] S1. Collect feedback data: record the characteristics of each experiment, the selected methods and the evaluation of the experimental results;

[0026] S2, update training data: add feedback data to the training set;

[0027] S3. Model update: Retrain the model using the expanded training set or continue training based on the original model.

[0028] Preferably, the automatic selection system workflow is as follows:

[0029] S1. The user inputs the parameters of the experimental design problem (number of factors, number of levels, etc.);

[0030] S2, feature extraction module extracts problem features;

[0031] S3, weight calculation module calculates the weight of each method;

[0032] S4, the method selection module selects the best method;

[0033] S5. Execute the selected method to generate an experimental design plan;

[0034] S6. The feedback learning module records the results and updates the model.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] In the present invention, through machine learning algorithms and dynamic weight mechanisms, intelligent selection of experimental design methods is achieved, the most suitable experimental design method is automatically selected, the time for manual judgment is reduced, the appropriate method is selected according to the experimental characteristics, and waste of resources is avoided. The system has good scalability and adaptability, and can meet the experimental design needs of different scenarios. Through the feedback learning mechanism, the system can continuously optimize the selection strategy, can handle experimental design needs of different scales and complexities, and improve the quality and efficiency of experimental design. Based on the automatic experimental design selection system of the present invention, the system architecture supports the addition of new experimental design methods and features, and can be widely used in industrial experiments, medical research and development, materials science, agricultural scientific research and other fields, significantly improving experimental efficiency and reducing experimental costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the method flow of the present invention;

[0038] Figure 2 This is the conditional weight adjustment rule diagram of the present invention. DETAILED DESCRIPTION

[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0040] See also Figure 1 , an automatic selection system for experimental design methods based on machine learning, including:

[0041] Feature extraction module: extract key features from experimental design problems;

[0042] Model training module: Use CatBoost algorithm training method to select the model;

[0043] Weight calculation module: calculate the weight of each experimental design method;

[0044] Dynamic selection module: select the best experimental design method based on weights;

[0045] Feedback learning module: updates the model based on experimental results.

[0046] The automatic selection system also includes: an experimental design generation module, a result analysis module, and a visualization module.

[0047] The experimental design features include the number of factors (n), the number of levels (L), the number of samples (s), the adequacy of computing resources (r), and the optimization requirements (o), and the feature vector in the feature extraction module is represented as: X = [n, L, s, r, o];

[0048] First, historical experimental data are collected, including the number of factors, the number of levels, the experimental design method, and the experimental results. The data are preprocessed and feature extracted to construct a training dataset. The experimental design features include the number of factors, the number of levels, the number of samples, the adequacy of computing resources, and the optimization requirements.

[0049] As a preferred embodiment of the present invention: the CatBoost training model training steps include:

[0050] S1. Prepare training data: collect historical experimental data, extract features and labels;

[0051] S2, model configuration: iterations: 100, learning rate: 0.1, tree depth: 6, loss function: MultiClass;

[0052] S3, model training: fit the model using training data;

[0053] S4. Model evaluation: Model performance was evaluated using cross-validation.

[0054] Mathematical representation of model training:

[0055] Given a training set where X i is the eigenvector, y i is the corresponding experimental design method label. The CatBoost model is trained by minimizing the following loss function:

[0056] L=∑i=1 N LMultiClass(y i ,F(X i ))

[0057] Where F(X i ) is the prediction result of the model, L MultiClass is the multi-classification loss function.

[0058] The weight calculation module is the core of the system and consists of three parts: basic weight, conditional weight and complexity penalty;

[0059] The base weight is predicted by the CatBoost model and ranges from 0-40%. The model outputs the probability distribution of each method, which needs to be converted into a percentage weight: BaseWeightm = Probm × 40%, where m represents different experimental design methods, Prob m is the probability of method m predicted by the model; the conditional weights are adjusted according to the characteristics of the problem; the computational complexity penalty is calculated based on the number of factors and the number of levels: ComplexityPenalty=0.1×n×L;

[0060] The final weights of each method are:

[0061] TotalWeightm=max(0,ConditionalWeightm-ComplexityPenalty), the final weight is guaranteed not to be a negative value.

[0062] The dynamic selection module selects the best method based on the calculated weights. The algorithm steps are as follows:

[0063] S1. Normalized weight: Calculate the total weight S = ∑mTotalWeightm, normalized weight NormalizedWeightm = TotalWeightm / S;

[0064] S2. Threshold determination: Calculate the maximum weight Max=maxmTotalWeightm and the minimum weight Min=minmTotalWeightm. If Max-Min<5%, randomly select a method; otherwise, proceed to S3.

[0065] S3, Probability selection: According to the weighted probability selection method, the probability of method m being selected is NormalizedWeight m , mathematical expression: Where P(m) represents the probability of selecting method m;

[0066] Total weight calculation = basic weight + conditional weight - computational complexity penalty, and the random sampling probability is proportional to the total weight. Set the selection threshold: when the weight difference is <5%, the two methods are randomly selected.

[0067] The feedback learning module continuously optimizes the model by recording experimental results. The steps are as follows:

[0068] S1. Collect feedback data: record the characteristics of each experiment, the selected methods and the evaluation of the experimental results;

[0069] S2, update training data: add feedback data to the training set;

[0070] S3, model update: retrain the model using the expanded training set or continue training based on the original model;

[0071] Record each selection result and actual performance, update the model regularly, adjust the basic weight calculation, and continuously optimize the selection strategy based on the feedback of experimental results. The mathematical expression of model update is: given the original model F0 and the new training data Dnew = (X i ,y i )i=1 M , the updated model is: Fnew = Fit(F0, Dnew), where Fit represents the function that continues training based on the original model.

[0072] As a preferred embodiment of the present invention, the automatic selection system workflow is as follows:

[0073] S1. The user inputs the parameters of the experimental design problem (number of factors, number of levels, etc.);

[0074] S2, feature extraction module extracts problem features;

[0075] S3, weight calculation module calculates the weight of each method;

[0076] S4, the method selection module selects the best method;

[0077] S5. Execute the selected method to generate an experimental design plan;

[0078] S6, the feedback learning module records the results and updates the model;

[0079] The following table compares the characteristics of different experimental design methods:

[0080]

[0081]

[0082] Example 1: Basic system architecture

[0083]

[0084] Example 2: Weight calculation system

[0085]

[0086]

[0087] Example 3: Dynamic selection mechanism

[0088]

[0089]

[0090] Example 4: Feedback Learning Mechanism

[0091]

[0092] Example 1: Simple scenario (3 factors, 2 levels)

[0093] Input parameters:

[0094] Number of factors: 3

[0095] Number of levels: 2

[0096] Sample size: 100

[0097] Computing resource sufficiency: 0.5

[0098] Whether optimization is needed: No

[0099] Calculation process:

[0100] Basic weights (assuming model prediction results):

[0101] full_factorial:15%

[0102] orthogonal: 10%

[0103] lhs:8%

[0104] Sobol: 5%

[0105] Bayesian: 2%

[0106] Condition weight adjustment:

[0107] full_factorial:15%+30%=45%

[0108] Other methods remain unchanged

[0109] Computational complexity penalty:

[0110] 0.1×3×2=0.6%

[0111] Final weight:

[0112] full_factorial: 44.4%

[0113] orthogonal: 9.4%

[0114] lhs:7.4%

[0115] Sobol: 4.4%

[0116] Bayesian: 1.4%

[0117] Selection result: full_factorial (full factorial design)

[0118] Analysis: Because the number of factors is small (3), the full factorial design obtains additional conditional weights and has lower computational complexity, so the system chooses the full factorial design.

[0119] Example 2: High-dimensional scenario (12 factors, 4 levels)

[0120] Input parameters:

[0121] Number of factors: 12

[0122] Number of levels: 4

[0123] Sample size: 500

[0124] Computing resource adequacy: 0.9

[0125] Whether optimization is needed: Yes

[0126] Calculation process:

[0127] Basic weights (assuming model prediction results):

[0128] full_factorial: 2%

[0129] orthogonal: 5%

[0130] lhs:10%

[0131] Sobol: 15%

[0132] Bayesian: 8%

[0133] Condition weight adjustment:

[0134] Sobol: 15% + 25% = 40%

[0135] Bayesian: 8% + 20% = 28%

[0136] Other methods remain unchanged

[0137] Computational complexity penalty:

[0138] 0.1×12×4=4.8%

[0139] Final weight:

[0140] full_factorial: 0% (negative values ​​are taken as 0)

[0141] orthogonal: 0.2%

[0142] lhs:5.2%

[0143] Sobol: 35.2%

[0144] Bayesian: 23.2%

[0145] Selection result: sobol (Sobol sequence)

[0146] Analysis: Because there are a large number of factors (12), the Sobol sequence obtains additional conditional weights. Although Bayesian optimization also obtains conditional weights, the Sobol sequence has a higher base weight, so the system chooses the Sobol sequence.

[0147] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be included therein. Any reference sign in a claim should not be construed as limiting the claim to which it relates.

Claims

1. An automatic selection system for experimental design methods based on machine learning, characterized in that: include: Feature extraction module: extract key features from experimental design problems; Model training module: Use CatBoost algorithm training method to select the model; Weight calculation module: calculate the weight of each experimental design method; Dynamic selection module: select the best experimental design method based on weights; Feedback learning module: updates the model based on experimental results.

2. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The experimental design features include the number of factors (n), the number of levels (L), the number of samples (s), the adequacy of computing resources (r), and the optimization requirement (o), and the feature vector in the feature extraction module is represented as: X = [n, L, s, r, o].

3. The automatic selection system for experimental design methods based on machine learning according to claim 2, characterized in that: The number of factors is the number of independent variables in the experiment, the number of levels is the number of possible values ​​for each factor, the number of samples is the maximum number of experiments allowed, the computing resource sufficiency is a value between 0 and 1, indicating the available computing resources, and the optimization requirement is a Boolean value indicating whether the objective function needs to be optimized.

4. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The CatBoost algorithm is a machine learning algorithm based on a gradient boosting decision tree, which is suitable for processing classification problems. The CatBoost training model regards the selection of the experimental design method as a classification problem.

5. The automatic selection system for experimental design methods based on machine learning according to claim 4, characterized in that: The CatBoost training model training steps include: S1. Prepare training data: collect historical experimental data, extract features and labels; S2, model configuration: iterations: 100, learning rate: 0.1, tree depth: 6, loss function: MultiClass; S3, model training: fit the model using training data; S4. Model evaluation: Model performance was evaluated using cross-validation.

6. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The weight calculation module is the core of the system and includes three parts: basic weight, conditional weight and complexity penalty.

7. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The dynamic selection module selects the best method based on the calculated weights, and the algorithm steps are as follows: S1. Normalized weight: Calculate the total weight S = ∑mTotalWeightm, normalized weight NormalizedWeightm = TotalWeightm / S; S2. Threshold determination: Calculate the maximum weight Max=maxmTotalWeightm and the minimum weight Min=minmTotalWeightm. If Max-Min<5%, randomly select a method; otherwise, proceed to S3. S3, Probability selection: According to the weighted probability selection method, the probability of method m being selected is NormalizedWeight m , mathematical expression: Where P(m) represents the probability of selecting method m.

8. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The feedback learning module continuously optimizes the model by recording experimental results. The steps are as follows: S1. Collect feedback data: record the characteristics of each experiment, the selected methods and the evaluation of the experimental results; S2, update training data: add feedback data to the training set; S3. Model update: Retrain the model using the expanded training set or continue training based on the original model.

9. The automatic selection system for experimental design methods based on machine learning according to claim 1, characterized in that: The automatic selection system workflow is as follows: S1. The user inputs the parameters of the experimental design problem (number of factors, number of levels, etc.); S2, feature extraction module extracts problem features; S3, weight calculation module calculates the weight of each method; S4, the method selection module selects the best method; S5. Execute the selected method to generate an experimental design plan; S6. The feedback learning module records the results and updates the model.