Method and apparatus for experimental design of multi-objective variables
By reducing the sample space through pre-sampling and cluster analysis and combining it with the Bayesian optimization model, the application difficulty of Bayesian optimization in large sample spaces was solved, and efficient experimental design was achieved.
Patent Information
- Application Number
- CN202410302493.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-03-15
AI Technical Summary
The huge sample space generated by Bayesian optimization experimental design under conditions of high factors and high level values leads to storage and computational burdens, making it difficult to apply to experimental design with large sample spaces.
By obtaining multiple characteristic variables to form the original sample space, pre-sampling and cluster analysis are performed to narrow the sample space, and the target sample space is generated using the pre-stored sampling algorithm. The experimental point combination optimization is performed in combination with the Bayesian optimization model.
The sample space is effectively reduced, making it possible to apply the Bayesian optimization method to experimental design with large sample spaces, thereby improving the efficiency and accuracy of experimental design.
Smart Images

Figure CN118098402B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of formula experiment, in particular to a method and device for designing an experiment scheme with multiple target variables. BACKGROUND
[0002] Bayesian optimization experiment design is an experiment design technique based on Bayesian optimization method. It uses Bayesian statistical method to optimize the experiment process, uses the information predicted by the prior model to guide the experiment, and finds the best parameter combination by continuously iterating and exploring the observation results, so as to more efficiently explore the parameter space. Compared with other experiment design methods, Bayesian optimization has the advantages of high efficiency, support for continuous and discrete parameters, and independence on parameter correlation. However, if the types of experimental factors and the level values of factors are too high, a huge sample space will be generated, which may cause the sample space to be unable to be stored and the Bayesian optimization process to be unable to be performed.
[0003] Therefore, it is an urgent problem for those skilled in the art to provide a method and device for designing an experiment scheme with multiple target variables, so as to provide data support for experiment scheme design by using a sample data set with a reduced space, and make it possible to apply the Bayesian optimization experiment method to a large sample space. SUMMARY
[0004] To this end, the embodiments of the present application provide a method and device for designing an experiment scheme with multiple target variables, so as to provide data support for experiment scheme design by using a sample data set with a reduced space, and make it possible to apply the Bayesian optimization experiment method to a large sample space.
[0005] In order to achieve the above-mentioned purpose, the embodiments of the present application provide the following technical solutions:
[0006] The present application provides a method for designing an experiment scheme with multiple target variables, the method comprising:
[0007] Obtaining multiple characteristic variables constituting an experiment scheme, and using all the characteristic variables to constitute an original sample space;
[0008] Pre-sampling the values of each characteristic variable in the original sample space to obtain a statistical sampling result;
[0009] Based on the statistical sampling result, performing cluster analysis on each characteristic variable in the pre-processed sample space to obtain target sampling points, and using all the target sampling points to constitute a pre-processed sample space;
[0010] Using a pre-stored sampling algorithm to sample the pre-processed sample space to obtain a target sample space, and the samples in the target sample space are used as experimental samples in the experiment scheme design.
[0011] In some embodiments, the values of each feature variable are sampled to obtain statistical sampling results, specifically including:
[0012] In the case that the feature variable is a continuous feature variable, the values of each feature variable are sampled to obtain statistical sampling results, the statistical sampling results at least including the mean, extreme value, median and mode of all continuous feature variables.
[0013] In some embodiments, based on the statistical sampling results, a cluster analysis is performed on each feature variable in the pre-processed sample space to obtain target sampling points, specifically including:
[0014] Each feature variable is subjected to cluster analysis, and the cluster centers of each cluster result are taken as the target sampling points according to the dispersion degree of data.
[0015] In some embodiments, the pre-stored sampling algorithm is a Latin square sampling algorithm or a random sampling algorithm.
[0016] In some embodiments, the samples in the target sample space are used as experimental samples for the verification process of the experimental scheme design, specifically including:
[0017] A high-throughput full-factor virtual data set is constructed, the high-throughput full-factor virtual data set including all sample points in the space and the real experimental results corresponding to each sample point of the sample points;
[0018] The high-throughput full-factor virtual data set is pre-processed to obtain a target sample space;
[0019] The sample data in the target sample space are taken as an initial data set to iteratively optimize a Bayesian initial model to obtain a Bayesian optimization model;
[0020] An experimental point combination is obtained based on the Bayesian optimization model.
[0021] In some embodiments, the high-throughput full-factor virtual data set includes 5 independent variables, 3 target variables, and a total sample size of 150W.
[0022] In some embodiments, based on the Bayesian optimization model, the experimental point combination is obtained, and then further including:
[0023] The experimental point combination is verified by experiments, and in the case that the verification result meets the preset condition, the iteration of the Bayesian optimization model is ended.
[0024] The application also provides a multi-target variable experimental scheme design device, the device including:
[0025] The feature extraction unit is configured to acquire a plurality of characteristic variables constituting an experimental scheme, and to form an original sample space by using all the characteristic variables.
[0026] The pre-sampling unit is configured to pre-sample the values of each of the characteristic variables in the original sample space to obtain a statistical quantity sampling result.
[0027] The clustering analysis unit is configured to perform clustering analysis on each of the characteristic variables in the pre-processed sample space based on the statistical quantity sampling result to obtain target sampling points, and to form a pre-processed sample space by using all the target sampling points.
[0028] The result generation unit is configured to sample the pre-processed sample space by using a pre-stored sampling algorithm to obtain a target sample space, and to use the samples in the target sample space as experimental samples in the experimental scheme design.
[0029] The present application also provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method as described above when executing the program.
[0030] The present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executable on a processor to implement the steps of the method as described above.
[0031] The experimental scheme design method for multiple target variables provided by the present application acquires a plurality of characteristic variables constituting an experimental scheme, forms an original sample space by using all the characteristic variables, pre-samples the values of each of the characteristic variables in the original sample space to obtain a statistical quantity sampling result, performs clustering analysis on each of the characteristic variables in the pre-processed sample space based on the statistical quantity sampling result to obtain target sampling points, forms a pre-processed sample space by using all the target sampling points, samples the pre-processed sample space by using a pre-stored sampling algorithm to obtain a target sample space, and uses the samples in the target sample space as experimental samples in the experimental scheme design. The method and device provided by the present application can reduce the sample space by using a unique statistical quantity clustering pre-sampling method, and make the feasibility of Bayesian optimization in a large sample space possible. The sample data set with the reduced space can provide data support for the experimental scheme design, and make it possible to apply the experimental method of Bayesian optimization to a large sample space. BRIEF DESCRIPTION OF DRAWINGS
[0032] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required in the embodiments or prior art description. Obviously, the drawings in the following description are only exemplary, and for those skilled in the art, other drawings can also be obtained from the provided drawings without creative labor.
[0033] The structures, proportions, sizes, etc. shown in the specification are only used to cooperate with the content disclosed in the specification, to be understood and read by those skilled in the art, and are not used to limit the conditions that the present application can be implemented, so they do not have technical significance. Any modification of structure, change of proportion relationship or adjustment of size, without affecting the effects and purposes that the present application can produce, should still fall within the scope of the technical content disclosed by the present application.
[0034] Figure 1 One of the flowcharts of the experimental scheme design method of multiple target variables provided by the present application;
[0035] Figure 2 A data sampling effect diagram;
[0036] Figure 3 The second flowchart of the experimental scheme design method of multiple target variables provided by the present application;
[0037] Figure 4 The third flowchart of the experimental scheme design method of multiple target variables provided by the present application;
[0038] Figure 5 A chart of 10 data in Ds;
[0039] Figure 6 A chart of part of the data set of the high-throughput full-factor virtual data set Dv;
[0040] Figure 7 A schematic diagram of 9 sample points recommended by Bayesian optimization;
[0041] Figure 8 A schematic diagram of experimental conditions and value levels of direct arylation reaction of imidazole organic chemistry;
[0042] Figure 9 A schematic diagram of experimental results of product yield and experimental cost in 10 rounds of iteration;
[0043] Figure 10 A structural block diagram of the experimental scheme design device of multiple target variables provided by the present application;
[0044] Figure 11 A structural block diagram of a computer device provided by the present application. DETAILED DESCRIPTION
[0045] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0046] Please refer to Figure 1 , Figure 1 This is one of the flow charts of the experimental scheme design method for multiple objective variables provided by the present invention.
[0047] In a specific embodiment, the experimental design method for multiple objective variables provided by the present invention includes the following steps:
[0048] S110: Acquire multiple characteristic variables that constitute the experimental plan, and use all of these characteristic variables to form the original sample space. It should be understood that characteristic variables in the experimental plan refer to experimental factors, which are provided by the specific experiment being studied, such as the experimental temperature, experimental humidity, etc. In fact, characteristic variables are not necessarily continuous. For example, the types of acids include formic acid, acetic acid, hydrochloric acid, sulfuric acid, etc., which are categorized variables. Characteristic variables are experimental factors, and target variables are the variables that need to be optimized in the experiment and are the variables that the experiment aims to obtain, such as the yield of organic chemistry and the experimental cost.
[0049] S120: Pre-sample the values of each of the characteristic variables in the original sample space to obtain statistical sampling results; specifically, when the characteristic variable is a continuous characteristic variable, sample the values of each of the characteristic variables to obtain statistical sampling results, and the statistical sampling results at least include the mean, extreme value, median and mode of all continuous characteristic variables. In a specific usage scenario, statistical sampling is performed on the values of the continuous characteristic variable A, and its conventional statistics are taken, which mainly include the mean, extreme value, median, mode and other statistical results to represent the information contained in the variable. In this embodiment, the continuous characteristic variable A refers to the experimental factor, which can be understood as the temperature of the experiment.
[0050] S130: Based on the statistical sampling results, cluster analysis is performed on each of the characteristic variables in the preprocessed sample space to obtain target sampling points, and all of the target sampling points are used to form the preprocessed sample space. Specifically, cluster analysis is performed on each of the characteristic variables, clustering is performed according to the degree of dispersion of the data, and the centroid of each cluster result is taken as the target sampling point. In the above specific usage scenario, cluster analysis is performed on the continuous characteristic variable A, and the continuous characteristic variable A is clustered according to the degree of dispersion of the data, and the centroid of the cluster is taken as the sampling point.
[0051] S140: Using a pre-stored sampling algorithm, the pre-processed sample space is sampled to obtain a target sample space. The samples in the target sample space are used as experimental samples in the experimental design. The pre-stored sampling algorithm is a Latin square sampling algorithm or a random sampling algorithm. Both the Latin square sampling algorithm and the random sampling algorithm are conventional algorithms and are not described in detail.
[0052] In the above specific usage scenario, for the continuous feature variable A, the factor level values represented by the mean, extreme value, median, mode, cluster centroid and other features are obtained. Figure 2 As shown, if the sample space that has not been sampled is recorded as D Ocean , after the statistical clustering pre-sampling method, we can get a Ocean The sample space is small and can be explicitly stored, denoted as D Lake Since the acquisition function used in the Bayesian optimization process consumes a lot of resources, Lake On this basis, conventional sampling functions, such as Latin square sampling and random sampling, are used to further narrow the search space of samples. The obtained sample space is recorded as D Pool .
[0053] In principle, the statistical clustering pre-sampling is to reduce the sample size of the multi-objective Bayesian optimization of statistical clustering pre-sampling method (MOBOSCP). In order to verify the effectiveness of the MOBOSCP in the condition of continuous large sample experimental space, a high-throughput virtual data set containing 150w full-factor 2 target variables is constructed, and the MOBOSCP is used to search for an experimental combination better than the current best result. In addition, in order to verify the scalability of the MOBOSCP in the application field, the MOBOSCP is applied to an organic chemical reaction containing continuous and discrete characteristic variables, and an experimental factor combination with high yield and low cost is searched by using the MOBOSCP.
[0054] In some embodiments, the samples in the target sample space are used as experimental samples for the verification process of the experimental scheme design, as shown in Figure 3 The method comprises the following steps:
[0055] S310: A high-throughput full-factor virtual data set is constructed, which includes all sample points in the space and the true experimental results corresponding to each sample point. The high-throughput full-factor virtual data set includes 5 independent variables, 3 target variables, and a total sample size of 150W.
[0056] S320: The high-throughput full-factor virtual data set is preprocessed to obtain a target sample space.
[0057] S330: The sample data in the target sample space is used as an initial data set to iteratively optimize the Bayesian initial model to obtain a Bayesian optimization model.
[0058] S340: An experimental point combination is obtained based on the Bayesian optimization model.
[0059] S350: The experimental point combination is verified by experiments, and in the case that the verification result meets the preset condition, the iteration of the Bayesian optimization model is ended.
[0060] In a specific use scenario, as shown in Figure 4 The MOBOSCP implementation process comprises the following steps:
[0061] (1) Prepare the input data; as shown in Figure 5 The data set mainly includes experimental factors, experimental factor level values, target variable values, and the expected optimization direction of the target variable, for example,Figure 4 Target1 is large, while target2 is small.
[0062] (2) Input the data in step (1) into MOBOSCP for Bayesian optimization to recommend a potentially better experimental point combination;
[0063] (3) Conduct experiments based on the experimental point combination recommended in step (2) to obtain the target variables and see whether the values of these target variables meet the experimental requirements. If they do, exit the loop iteration; otherwise, continue the loop iteration and continue to use MOBOSCP to recommend new experimental points.
[0064] For ease of understanding, the following briefly describes the implementation process and technical effects of the method provided by the present invention using three specific usage scenarios as examples.
[0065] In Example 1, a Bayesian optimization experiment based on statistical clustering pre-sampling in a large sample space to optimize multiple objective variables is taken as an example.
[0066] Since full-factor high-throughput experimental data are often difficult to obtain in reality, in order to verify the effectiveness of the statistical clustering pre-sampling strategy, a high-throughput full-factor virtual data set with 5 independent variables, 3 target variables and a total sample size of 1.5 million was constructed, denoted as Dv. Some data of Dv are as follows Figure 6 shown.
[0067] This experiment simulates a real-world scenario: optimizing the strength and toughness of a high-entropy alloy. Dv includes all sample points in space (assuming there are 150,000 sample points in the entire experimental space) and their corresponding real-world experimental results (strength and toughness), while Ds represents the 10 data points obtained by the research team's current experiments.
[0068] exist Figure 6 In the example, the maximum value of target1, target1max, is 24, and the minimum value of target2, target2min, is 3838. The optimization goal is to find at least one set of sample points that simultaneously satisfy target1>=24 & target2<=3838. Figure 7 As shown in the figure, Ds is first used as the initial Bayesian optimization modeling dataset into MOBOSCP to train the initial model. In the first round of nine recommended test points given by the model, it can be found that the target 1 and target 2 corresponding to sample points 2 to 9 in the Dv dataset already meet the experimental objectives.
[0069] Following this approach, we can repeatedly iterate experimental data and models based on MOBOSCP, obtaining better experimental plans or materials through limited experiments. It should be noted that the goal of the experiment is not to find the global optimal solution in Dv. Even using Bayesian optimization or other related optimization methods, it is impossible to find or approximate the global optimal solution in Dv with just dozens or even hundreds of sample points. A more realistic goal is to find a point that is better than the current point and break through the existing level.
[0070] In Example 2, the product yield and experimental cost of the direct arylation organic chemical reaction of imidazole are optimized using Bayesian optimization as an example.
[0071] like Figure 8 As shown in the figure, in the direct arylation reaction of imidazole, the experimental factors that need to be optimized are 4 bases, 12 ligands, 4 solvents, 3 solvent concentrations, and 3 temperatures. The high-throughput full-factor experimental space formed by these experimental conditions and level values is 1728. These 1728 sample points have been completed by high-throughput experiments. The target variables that need to be optimized are product yield and experimental cost. Among them, the yield value range is (0-100), and the experimental cost is mainly determined by the type and content of the experimental drug, and its value range is (0.02-0.48). The following are the results of the optimization using MOBOSCP. Figure 9 As shown, after 5 rounds, 5 times each round, and 25 experimental points in total, we found the experimental sample point with a yield of 99% and a cost of 0.03, and completed the optimization task very well.
[0072] The verification results show that the present invention uses a unique statistical clustering pre-sampling method to reduce the experimental sample space and proposes a multi-objective variable Bayesian optimization experimental design method based on the statistical clustering pre-sampling method, MOBOSCP, which effectively addresses the problem of continuous large sample space and successfully optimizes the experimental results of high-throughput virtual data sets and organic chemical reactions, finding better experimental points than existing data sets. This method has a wide range of applications and feasibility.
[0073] In Example 3, the specific scenario of the multi-objective variable Bayesian optimization algorithm based on the statistical clustering pre-sampling method can be used to design a solid coating experimental formula. The coating formula components mainly include experimental components such as resin, filler, and curing agent. In Example 3, there are mainly 11 experimental factors, namely A1 to A9, B1 to B2, and the specific components are shown in Table 1. The purpose is to find a coating formula that meets the requirements of good impact resistance (IMPACT), wear resistance (ABR), and electrical properties (ELECD). The specific original data is shown in Table 2.
[0074] Table 1 Experimental component table of paint
[0075] Component Class A1 Resin A2 Resin A3 Resin A4 Filler (solid) A5 Filler (solid) A6 Filler (solid) A7 Filler (solid) A8 Resin A9 Resin B1 Curing agent B2 Curing agent
[0076] Table 2 Original paint component parameter and its paint performance data table
[0077]
[0078]
[0079] In this embodiment 3, the method for paint formulation optimization using the multi-target variable experimental scheme design method provided by the present application includes the following steps:
[0080] S1: Obtain a plurality of characteristic variables constituting the experimental scheme, and use all the characteristic variables to constitute an original sample space; in this embodiment, specifically, 11 experimental characteristic variables A1-A9, B1-B2 are arranged and combined to obtain an experimental space of 46 billion original spaces, denoted as D raw .
[0081] S2: Pre-sample the values of each of the characteristic variables in the original sample space to obtain a statistical sampling result; in this embodiment, specifically, calculate the mean, extreme value, mode, quantile, etc. of the characteristic variables D raw , and the sampling results of these statistics after permutation and combination are denoted as D sample1 .
[0082] S3: Based on the statistical sampling result, perform cluster analysis on each of the characteristic variables in the pre-processed sample space to obtain target sampling points, and use all the target sampling points to constitute a pre-processed sample space; in this embodiment, specifically, perform Kmeans algorithm clustering on the statistics D sample1 , wherein the number of clusters is defined according to the size of the sample space, and in this example, the number of clusters is set to 10, and the sample space D sample2 .
[0083] S4: Sample the pre-processed sample space using a pre-stored sampling algorithm to obtain a target sample space, and the samples in the target sample space are used for experimental samples in the experimental scheme design. In this embodiment, the experimental samples are specifically D sample2 .
[0084] In this way, when the method provided by the present invention is used to design the experiment of the paint formula, the subsequent Bayesian optimization cannot be performed without using the statistical clustering pre-sampling method to form a sample space of up to 46 billion experimental point combinations. After pre-sampling using the statistical clustering pre-sampling method combined with Bayesian optimization, an experimental formula that meets multi-objective performance is found.
[0085] In the above specific embodiment, the experimental design method for multiple target variables provided by the present invention obtains multiple characteristic variables constituting the experimental plan, and uses all the characteristic variables to form an original sample space; pre-samples the values of each characteristic variable in the original sample space to obtain a statistical sampling result; based on the statistical sampling result, cluster analysis is performed on each characteristic variable in the pre-processed sample space to obtain a target sampling point, and all the target sampling points are used to form a pre-processed sample space; the pre-processed sample space is sampled using a pre-stored sampling algorithm to obtain a target sample space, and the samples in the target sample space are used as experimental samples in the experimental design. The method and device provided by the present invention, in response to the problem that there are many levels of experimental factors and their values, resulting in a huge sample space, adopt a unique statistical clustering pre-sampling method to reduce the sample space, thereby achieving the feasibility of Bayesian optimization of large sample spaces; the sample data set with reduced space is used to provide data support for experimental design, making it possible to apply the Bayesian optimization experimental method to large sample spaces.
[0086] In addition to the above method, the present invention also provides a multi-objective variable experimental design device, such as Figure 10 As shown, the device includes:
[0087] A feature extraction unit 1010 is used to obtain multiple feature variables constituting the experimental plan, and use all of the feature variables to form an original sample space;
[0088] A pre-sampling unit 1020 is configured to pre-sample the values of each of the characteristic variables in the original sample space to obtain a statistical sampling result;
[0089] A cluster analysis unit 1030 is configured to perform cluster analysis on each of the characteristic variables in the preprocessed sample space based on the statistical sampling result to obtain target sampling points, and use all the target sampling points to form the preprocessed sample space;
[0090] The result generating unit 1040 is configured to sample the pre-processed sample space using a pre-stored sampling algorithm to obtain a target sample space. The samples in the target sample space are used as experimental samples in the experimental scheme design.
[0091] In some embodiments, the values of each of the feature variables are sampled to obtain statistical sampling results, specifically including:
[0092] In the case that the feature variables are continuous feature variables, the values of each of the feature variables are sampled to obtain statistical sampling results, the statistical sampling results at least including the mean, extreme value, median and mode of all continuous feature variables.
[0093] In some embodiments, based on the statistical sampling results, a cluster analysis is performed on each of the feature variables in the pre-processed sample space to obtain target sampling points, specifically including:
[0094] Each of the feature variables is subjected to cluster analysis, and the cluster is performed according to the dispersion degree of data, and the centroid of each cluster result is taken as the target sampling point.
[0095] In some embodiments, the pre-stored sampling algorithm is a Latin square sampling algorithm or a random sampling algorithm.
[0096] In some embodiments, the samples in the target sample space are used as experimental samples for the verification process of the experimental scheme design, specifically including:
[0097] A high-throughput full-factor virtual data set is constructed, the high-throughput full-factor virtual data set including all sample points in the space and the real experimental results corresponding to each sample point of the sample points;
[0098] The high-throughput full-factor virtual data set is pre-processed to obtain a target sample space;
[0099] The sample data in the target sample space are taken as an initial data set to perform iterative optimization on a Bayesian initial model to obtain a Bayesian optimization model;
[0100] An experimental point combination is obtained based on the Bayesian optimization model.
[0101] In some embodiments, the high-throughput full-factor virtual data set includes 5 independent variables, 3 target variables, and a total sample size of 150W.
[0102] In some embodiments, based on the Bayesian optimization model to obtain the experimental point combination, further including:
[0103] The experimental point combination is verified by experiments, and in the case that the verification result meets the preset condition, the iteration of the Bayesian optimization model is ended.
[0104] In the foregoing specific embodiments, the multi-target variable experimental scheme design device provided by the application obtains a plurality of characteristic variables constituting an experimental scheme, uses all the characteristic variables to constitute an original sample space, pre-samples the values of each characteristic variable in the original sample space to obtain a statistical quantity sampling result, performs cluster analysis on each characteristic variable in the pre-processed sample space based on the statistical quantity sampling result to obtain target sampling points, uses all the target sampling points to constitute a pre-processed sample space, and uses a pre-stored sampling algorithm to sample the pre-processed sample space to obtain a target sample space, wherein the samples in the target sample space are used as experimental samples in experimental scheme design. The device provided by the application can reduce the sample space by using a unique statistical quantity cluster pre-sampling method, and can realize the feasibility of Bayesian optimization of a large sample space. The reduced sample data set can provide data support for experimental scheme design, and can make it possible to apply the Bayesian optimization experimental method to a large sample space.
[0105] In one embodiment, a computer device, which can be a server, has an internal structure diagram as shown in Figure 11 The computer device includes a processor, a memory and a network interface connected by a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a model prediction. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The model prediction of the computer device is configured to store static information and dynamic information data. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the steps in the above method embodiments.
[0106] Those skilled in the art can understand that Figure 11 The structure shown in the above embodiment is only a block diagram of part of the structure related to the application scheme, and does not limit the computer device to which the application scheme is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0107] Corresponding to the above embodiments, the application also provides a computer storage medium containing one or more program instructions. The one or more program instructions are used to execute the method described above.
[0108] The application further provides a computer program product, which comprises a computer program, the computer program being stored in a non-transitory computer readable storage medium, and the computer program being capable of executing the above method when executed by a processor.
[0109] In the embodiments of the present application, the processor can be an integrated circuit chip with a processing capability of signals. The processor can be a general processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0110] The disclosed methods, steps and logic block diagrams in the embodiments of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register or other mature storage medium in the art. The processor reads the information in the storage medium and combines the hardware to complete the steps of the above method.
[0111] The storage medium can be a memory, for example, can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
[0112] Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM) or a flash memory.
[0113] The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example, and not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The below-described subject matter can be implemented with computers using one or more of the above or any other suitable memory.
[0114] The storage media described in the embodiments of the present application is intended to include, but not be limited to, these and any other suitable types of memory.
[0115] Those skilled in the art should be aware that the functions described in the embodiments of the present application can be implemented in combination of hardware and software in one or more of the above examples. When the software is applied, the corresponding functions can be stored in a computer readable medium or transmitted as one or more instructions or codes on the computer readable medium. The computer readable medium includes a computer storage medium and a communication medium, wherein the communication medium includes any medium that facilitates the transfer of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special purpose computer.
[0116] The above detailed description is further intended to serve as a purpose, technical solutions and beneficial effects of the present application. It should be understood that the above is only a specific embodiment of the present application, and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the present application shall be included in the protection scope of the present application.
Claims
1. A method for designing a multi-objective experimental scheme, characterized in that: The experimental scheme design method is used for coating formula optimization, and the method comprises the following steps: A plurality of characteristic variables constituting an experimental scheme are acquired, and an original sample space is constituted using all of the characteristic variables; the characteristic variables include original paint component parameters, and the original sample space is obtained by arranging and combining all of the characteristic variables, denoted as D raw ; The value of each feature variable in the original sample space is pre-sampled to obtain a statistical quantity sampling result; specifically, D raw The statistical quantity of the feature variable, and the arrangement and combination of all statistical quantities to obtain a statistical quantity sampling result, denoted as D sample1 The statistical quantity includes mean, extreme value, median, and mode; Based on the statistical sampling results, each of the feature variables in the pre-processed sample space is subjected to cluster analysis to obtain target sampling points, and all the target sampling points are used to form the pre-processed sample space; specifically, D sample1 The statistical quantity is subjected to Kmeans algorithm clustering, wherein the number of clusters is defined according to the size of the sample space, and the sample space D sample2 is obtained. Using the pre-stored sampling algorithm, the pre-processed sample space D sample2 performing sampling to obtain a target sample space, wherein samples within the target sample space are used as experimental samples in an experimental design of a coating formulation; The samples in the target sample space are used as experimental samples for the verification process of the experimental scheme design, and specifically comprises the following steps: A high-throughput full-factor virtual data set is constructed, which comprises all sample points in the space and the real experimental results corresponding to each sample point; The high-throughput full-factor virtual data set is preprocessed to obtain a target sample space; The sample data in the target sample space are used as an initial data set to iteratively optimize a Bayesian initial model to obtain a Bayesian optimization model; Based on the Bayesian optimization model, an experimental point combination is obtained; In the coating formula optimization process, the implementation steps of the Bayesian optimization model comprise the following steps: (1) preparing input data; the data set mainly comprises experimental factors, experimental factor level values, target variable values, and the expected optimization direction of the target variable; (2) inputting the data in step (1) into the Bayesian optimization model to perform Bayesian optimization and recommend a potential optimal experimental point combination; (3) performing experiments according to the experimental point combination recommended in step (2) to obtain target variables, and checking whether the values of the target variables meet the experimental requirements; if yes, the loop iteration is exited; otherwise, the loop iteration is continued, and the Bayesian optimization model is used to recommend a new experimental point.
2. The method of designing an experiment for multiple objectives and variables of claim 1, wherein, Based on the statistical quantity sampling result, each feature variable in the preprocessed sample space is subjected to cluster analysis to obtain a target sampling point, and specifically comprises the following steps: Each feature variable is subjected to cluster analysis, and the cluster results are clustered according to the dispersion degree of the data, and the centroids of the cluster results are taken as the target sampling points.
3. The method of designing an experiment for multiple objectives and variables of claim 1, wherein, The pre-stored sampling algorithm is a Latin square sampling algorithm or a random sampling algorithm.
4. The method of designing an experiment for multiple objectives and variables of claim 1, wherein, The high-throughput full-factor virtual data set comprises five independent variables, three target variables, and a total sample size of 150W.
5. The method of designing an experiment protocol for multiple objectives and variables as recited in claim 1, wherein, Based on the Bayesian optimization model, an experimental point combination is obtained, and then the following steps are further included: The experimental point combination is verified by experiments, and the iteration of the Bayesian optimization model is ended when the verification result meets the preset condition.
6. A multi-objective variable experimental design apparatus for implementing the method of any one of claims 1-5, characterized by, The device comprises: a feature extraction unit configured to acquire a plurality of feature variables constituting an experimental scheme, and to use all the feature variables to constitute an original sample space; a pre-sampling unit configured to pre-sample the values of each feature variable in the original sample space to obtain a statistical quantity sampling result; a cluster analysis unit configured to, based on the statistical quantity sampling result, perform cluster analysis on each feature variable in the preprocessed sample space to obtain a target sampling point, and to use all the target sampling points to constitute a preprocessed sample space; a result generation unit configured to use a pre-stored sampling algorithm to sample the preprocessed sample space to obtain a target sample space, and to use the samples in the target sample space as experimental samples in the experimental scheme design.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the method of any one of claims 1-5 when executing the program.
8. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1-5.
Citation Information
Patent Citations
Hyper-parameter optimization method, related device and storage medium
CN116029368A
Experimental design generation method and device and computer equipment
CN116384034A