Multi-fidelity modeling method and device for improving material formula performance prediction
By dividing the experimental data into low-fidelity and high-fidelity data sets, Gaussian process regression models were established and multi-fidelity fusion learning was carried out, which solved the data scarcity and model robustness problems in material formulation performance prediction, and achieved high-precision multi-performance index prediction.
Patent Information
- Application Number
- CN202510356108.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art has problems of data scarcity, model overfitting and poor robustness in material formulation performance prediction, and it is difficult to effectively predict multiple mutually restricted performance indicators, resulting in lack of reference significance for the prediction results.
By dividing the experimental data into low-fidelity and high-fidelity data sets, Gaussian process regression models are established separately, and multi-fidelity fusion learning is carried out to generate multi-fidelity models to improve the prediction effect.
It significantly improves the prediction accuracy and robustness of the new sample, can effectively avoid the impact of data errors when the data volume is scarce, and achieve high-precision multi-performance indicator prediction.
Smart Images

Figure CN120299580A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-fidelity modeling, and particularly to a multi-fidelity modeling method and device for improving the performance prediction of material formulations. Background Art
[0002] In people's daily lives, the use of various chemical materials is indispensable. And preparing reasonable and effective materials requires great efforts from scientific researchers, who need to conduct a large number of attempts on the combinations of various formula raw materials to ensure that the prepared materials meet various excellent properties. However, due to the huge search space, it is impossible to try them one by one. Therefore, deriving the properties that the untested raw material combinations should have from the existing solutions, that is, performance prediction, can well assist scientific researchers in material preparation work. Usually, scientific researchers will conduct multiple rounds of experimental verification on several promising formulas respectively to obtain the performance index values of multiple groups of experiments. These values may vary greatly in some cases, resulting in too large a fluctuation range of the performance index values in the test data of different batches of the same formula. The effect of the conventional method of establishing a machine learning model to fit all the data simultaneously is not ideal, and the prediction effect on new formulas is also not satisfactory. There is a large deviation between the prediction result and the real experimental value, and it cannot effectively provide reference for relevant personnel.
[0003] Currently, due to the consumption of manpower, material resources and financial resources in the material preparation process, it is difficult to easily obtain a large amount of high-quality real experimental data. Usually, only a small part of experimental data can be accumulated in a long period of time. The problem of scarcity of experimental data widely exists in the chemical industry. At the same time, using machine learning algorithms to fit all historical data at once for modeling may lead to overfitting problems of the model. And when there are a large number of outliers or abnormal values in the data, it is easy to interfere with the learning process of the model, and there will be a large deviation when predicting new samples, and the robustness of the model is poor. In addition, usually a certain material needs to meet multiple different performance indicators at the same time, and these indicators are very likely to restrict each other, which poses a huge challenge to the performance prediction task. The existing methods cannot effectively predict multiple properties at the same time, resulting in the lack of reference significance and value of the prediction results.
[0004] Therefore, how to invent a multi-fidelity modeling method to reasonably divide historical data according to different fidelity levels and achieve high-precision and high-value performance prediction has become an urgent problem to be solved. Summary of the Invention
[0005] To this end, the present invention provides a multi-fidelity modeling method and device for improving the prediction of material formulation performance. By dividing experimental data into a low-fidelity data set and a high-fidelity data set, Gaussian process regression models are respectively established for fitting, and then the two models are subjected to multi-fidelity fusion learning (Multi-fidelity learning) to generate a multi-fidelity model, which can significantly improve the prediction effect on new samples.
[0006] To achieve the above object, the present invention provides the following technical solutions: A multi-fidelity modeling method for improving the prediction of material formulation performance, comprising:
[0007] By collecting the formulation and proportion data of several materials and annotating them with formulation numbers, a formulation data set is generated; by collecting several test samples corresponding to the formulations in the formulation data set, an experimental data set is generated;
[0008] Preprocess the experimental data in the experimental data set to obtain preprocessed experimental data; group and aggregate the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0009] In the experimental data groups, calculate the performance standard deviation of the formulations respectively to obtain standard deviation values; according to the magnitudes of the standard deviation values, divide the experimental data into low-fidelity data and high-fidelity data;
[0010] According to the low-fidelity data and high-fidelity data, respectively construct a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model through a Gaussian process regression strategy;
[0011] Perform multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model;
[0012] Predict a target formulation sample through the multi-fidelity model to obtain a target prediction performance.
[0013] As a preferred scheme of a multi-fidelity modeling method for improving the prediction of material formulation performance, in the process of preprocessing the experimental data in the experimental data set, the experimental data is preprocessed by standardization and normalization using a tool kit in python.
[0014] As a preferred scheme of a multi-fidelity modeling method for improving the prediction of material formulation performance, in the process of calculating the performance standard deviation of the formulations, the calculation formula for the standard deviation is:
[0015]
[0016] Wherein, S is the standard deviation value; xi is each observed value; xm is the average value of each group of observed values; n is the number of observed values.
[0017] As an optimal solution of a multi-fidelity modeling method for improving the prediction of material formulation performance, in the process of performing multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, multi-fidelity fusion learning is carried out through the multi-fidelity data aggregation strategy of the convolutional neural network.
[0018] As an optimal solution of a multi-fidelity modeling method for improving the prediction of material formulation performance, in the process of performing multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, the expression of multi-fidelity fusion learning is:
[0019] Y M (x) = pY L (x) + Y H (x)
[0020] Wherein, Y M (x) is the multi-fidelity model; Y L (x) is the low-fidelity Gaussian process regression model; Y H (x) is the high-fidelity Gaussian process regression model; p is an adjustable constant factor.
[0021] The present invention also provides a multi-fidelity modeling device for improving the prediction of material formulation performance. Based on the above-mentioned multi-fidelity modeling method for improving the prediction of material formulation performance, it includes:
[0022] The formulation dataset and experimental dataset generation module is used to generate a formulation dataset by collecting the formulation and proportion data of several materials and annotating them with formulation numbers; and generate an experimental dataset by collecting several test samples corresponding to the formulations in the formulation dataset;
[0023] The experimental data grouping and aggregation module is used to preprocess the experimental data in the experimental dataset to obtain preprocessed experimental data; group and aggregate the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0024] The low-fidelity data and high-fidelity data division module is used to calculate the performance standard deviation of the formulation in the experimental data group respectively to obtain the standard deviation value; and divide the experimental data into low-fidelity data and high-fidelity data according to the size of the standard deviation value;
[0025] The low / high-fidelity Gaussian process regression model construction module is used to construct a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through the Gaussian process regression strategy according to the low-fidelity data and high-fidelity data;
[0026] A multi-fidelity model acquisition module, configured to perform multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model;
[0027] A multi-fidelity model prediction module, configured to predict a target formulation sample through the multi-fidelity model to obtain a target prediction performance.
[0028] As a preferred solution of a multi-fidelity modeling device for improving the prediction of material formulation performance, in the experimental data grouping and aggregation module, during the preprocessing of the experimental data in the experimental data set, the experimental data is preprocessed by standardization and normalization through a toolkit in Python.
[0029] As a preferred solution of a multi-fidelity modeling device for improving the prediction of material formulation performance, in the low-fidelity data and high-fidelity data division module, during the calculation of the performance standard deviation of the formulation, the calculation formula of the standard deviation is:
[0030]
[0031] In the formula, S is the standard deviation value; xi is each observation value; xm is the average value of each group of observation values; n is the number of observation values.
[0032] As a preferred solution of a multi-fidelity modeling device for improving the prediction of material formulation performance, in the multi-fidelity model acquisition module, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, multi-fidelity fusion learning is performed through the multi-fidelity data aggregation strategy of a convolutional neural network.
[0033] As a preferred solution of a multi-fidelity modeling device for improving the prediction of material formulation performance, in the multi-fidelity model acquisition module, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, the expression of multi-fidelity fusion learning is:
[0034] Y M (x) = pY L (x) + Y H (x)
[0035] In the formula, Y M (x) is the multi-fidelity model; Y L (x) is the low-fidelity Gaussian process regression model; Y H (x) is the high-fidelity Gaussian process regression model; p is an adjustable constant factor.
[0036] The present invention has the following advantages: The present invention collects the formula and proportion data of several materials, marks them with formula numbers, and generates a formula dataset; collects several test samples corresponding to the formulas in the formula dataset to generate an experimental dataset; preprocesses the experimental data in the experimental dataset to obtain preprocessed experimental data; groups and aggregates the preprocessed experimental data according to the formula numbers to obtain experimental data groups; in the experimental data groups, calculates the performance standard deviation of the formulas respectively to obtain standard deviation values; divides the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values; constructs a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy based on the low-fidelity data and the high-fidelity data; performs multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model; predicts a target formula sample through the multi-fidelity model to obtain a target predicted performance. In the present invention, the performance differences between different batches of test data under the same formula are relatively large and belong to low-fidelity data. Although it is not conducive to directly establishing a model for fitting and learning, it can provide some additional trend information and prior knowledge for the relatively small amount of high-fidelity data, thus becoming a beneficial supplement to improving the effect of the performance prediction model. Conventional machine learning algorithms are prone to overfitting when the amount of data is small and are easily affected by outliers or abnormal points in the dataset, resulting in an excessive overall deviation of the fitting curve. However, the distribution containing deviations fitted by the Gaussian process regression in the present invention can better take into account a reasonable deviation range, can also play an advantage when the amount of data is small, has better robustness, and compared with other machine learning algorithms, can perform automatic iterative optimization, is easy to implement multi-model fusion learning, and is more suitable for dealing with the situation of scarce data. The present invention divides historical data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation, and respectively establishes corresponding models for learning, which can effectively avoid the influence of data itself errors and make the mapping relationship learned by the model more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be obtained based on the provided drawings.
[0038] The structures, proportions, sizes, etc. shown in this specification are only used to match the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the implementation conditions of the present invention. Therefore, they do not have substantial technical significance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the efficacy that the present invention can produce and the purpose that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0039] Figure 1 It is a schematic flowchart of a multi-fidelity modeling method for improving the performance prediction of material formulations provided in Embodiment 1 of the present invention;
[0040] Figure 2 It is a schematic flowchart of the specific implementation process of a multi-fidelity modeling method for improving the performance prediction of material formulations provided in Embodiment 1 of the present invention;
[0041] Figure 3 It is a schematic diagram of the architecture of a multi-fidelity modeling device for improving the performance prediction of material formulations provided in Embodiment 2 of the present invention. Specific implementation manners
[0042] The following specific embodiments illustrate the implementation manners of the present invention. Those familiar with this technology can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0043] Embodiment 1
[0044] Refer to Figure 1 and Figure 2 Embodiment 1 of the present invention provides a multi-fidelity modeling method for improving the performance prediction of material formulations, including the following steps:
[0045] S1. By collecting the formulation and proportion data of several materials and labeling them with formulation numbers, a formulation data set is generated; by collecting several test samples corresponding to the formulations in the formulation data set, an experimental data set is generated;
[0046] S2. Preprocess the experimental data in the experimental data set to obtain preprocessed experimental data; group and aggregate the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0047] S3. In the experimental data groups, calculate the performance standard deviation of the formulations respectively to obtain standard deviation values; divide the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values;
[0048] S4. Based on the low-fidelity data and high-fidelity data, construct a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy;
[0049] S5. Perform multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model;
[0050] S6. Predict the target formulation sample through the multi-fidelity model to obtain the target prediction performance.
[0051] In this embodiment, in step S1, by collecting the formulations and proportioning data of a number of materials and annotating them with formulation numbers, a formulation data set is generated; by collecting a number of test samples corresponding to the formulations in the formulation data set, an experimental data set is generated;
[0052] Specifically, collect the respective formulations and their proportioning data of a variety of different materials, annotate them with different formulation numbers as the formulation data set; collect a number of test samples corresponding to each formulation as the experimental data set;
[0053] Among them, the formulation data set is shown in Table 1:
[0054] Formulation Number Formulation Ratio 1 Raw Material 1 66 1 Raw Material 2 2 1 Raw Material 3 20 1 Raw Material 4 8 1 Raw Material 5 3 2 Raw Material 9 50 2 Raw Material 2 50 3 Raw Material 13 22 3 Raw Material 7 48 3 Raw Material 1 30 ... ... ...
[0055] Table 1 Formulation data set
[0056] The experimental data set is shown in Table 2:
[0057] Formulation Number Performance Index 1 Performance Index 2 Performance Index 3 ... 1 27.9 106 134 ... 1 26.5 31 99 ... 1 27.4 85 40 ... ... ... ... ... ... 2 26.4 131 33.7 ... 2 25.8 131 32.7 ... 2 26.8 145 33.5 ... 2 26.4 176 33.6 ... ... ... ... ... ...
[0058] Table 2 Experimental data set
[0059] In this embodiment, in step S2, preprocess the experimental data in the experimental data set to obtain preprocessed experimental data; group and aggregate the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0060] Specifically, use the toolkits in python to process problems such as missing values and outliers in the experimental data, and perform data preprocessing operations such as data normalization and standardization on the data to obtain preprocessed experimental data; group and aggregate all the data in the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0061] Taking the datasets in Table 1 and Table 2 as examples, first count the types of raw materials included in all the formulations in Table 1, and use them as attribute columns, while the number of rows is the number of rows in the experimental dataset, thus forming a multi-dimensional matrix. If a certain raw material appears in the formulation, it is filled according to its corresponding ratio value; otherwise, it is filled with a value of 0, indicating that the raw material is not included in the formulation. In this way, a sparse matrix can be formed, and then it is combined with the data in Table 2 according to the formulation number to generate a dataset containing performance indicators.
[0062] In this embodiment, in step S3, in the experimental data group, calculate the performance standard deviation of the formulations respectively to obtain the standard deviation value; according to the magnitude of the standard deviation value, divide the experimental data into low-fidelity data and high-fidelity data;
[0063] Specifically, within each group, calculate the standard deviation between each performance in different formulations respectively, and divide the data into two parts: low-fidelity data and high-fidelity data according to the magnitude of the standard deviation;
[0064] Among them, the calculation formula of the standard deviation is:
[0065]
[0066] In the formula, S is the standard deviation value; xi is each observed value; xm is the average value of each group of observed values; n is the number of observed values.
[0067] In this embodiment, the low-fidelity data refers to the test data with a relatively large standard deviation within the group, indicating that although the same formulation ratio is adopted, the fluctuation range of each test result is relatively large. If such data is directly used for machine learning algorithm fitting, it will cause the model to be difficult to learn the correct mapping relationship, which is not conducive to subsequent prediction tasks; the remaining data with a relatively small standard deviation within the group and only one test record can be regarded as high-fidelity data, indicating that the performance of these formulations has a relatively small fluctuation range in different test batches and the overall performance is stable.
[0068] Among them, the threshold of the standard deviation can be set manually according to the actual business scenario and problem type; the processed formulation datasets are respectively merged with the low-fidelity dataset and the high-fidelity dataset to form the final low-fidelity dataset and high-fidelity dataset for Gaussian process regression; the specific merging steps are as follows:
[0069] First, count the types of raw materials included in all formulations in the formulation dataset, and use them as attribute columns, while the number of rows is the number of rows in the experimental dataset, thus forming a multi-dimensional matrix. If a certain raw material appears in the formulation, it is filled according to its corresponding formulation ratio; otherwise, it is filled with a value of 0, indicating that the raw material is not included in the formulation. In this way, a sparse matrix can be formed. Then, the low-fidelity dataset and the high-fidelity dataset are respectively merged with this sparse matrix according to the formulation number to form the final dataset for the subsequent steps.
[0070] In this embodiment, in step S4, according to the low-fidelity data and the high-fidelity data, a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model are respectively constructed by means of a Gaussian process regression strategy;
[0071] Specifically, two models are respectively established for the low-fidelity data and the high-fidelity data by using the Gaussian process regression strategy to fit the mapping relationship between the raw materials and performance indicators of different formulations, and a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model are generated; the main advantage of Gaussian process regression is that it can provide an estimate of the uncertainty of the prediction and can work under the condition of small samples, so it is effective for scarce experimental data.
[0072] In this embodiment, in step S5, the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model are subjected to multi-fidelity fusion learning to obtain a multi-fidelity model;
[0073] Specifically, the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model are subjected to multi-fidelity fusion learning according to the fusion formula, and the methods for performing multi-fidelity fusion learning include but are not limited to multi-fidelity data aggregation using a convolutional neural network (CNN); this method includes steps such as data compilation, convolutional layer processing, and deep neural network processing, and can utilize all available low-fidelity data to learn the potential relationship between the low-fidelity data and the high-fidelity data.
[0074] Among them, the expression for multi-fidelity fusion learning is:
[0075] Y M (x) = pY L (x) + Y H (x)
[0076] In the formula, Y M (x) is the multi-fidelity model; Y L (x) is the low-fidelity Gaussian process regression model; Y H (x) is the high-fidelity Gaussian process regression model; p is an adjustable constant factor.
[0077] In this embodiment, in step S6, the multi-fidelity model is used to predict the target formulation sample to obtain the target prediction performance.
[0078] Specifically, the multi-fidelity model is used to predict the performance achieved by the untried formulation combinations to obtain the target prediction performance.
[0079] In summary, the present invention collects the formulation and proportion data of several materials, annotates them with formulation numbers to generate a formulation dataset; collects several test samples corresponding to the formulations in the formulation dataset to generate an experimental dataset; preprocesses the experimental data in the experimental dataset to obtain preprocessed experimental data; groups and aggregates the preprocessed experimental data according to the formulation numbers to obtain experimental data groups; in the experimental data groups, calculates the performance standard deviation of the formulations respectively to obtain standard deviation values; divides the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values; constructs a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy according to the low-fidelity data and the high-fidelity data; performs multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model; uses the multi-fidelity model to predict the target formulation sample to obtain the target prediction performance. In the present invention, the performance differences between different batches of test data under the same formulation are relatively large and belong to low-fidelity data. Although it is not conducive to directly establishing a model for fitting and learning, it can provide some additional trend information and prior knowledge for the relatively small amount of high-fidelity data, thus becoming a beneficial supplement to improving the effect of the performance prediction model. Conventional machine learning algorithms are prone to overfitting when the amount of data is small and are easily affected by outliers or abnormal points in the dataset, resulting in an overly large overall deviation of the fitting curve. However, the present invention fits a distribution containing deviations through Gaussian process regression, which can better take into account a reasonable deviation range, and can also play an advantage when the amount of data is small, has better robustness, and compared with other machine learning algorithms, can perform automatic iterative optimization, is easy to implement multi-model fusion learning, and is more suitable for dealing with the situation of scarce data. The present invention divides the historical data into low-fidelity data and high-fidelity data according to the magnitude of the standard deviation, and establishes corresponding models for learning respectively, which can effectively avoid the influence of data itself errors and make the mapping relationship learned by the model more accurate.
[0080] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0081] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0082] Embodiment 2
[0083] See Figure 3 , Embodiment 2 of the present invention also provides a multi-fidelity modeling device for improving the prediction of material formulation performance, including:
[0084] A formulation dataset and experimental dataset generation module 001, configured to generate a formulation dataset by collecting the formulation and proportion data of several materials and annotating them with formulation numbers; and generate an experimental dataset by collecting several test samples corresponding to the formulations in the formulation dataset;
[0085] An experimental data grouping and aggregation module 002, configured to preprocess the experimental data in the experimental dataset to obtain preprocessed experimental data; group and aggregate the preprocessed experimental data according to the formulation numbers to obtain experimental data groups;
[0086] A low-fidelity data and high-fidelity data division module 003, configured to calculate the performance standard deviation of the formulations in the experimental data groups respectively to obtain standard deviation values; divide the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values;
[0087] A low / high-fidelity Gaussian process regression model construction module 004, configured to construct a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy according to the low-fidelity data and the high-fidelity data;
[0088] A multi-fidelity model acquisition module 005, configured to perform multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model;
[0089] The multi-fidelity model prediction module 006 is configured to predict a target formulation sample through the multi-fidelity model to obtain a target prediction performance.
[0090] In this embodiment, in the experimental data grouping and aggregating module 002, during the preprocessing of the experimental data in the experimental dataset, the experimental data is preprocessed by standardization and normalization using a toolkit in Python.
[0091] In this embodiment, in the low-fidelity data and high-fidelity data partitioning module 003, during the calculation of the performance standard deviation of the formulation, the calculation formula for the standard deviation is:
[0092]
[0093] In the formula, S is the standard deviation value; xi is each observed value; xm is the average value of each group of observed values; n is the number of observed values.
[0094] In this embodiment, in the multi-fidelity model acquisition module 005, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, multi-fidelity fusion learning is performed through the multi-fidelity data aggregation strategy of the convolutional neural network.
[0095] In this embodiment, in the multi-fidelity model acquisition module 005, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, the expression for multi-fidelity fusion learning is:
[0096] Y M (x) = pY L (x) + Y H (x)
[0097] In the formula, Y M (x) is the multi-fidelity model; Y L (x) is the low-fidelity Gaussian process regression model; Y H (x) is the high-fidelity Gaussian process regression model; p is an adjustable constant factor.
[0098] It should be noted that for the information interaction, execution process, etc. between the above system modules, since they are based on the same concept as the method embodiment in Embodiment 1 of the present application, the technical effects brought by them are the same as those of the method embodiment of the present application. For specific content, reference can be made to the description in the method embodiment shown above in the present application, and details will not be repeated here.
[0099] Embodiment 3
[0100] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which program code for a multi-fidelity modeling method for improving the performance prediction of a material formulation is stored. The program code includes instructions for executing the multi-fidelity modeling method for improving the performance prediction of a material formulation according to Embodiment 1 or any possible implementation thereof.
[0101] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)), etc.
[0102] Embodiment 4
[0103] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0104] The processor and the memory communicate with each other through a bus; the memory stores program instructions executable by the processor, and the processor can execute a multi-fidelity modeling method for improving the performance prediction of a material formulation according to Embodiment 1 or any possible implementation thereof by calling the program instructions.
[0105] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in the memory. The memory can be integrated in the processor or can be located outside the processor and exist independently.
[0106] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (e.g., infrared, wireless, microwave, etc.).
[0107] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing system. They can be concentrated on a single computing system or distributed over a network composed of multiple computing systems. Optionally, they can be implemented by program codes executable by the computing system. Thus, they can be stored in a storage system and executed by the computing system. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.
[0108] Although the present invention has been described in detail with general descriptions and specific embodiments above, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of the present invention claimed.
Claims
1. A multi-fidelity modeling method for improving the performance prediction of material formulations, characterized in that, Including: Generating a recipe dataset by collecting the recipe and proportion data of several materials and annotating with recipe numbers; Generating an experimental dataset by collecting several test samples corresponding to the recipes in the recipe dataset; Preprocessing the experimental data in the experimental dataset to obtain preprocessed experimental data; Grouping and aggregating the preprocessed experimental data according to the recipe numbers to obtain experimental data groups; In the experimental data groups, calculating the performance standard deviation of the recipes respectively to obtain standard deviation values; dividing the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values; Constructing a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy according to the low-fidelity data and high-fidelity data; Performing multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model; Predicting a target recipe sample through the multi-fidelity model to obtain a target predicted performance.
2. A multi-fidelity modeling method for predicting the performance of a lifting material formula according to claim 1, characterized in that During the process of preprocessing the experimental data in the experimental dataset, standardizing and normalizing the experimental data through a toolkit in Python.
3. A multi-fidelity modeling method for predicting the performance of a lifting material formula according to claim 2, characterized in that, During the process of calculating the performance standard deviation of the recipes, the calculation formula of the standard deviation is: In the formula, S is the standard deviation value; xi is each observation value; xm is the average value of each group of observation values; n is the number of observation values.
4. A multi-fidelity modeling method for improving the performance prediction of material formulas according to claim 3, characterized in that During the process of performing multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, performing multi-fidelity fusion learning through a multi-fidelity data aggregation strategy of a convolutional neural network.
5. A multi-fidelity modeling method for improving the performance prediction of a material formula according to claim 4, characterized in that During the process of performing multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, the expression of multi-fidelity fusion learning is: Y M f(x) = pY L f(x)+Y H f(x) where Y M (x) is a multi-fidelity model; Y L (x) is a low-fidelity Gaussian process regression model; Y H (x) is a high-fidelity Gaussian process regression model; p is an adjustable constant factor.
6. A multi-fidelity modeling device for improving the performance prediction of material formulations, which adopts a multi-fidelity modeling method for improving the performance prediction of material formulations according to any one of claims 1-5, characterized in that, Including: A recipe dataset and experimental dataset generation module, which is used to generate a recipe dataset by collecting the recipe and proportion data of several materials and annotating with recipe numbers; Generating an experimental dataset by collecting several test samples corresponding to the recipes in the recipe dataset; An experimental data grouping and aggregating module, which is used to preprocess the experimental data in the experimental dataset to obtain preprocessed experimental data; grouping and aggregating the preprocessed experimental data according to the recipe numbers to obtain experimental data groups; A low-fidelity data and high-fidelity data division module, which is used to calculate the performance standard deviation of the recipes respectively in the experimental data groups to obtain standard deviation values; Dividing the experimental data into low-fidelity data and high-fidelity data according to the magnitudes of the standard deviation values; A low / high-fidelity Gaussian process regression model construction module, which is used to construct a low-fidelity Gaussian process regression model and a high-fidelity Gaussian process regression model respectively through a Gaussian process regression strategy according to the low-fidelity data and high-fidelity data; A multi-fidelity model acquisition module, which is used to perform multi-fidelity fusion learning on the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model to obtain a multi-fidelity model; The multi-fidelity model prediction module is used to predict the target formulation sample through the multi-fidelity model to obtain the target prediction performance.
7. A multi-fidelity modeling device for improving the performance prediction of a material formula according to claim 6, characterized in that, In the experimental data grouping and aggregation module, during the preprocessing of the experimental data in the experimental dataset, the experimental data is preprocessed by standardization and normalization using the toolkits in Python.
8. The multi-fidelity modeling device for predicting the performance of an enhanced material formulation according to claim 7, wherein, In the low-fidelity data and high-fidelity data division module, during the calculation of the performance standard deviation of the formulation, the calculation formula for the standard deviation is: Where S is the standard deviation value; xi is each observed value; xm is the average value of each group of observed values; n is the number of observed values.
9. The multi-fidelity modeling device for predicting the performance of a lifting material formula according to claim 8, wherein, In the multi-fidelity model acquisition module, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, multi-fidelity fusion learning is performed through the multi-fidelity data aggregation strategy of the convolutional neural network.
10. A multi-fidelity modeling device for improving the performance prediction of a material formulation, as claimed in claim 9, wherein, In the multi-fidelity model acquisition module, during the multi-fidelity fusion learning of the low-fidelity Gaussian process regression model and the high-fidelity Gaussian process regression model, the expression for multi-fidelity fusion learning is: Y M f(x) = pY L f(x) + Y H f(x) where Y M (x) is a multi-fidelity model; Y L (x) is a low-fidelity Gaussian process regression model; Y H (x) is a high-fidelity Gaussian process regression model; p is an adjustable constant factor.
Citation Information
Patent Citations
Structural member complexity performance prediction method based on multi-fidelity data fusion driving
CN116306259A
Material formula design and performance prediction method
CN116434887A
Construction method of material performance prediction model, material performance prediction method and system
CN116562159A
Hypersonic aircraft overall control collaborative design method based on multi-fidelity data fusion
CN117634020A
Multi-fidelity data fusion method for electromagnetic field prediction of linear propulsion electromagnetic energy equipment
CN119004373A