Quantum annealing-based device and method for searching for novel material
The quantum annealing-based method for new material search addresses the inefficiencies of existing methods by varying uncertainty parameters in a predictive model, enabling a global exploration of chemical space and enhancing the likelihood of discovering materials with desired properties.
Patent Information
- Application Number
- PCT/KR2024/016529
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-30
- Filing Date
- 2024-10-28
- Publication Date
- 2025-05-08
AI Technical Summary
Existing methods for developing new chemical materials are time-consuming and inefficient due to the need to review a vast combination of microscopic factors in chemical space, often focusing on local areas rather than exploring the entire global chemical space.
A quantum annealing-based new material search device and method that varies the uncertainty parameter value of a predictive model, allowing for the extraction of sample data corresponding to desired characteristics across the entire chemical space, rather than being limited to local areas.
This approach effectively expands the search for new chemical materials to the global chemical space, increasing the probability of discovering materials with desired properties by exploring a broader range of combinations.
Smart Images

Figure KR2024016529_08052025_PF_FP_ABST
Abstract
Description
Quantum annealing-based novel material exploration device and method
[0001] The present disclosure relates to a quantum annealing-based novel material exploration device and method that can effectively explore new materials satisfying desired characteristics using a quantum annealing-based quantum computing device.
[0002] Recently, with the demand for high functionality and diversification of chemical materials, there is a need for the development of new chemical materials with properties and functions that did not previously exist.
[0003] However, since the properties of chemical materials depend on many microscopic factors, a review of the vast combinations in chemical space is necessary.
[0004] Due to these factors, developing new chemical materials requires a lot of time and effort, and there are many difficulties in finding the optimal solution.
[0005] Recently, a method has been developed that uses an algorithm to search for chemical materials that satisfy desired properties in order to shorten the development time of chemical materials.
[0006] However, although this method can interpret molecules using algorithms, it still has the problem of taking a lot of time to search for the target chemical material.
[0007] Meanwhile, among the algorithms that can explore chemical materials, the method using a quantum annealer is widely used in the exploration of new materials because it has a high probability of obtaining a combination that realizes the lowest energy in the chemical space.
[0008] However, this method also had the problem that the probability of discovering new chemical materials was low because the chemical material exploration was focused only on a few local areas in the entire chemical space.
[0009] Therefore, in the future, it is necessary to develop a novel material exploration device that can effectively explore chemical materials by expanding the entire global area rather than being limited to a few local areas in the chemical space, thereby increasing the probability of discovering new materials.
[0010] The present disclosure aims to solve the above-mentioned problems and other problems.
[0011] The present disclosure aims to provide a novel material exploration device and method capable of increasing the probability of new material exploration by varying the uncertainty parameter value of a prediction model using a quantum annealing method, inputting a dataset into the prediction model with the varied uncertainty parameter value, and extracting sample data corresponding to the predicted target characteristic, thereby effectively exploring chemical materials by expanding the entire global region rather than being limited to some local regions in the chemical space.
[0012] A novel material exploration device according to one embodiment of the present disclosure includes a database storing datasets of chemical materials, and a processor for exploring a target material from the database, wherein the processor inputs a dataset into a prediction model including an uncertainty parameter to extract sample data corresponding to a predicted target characteristic, varies an uncertainty parameter value of the prediction model based on the extracted sample data, inputs a dataset into the prediction model with the varied uncertainty parameter value to additionally extract sample data corresponding to the predicted target characteristic, and explores the target material from the extracted sample data.
[0013] A method for exploring a new material of a new material exploration device according to one embodiment of the present disclosure may include a step of inputting a dataset of chemical materials into a prediction model including uncertainty parameters to extract sample data corresponding to a prediction target characteristic, a step of varying an uncertainty parameter value of the prediction model based on the extracted sample data, a step of inputting a dataset into the prediction model with the varied uncertainty parameter value to additionally extract sample data corresponding to the prediction target characteristic, and a step of exploring a target material from the extracted sample data.
[0014] According to one embodiment of the present disclosure, a novel material exploration device can effectively explore chemical materials by expanding the entire global region rather than being limited to some local regions in chemical space by varying the uncertainty parameter value of a prediction model using a quantum annealing method and extracting sample data corresponding to the predicted target characteristic by inputting a dataset into the prediction model with the varied uncertainty parameter value, thereby increasing the probability of exploring new materials.
[0015] FIG. 1 is a drawing for explaining a novel material exploration device according to one embodiment of the present disclosure.
[0016] FIG. 2 is a drawing for explaining the operation of a novel material exploration device according to an embodiment of the present disclosure.
[0017] FIGS. 3 to 6 are drawings for explaining a prediction model of a novel material exploration device according to an embodiment of the present disclosure.
[0018] FIGS. 7 to 11 are diagrams for explaining performance evaluation of a prediction model of a novel material exploration device according to an embodiment of the present disclosure.
[0019] FIG. 12 is a drawing for explaining a novel material search method of a novel material search device according to an embodiment of the present disclosure.
[0020] Hereinafter, embodiments disclosed in this specification will be described in detail with reference to the attached drawings. Regardless of the drawing numbers, identical or similar components will be given the same reference numbers and redundant descriptions thereof will be omitted. The suffixes "module" and "part" used for components in the following description are assigned or used interchangeably only for the convenience of writing the specification, and do not in themselves have distinct meanings or roles. In addition, when describing the embodiments disclosed in this specification, if it is determined that a specific description of a related known technology may obscure the gist of the embodiments disclosed in this specification, a detailed description thereof will be omitted. In addition, the attached drawings are only intended to facilitate easy understanding of the embodiments disclosed in this specification, and the technical ideas disclosed in this specification are not limited by the attached drawings, and should be understood to include all modifications, equivalents, and substitutes included in the spirit and technical scope of the present disclosure.
[0021] Terms that include ordinal numbers, such as first, second, etc., may be used to describe various components, but the components are not limited by these terms. These terms are used solely to distinguish one component from another.
[0022] When a component is referred to as being "connected" or "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but that there may be other components intervening. Conversely, when a component is referred to as being "directly connected" or "connected" to another component, it should be understood that there are no other components intervening.
[0023] FIG. 1 is a drawing for explaining a novel material exploration device according to one embodiment of the present disclosure.
[0024] The new material search device (100) can be implemented as a fixed device or a movable device, such as a TV, a projector, a mobile phone, a smart phone, a desktop computer, a laptop, a digital broadcasting terminal, a PDA (personal digital assistant), a PMP (portable multimedia player), a navigation device, a tablet PC, a wearable device, a set-top box (STB), a DMB receiver, a radio, a washing machine, a refrigerator, a desktop computer, digital signage, a robot, a vehicle, etc.
[0025] Referring to FIG. 1, a new material exploration device (100) may include a communication unit (110), an input unit (120), a learning processor (130), a sensing unit (140), an output unit (150), a memory (170), and a processor (180).
[0026] The communication unit (110) can transmit and receive data with external devices such as other AI devices or AI servers using wired or wireless communication technology. For example, the communication unit (110) can transmit and receive sensor information, user input, learning models, control signals, etc. with external devices.
[0027] At this time, the communication technologies used by the communication unit (110) include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc.
[0028] The input unit (120) can obtain various types of data.
[0029] At this time, the input unit (120) may include a camera for inputting a video signal, a microphone for receiving an audio signal, a user input unit for receiving information from a user, etc. Here, the camera or microphone may be treated as a sensor, and a signal obtained from the camera or microphone may be referred to as sensing data or sensor information.
[0030] The input unit (120) can obtain input data to be used when obtaining output using learning data and learning models for model learning. The input unit (120) can also obtain unprocessed input data, in which case the processor (180) or learning processor (130) can extract input features as preprocessing for the input data.
[0031] The learning processor (130) can train a model composed of an artificial neural network using learning data. Here, the trained artificial neural network may be referred to as a learning model. The learning model can be used to infer result values for new input data other than the learning data, and the inferred values can be used as a basis for making decisions regarding certain actions.
[0032] At this time, the running processor (130) can perform AI processing together with the running processor of the AI server.
[0033] At this time, the running processor (130) may include a memory integrated or implemented in the new material exploration device. Alternatively, the running processor (130) may be implemented using a memory (170), an external memory directly coupled to the new material exploration device (100), or a memory maintained in an external device.
[0034] The sensing unit (140) can obtain at least one of internal information of the new material exploration device (100), environmental information surrounding the new material exploration device (100), and user information using various sensors.
[0035] At this time, the sensors included in the sensing unit (140) include a proximity sensor, a light sensor, an acceleration sensor, a magnetic sensor, a gyro sensor, an inertial sensor, an RGB sensor, an IR sensor, a fingerprint recognition sensor, an ultrasonic sensor, a light sensor, a microphone, a lidar, a radar, etc.
[0036] The output unit (150) can generate output related to vision, hearing, or touch.
[0037] At this time, the output unit (150) may include a display unit that outputs visual information, a speaker that outputs auditory information, a haptic module that outputs tactile information, etc.
[0038] The memory (170) can store data that supports various functions of the new material exploration device (100). For example, the memory (170) can store input data, learning data, learning models, learning history, etc. acquired from the input unit (120).
[0039] The processor (180) may include a quantum processor (QPU) that executes a multidimensional quantum algorithm using qubits, and may also include simulated annealing, simulated quantum annealing, or CMOS annealing that implements operations similar to those of the QPU using non-quantum technology.
[0040] The processor (180) may determine at least one executable operation of the new material exploration device (100) based on information determined or generated using a data analysis algorithm or a machine learning algorithm. Then, the processor (180) may control components of the new material exploration device (100) to perform the determined operation.
[0041] To this end, the processor (180) can request, search, receive or utilize data from the running processor (130) or memory (170), and control components of the new material search device (100) to execute at least one of the executable operations, a predicted operation or an operation determined to be desirable.
[0042] At this time, if connection of an external device is required to perform a determined operation, the processor (180) can generate a control signal for controlling the external device and transmit the generated control signal to the external device.
[0043] The processor (180) can obtain intent information for user input and determine the user's requirement based on the obtained intent information.
[0044] At this time, the processor (180) can obtain intent information corresponding to the user input by using at least one of an STT (Speech To Text) engine for converting voice input into a string or a natural language processing (NLP) engine for obtaining intent information of natural language.
[0045] At this time, at least one of the STT engine or the NLP engine may be configured with an artificial neural network, at least in part, trained according to a machine learning algorithm. Furthermore, at least one of the STT engine or the NLP engine may be trained by the learning processor (130), by the learning processor of the AI server, or by distributed processing thereof.
[0046] The processor (180) can collect history information, including the operation details of the new material exploration device (100) or user feedback on the operation, and store the information in the memory (170) or the learning processor (130), or transmit the information to an external device such as an AI server. The collected history information can be used to update the learning model.
[0047] The processor (180) can control at least some of the components of the new material search device (100) to drive an application program stored in the memory (170). Furthermore, the processor (180) can operate two or more of the components included in the new material search device (100) in combination to drive the application program.
[0048] FIG. 2 is a drawing for explaining the operation of a novel material exploration device according to an embodiment of the present disclosure.
[0049] As illustrated in FIG. 2, the novel material exploration device (100) of the present disclosure may include a database (500) that stores datasets of chemical materials, and a processor (180) that explores a target material from the database (500).
[0050] Here, the database (500) may include datasets corresponding to molecular structures of chemical materials.
[0051] For example, a dataset may include information about molecules having at least one atom and having at least one of geometric, thermodynamic, and electronic properties.
[0052] And, the processor (180) inputs a data set into a prediction model (190) including an uncertainty parameter to extract sample data corresponding to a prediction target characteristic, varies the uncertainty parameter value of the prediction model (190) based on the extracted sample data, inputs the data set into the prediction model (190) with the varied uncertainty parameter value to additionally extract sample data corresponding to the prediction target characteristic, and can search for a target material from the extracted sample data.
[0053] Here, the processor (180) can estimate uncertainty through a prediction model (190) that produces a variant covariant matrix based on the regression equation of the linear regression model.
[0054] That is, the processor (180) can estimate uncertainty through a prediction model (190) that calculates regression coefficients that minimize RMSE (Root Mean Square Error) based on the regression equation of a linear regression model and maximizes likelihood by applying a Gaussian distribution to the error of the regression equation.
[0055] For example, the processor (180) can calculate a regression coefficient that minimizes RMSE based on the following mathematical expressions 1 to 5 through a prediction model (190).
[0056]
[0057] Here, Y is the predicted value, a is the coefficient, b is the intercept, and ε is the error.
[0058]
[0059]
[0060]
[0061] Here, is X i is the average of, is Y i is the average.
[0062]
[0063] Here, is X i is the average of, is Y i is the average.
[0064] That is, the processor (180) calculates X through the equation of mathematical formula 2 in the regression equation of the linear regression model consisting of mathematical formula 1. i The value can be predicted, the sum of squared errors can be minimized using the formula in Equation 3, and the regression coefficient b can be calculated using the formulas in Equation 4 and Equation 5.
[0065] As another example, the processor (180) can maximize the likelihood by applying a Gaussian distribution to the error of the regression equation based on the following mathematical equations 6 to 9 through the prediction model (190).
[0066]
[0067]
[0068] Here, the mean is μ = a + bx, and the variance is σ.
[0069]
[0070]
[0071] That is, the processor (180) can estimate the error ε of the linear regression model of the equation 1 by applying the Gaussian distribution of the equation 6 to the error ε of the equation 7, and can estimate the uncertainty by maximizing the likelihood through the equation 9 by taking the log of the likelihood equation of the equation 8.
[0072] Here, the processor (180) can calculate a regression coefficient b that is identical to the regression coefficient that minimizes RMSE through the equation formed by Equation 4 and the equation formed by Equation 5 when maximizing the likelihood.
[0073] Additionally, the processor (180) may determine the coefficient a based on mathematical expressions 10 to 13 through the prediction model (190), and may include an uncertainty parameter σ.
[0074]
[0075] Here, n is the number of data in the dataset.
[0076]
[0077]
[0078]
[0079] That is, the processor (180) minimizes the loss function through the equation of mathematical formula 10 so that the predicted value aX approximates the actual cost function value y in the regression equation Y = aX of the linear regression model, calculates the transformed covariance matrix of mathematical formula 12 through the equation of mathematical formula 11, determines the coefficient a through the equation of mathematical formula 13, and may include the uncertainty parameter σ.
[0080] Next, the processor (180) can convert the molecular structure corresponding to the datasets of chemical materials into a fingerprint by encoding it into binary before extracting the sample data.
[0081] Here, the processor (180) can convert the molecular structure corresponding to each dataset into a fingerprint by encoding it into a series of binary numbers indicating the presence or absence of a substructure within the molecule.
[0082] For example, the processor (180) can convert all datasets stored in the database (500) into fingerprints.
[0083] Next, the processor (180) can pre-train a prediction model (190) to predict data characteristics corresponding to the characteristic conditions based on fingerprints of training data and test data when characteristic conditions of a chemical material to be explored are input before extracting sample data.
[0084] Here, the characteristic conditions of the chemical material may include the target characteristic of the chemical material to be explored and the target value of the target characteristic.
[0085] And, when extracting sample data, the processor (180) can select at least one uncertainty parameter value and input a data set into a prediction model (190) including the selected uncertainty parameter value to extract sample data corresponding to the prediction target characteristic.
[0086] Here, the processor (180) may include one selected uncertainty parameter value in the prediction model (190) if there is one selected uncertainty parameter value, and may sequentially include multiple selected uncertainty parameter values in the prediction model (190) if there are multiple different selected uncertainty parameter values.
[0087] For example, when sequentially including a plurality of selected uncertainty parameter values, the processor (180) may sequentially include the plurality of selected uncertainty parameter values in the prediction model (190) from the lowest value to the highest value.
[0088] For example, the processor (180) selects a plurality of uncertainty parameter values such that σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 If we include the uncertainty parameter value σ as 0, 4 × 10 -3 , 8 × 10 -3 , 12 × 10 -3 can be included in the prediction model (190) in the order of .
[0089] As another example, when sequentially including a plurality of selected uncertainty parameter values, the processor (180) may sequentially include the plurality of selected uncertainty parameter values in the prediction model (190) from the highest value to the lowest value.
[0090] For example, the processor (180) selects a plurality of uncertainty parameter values such that σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 If we include the uncertainty parameter value σ, we get 12 × 10 -3 , 8 × 10 -3 , 4 × 10 -3 , can be included in the prediction model (190) in the order of 0.
[0091] As another example, when sequentially including a plurality of selected uncertainty parameter values, the processor (180) may first include the lowest value among the plurality of selected uncertainty parameter values in the prediction model if the search condition of the target material is search for a material having a specific characteristic, and may first include the highest value among the plurality of selected uncertainty parameter values in the prediction model if the search condition of the target material is search for a material having a new characteristic.
[0092] For example, when the search condition of the target material is to search for a material having specific characteristics, the processor (180) selects a plurality of uncertainty parameter values such that σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 Among the uncertainty parameter values, σ = 0 is included first in the prediction model (190), and when the search condition for the target material is to search for a material with new characteristics, the selected multiple uncertainty parameter values are σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 Among the uncertainty parameter values, σ = 12 × 10 -3 can be included first in the prediction model (190).
[0093] As another example, the processor (180) may include one selected uncertainty parameter value in one prediction model if there is one selected uncertainty parameter value, and may include multiple selected uncertainty parameter values in multiple prediction models in a one-to-one correspondence at the same time if there are multiple different selected uncertainty parameter values.
[0094] Here, the processor (180) can generate the same number of prediction models so that, if there are multiple different selected uncertainty parameter values, the multiple selected uncertainty parameter values are included in a one-to-one correspondence.
[0095] For example, the processor (180) selects a plurality of uncertainty parameter values such that σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 The first prediction model includes uncertainty parameter value σ = 0, uncertainty parameter value σ = 4 × 10 -3 A second prediction model with uncertainty parameter value σ = 8 × 10 -3 The third prediction model, which includes uncertainty parameter value σ = 12 × 10 -3 A fourth prediction model can be created that includes .
[0096] Next, when extracting sample data, the processor (180) checks whether the number of sample data to be extracted is preset, and if the number of sample data is set, the processor can extract the preset number of sample data based on the cost function of the prediction model (190).
[0097] Here, when the processor (180) checks whether the number of sample data is preset, if the number of sample data is not preset, all sample data generated based on the cost function of the prediction model can be extracted.
[0098] In some cases, when checking whether the number of sample data is preset, the processor (180) may request a user input corresponding to the sample data number setting if the number of sample data is not preset, and when a user input corresponding to the sample data number setting is received, the processor may extract sample data in the set number corresponding to the user input.
[0099] Here, the processor (180) can extract all sample data generated based on the optimized cost function if a user input corresponding to the sample data number setting is not received within a predetermined time.
[0100] Next, when varying the uncertainty parameter value of the prediction model (190), the processor (180) can check the first uncertainty parameter value included in the prediction model (190), determine a second uncertainty parameter value that is greater than or less than the first uncertainty parameter value based on the extracted sample data, and vary the uncertainty parameter value of the prediction model (190) based on the determined second uncertainty parameter value.
[0101] Here, when determining the second uncertainty parameter value, the processor (180) may determine a second uncertainty parameter value that is greater than the first uncertainty parameter value if the extracted sample data is data that is concentrated on a specific characteristic, and may determine a second uncertainty parameter value that is greater or smaller than the first uncertainty parameter value depending on the undistributed area of the data if the extracted sample data is data that has a new characteristic.
[0102] For example, if there are multiple first uncertainty parameter values included in the prediction model (190), the processor (180) can determine multiple second uncertainty parameter values that do not overlap with the multiple first uncertainty parameter values.
[0103] Additionally, when determining a plurality of second uncertainty parameter values that do not overlap with a plurality of first uncertainty parameter values, the processor (180) may determine the second uncertainty parameter values in the same number as the first uncertainty parameter values included in the prediction model (190).
[0104] And, when searching for a target material, the processor (180) can evaluate feature importance from extracted sample data, select upper-level features based on the feature importance, and search for a target material based on the selected upper-level features.
[0105] Here, when evaluating feature importance, the processor (180) can evaluate the feature importance of each sample data from the frequency obtained from the sample data.
[0106] Additionally, the processor (180) can sequentially list features in order of high feature importance when the feature importance of each sample data is evaluated.
[0107] In addition, when selecting upper-level features, the processor (180) can check whether a reference value for feature selection has been preset, and if the reference value for feature selection has been preset, select upper-level features having a feature importance higher than the reference value based on the preset reference value.
[0108] Here, if a reference value for feature selection is not set, the processor (180) can select a preset number of features belonging to a higher level from among features arranged in order of high feature importance.
[0109] For example, the processor (180) can select from a first-order level feature with the highest feature importance to a specific order level feature corresponding to a preset number.
[0110] Additionally, the processor (180) can search for target materials that optimize new substituent combinations attached to a specific base structure based on selected high-level features when searching for target materials.
[0111] FIGS. 3 to 6 are drawings for explaining a prediction model of a novel material exploration device according to an embodiment of the present disclosure.
[0112] As illustrated in FIGS. 3 to 6, the prediction model of the present disclosure can estimate uncertainty by calculating a variant covariant matrix based on the regression equation of the linear regression model.
[0113] Here, the prediction model can estimate uncertainty by calculating regression coefficients that minimize RMSE (Root Mean Square Error) based on the regression equation of the linear regression model and applying a Gaussian distribution to the error of the regression equation to maximize the likelihood.
[0114] As shown in Fig. 3, the prediction model is a linear regression model consisting of Equation 1, and X is calculated using Equation 2. i The value can be predicted, the sum of squared errors can be minimized using the formula in Equation 3, and the regression coefficient b can be calculated using the formulas in Equation 4 and Equation 5.
[0115] And, as shown in Fig. 4, the prediction model can maximize the likelihood by applying a Gaussian distribution to the error of the regression equation based on mathematical equations 6 to 9.
[0116] That is, the prediction model estimates the error ε of Equation 7 by applying the Gaussian distribution of Equation 6 to the error ε in the regression equation of the linear regression model of Equation 1, and estimates the uncertainty by maximizing the likelihood through the equation of Equation 9 by taking the logarithm of the likelihood equation of Equation 8.
[0117] Here, the prediction model can produce a regression coefficient b that is identical to the regression coefficient that minimizes RMSE through the equations formed by Equation 4 and Equation 5 when maximizing the likelihood.
[0118] Next, as shown in FIG. 5, the prediction model may determine the coefficient a based on mathematical expressions 10 to 13 and include an uncertainty parameter σ.
[0119] That is, the prediction model minimizes the loss function through the equation of mathematical formula 10 so that the predicted value aX approximates the actual cost function value y in the regression equation of the linear regression model Y = aX, calculates the transformed covariance matrix of mathematical formula 12 through the equation of mathematical formula 11, determines the coefficient a through the equation of mathematical formula 13, and may include the uncertainty parameter σ.
[0120] Next, the prediction model can repeatedly extract sample data by varying the uncertainty parameter value and search for target materials from the extracted sample data.
[0121] As shown in Fig. 6, the prediction model can search for target materials that optimize new substituent combinations attached to a specific base structure based on high-level features among sample data.
[0122] Here, the prediction model may have a higher probability of extracting sample data concentrated on a specific substituent than of extracting sample data containing a new substituent when the uncertainty parameter value is low, and may have a higher probability of extracting sample data containing a new substituent than of extracting sample data concentrated on a specific substituent when the uncertainty parameter value is high.
[0123] FIGS. 7 to 11 are diagrams for explaining performance evaluation of a prediction model of a novel material exploration device according to an embodiment of the present disclosure.
[0124] Figure 7 shows the coefficient of determination R according to the loop number for each sample data of a substituent corresponding to different uncertainty parameter values. 2 This is a graph showing the changes.
[0125] As shown in Fig. 7, the prediction model of the present disclosure is tested with about 1000 data and 4 uncertainty parameter values σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 , σ = 12 × 10 -3 As a result of performance evaluation using uncertainty parameter values σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 When applied to this prediction model, the coefficient of determination of the prediction model was found to operate at approximately 0.1 or higher, and the uncertainty parameter value σ = 12 × 10 -3 When applied to this prediction model, it can be seen that the coefficient of determination of the prediction model operates at approximately 0.1 or less.
[0126] Additionally, it can be seen that the prediction model of the present disclosure has an increasing coefficient of determination as the number of loops increases.
[0127] Therefore, the prediction model of the present disclosure can effectively search for chemical materials by expanding to the entire global region rather than being limited to a local region of the chemical space, with improved performance in new material search.
[0128] Figures 8a and 8b are histograms showing sample data of substituents extracted according to uncertainty parameter values.
[0129] Figure 8a is a histogram showing sample data R3 and R4 extracted by applying uncertainty parameter value σ = 0 to the prediction model, and Figure 8b is a histogram showing sample data R3 and R4 extracted by applying uncertainty parameter value σ = 4 × 10 -3 This is a histogram showing the sample data R3 and R4 extracted by applying the prediction model.
[0130] As shown in Fig. 8a, when the uncertainty parameter value σ = 0 is applied to the prediction model, it can be seen that sample data concentrated on a specific substituent is extracted.
[0131] And, as shown in Fig. 8b, the uncertainty parameter value σ = 4 × 10 -3 When applied to the prediction model, it can be seen that various sample data containing new substituents are extracted.
[0132] In this way, the prediction model of the present disclosure can extract various sample data including new substituents by expanding to the entire global region rather than being limited to a local region of the chemical space through changes in the uncertainty parameter values.
[0133] Here, various sample data containing the extracted new substituents can be used to strengthen the database dataset and update the prediction model.
[0134] Figure 9 shows sample data R3 and R4 extracted through approximately 20 loops by applying uncertainty parameter value σ = 0 to the prediction model and uncertainty parameter value σ = 4 × 10 -3 Histogram showing sample data R3 and R4 extracted through approximately 20 loops by applying to the prediction model.
[0135] As shown in Fig. 9, when the uncertainty parameter value σ = 0 is applied to the prediction model, it can be seen that sample data concentrated on a specific substituent are extracted, and sample data corresponding to new substituents to be added to the database are hardly extracted.
[0136] In contrast, the uncertainty parameter value σ = 4 × 10 -3 When applied to a prediction model, it can be seen that various sample data containing new substituents to be added to the database are extracted.
[0137] Figure 10 is a graph showing the change in material properties according to the loop number for each sample data of a substituent corresponding to the uncertainty parameter value.
[0138] As shown in Figure 10, the uncertainty parameter values σ = 0, σ = 4 × 10 -3 , σ = 8 × 10 -3 When applied to this prediction model, we can see that the average attribute y of the top 10 sample data increases as the number of loops increases.
[0139] This is because as the number of loops increases, the extracted data strengthens the database, which in turn updates the prediction model and improves the exploration accuracy.
[0140] And, the uncertainty parameter value σ = 12 × 10 -3 When applied to this prediction model, the average attribute y of the top 10 sample data does not change significantly as the number of loops increases, but by exploring various substituents, the probability of finding materials with high attributes may be lower than when other uncertainty parameter values are applied.
[0141] In this way, the prediction model of the present disclosure can effectively explore new materials with high properties.
[0142] Figure 11 is a diagram showing the results of a search for new materials with high properties for each sample data of a substituent corresponding to the uncertainty parameter value.
[0143] As shown in Fig. 11, the prediction model with the uncertainty parameter value σ = 0 can obtain about 25 data sets of new substituent combinations with high attributes, such as y values of 88% or more, and the uncertainty parameter value σ = 4 × 10 -3 This applied prediction model can obtain about 11 data sets of new substituent combinations with high properties, with y values of 88% or more, and the uncertainty parameter value σ = 8 × 10 -3This applied prediction model can obtain about 6 data sets of new substituent combinations with high properties, such as y values of 88% or more, and the uncertainty parameter value σ = 12 × 10 -3 This applied prediction model can obtain about 1 data set of new substituent combinations with high attributes, such as y-values of 88% or higher.
[0144] In this way, the present disclosure can effectively explore new materials with high properties by changing the uncertainty parameter value.
[0145] FIG. 12 is a drawing for explaining a novel material search method of a novel material search device according to an embodiment of the present disclosure.
[0146] As illustrated in FIG. 12, the present disclosure can input a dataset of chemical materials into a prediction model including uncertainty parameters to extract sample data corresponding to predicted target characteristics (S10).
[0147] Here, the present disclosure can convert molecular structures corresponding to datasets of chemical materials into fingerprints by encoding them in binary before extracting sample data.
[0148] In addition, the present disclosure can pre-train a prediction model to predict data characteristics corresponding to the characteristic conditions based on fingerprints of training data and test data when characteristic conditions of a chemical material to be explored are input before extracting sample data.
[0149] The present disclosure can extract sample data corresponding to a prediction target characteristic by selecting at least one uncertainty parameter value and inputting a dataset into a prediction model including the selected uncertainty parameter value.
[0150] For example, the present disclosure may include one selected uncertainty parameter value in the prediction model if there is one selected uncertainty parameter value, and may sequentially include multiple selected uncertainty parameter values in the prediction model if there are multiple different selected uncertainty parameter values.
[0151] In some cases, when the present disclosure sequentially includes uncertainty parameter values, if the search condition for the target material is to search for a material having a specific characteristic, the lowest value among the selected plurality of uncertainty parameter values may be included first in the prediction model, and if the search condition for the target material is to search for a material having a new characteristic, the highest value among the selected plurality of uncertainty parameter values may be included first in the prediction model.
[0152] As another example, the present disclosure may include one selected uncertainty parameter value in one prediction model if there is one selected uncertainty parameter value, and may include multiple selected uncertainty parameter values in multiple prediction models in a one-to-one correspondence at the same time.
[0153] Here, the present disclosure can generate the same number of prediction models so that the plurality of selected uncertainty parameter values are included in a one-to-one correspondence when the plurality of selected uncertainty parameter values are different from each other.
[0154] In addition, the present disclosure can check whether the number of sample data to be extracted is preset, and if the number of sample data is set, can extract the preset number of sample data based on the cost function of the prediction model.
[0155] And, the present disclosure can vary the uncertainty parameter value of the prediction model based on the extracted sample data (S20).
[0156] Here, the present disclosure can verify a first uncertainty parameter value included in a prediction model, determine a second uncertainty parameter value that is greater than or less than the first uncertainty parameter value based on extracted sample data, and vary the uncertainty parameter value of the prediction model based on the determined second uncertainty parameter value.
[0157] In the present disclosure, when determining the second uncertainty parameter value, if the extracted sample data is data concentrated on a specific characteristic, a second uncertainty parameter value greater than the first uncertainty parameter value can be determined, and if the extracted sample data is data having a new characteristic, a second uncertainty parameter value greater than or less than the first uncertainty parameter value can be determined depending on the non-distribution area of the data.
[0158] For example, the present disclosure can determine a plurality of second uncertainty parameter values that do not overlap with the plurality of first uncertainty parameter values when there are a plurality of first uncertainty parameter values included in the prediction model.
[0159] Here, the present disclosure can determine the second uncertainty parameter values in the same number as the first uncertainty parameter values included in the prediction model.
[0160] Next, the present disclosure can additionally extract sample data corresponding to the prediction target characteristic by inputting a dataset into a prediction model with a variable uncertainty parameter value (S30).
[0161] Next, the present disclosure can search for target materials from extracted sample data (S40).
[0162] Here, the present disclosure can evaluate feature importance from extracted sample data, select upper-level features based on the feature importance, and search for target materials based on the selected upper-level features.
[0163] For example, the present disclosure can evaluate the feature importance of each sample data from the frequency obtained from the sample data.
[0164] And, in the present disclosure, when the feature importance of each sample data is evaluated, the features can be sequentially listed in order of high feature importance levels.
[0165] Additionally, the present disclosure can explore target materials that optimize new substituent combinations attached to a specific base structure based on selected high-level features.
[0166] In this way, the present disclosure can effectively explore chemical materials by expanding the entire global region rather than being limited to some local regions in the chemical space by varying the uncertainty parameter value of a prediction model using a quantum annealing method and inputting a dataset into the prediction model with the varied uncertainty parameter value to extract sample data corresponding to the prediction target characteristic, thereby increasing the probability of exploring new materials.
[0167] The above-described present disclosure can be implemented as computer-readable code on a program-recorded medium. The computer-readable medium includes all types of recording devices that store data that can be read by a computer system. Examples of computer-readable media include hard disk drives (HDDs), solid-state disk drives (SSDs), silicon disk drives (SDDs), read-only memory (ROM), random access memory (RAM), CD-ROMs, magnetic tapes, floppy disks, and optical data storage devices. In addition, the computer may include a processor (180) of an artificial intelligence device.
[0168] According to the novel material exploration device according to the present disclosure, by varying the uncertainty parameter value of a prediction model using a quantum annealing method, inputting a dataset into the prediction model with the varied uncertainty parameter value, and extracting sample data corresponding to the predicted target characteristic, the probability of exploring new materials can be increased, and thus the possibility of industrial application is remarkable.
Claims
1. A database that stores datasets of chemical materials; and, A processor for searching for a target material from the database, The above processor, A novel material exploration device characterized in that the data set is input into a prediction model including an uncertainty parameter to extract sample data corresponding to a prediction target characteristic, the uncertainty parameter value of the prediction model is varied based on the extracted sample data, the data set is input into the prediction model with the varied uncertainty parameter value to additionally extract sample data corresponding to the prediction target characteristic, and the target material is explored from the extracted sample data.
2. In paragraph 1, The above prediction model is, A novel material exploration device characterized by estimating uncertainty by calculating a variant covariant matrix based on the regression equation of a linear regression model.
3. In paragraph 2, The above prediction model is, A novel material exploration device characterized in that it calculates a regression coefficient that minimizes RMSE (Root Mean Square Error) based on the regression equation of the linear regression model, and estimates uncertainty by applying a Gaussian distribution to the error of the regression equation to maximize the likelihood.
4. In paragraph 1, The above processor, A novel material exploration device characterized in that, when extracting the sample data, at least one uncertainty parameter value is selected, and the dataset is input into a prediction model including the selected uncertainty parameter value to extract sample data corresponding to the prediction target characteristic.
5. In paragraph 4, The above processor, If the above-mentioned selected uncertainty parameter value is 1, the above-mentioned selected uncertainty parameter value is included in the above-mentioned prediction model, A novel material exploration device characterized in that, if the above-mentioned selected uncertainty parameter values are different from each other, the above-mentioned selected uncertainty parameter values are sequentially included in the prediction model.
6. In paragraph 5, The above processor, When sequentially including the above-selected multiple uncertainty parameter values, if the search condition for the target material is to search for a material with specific characteristics, the lowest value among the above-selected multiple uncertainty parameter values is included first in the prediction model, A novel material exploration device characterized in that, if the search condition for the above target material is to search for a material having new characteristics, the highest value among the selected plurality of uncertainty parameter values is first included in the prediction model.
7. In paragraph 4, The above processor, If the above-mentioned selected uncertainty parameter value is 1, then 1 prediction model includes the above-mentioned selected uncertainty parameter value, A novel material exploration device characterized in that, if the above-mentioned selected uncertainty parameter values are different from each other, the selected uncertainty parameter values are simultaneously and collectively included in multiple prediction models in a one-to-one correspondence.
8. In paragraph 7, The above processor, A novel material exploration device characterized in that, if the above-mentioned plurality of different uncertainty parameter values are selected, an equal number of prediction models are generated so that the selected plurality of uncertainty parameter values are included in a one-to-one correspondence.
9. In paragraph 1, The above processor, A novel material exploration device characterized in that, when extracting the above sample data, it is checked whether the number of sample data to be extracted is preset, and if the number of sample data is preset, the preset number of sample data is extracted based on the cost function of the prediction model.
10. In paragraph 1, The above processor, A novel material exploration device characterized in that when varying the uncertainty parameter value of the prediction model, a first uncertainty parameter value included in the prediction model is checked, a second uncertainty parameter value greater than or less than the first uncertainty parameter value is determined based on the extracted sample data, and the uncertainty parameter value of the prediction model is varied based on the determined second uncertainty parameter value.
11. In paragraph 10, The above processor, When determining the second uncertainty parameter value, if the extracted sample data is data that is concentrated on a specific characteristic, a second uncertainty parameter value greater than the first uncertainty parameter value is determined, A novel material exploration device characterized in that, if the extracted sample data is data having new characteristics, a second uncertainty parameter value that is greater or smaller than the first uncertainty parameter value is determined according to a non-distribution area of the data.
12. In paragraph 10, The above processor, A novel material exploration device characterized in that, when there are multiple first uncertainty parameter values included in the above prediction model, multiple second uncertainty parameter values that do not overlap with the multiple first uncertainty parameter values are determined.
13. In paragraph 1, The above processor, A novel material exploration device characterized in that, when exploring the target material, feature importance is evaluated from the extracted sample data, upper-level features are selected based on the feature importance, and the target material is explored based on the selected upper-level features.
14. In paragraph 13, The above processor, A novel material exploration device characterized in that, when exploring the above target material, it explores the target material that optimizes a new combination of substituents attached to a specific base structure based on the above-selected upper-level features.
15. A step of inputting a dataset of chemical materials into a prediction model including uncertainty parameters to extract sample data corresponding to the prediction target characteristic; A step of varying the uncertainty parameter value of the prediction model based on the extracted sample data; A step of inputting the dataset into a prediction model in which the uncertainty parameter value is varied to additionally extract sample data corresponding to the prediction target characteristic; and A novel material exploration method, characterized by including a step of exploring the target material from the extracted sample data.
Citation Information
Patent Citations
Apparatus And Method For Generating Immersive Advertisement Contents Through One-Shot Face Swap
KR1020240081602A
Stochastic Model-Predictive Control of Uncertain System
US20220187793A1
Cited By
Model training method and related equipment
CN120632469A