Physical property prediction method, information processing device, and program

The method generates descriptors for mixtures by multiplying compound descriptors by composition ratios and using learning models to predict properties, addressing the inadequacies of existing technologies and enhancing prediction accuracy for liquid crystal materials.

JP7798221B1Active Publication Date: 2026-01-14DIC CORP
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2025068948
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2026-01-14
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing property prediction technologies using machine learning do not adequately consider compositions containing multiple compounds, such as liquid crystal materials, leading to insufficient prediction of physical properties.

Method used

A method that generates descriptors for mixtures by multiplying each element of a compound's descriptor by its composition ratio and summing the results, using Morgan Fingerprint and Mordred as physical property descriptors, and inputs these into a learning model to predict physical properties.

Benefits of technology

Improves the accuracy of property prediction for mixtures by reflecting intermolecular bonding information, reducing experimental costs, and enabling consistent predictions for varying compositions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007798221000001_ABST
    Figure 0007798221000001_ABST
Patent Text Reader

Abstract

A physical property prediction method, an information processing device, and a program are provided that improve physical property prediction technology using machine learning. [Solution] A physical property prediction method executed by an information processing device, the information processing device includes generating a descriptor for a mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and adding up the results, inputting the descriptor for the mixture into a learning model as an explanatory variable, predicting the physical properties of the mixture based on the output from the learning model, and outputting the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a property prediction method, an information processing device, and a program. [Background technology]

[0002] Conventionally, molecular structures have been converted into vector data, and the converted information has been used to predict physical properties based on machine learning models. Furthermore, studies have been conducted on appropriately expressing the molecular structures of synthetic compounds, which are synthesized from multiple types of structural units, in data (e.g., Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-81920 Summary of the Invention [Problem to be solved by the invention]

[0004] Although the technology of Patent Document 1 considers appropriately representing the molecular structure of a synthetic compound synthesized from multiple types of structural units in data, it does not consider generating descriptors corresponding to compositions containing multiple compounds, such as liquid crystal materials. Therefore, consideration of predicting physical properties of compositions containing multiple compounds is insufficient. As such, there is room for improvement in property prediction technologies using machine learning.

[0005] In view of the above circumstances, an object of the present disclosure is to improve property prediction techniques using machine learning. [Means for solving the problem]

[0006] (1) A physical property prediction method according to an embodiment of the present disclosure includes: A physical property prediction method executed by an information processing device, comprising: generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors of the mixture into a learning model as explanatory variables; predicting physical properties of the mixture based on output from the learning model; and Includes.

[0007] (2) A physical property prediction method according to an embodiment of the present disclosure includes: The physical property prediction method according to (1), The descriptors include physical property descriptors.

[0008] (3) A physical property prediction method according to an embodiment of the present disclosure includes: (2) The physical property prediction method according to The physical property descriptor is either Mordred or RDKIT.

[0009] (4) A physical property prediction method according to an embodiment of the present disclosure includes: The physical property prediction method according to (2) or (3), The mixture is a liquid crystal, and the physical property descriptors include elements related to HSP.

[0010] (5) A physical property prediction method according to an embodiment of the present disclosure includes: The physical property prediction method according to any one of (1) to (4), The descriptors include structural descriptors.

[0011] (6) A physical property prediction method according to an embodiment of the present disclosure includes: (5) A physical property prediction method according to The structural descriptor is the Morgan Fingerprint.

[0012] (7) A physical property prediction method according to an embodiment of the present disclosure includes: (6) A physical property prediction method according to the present invention, Each element of the Morgan Fingerprint is a count-based value.

[0013] (8) A physical property prediction method according to an embodiment of the present disclosure includes: The physical property prediction method according to (6) or (7), The radius of the Morgan Fingerprint is 3 or greater.

[0014] (9) An information processing device according to an embodiment of the present disclosure includes: An information processing device including a control unit, The control unit generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors related to the mixture into a learning model as explanatory variables; Based on the output from the learning model, the physical properties of the mixture are predicted.

[0015] (10) A program according to an embodiment of the present disclosure includes: On the computer, generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors of the mixture into a learning model as explanatory variables; predicting physical properties of the mixture based on output from the learning model; and Execute the following. [Effects of the Invention]

[0016] According to one embodiment of the present disclosure, a property prediction technique using machine learning is improved. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a block diagram illustrating a schematic configuration of an information processing device according to an embodiment of the present disclosure. [Figure 2] 1 is a schematic diagram of the count-based Morgan Fingerprint. [Figure 3]FIG. 1 illustrates the radius of the Morgan Fingerprint. [Figure 4] FIG. 1 shows the Morgan Fingerprint of ibuprofen for a radius of 0. [Figure 5A] FIG. 1 shows the Morgan Fingerprint of ibuprofen for a radius of 1. [Figure 5B] FIG. 1 shows the Morgan Fingerprint of ibuprofen for a radius of 2. [Figure 5C] FIG. 1 shows the Morgan Fingerprint of ibuprofen for a radius of 3. [Figure 6] FIG. 1 shows the Morgan Fingerprint of C16H18FN when the radius is 0. [Figure 7A] FIG. 1 shows the Morgan Fingerprint of C16H18FN for a radius of 1. [Figure 7B] FIG. 1 shows the Morgan Fingerprint of C16H18FN for a radius of 2. [Figure 7C] FIG. 1 shows the Morgan Fingerprint of C16H18FN for a radius of 3. [Figure 8] FIG. 1 is a conceptual diagram illustrating a method for generating a descriptor for a mixture according to an embodiment of the present disclosure. [Figure 9] FIG. 1 is a conceptual diagram illustrating an example of a descriptor for a mixture according to an embodiment of the present disclosure. [Figure 10] 10 is a flowchart illustrating an operation of an information processing device according to an embodiment of the present disclosure. [Figure 11] 10 is a graph showing prediction results of a property prediction method according to a comparative example. [Figure 12] 1 is a graph showing prediction results of a property prediction method according to an embodiment of the present disclosure. [Figure 13] 1 is a graph showing prediction results of a property prediction method according to an embodiment of the present disclosure. [Figure 14] 1 is a graph showing prediction results of a property prediction method according to an embodiment of the present disclosure. [Figure 15] 1 is a graph showing prediction results of a property prediction method according to an embodiment of the present disclosure. [Figure 16] 10 is a graph showing prediction results of a property prediction method according to a comparative example. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, embodiments of the present disclosure will be described.

[0019] (Outline of the embodiment) First, an overview of this embodiment will be described, and details will be provided later. The property prediction method according to this embodiment is executed by an information processing device 10. First, a descriptor for the mixture is generated by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results. The information processing device 10 inputs the descriptor for the mixture as an explanatory variable into a machine learning model. Then, the information processing device 10 predicts the properties of the mixture based on the output from the machine learning model.

[0020] As described above, according to this embodiment, a descriptor for a mixture is generated by multiplying each element of a descriptor for each compound in a mixture of liquid crystal materials or the like by the composition ratio and adding up the results, and the physical properties of the mixture can be predicted using the descriptor for the mixture and a machine learning model, thereby improving the property prediction technology using machine learning.

[0021] Next, each component of the information processing device will be described in detail.

[0022] (Configuration of information processing device) As shown in FIG. 1, the information processing device 10 includes a control unit 11, a storage unit 12, an input unit 13, and an output unit 14.

[0023] The control unit 11 includes at least one processor, at least one dedicated circuit, or a combination thereof. The processor is a general-purpose processor such as a central processing unit (CPU) or a graphics processing unit (GPU), or a dedicated processor specialized for a specific process. The dedicated circuit is, for example, a field-programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The control unit 11 executes processes related to the operation of the information processing device 10 while controlling each unit of the information processing device 10.

[0024] The storage unit 12 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory is, for example, a random access memory (RAM) or a read only memory (ROM). The RAM is, for example, a static random access memory (SRAM) or a dynamic random access memory (DRAM). The ROM is, for example, an electrically erasable programmable read only memory (EEPROM). The storage unit 12 functions as, for example, a main storage device, an auxiliary storage device, or a cache memory. The storage unit 12 stores data used in the operation of the information processing device 10 and data obtained by the operation of the information processing device 10.

[0025] The input unit 13 includes at least one input interface. The input interface is, for example, a physical key, a capacitance key, a pointing device, or a touch screen integrated with a display. The input interface may also be, for example, a sound sensor that accepts voice input, or a camera that accepts gesture input. The input unit 13 accepts an operation to input data used for the operation of the information processing device 10. The input unit 13 may be connected to the information processing device 10 as an external input device instead of being provided in the information processing device 10. Any connection method may be used, for example, a Universal Serial Bus (USB), a High-Definition Multimedia Interface (HDMI) (registered trademark), or Bluetooth (registered trademark).

[0026] The output unit 14 includes at least one output interface. The output interface is, for example, a display that outputs information as a video, or a speaker that outputs information as a sound. The display is, for example, an LCD (liquid crystal display) or an organic EL (electro luminescence) display. The output unit 14 outputs data obtained by the operation of the information processing device 10. The output unit 14 may be connected to the information processing device 10 as an external output device instead of being provided in the information processing device 10. Any connection method can be used, for example, USB, HDMI (registered trademark), or Bluetooth (registered trademark).

[0027] The functions of the information processing device 10 are realized by executing a program according to this embodiment on a processor corresponding to the control unit 11. That is, the functions of the information processing device 10 are realized by software. The program causes a computer to execute the operations of the information processing device 10, thereby causing the computer to function as the information processing device 10. That is, the computer functions as the information processing device 10 by executing the operations of the information processing device 10 in accordance with the program.

[0028] In this embodiment, the program can be recorded on a computer-readable recording medium. The computer-readable recording medium includes non-transitory computer-readable media, such as a magnetic recording device, an optical disc, a magneto-optical recording medium, or a semiconductor memory. The program can be distributed, for example, by selling, transferring, or lending a portable recording medium, such as a DVD (digital versatile disc) or a CD-ROM (compact disc read only memory), on which the program is recorded. The program can also be distributed by storing the program in the storage of an external server and transmitting the program from the external server to another computer. The program can also be provided as a program product.

[0029] Some or all of the functions of the information processing device 10 may be implemented by a dedicated circuit equivalent to the control unit 11. In other words, some or all of the functions of the information processing device 10 may be implemented by hardware.

[0030] (Generation of descriptors for mixtures) 2 to 7C, a description will be given of a method for generating a descriptor for a mixture according to this embodiment. As described above, a descriptor for a mixture is generated by multiplying each element associated with each compound in the mixture by its composition ratio and summing the results.

[0031] Here, the descriptor includes at least one of a physical property descriptor and a structural descriptor. The physical property descriptor may be either Mordred or RDKIT. The structural descriptor may be a Morgan Fingerprint, Feature Morgan Fingerprint, Topological Fingerprint, Atom-pair Fingerprint, Topological Torsion Fingerprint, E-state Fingerprint, MACCS keys, etc. Each element of the Morgan Fingerprint may be a count-based value. Figure 2 shows an overview of the count-based Morgan Fingerprint. As shown in Figure 2, in the count-based Morgan Fingerprint, the target molecule is described by the number of substructures. The target molecule shown in Figure 2 includes substructures 201 to 226. These substructures are classified into eight types: MF0014, MF0001, MF0006, MF0027, MF0015, MF0005, MF0016, and MF0020. Substructure 201 is MF0014. Substructures 202 and 203 are MF0001. Substructures 204 and 205 are MF0006. Substructures 206 and 207 are MF0027. Substructures 208, 209, and 210 are MF0015. Substructures 211 to 214 are MF0005. Substructures 215 to 220 are MF0016. Substructures 221 to 226 are MF0014. Thus, the molecule of interest contains 1, 2, 2, 2, 3, 4, 6, and 6 of MF0014, MF0001, MF0006, MF0027, MF0015, MF0005, MF0016, and MF0020, respectively. Therefore, the molecule is represented by count-based representation 300. The adoption of the count-based Morgan Fingerprint makes it possible to quantitatively indicate the frequency of occurrence of specific molecular structures, allowing for a more detailed reflection of the overall molecular characteristics. The adoption of the count-based method also makes it easier to reflect the contribution of chemical or physical characteristics of the structure in the model, which is expected to improve prediction accuracy.

[0032] The advantages of adopting a count-based approach rather than a binary approach to Morgan Fingerprints are as follows. First, Morgan Fingerprints are often provided as binary data indicating the presence or absence of a structure, including the bonding state, of all atoms in a molecule except for hydrogen atoms. Binary data describes only the presence or absence of a specific structure. For example, if the characteristics of a compound are determined solely by the presence or absence of a structure of interest, there is little need to consider the number of structures (functional groups). However, for compounds with continuous structures, such as polymer compounds, or compounds in which specific functional groups exist within the same molecule but the physical properties of the compound differ depending on the number of functional groups, binary information alone may be insufficient. In other words, to analyze the relationship between structure and physical properties in more detail, it is preferable to understand the number of substructures present. In particular, in this embodiment, descriptors are proportionally divided and totaled according to the abundance ratio of each compound constituting the composition, so it is preferable to clearly define the number of substructures present.

[0033] Figure 3 shows the radius of the Morgan Fingerprint. The radius of the Morgan Fingerprint can be 3 or greater. For example, by setting the radius to 3, the entire six-membered ring is included within the radius, allowing for an appropriate description of the mixture. Setting the radius in this way allows for a more accurate reflection of the overall characteristics of the chemical structure and improved predictive model performance. Here, the carbon next to a given carbon is called the α-carbon, and the carbon two carbons away is called the β-carbon. The positions themselves are also referred to as the α-position and β-position. In organic compounds, the influence of a carbon on that carbon is greatest up to the β-position, while the influence of a distance of three or more carbons is extremely small. For this reason, a radius of 2 is generally considered sufficient for representing organic compounds with a fingerprint. On the other hand, for benzene, one of the most useful organic compounds, a radius of 2 is considered insufficient. Benzene has numerous derivatives (e.g., aniline, phenol, benzoic acid, etc.), and the six carbons in benzene form a hexagon. When a substituent is located at the α-position of a given carbon, it is called ortho, and when it is located at the β-position, it is called meta. When a substituent is located three carbons away, i.e., on the diagonal, it is called para. Many benzene ring derivatives have characteristic physical properties even in para-position compounds. Therefore, to accurately analyze the properties of six-membered ring compounds, such as benzene, it is necessary to incorporate descriptors that represent the entire six-membered ring into the explanatory variables. In other words, the Morgan Fingerprint radius must be 3. For example, in the analysis of liquid crystal compounds, the coefficient of determination and error value in the analysis tend to improve as the radius is increased, and accuracy improves at a radius of 3 or 4. On the other hand, when the radius is set to 2, the coefficient of determination decreases and the error increases, which is considered to be supporting evidence that the descriptors that represent the six-membered ring are necessary for the analysis.

[0034] Specifically, you can obtain the Morgan Fingerprint using the Python script shown below. fp = AllChem.GetMorganFingerprint(mol, radius , useCounts=True) In the above script, radius is a parameter that specifies how many bonds from the central atom should be reflected in the fingerprint. For example, if radius=2, the partial structure up to two bonds from the central atom will be the subject of the fingerprint. Furthermore, by setting useCounts=True, it is possible to obtain a count-based fingerprint that reflects the number of times the same feature exists multiple times within a molecule, rather than simply expressing the presence or absence of a feature with bits. The Morgan Fingerprint obtained in this way can efficiently and comprehensively capture the features of molecular structure, making it effective for various analyses.

[0035] Figure 4 shows the Morgan Fingerprint generated for ibuprofen with radius=0. When radius=0, the area around each atom is not traced at all, and the characteristics corresponding to the properties of the atom alone are reflected. For example, the carbon with a double bond (C(sp 2 )) 、 Single bond carbon (C(sp 3 )), CH3, CH2, CH, C 、 Oxygen atom (O) 、 Hydroxy group (OH) 、 Carbonyl group (=O) etc. Hash values ​​(bits) are assigned based on information such as "atom type, number of bonds, electronic state." In table 400 in Figure 4, the Bit column shows numerical values ​​such as "864662311" and "864942730." These are unique hash values ​​obtained during the Morgan Fingerprint calculation process. The same hash value is obtained for the same atom type and bonding state. In this way, the characteristics of the molecule are encoded. The Structure column in table 400 shows in easy-to-understand text the specific structure represented by the Bit hash value. Morgan Fingerprint does not store this text information as is, but replaces the elements obtained by the calculation with hash values. The Count column shows how many of each atomic environment exist within the molecule. For example, "CH(sp2 If the line "Count=2" in the table reads "2", it means that the ibuprofen molecule contains two atoms with this property. If you specify "useCounts=True" when generating a Morgan Fingerprint, a count-based fingerprint will be generated that reflects the number of times the same feature is found, in cases where multiple instances of the same feature exist.

[0036] Figures 5A to 5C show Morgan Fingerprints generated for ibuprofen with a radius of 1 to 3. Figure 5A corresponds to a radius of 1, Figure 5B corresponds to a radius of 2, and Figure 5C corresponds to a radius of 3. When radius = 1, structural information from the central atom to its directly bonded atoms is obtained. When radius = 2, structural information from the central atom to the atoms directly bonded to it is obtained. When radius = 3, structural information from the central atom to the atoms directly bonded to it is obtained. The larger the radius, the wider the range of structural elements (substituents, parts where multiple atoms are connected, etc.) that can be reflected in the fingerprint. As shown in Figure 5A, when radius = 1, information from the central atom to the atoms directly bonded to it is reflected as bits (hash values). As shown in Table 501, each local structure is listed. When radius = 2, the fingerprint is more diverse than when radius = 1 because it considers the atoms directly bonded to it (e.g., the carbon with the substituent, the carbon beyond that, and the oxygen). Specifically, the number of newly detected hash values ​​increases, and the ring structure and substituent structure two bonds away are added as features. More specifically, the bits shown in table 502 (e.g., 602262358, 2279287705, etc.) indicate the inclusion of a portion of a continuous carbon chain or a group outside the ring, which could not be covered by a single bond alone. Furthermore, when radius = 3, information on the portion three bonds away from the central atom, such as a large substituent at the end of the molecule or a functional group reached via an aromatic ring, can be reflected. Note that in Figures 5B and 5C, fingerprints centered on terminal atoms are not generated when radius = 2 (radius = 2), and on atoms one bond further back are not generated when radius = 3 (radius = 3). This is because even if radius 2 or 3 is set around a terminal atom (such as a terminal carbon atom), there are no structural elements beyond that, so the fingerprint may not be counted as a new fingerprint (bit). Such cases are indicated by an "x" in Figures 5B and 5C. Furthermore, in the case of a radius of 3 (radius=3), there are multiple shaded bits (hash values) in table 503, but the count is 1.These are due to the specifications of the hash calculation used in Morgan Fingerprint and the handling of terminal atoms. However, when radius = 3 or more is set, information below that radius, i.e., information from radius = 0 to 2, is automatically included, so there is no need to worry about missing information.

[0037] Figure 6 shows chemical formula C 16 H 18 The Morgan Fingerprint generated for the compound FN (hereinafter referred to as this compound) with radius = 0 is shown. Referring to table 600 in FIG. 6, atoms such as N and F are present, and carbon (C(sp 2 ), C(sp 3 )) and that they form a different bond state than carbonyl groups, etc. These are assigned corresponding hash values ​​(bits) based on the characteristics (atom type, number of bonds, electronic state) when the central atom is not traced at all (radius = 0). For example, the hash value corresponding to an N atom is "847433064," and the hash value corresponding to an F atom is "882399112." The same bit is obtained for the same atom type and bond state. The Count column in Table 600 indicates how many of each characteristic are present in the molecule. By specifying useCounts=True, the count value is reflected when there are multiple identical environments.

[0038] Figures 7A to 7C show Morgan Fingerprints generated for this compound with radius = 1, 2, and 3, respectively. Figure 7A corresponds to radius = 1, Figure 7B corresponds to radius = 2, and Figure 7C corresponds to radius = 3. With radius = 1, the structure from the central atom to the next bond is captured as a feature. As shown in Table 701, C(sp 2 ), C(sp 3), as well as bond information with carbons adjacent to the F and N atoms, are listed as bits (hash values). With radius=2, the structure is considered up to two bonds from the central atom, resulting in an increased hash value, as shown in Table 702. This allows for more detailed capture of aromatic rings and linked substituents. For example, bits such as "973199101" and "714711401" reflect parts of the carbon chain or cyclic structure that cannot be fully expressed with a single bond. Note that if the structural expansion is interrupted when the terminal atom is the center, a new bit is not generated, and an event such as the "x" in the figure may occur. With radius=3, the structure is considered up to three bonds from the central atom, resulting in more newly recognized structural elements, such as "2955280293" and "3722802267" shown in Table 703 in Figure 7C. These are the results of reflecting the substituents at the ends of the molecule and the functional groups that exist beyond the boundaries of multiple carbon atoms. However, as with the example of ibuprofen, there are cases where the count is 1 for multiple identical partial structures.

[0039] Table 100 shown in FIG. 8 is a table that schematically shows the relationship between the compounds that make up mixture A, their composition ratios, and each element of the descriptor for each compound. Here, mixture A is assumed to be a mixture of compounds 1 to n. As shown in table 100, the composition ratios for compounds 1 to n are assumed to be 35%, 5%, 14%, . . . , 2%, respectively. The descriptor for each compound is assumed to be a vector composed of element 1 to element k. Specifically, d 11 , d 12 , d 13 , , d 1k are the values ​​of element 1, element 2, element 3, ..., and element k of the descriptor of compound 1, respectively. Similarly, d 21 , d 22 , d 23 , , d 2k are the values ​​of element 1, element 2, element 3, ..., and element k of the descriptor of compound 2, respectively. Similarly, d 31 , d 32 , d 33 , , d 3kare the values ​​of element 1, element 2, element 3, . . . and element k of the descriptor of compound 3, respectively. Similarly, d n1 , d n2 , d n3 , , d nk are the values ​​of element 1, element 2, element 3, . . . , and element k of the descriptor of compound n, respectively.

[0040] Table 101 shows the relationship between the compounds constituting mixture A and the values ​​obtained by multiplying each element of the descriptor for each compound in the mixture by its composition ratio. Table 101 is generated from table 100. As shown in table 101, each element of the descriptor for compound 1 is multiplied by the composition ratio of compound 1. Specifically, for compound 1, elements 1, 2, . . . , and k of the descriptor are multiplied by 0.35, which is the composition ratio of compound 1. Similarly, for compound 2, elements 1, 2, . . . , and k of the descriptor are multiplied by 0.05, which is the composition ratio of compound 2. Similarly, for compound 3, elements 1, 2, . . . , and k of the descriptor are multiplied by 0.14, which is the composition ratio of compound 3. Similarly, for compound n, elements 1, 2, . . . , and k of the descriptor are multiplied by 0.02, which is the composition ratio of compound n.

[0041] Table 102 is a table obtained by summing up the values ​​obtained by multiplying each element of the descriptor for each compound by the composition ratio. Table 102 is generated from table 101. Here, the composition ratio of each compound is expressed as A i where i is an integer from 1 to n corresponding to the compound. In this way, element 1, element 2, ..., element k of the descriptor of mixture A (in other words, the descriptor of mixture A) are expressed by the following formula (1).

[0042]

number

[0043] Table 110 shown in FIG. 9 is a table showing the composition ratios of each compound in a plurality of mixtures A to J. Table 111 is generated from table 110 by the same process as in FIG. 8. As shown in table 111, for example, the descriptors of mixtures B, C, and J are expressed by the following formulas (2), (3), and (4), respectively. Here, the composition ratios of each compound constituting mixtures B, C, and J are expressed as B i , C i , and J i where i is an integer from 1 to n corresponding to each compound.

[0044]

number

[0045] In this embodiment, the properties of a mixture are predicted using the mixture-related descriptors generated by the above-described method and a learning model.

[0046] (Operation of information processing device) The operation of the information processing device 10 according to this embodiment will be described with reference to FIG.

[0047] Step S1: The control unit 11 of the information processing device 10 generates a descriptor for the mixture by multiplying each element of the descriptor for each compound in the mixture by the composition ratio and adding up the results.

[0048] Step S2: The control unit 11 inputs the descriptors of the mixture as explanatory variables into the learning model. In this embodiment, the learning model is a model that uses the descriptors of the mixture as explanatory variables and the physical properties of the mixture as target variables. In other words, the learning model is a machine learning model trained using the descriptors of the mixture as features and the actual values ​​of the physical properties of the mixture as labels. The learning model may be a model created by machine learning using a machine learning algorithm. The learning model may be, for example, a machine learning model built based on a decision tree. Examples of machine learning models built based on a decision tree include, but are not limited to, Light GBM and XGBoost. Alternatively, the learning model may be a model generated based on a machine learning algorithm such as a convolutional neural network (CNN), a recurrent neural network (RNN), or other deep learning.

[0049] Step S3: The control unit 11 predicts the physical properties of the mixture based on the output from the learning model.

[0050] Step S4: The control unit 11 outputs the prediction result. Any method can be used for outputting the prediction result. For example, the control unit 11 may cause the output unit 14 to display and output information.

[0051] As described above, the information processing device 10 according to this embodiment generates a descriptor for a mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results. The information processing device 10 inputs the descriptor for the mixture as an explanatory variable into a machine learning model. The information processing device 10 then predicts the physical properties of the mixture based on the output from the machine learning model.

[0052] This configuration improves machine learning-based property prediction technology by generating a descriptor for a mixture by multiplying each element of a descriptor for each compound in a mixture, such as a liquid crystal material, by its composition ratio and summing the results. The descriptor for the mixture can then be used to predict the properties of the mixture using the descriptor and a machine learning model. Furthermore, the use of well-known molecular descriptors, such as Morgan Fingerprint and Mordred, maintains versatility, and by combining physical property descriptors with structural descriptors, prediction results that reflect intermolecular bonding information can be obtained. In other words, even those unfamiliar with machine learning can easily obtain highly accurate prediction results, reducing the effort and cost of experimental property measurements. Furthermore, consistent predictions can be made for mixtures of different compositions, making the method useful for the design of new materials.

[0053] Referring to Figures 11 to 16, the accuracy of the prediction results for the physical properties of mixtures is shown. In this example, liquid crystals were used as the mixture, and the transition point was predicted and measured as a physical property of the mixture. The transition point is the temperature at which the liquid crystal state changes to a different phase (state). In this case, if the descriptor includes a physical property descriptor, the physical property descriptor includes elements related to HSP (Hansen Solubility Parameters). By including elements related to HSP in the physical property descriptor, predictions that take into account intermolecular interactions, the solubility of the mixture, and interaction patterns are expected. This can improve the accuracy of transition point predictions.

[0054] FIG. 11 is a graph showing a comparison between the predicted transition point results (horizontal axis: Predicted) by a learning model according to a comparative example and the actual measurement results of the transition point (vertical axis: Actual). The explanatory variables of the machine learning model in the comparative example are the compositions that make up the mixture and the ratios of those compositions. In the comparative example, an Elastic-Net Regressor was used as the machine learning model. FIG. 11 shows two results: one in which data was divided into five parts using Cross-Validation (CV) to evaluate the performance of the learning model, and the other in which a portion of the data (here, 20%) was set aside for testing using Hold Out (HO). R 2 is the coefficient of determination, and the closer it is to 1, the higher the accuracy. Root Mean Squared Error (RMSE) is the root mean square error. The smaller the RMSE value, the higher the accuracy. As shown in Figure 11, there is a relatively high correlation between the prediction results of the physical properties by the learning model and the actual measurement results of the physical properties. In CV, R 2 , and RMSE are 0.941 and 4.942, respectively. 2 , and RMSE are 0.966 and 3.937, respectively.

[0055] FIG. 12 is a graph showing a comparison between the predicted transition point results (horizontal axis: Predicted) obtained by employing the method for generating descriptors for mixtures according to this embodiment and inputting the descriptors into a learning model, and the actual measurement results of the transition point (vertical axis: Actual). Here, the descriptor for the mixture was generated by multiplying each element of the property descriptor for each compound in the mixture by the composition ratio and summing the results. Here, the property descriptor is RDKit. Elastic-Net Regressor was used as the machine learning model. FIG. 12 shows two results: one in which data was divided into five parts using CV to evaluate the performance of the learning model, and the other in which a portion of the data (here, 20%) was set aside for testing using HO. As shown in FIG. 12, there is a high correlation between the predicted property results by the learning model and the actual measurement results of the physical properties. In CV, R 2, and RMSE are 0.960 and 4.015, respectively, which shows that the prediction is more accurate than the comparative example. 2 The RMSE and RMSE were 0.965 and 3.732, respectively, which indicates that the prediction was more accurate than in the comparative example.

[0056] Figure 13 is a graph showing a comparison of the predicted transition points by the learning model (horizontal axis: Predicted) and the actual measured transition points (vertical axis: Actual) when the method for generating descriptors for mixtures is changed. Here, the descriptor for the mixture is generated by multiplying each element of the property descriptor for each compound in the mixture by the composition ratio and adding them up. Here, the property descriptor is Mordred. Elastic-Net Regressor was adopted as the machine learning model. Figure 13 shows two results: one in which data was divided into five parts using CV to evaluate the performance of the learning model, and the other in which a portion of the data (20% in this case) was set aside for testing using HO. As shown in Figure 13, there is a high correlation between the predicted property results by the learning model and the actual measured property results. In CV, R 2 , and RMSE are 0.983 and 2.569, respectively, which shows that the prediction is more accurate than the comparative example. 2 The RMSE and RMSE were 0.988 and 2.117, respectively, which indicates that the prediction was more accurate than in the comparative example.

[0057] Figure 14 is a graph showing a comparison between the predicted transition points (horizontal axis: Predicted) by a learning model and the actual measured transition points (vertical axis: Actual) when a descriptor for a mixture is generated by multiplying each element of the structural descriptor for each compound in the mixture by the composition ratio and adding them up. Here, the structural descriptor is Morgan Fingerprint, and Elastic-Net Regressor is used as the machine learning model. Figure 14 shows two results: one in which data was divided into five parts using CV to evaluate the performance of the learning model, and the other in which a portion of the data (20% in this case) was set aside for testing using HO. As shown in Figure 14, there is a high correlation between the predicted physical properties by the learning model and the actual measured physical properties. Furthermore, in CV, R 2 , and RMSE are 0.987 and 2.261, respectively, which shows that the prediction is more accurate than the comparative example. 2 The RMSE and RMSE were 0.990 and 2.021, respectively, which indicates that the prediction was more accurate than in the comparative example.

[0058] Figure 15 is a graph comparing the predicted transition points (horizontal axis: Predicted) by the learning model with the actual measured transition points (vertical axis: Actual) when the method for generating descriptors for mixtures was changed. Here, the descriptor for the mixture was generated by multiplying each element of the physical property descriptor for each compound in the mixture by the composition ratio and summing the results, and by multiplying each element of the structural descriptor for each compound in the mixture by the composition ratio and summing the results. Here, the physical property descriptor is RDKit. The structural descriptor is Morgan Fingerprint. Elastic-Net Regressor was used as the machine learning model. Figure 15 shows two results: one in which data was divided into five parts using CV to evaluate the performance of the learning model, and the other in which a portion of the data (20% in this case) was set aside for testing using HO. As shown in Figure 15, there is a high correlation between the predicted physical properties by the learning model and the actual measured physical properties. In CV, R 2The RMSE and RMSE were 0.992 and 1.790, respectively, which shows that the prediction was more accurate than the results shown in the comparative example. 2 The σ and RMSE were 0.992 and 1.776, respectively, which indicates that the prediction was made with higher accuracy than the results shown in the comparative example.

[0059] Figure 16 is a graph comparing the predicted transition points (horizontal axis: Predicted) by the learning model with the actual measured transition points (vertical axis: Actual) when the method for generating descriptors for mixtures was changed. Here, the descriptor for the mixture was generated by multiplying each element of the physical property descriptor for each compound in the mixture by the composition ratio and summing the results, and also by multiplying each element of the structural descriptor for each compound in the mixture by the composition ratio and summing the results. Here, the physical property descriptor is Mordred. The structural descriptor is Morgan Fingerprint. Elastic-Net Regressor was used as the machine learning model. Figure 16 shows two results: one in which the data was divided into five parts using CV to evaluate the performance of the learning model, and one in which a portion of the data (20% in this case) was set aside for testing using HO. As shown in Figure 16, there is a high correlation between the predicted physical properties by the learning model and the actual measured physical properties. In CV, R 2 The RMSE and RMSE were 0.992 and 1.798, respectively, which shows that the prediction was more accurate than the results shown in the comparative example. 2 The RMSE and RMSE were 0.992 and 1.745, respectively, which indicates that the prediction was performed with a high degree of accuracy comparable to the results shown in the comparative example.

[0060] Although the present disclosure has been described based on the drawings and examples, it should be noted that those skilled in the art may make various modifications and alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included in the scope of the present disclosure. For example, the functions included in each component or step can be rearranged so as not to be logically inconsistent, and multiple components or steps can be combined or divided into one.

[0061] For example, in the above-described embodiment, the configuration and operation of the information processing device 10 may be distributed among a plurality of computers that can communicate with each other. [Explanation of symbols]

[0062] 1 System 10. Information processing equipment 11 Control section 12 Storage section 13 Input section 14 Output section Tables 100, 101, 102, 110, and 111 201~226 Partial structure 300 Expression format 400, 501, 502, 503, 600, 701, 702, 703 tables

Claims

1. A physical property prediction method executed by an information processing device, comprising: generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors of the mixture into a learning model as explanatory variables; predicting physical properties of the mixture based on output from the learning model; and Including, the descriptors include structural descriptors; each element of the structural descriptor is a count-based value; the radius of the structural descriptor is 3 or greater; Physical property prediction methods.

2. 2. The physical property prediction method according to claim 1, The method for predicting physical properties, wherein the descriptors include physical property descriptors.

3. 3. The physical property prediction method according to claim 2, The physical property prediction method, wherein the physical property descriptor is either Mordred or RDKIT.

4. 3. The physical property prediction method according to claim 2, A physical property prediction method, wherein the mixture is a liquid crystal and the physical property descriptor includes an element related to HSP.

5. 2. The physical property prediction method according to claim 1, The method for predicting physical properties, wherein the structural descriptor is a Morgan Fingerprint.

6. An information processing device including a control unit, The control unit generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors related to the mixture into a learning model as explanatory variables; predicting physical properties of the mixture based on an output from the learning model; the descriptors include structural descriptors; each element of the structural descriptor is a count-based value; the radius of the structural descriptor is 3 or greater; Information processing device.

7. On the computer, generating a descriptor for the mixture by multiplying each element of a descriptor for each compound in the mixture by its composition ratio and summing the results; inputting the descriptors of the mixture into a learning model as explanatory variables; predicting physical properties of the mixture based on output from the learning model; and Execute the descriptors include structural descriptors; each element of the structural descriptor is a count-based value; the radius of the structural descriptor is 3 or greater; program.

Citation Information

Patent Citations

  • Liquid crystal panel, liquid crystal display device and electronic apparatus

    JP2018077305A

  • Molecule descriptor creation system, molecule descriptor creation method and molecule descriptor creation program

    JP2021081920A

  • Method for searching for thermosetting epoxy resin composition, information processing apparatus, and program

    JP2023183286A

  • Prediction method, information processing apparatus, computer program, substance selection method, and substance manufacturing method

    JP2024145640A

  • Prediction system

    JP2024171479A