Method for designing organic compound and system for designing organic compound
The method and system for designing organic compounds using autoencoders and regression models address the inefficiencies of traditional methods by predicting desired physical properties, thereby enhancing development speed and reducing costs.
Patent Information
- Application Number
- PCT/IB2024/061764
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-05
AI Technical Summary
Current methods for designing organic compounds with desired physical properties are inefficient and require significant time and cost, as they rely on direct synthesis and measurement, which are not easily scalable or accessible to all experts.
A method and system that utilize autoencoders and regression models to predict the physical properties of organic compounds, allowing for the design of compounds with specific molecular orientations and properties, such as luminescence spectra and transition dipole moments, through a process involving learning datasets and neural networks.
This approach enables the easy and accurate prediction of organic compounds with desired physical properties, significantly improving development speed and reducing costs associated with traditional synthesis and measurement methods.
Smart Images

Figure IB2024061764_05062025_PF_FP_ABST
Abstract
Description
Organic compound design method and organic compound design system
[0001] One aspect of the present invention relates to a method and system for designing novel organic compounds that have predictable desired physical properties.
[0002] In research or development using organic compounds, organic compounds with corresponding physical properties are selected and used according to the required characteristics. However, the physical properties of an organic compound can only be known by synthesizing the target substance and directly measuring it. Therefore, when synthesizing a new organic compound, it is necessary to predict whether the organic compound will have the desired physical properties from known substances. Furthermore, the physical properties of an organic compound are determined by the molecular structure of the organic compound, but because there are so many candidate organic compounds, it is difficult even for experienced researchers to identify them.
[0003] Furthermore, it is possible to predict the approximate values of the physical properties of an organic compound with a certain molecular structure by using accumulated data, scientific calculations such as first-principles calculations, or simulations. However, not everyone can make accurate predictions, and currently simulations are extremely costly and time-consuming.
[0004] In recent years, methods for classification, estimation, and prediction using techniques such as machine learning have undergone significant advances. In particular, the performance of selection and prediction using deep learning with convolutional neural networks has improved significantly, achieving excellent results in a variety of fields. However, in the field of organic compounds, there are still few sufficient methods for describing organic compounds that can accurately understand their structure and extract characteristics related to their physical properties, while providing a manageable amount of information. Therefore, a design method and system that allows anyone to easily and accurately predict organic compounds with desired physical properties has not yet been realized.
[0005] Patent Document 1 and Non-Patent Document 1 disclose a method and an apparatus for searching for new substances using machine learning.
[0006] JP 2017-91526 A
[0007] Takahiro Ohkoshi, Masaki Ohkoshi, Yoshiaki Nagao, Satoshi Yotsuhashi, "Exploring new structures of optically functional organic materials using high-throughput calculations and machine learning," Panasonic Technical Journal Vol. 67 No. 2 Nov. 2021 p. 116, P. Liehm, and five others, "Applied Physics Letters," 101, 253304 (2012)
[0008] An object of one embodiment of the present invention is to provide a method for designing an organic compound that allows anyone to easily and accurately design an unknown organic compound having desired physical properties. Another object is to provide a system for designing an organic compound that allows anyone to easily and accurately design an organic compound. Another object of one embodiment of the present invention is to provide a method for generating a training dataset that can be used to design an organic compound having desired physical properties or to predict the physical properties of an organic compound.
[0009] Note that the description of these problems does not preclude the existence of other problems. Note that one embodiment of the present invention does not necessarily solve all of these problems. Note that problems other than these will become apparent from the description of the specification, drawings, claims, etc., and it is possible to extract other problems from the description of the specification, drawings, claims, etc.
[0010] One aspect of the present invention is a method for designing an organic compound, the method comprising the steps of: setting a first set of explanatory variables for a first group of organic compounds from a first set of explanatory variables; causing an autoencoder to learn the molecular structures of the first group of organic compounds and acquiring a first set of latent variables corresponding to the molecular structure of the first group of organic compounds; causing a regression model to learn the correlation between the first set of latent variables and the first set of explanatory variables; generating a second set of latent variables using random numbers; acquiring the molecular structures of the second group of organic compounds from the second latent variable set using the autoencoder; acquiring predicted values of the second set of explanatory variables from the second latent variable set using the regression model; selecting a third group of organic compounds from the second group of organic compounds using the predicted values of the second set of explanatory variables; and calculating a second set of explanatory variables for the third group of organic compounds, wherein the first set of explanatory variables and the second set of explanatory variables contain information about molecular orientation.
[0011] In the above, the first group of explanatory variables and the second group of explanatory variables are a method for designing an organic compound, each of which includes the length of the major axis of a molecule in the organic compound, the magnitude of the transition dipole moment in the organic compound, and the angle between the major axis and the transition dipole moment.
[0012] In the above, the first to third organic compound groups are methods for designing organic compounds that are organometallic complexes in which the coordination number of the central metal is four.
[0013] One aspect of the present invention includes an input unit, a calculation unit, and an output unit. The input unit has a function of inputting a molecular structure of a first group of organic compounds and a first group of explanatory variables. The calculation unit has a function of setting a first group of dependent variables from the first group of explanatory variables for the first group of organic compounds, a function of training an autoencoder to learn the molecular structure of the first group of organic compounds and acquiring a first group of dependent variables corresponding to the molecular structure of the first group of organic compounds, a function of training a regression model to learn the correlation between the first group of dependent variables and the first group of dependent variables, a function of generating a second group of dependent variables using random numbers, and a function of training an autoencoder to learn the correlation between the first group of dependent variables and the first group of dependent variables. a function of acquiring a molecular structure of a second group of organic compounds from the first group of latent variables, a function of acquiring a predicted value of a second group of objective variables from the second group of latent variables using a regression model, a function of selecting a third group of organic compounds from the second group of organic compounds using the predicted value of the second group of objective variables, and a function of calculating a second group of explanatory variables for the third group of organic compounds; an output unit has a function of outputting the molecular structure of the second organic compound or a physical property value related to the molecular orientation of the second group of organic compounds, and the first group of explanatory variables and the second group of explanatory variables have information related to molecular orientation.
[0014] In the above, the first group of explanatory variables and the second group of explanatory variables are a design system for organic compounds that respectively include the length of the long axis of the molecule in the organic compound, the magnitude of the transition dipole moment in the organic compound, and the angle between the long axis and the transition dipole moment.
[0015] In the above, the autoencoder is a design system for organic compounds, which is JT-VAE.
[0016] In the above, the organic compound is a design system of an organic compound that is an organometallic complex in which the coordination number of the central metal is four.
[0017] In another embodiment of the present invention, in the above-described structure, the physical property values of the organic compound are one or more of an emission spectrum, a half-width, an emission energy, an excitation spectrum, an absorption spectrum, a transmission spectrum, a reflection spectrum, a molar absorption coefficient, an excitation energy, a transient emission lifetime, a transient absorption lifetime, an S1 level, a T1 level, an Sn level, a Tn level, a Stokes shift value, an emission quantum yield, an oscillator strength, an oxidation potential, a reduction potential, a HOMO level, a LUMO level, a glass transition point, a melting point, a crystallization temperature, a decomposition temperature, a boiling point, a sublimation temperature, a carrier mobility, a refractive index, a molecular orientation parameter, a mass-to-charge ratio, a spectrum in an NMR measurement, a chemical shift value and the number of elements or a coupling constant thereof, a spectrum in an ESR measurement, a g factor, a D value, and an E value.
[0018] One embodiment of the present invention provides an organic compound design method capable of designing an organic compound having unknown but desired physical properties. Another embodiment of the present invention provides an organic compound design system capable of designing an organic compound having unknown but desired physical properties. Another embodiment of the present invention provides a method for generating training data applicable to the organic compound design method and the organic compound design system. Use of these organic compound design methods or systems facilitates the selection of organic compounds having desired physical properties, significantly improving the development speed.
[0019] The effects of one embodiment of the present invention are not limited to the effects listed above. The effects listed above do not preclude the existence of other effects. Note that the other effects are effects not mentioned in this section, which will be described below. Effects not mentioned in this section can be derived by a person skilled in the art from the description in the specification, drawings, etc., and can be extracted as appropriate from these descriptions. Note that one embodiment of the present invention has at least one of the effects listed above and / or other effects. Therefore, one embodiment of the present invention may not have the effects listed above in some cases.
[0020] FIG. 1 is a flowchart illustrating one embodiment of the present invention. FIGS. 2A and 2C are diagrams illustrating a training data set, and FIG. 2B is a diagram illustrating an autoencoder. FIGS. 3A and 3C are diagrams illustrating a regression model, and FIG. 3B is a diagram illustrating a decoder. FIGS. 4A and 4B are diagrams illustrating the configuration of a neural network. FIGS. 5A to 5C are diagrams illustrating a training data set. FIG. 6 is a diagram illustrating a method for converting molecular structures using fingerprinting. FIGS. 7A to 7D are diagrams illustrating types of fingerprinting. FIG. 8 is a diagram illustrating the process of converting from SMILES notation to fingerprinting notation. FIG. 9 is a diagram illustrating types of fingerprinting and notation overlap. FIGS. 10A and 10B are diagrams illustrating an example of a molecular structure notation using multiple fingerprinting methods. FIG. 11 is a diagram illustrating a physical property prediction system according to one embodiment of the present invention. FIG. 12A is a schematic diagram of a light-emitting device according to one embodiment of the present invention, and FIG. 12B is a diagram illustrating angles formed with each vector in an organic compound according to one embodiment of the present invention. FIG. 13 is a diagram for explaining the relationship between the observation direction of a measuring device in measuring the spatial distribution of emission intensity and each vector component of the transition dipole moment on the substrate.
[0021] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the present invention is not limited to the following description, and it will be readily understood by those skilled in the art that various changes in form and details can be made without departing from the spirit and scope of the present invention. Therefore, the present invention should not be interpreted as being limited to the description of the embodiments shown below.
[0022] Embodiment 1 A method for designing an organic compound according to one embodiment of the present invention will be described below. The method for designing an organic compound according to one embodiment of the present invention can be represented, for example, by a flowchart F100 shown in FIG.
[0023] 1, a data set 100 shown in Fig. 2A is prepared in advance. The data set 100 includes molecular structures, explanatory variables, and response variables of n organic compounds (hereinafter also referred to as an organic compound group), that is, organic compounds M1 to Mn (n is a natural number).
[0024] As explanatory variables included in the data set 100, for example, physical property values of organic compounds M1 to Mn (n is a natural number) can be used.
[0025] Specifically, it is preferable to use, as the physical property values, those related to molecular orientation (molecular orientation parameters), such as the length of the molecular major axis (X), the magnitude of the transition dipole moment (D), and the angle between the molecular major axis and the transition dipole moment (θ). Note that molecular orientation parameters will be described later in other embodiments.
[0026] Other physical property values include, for example, energy levels such as the lowest triplet excitation energy (T1) level, the lowest singlet excitation energy (S1) level, the HOMO (Highest Occupied Molecular Orbital) level, and the LUMO (Lowest Unoccupied Molecular Orbital) level, as well as an energy level difference between excited states and a rearrangement energy (λ).
[0027] Other physical property values that can be used include emission spectrum, half-width, emission energy, excitation spectrum, absorption spectrum, transmission spectrum, reflection spectrum, molar absorption coefficient, excitation energy, transient emission lifetime, transient absorption lifetime, Sn level, Tn level, Stokes shift value, emission quantum yield, oscillator strength, oxidation potential, reduction potential, glass transition point, melting point, crystallization temperature, decomposition temperature, boiling point, sublimation temperature, carrier mobility, refractive index, mass-to-charge ratio, spectrum in NMR measurement, chemical shift value and the number of elements or coupling constant thereof, spectrum in ESR measurement, g factor, D value or E value, etc.
[0028] The physical property values may be experimentally measured values, calculated values obtained by quantum chemical calculations, or predicted values using a regression model or the like.
[0029] Furthermore, the objective variable Y can be set from the explanatory variables included in the data set 100. The objective variable Y is set using an explanatory variable or a function using a mathematical model that reflects the explanatory variable as appropriate. For example, if one wishes to design an organic compound with excellent molecular orientation, a parameter related to the orientation can be selected as the explanatory variable.
[0030] Specifically, when X is the length of the long axis of the molecule, D is the magnitude of the transition dipole moment, and θ is the angle between the long axis of the molecule and the transition dipole moment, the inner product of the long axis vector of the molecule and the transition dipole moment can be set as the objective variable Y, as shown in equation (1).
[0031]
[0032] In formula (1), the larger the transition dipole moment is and the more it coincides with the direction of the major axis vector, the greater the response variable Y. In other words, the better the molecular orientation of the organic compound is, the greater the response variable Y tends to be.
[0033] Furthermore, for example, as shown in formula (2), a sigmoid function can be used for the objective variable Y. In formula (2), T1 represents the T1 level, λS0 represents the rearrangement energy, Ea represents the transition energy, and Sigmoid represents the sigmoid function.
[0034]
[0035] In formula (2), the objective variable Y tends to increase as the properties of the organic compound improve.
[0036] The objective variable Y is not limited to the above formula (1) or (2), and an appropriate function may be used as needed.
[0037] The number of compounds (n) included in the data set 100 is preferably as large as possible, and is preferably 100 or more, more preferably 500 or more, and even more preferably 1000 or more.
[0038] <<Method for Designing Organic Compounds>> Hereinafter, a description will be given according to the flowchart F100 in FIG.
[0039] Step S101: First, the first organic compound group M (organic compound M 1 ~Organic compound M n ), the first explanatory variables (the physical property values of the first organic compound group M), and the first objective variable are learned by an autoencoder (AE) 110.
[0040] <Step S102> Next, a data set 120 consisting of a first latent variable Z and a first objective variable Y is created using the data set 100.
[0041] 2B, the AE 110 is made up of an encoder 111 and a decoder 112. For example, a VAE (Variational Autoencoder) can be used as the AE 110. Among VAEs, it is preferable to use a JT (Junction Tree)-VAE.
[0042] The encoder 111 detects the organic compound M 1 ~Organic compound M n is given as an input, and the first latent variable Z (latent variable Z 1 or latent variable Z k (k is a natural number)). Meanwhile, the decoder 112 outputs the latent variable Z 1 or latent variable Z k is given as an input, and an organic compound M 1 ~Organic compound M n That is, the AE110 outputs the molecular structure of the latent variable Z 1 or latent variable Z k The dimension (k) of the latent variable Z can be selected to be a value between 10 and 100.
[0043] The molecular structure is preferably expressed in SMILES (Simplified Molecular Input Line Entry Specification Syntax) notation. In addition to SMILES, SMARTS notation, SMIRKS notation, InChI notation, WLN (Wiswesser Line Notation) notation, ROSDAL notation, SLN notation (Tripos), or names based on the IUPAC nomenclature can also be used.
[0044] As shown in FIG. 2C, as a result of learning by the AE 110, the organic compound M 1 ~Organic compound M n Each molecular structure is represented by a latent variable Z 1 or latent variable Z k That is, the AE 110 compresses the information of each organic compound into a k-dimensional value.
[0045] <Step S103> Next, the regression model 113 is trained on the data set 120.
[0046] As shown in FIG. 3A, the regression model 113 calculates the latent variable Z 1 or latent variable Z k When the input is made, the corresponding value of the objective variable Y (predicted value Y 1 or predicted value Y k ) is output. An appropriate regression model may be used as the regression model 113 as needed. Specifically, a linear regression model such as a Gaussian process regression model, or a nonlinear regression model such as a neural network may be used.
[0047] <Step S104> Next, a random number is generated, and the random number is used to generate a second latent variable Z′ (latent variable Z′ 1 to latent variable Z' j (j is an integer)
[0048] The larger the number (j) of newly generated organic compounds, the better, and it is preferably 1,000 or more, more preferably 5,000 or more, and even more preferably 10,000 or more. It is also preferable to adjust the random number so that the second latent variable Z' does not become a value that is significantly different from the first latent variable Z included in the dataset 120. Specifically, the maximum and minimum values of the first latent variable Z are determined, and the value of the random number is adjusted so that the second latent variable Z' does not deviate from that range.
[0049] <Step S105> Next, the second latent variable Z′ obtained in step S104 is input to the decoder 112 that has been trained in step S101, thereby generating a second organic compound M′ (organic compound M′ 1 to organic compound M' j ) to obtain the molecular structure (Figure 3B).
[0050] <Step S106> Next, the second latent variable Z′ obtained in step S104 is input to the regression model 113 that has been trained in step S103, and the value of the second objective variable Y′ (predicted value Y′ 1 to predicted value Y' j ) is obtained (Figure 3C).
[0051] For example, when a Gaussian process regression model is used as the regression model 113, the output average value corresponds to the predicted value Y'.
[0052] <<Step S107>> Next, the second organic compound M′ (organic compound M′ 1 to organic compound M' j ) in which the predicted value Y′ 1 to predicted value Y' j Selection is carried out by:
[0053] For example, Bayesian optimization may be used for the selection of organic compounds. For example, when a Gaussian process regression model is used for the regression model 113, a mean value and a variance are output, and an acquisition function may be generated from these values and used as a criterion for the selection of organic compounds. Examples of the acquisition function include the probability of improvement PI (Probability of Improvement) or the expected improvement EI (Expected Improvement).
[0054] Alternatively, for example, the organic compounds may be arranged in descending order of predicted value Y', and quantum chemical calculations may be performed on the top 100 compounds.
[0055] <Step S108> Next, the physical property values (calculated values) of the second organic compound M' selected in step S107 are obtained. The physical property values can be calculated by scientific calculation.
[0056] As the scientific calculation, quantum chemical calculation using software such as Gaussian, Jaguar, or GAMESS can be used.
[0057] In one aspect of the present invention, in step S107, the second organic compound M′ is selected, and the generated second organic compound M′ (organic compound M′ 1 to organic compound M' j This can reduce the enormous cost that would be incurred if quantum chemical calculations were performed on all of the molecules.
[0058] <<Step D101>> Next, the physical property values of the second organic compound M' obtained in step S108 are determined. If an organic compound having the obtained physical property values within the desired range is generated from the second organic compound M', the organic compound is output, and flow chart F100 ends. On the other hand, if an organic compound having the obtained physical property values within the desired range is not generated, the process proceeds to step S109.
[0059] <<Step S109>> If an organic compound within the desired range of physical property values is not obtained in step S101, the data set 120 is updated using information on the selected second organic compound M′. Specifically, the latent variable Z′ obtained in step S104 and the objective variable obtained from the physical property values (calculated values) obtained in step S108 are added to the data set 120.
[0060] In addition, the structural formula of the selected second organic compound M′ and the physical property values (calculated values) obtained in step S 108 may be added to the data set 120 .
[0061] Next, the process returns to step S103, and the regression model 113 is trained again using the updated data set 120. Since the amount of data in the data set 120 is greater than in the previous training, the regression model 113 can make predictions with higher accuracy than in the previous training.
[0062] In this embodiment, the data set 100 and the data set 120 may be treated as a single data set.
[0063] As described above, by repeating steps S101 to S109 and step D101 according to flow chart F100, an organic compound having desired physical property values can be designed.
[0064] As described above, by using the organic compound design system of one embodiment of the present invention, it is possible to efficiently design organic compounds with excellent molecular orientation. In addition, by using the organic compound design system of one embodiment of the present invention, it is possible to efficiently design organic compounds with desired physical properties.
[0065] <<Neural Network>> A predicted value by a neural network can be used as the regression model described above. In the following, a neural network that can be used for supervised learning will be described as an example of a neural network.
[0066] As shown in FIG. 4A, the neural network NN can be composed of an input layer IL, an output layer OL, and a hidden layer HL. Each of the input layer IL, output layer OL, and hidden layer HL has one or more neurons (units). The hidden layer HL may have one layer or two or more layers. A neural network having two or more hidden layers HL can also be called a deep neural network (DNN). Learning using a deep neural network can also be called deep learning.
[0067] Input data is input to each neuron in the input layer IL. An output signal from a neuron in the previous or next layer is input to each neuron in the hidden layer HL. An output signal from a neuron in the previous layer is input to each neuron in the output layer OL. Note that each neuron may be connected to all neurons in the previous or next layer (fully connected), or may be connected to only a portion of the neurons.
[0068] An example of a neuron operation is shown in Figure 4B. Here, neuron N and two neurons in the previous layer that output signals to neuron N are shown. Neuron N receives the output x of the previous layer neuron. 1 and the output of the previous layer neuron x 2 is input. Then, in neuron N, the output x 1 And weight lol 1 The multiplication result (x 1 w 1 ) and output x 2 And weight lol 2 The multiplication result (x 2 w 2 ) the sum x 1 w 1 +x 2 w 2 After is calculated, a bias b is added if necessary, and the value a = x 1 w 1 +x 2 w 2 +b is obtained. The value a is then transformed by the activation function h, and an output signal y=ah is output from the neuron N. As the activation function h, for example, a sigmoid function, a tanh function, a softmax function, a ReLU function, a threshold function, or the like can be used.
[0069] In this way, the operation by a neuron includes the operation of adding the product of the output of the neuron in the previous layer and the weight, that is, the product-sum operation (the above x 1 w 1 +x 2 w 2). This product-sum operation may be performed on software using a program, or may be performed by hardware. When the product-sum operation is performed by hardware, a product-sum operation circuit may be used. This product-sum operation circuit may be a digital circuit or an analog circuit. When an analog circuit is used for the product-sum operation circuit, it is possible to reduce the circuit size of the product-sum operation circuit or the number of memory accesses, thereby improving processing speed and reducing power consumption.
[0070] The product-sum operation circuit may be configured using a transistor containing silicon (such as single crystal silicon) in a channel formation region (hereinafter also referred to as a Si transistor) or a transistor containing an oxide semiconductor in a channel formation region (hereinafter also referred to as an OS transistor). OS transistors have extremely low off-state current and are therefore suitable as transistors constituting the analog memory of the product-sum operation circuit. The product-sum operation circuit may be configured using both Si transistors and OS transistors.
[0071] The above is a description of the neural network. In one aspect of the present invention, it is preferable to use deep learning. In other words, it is preferable to use a neural network having two or more hidden layers HL.
[0072] <<Learning Data Set>> Next, a learning data set used in supervised learning when carrying out the above-described regression model will be described.
[0073] 5A and 5B are diagrams showing the configuration of a training dataset 50. The training dataset 50 includes training data 51_1 to 51_m (m is an integer equal to or greater than 2). The training data 51_1 to 51_m include input data 52_1 to 52_m and teacher data 53_1 to 53_m, respectively. Note that the training data 51_1 to 51_m include information on organic compounds M_1 to M_m and data on the properties of the organic compounds M_1 to M_m, respectively.
[0074] The target of prediction in this embodiment is the properties of the organic compound M. Examples of the properties include emission spectrum, reflection spectrum, absorption spectrum, and transmission spectrum, as well as oxidation potential, reduction potential, glass transition point, melting point, crystallization temperature, and carrier mobility. The properties of the organic compound M can also include, for example, the lowest triplet excitation energy level (T 1 level), the lowest singlet excited energy level (S 1 The energy level of the HOMO is influenced by the energy levels of the HOMO (Highest Occupied Molecular Orbital) and LUMO (Lowest Unoccupied Molecular Orbital), as well as by the overlap of the HOMO and LUMO distributions, the difference in the excited state energy levels, and the rearrangement energy (λ). Therefore, it is preferable to include such data in the training dataset.
[0075] The data input to the autoencoder or the regression model includes information about the above-mentioned organic compound M, data about the properties of the above-mentioned organic compound M, etc. Therefore, the training dataset is generated by extracting, processing, or converting the data input to the autoencoder or the regression model.
[0076] It is preferable that the data contained in the training dataset used in supervised learning be quantified, since quantification of the data contained in the training dataset can prevent the machine learning model from becoming too complex, compared to when the training dataset contains non-numeric data.
[0077] In the training dataset 50 shown in FIG. 5A , input data 52_1 through input data 52_m include information about organic compounds M_1 through M_m, respectively. Furthermore, teacher data 53_1 through 53_m include data on the properties of organic compounds M_1 through M_m, respectively.
[0000] In the training dataset 50 shown in FIG. 5B , input data 52_1 through 52_m include, in addition to information about organic compounds M_1 through M_m, first properties of organic compounds M_1 through M_m, respectively. Furthermore, teacher data 53_1 through 53_m include, in addition to data on the properties of organic compounds M_1 through M_m, second properties of organic compounds M_1 through M_m, respectively.
[0078] As described above, the data on the properties of the organic compounds M_1 to M_m, the data on the first properties of the organic compounds M_1 to M_m, and the data on the second properties of the organic compounds M_1 to M_m are quantified data and can be included in the teacher data 53_1 to 53_m, respectively, without any particular conversion. Note that characteristic points of the data on the properties of the organic compounds M_1 to M_m may be extracted and included in the teacher data 53_1 to 53_m, respectively. Alternatively, points may be extracted so that the values of the control variables are equally spaced and included in the teacher data 53_1 to 53_m, respectively.
[0079] Here, information about the organic compounds M_1 to M_m will be described. First, information about the organic compound M_1 included in the input data 52_1 will be described with reference to FIG.
[0080] The information about the organic compound M includes the molecular structure of the organic compound M. The information about the molecular structure of the organic compound M includes, for example, the ligand, the bonding state, the hole distribution, the electron distribution, and the like.
[0081] As shown in FIG. 5C, for example, the input data 52_1 includes information about the structures of the organic compounds M_1 to M_m as information about the organic compounds M_1 to M_m.
[0082] It is preferable that the data contained in the learning dataset used in supervised learning be digitized.
[0083] Furthermore, in the data included in the training dataset, information about the organic compound M is often input as non-numeric data, such as a structural formula, a molecular structure expressed in SMILES notation (sometimes simply referred to as molecular structure), a name according to the nomenclature of compounds determined by IUPAC, etc. Therefore, it is preferable to digitize information about the organic compound that is input as non-numeric data (for example, molecular structure).
[0084] One way to quantify a molecular structure is to replace it with the physical properties of an organic compound represented by the molecular structure. Examples of the physical properties of an organic compound include an emission spectrum, an absorption spectrum, a transmission spectrum, a reflection spectrum, an S1 level, a T1 level, an oxidation potential, a reduction potential, a HOMO level, a LUMO level, a glass transition point, a melting point, a crystallization temperature, and carrier mobility. These physical properties can be treated as numerical values, but since measurements and simulations are required, generating a training dataset requires a great deal of effort.
[0085] Therefore, in one aspect of the present invention, the molecular structure that specifies the organic compound M is converted using a certain method. Note that the certain method only needs to be able to express molecular similarity. Well-known methods for expressing molecular similarity include quantitative structure-activity relationship (QSAR), fingerprinting, and graph structures. For example, it is preferable to digitize the molecular structure using linear notation or matrix notation. By digitizing the molecular structure that specifies the organic compound M, it can be used as training data.
[0086] When quantifying the structure of an inorganic compound, it is advisable to use descriptors such as the radial distribution function (RDF) and the orbital field matrix (OFM).
[0087] <<Example of Numerical Expression of Molecular Structure of Organic Compound>> Here, an example of numerical expression (mathematical expression) of the molecular structure of an organic compound will be described.
[0088] When information about organic compounds is input as non-numeric data other than SMILES notation, it is preferable to first convert it to SMILES notation. SMILES notation expresses organic compounds as a continuous string of characters, making it preferable for computer-handled data. SMILES notation and the fingerprinting method described below are both classified as linear notations, making them easy to convert between, making them preferable.
[0089] The open-source cheminformatics toolkit RDKit can be used to quantify molecular structures. RDKit can convert the SMILES notation of the input molecular structure into mathematical data (quantification) using the fingerprinting method.
[0090] In fingerprinting, a molecular structure is represented by assigning a partial structure (fragment) of the molecular structure to each bit, as shown in FIG. 6, and if the corresponding partial structure exists in the molecule, the bit is set to "1", and if not, "0". In other words, by using fingerprinting, it is possible to obtain a mathematical formula that extracts the characteristics of the molecular structure. Furthermore, the molecular structure formula represented by fingerprinting is generally several hundred to tens of thousands of bits long, which is an easy-to-handle size. Furthermore, by using fingerprinting to represent a molecular structure as a mathematical formula of 0s and 1s, it is possible to achieve extremely high-speed calculation processing.
[0091] In addition, there are many types of fingerprinting methods (those that take into account different bit generation algorithms, atom types, bond types, and aromaticity conditions, and those that dynamically generate bit lengths using hash functions, etc.), each with its own characteristics.
[0092] 7A to 7D show examples of types of fingerprinting. Typical types of fingerprinting include the circular type shown in Fig. 7A (surrounding atoms up to a specified radius from a starting atom are considered to be the partial structure), the path-based type shown in Fig. 7B (atoms from the starting atom up to a specified path length are considered to be the partial structure), the substance keys type shown in Fig. 7C (partial structures are defined for each bit), and the atom pair type shown in Fig. 7D (atom pairs generated for all atoms in a molecule are considered to be the partial structure). Fingerprints of each of these types are implemented in RDKit.
[0093] Figure 8 shows an example of the molecular structure of an organic compound actually expressed as a mathematical formula using the fingerprint method. In this way, the molecular structure can be converted into a SMILES notation and then converted into a fingerprint.
[0094] When expressing the molecular structure of an organic compound using a fingerprinting method, the resulting mathematical formula may be identical between different organic compounds with similar structures. As described above, there are several types of fingerprinting methods depending on the notation method, but the tendency for compounds to be identical varies depending on the notation method, as shown in Figure 9, which includes the circular type (Morgan Fingerprint), path-based type (RDK Fingerprint), substance keys type (Avalon Fingerprint), and atom pair type (Hash atom pair). In Figure 9, the molecules within each double arrow each represent the same mathematical formula (notation). Therefore, it is preferable to use a fingerprinting method in which, when the molecular structure of each organic compound to be learned is expressed using at least one of the fingerprinting methods, the notations for each organic compound are all different. In FIG. 9, it can be seen that compounds with different atom pair types can be expressed without overlap, but depending on the population of organic compounds to be learned, other expression methods may also be possible without overlap.
[0095] Therefore, when describing the molecular structure of an organic compound using a fingerprinting method, it is preferable to use multiple different types of fingerprinting methods. Any number of types may be used, but two or three types are preferred because they are easier to handle in terms of data volume. When using multiple types of fingerprinting methods, the molecular structure of an organic compound may be described by connecting a formula represented by one type of fingerprinting method to a formula represented by another type of fingerprinting method, or by describing a single organic compound as having multiple different formulas. Figures 10A and 10B show an example of a method for describing the molecular structure of an organic compound using multiple fingerprints of different types.
[0096] Fingerprints are a method of describing the presence or absence of substructures, and information about the overall molecular structure is lost. However, if a molecular structure is expressed mathematically using multiple fingerprints of different types, different substructures are generated for each fingerprint type, and information about the presence or absence of these substructures can be used to complement information about the overall molecular structure. If a feature that cannot be fully expressed by a fingerprint has a significant impact on the characteristics of a light-emitting device, it can be complemented by other fingerprints, making the method of describing a molecular structure using multiple fingerprints of different types effective.
[0097] As shown in FIG. 10A, when representation is performed using two types of fingerprint methods, it is preferable to use the Substructure Keys type and the Circular type, as this allows for accurate prediction of physical properties.
[0098] Furthermore, as shown in FIG. 10B , when notating using three types of fingerprint methods, it is preferable to use the Atom Pair type, the Circular type, and the Substructure Keys type, as these enable accurate prediction of physical properties.
[0099] Furthermore, when a circular fingerprinting method is used, the radius r is preferably 3 or more, and more preferably 5 or more. Note that the radius r is the number of atoms counted in connection from a certain atom that serves as the starting point, which is set as 0.
[0100] When selecting a fingerprinting method to be used, as mentioned above, it is preferable to select at least one method in which the molecular structures of the organic compounds are all represented differently.
[0101] By increasing the bit length (number of bits) used to represent a fingerprint, the possibility of generating descriptions that completely match between organic compounds can be reduced; however, increasing the bit length too much results in a trade-off in increased computational costs or database management costs. On the other hand, by simultaneously using multiple fingerprints for representation, even if there are multiple molecular structures whose representations are completely identical in a certain fingerprint type, combining different fingerprint types may result in an overall incomplete match. As a result, a state in which multiple organic compounds whose representations are completely identical in fingerprints can be generated with as small a bit length as possible. There is no particular limit to the bit length of the fingerprint to be generated; however, considering computational costs or database management costs, for molecules with molecular weights of up to about 2000, fingerprints that do not completely match between organic compounds can be generated with a bit length of 4096 or less, preferably 2048 or less, and in some cases 1024 or less for each fingerprint type.
[0102] Furthermore, the bit length of the fingerprints generated for each fingerprint type can be adjusted appropriately taking into account the characteristics of the type and the overall molecular structure, and does not need to be unified. For example, the bit length may be expressed as 1024 bits for the Atom Pair type and 2048 bits for the Circular type, and then these may be concatenated.
[0103] The above is an explanation of how to quantify the molecular structure of an organic compound.
[0104] As described above, one embodiment of the present invention can provide a learning method, a design method, and a design system for a design model of a molecular structure of an organic compound having desired properties.
[0105] One aspect of the present invention makes it possible to predict the properties of the molecular structure of a novel organic compound. Furthermore, by using past experimental data, it is possible to speed up the process by optimizing the molecular structure of an organic compound with desired properties. Even if a human observes the data and finds it to be non-interpolative, the non-linear or high-order representation of the machine learning model may make it so. Furthermore, by extracting and examining fragments of the representation obtained by the machine learning model, it is possible to discover patterns that would not have been previously noticed.
[0106] This embodiment can be carried out by combining parts thereof as appropriate.
[0107] Embodiment 2 In Embodiment 2, a system for designing an organic compound, which is one embodiment of the present invention, will be described.
[0108] A design system 20 for an organic compound according to one embodiment of the present invention includes at least an input unit, a learning unit, a prediction unit, an output unit, and a data server. These units may be incorporated into a single device, or may be separate devices, or some may be incorporated into the same device, or the data server may be a cloud, as long as they are capable of exchanging data. However, these units are collectively referred to as a property prediction system.
[0109] 11, one embodiment of the present invention will be described by taking as an example an organic compound design system comprising an information terminal having input means, learning means, prediction means, and output means, and a data server 28. The information terminal 21 has at least an input section, learning means, prediction means, and output section, and is capable of exchanging data with the separately installed data server 28. Note that the data server 28 may be installed within the information terminal 21.
[0110] The information terminal 21 mainly comprises an input unit 22, a calculation unit 24, and an output unit 26. The calculation unit 24 functions as both a learning means and a prediction means. Preferably, the calculation unit 24 also includes a neural network circuit. Data provided from the data server 28 is used as data for learning or prediction by the neural network circuit. By using a portion of this data as verification data and training data for the learned learning means, it is possible to update the weight coefficients in the neural network circuit and generate learned weight coefficients. This can further improve the accuracy of prediction.
[0111] 11, the flow of signals is shown by arrows (Din: data input, Dout: data output) in the order of the input unit 22, the calculation unit 24, the data server 28, and the output unit 26. In this specification, the term "signal" can be interpreted as "data" or "information" as appropriate.
[0112] The data server 28 provides the structures and physical property values of the organic compounds to be learned to the learning means of the calculation unit 24. The structures of the organic compounds provided are represented using two or more types of fingerprints. The learning means of the calculation unit 24 preferably has a neural network circuit.
[0113] The input unit 22 has a function for the user to input information. Specific examples of the input unit 22 include any input means such as a keyboard, a mouse, a touch panel, a pen tablet, a microphone, or a camera.
[0114] Input information D in is the data output from the input unit 22 to the calculation unit 24. in is information input by the user. For example, if the input unit 22 is a touch panel, it is information obtained by inputting characters by operating the touch panel. Alternatively, if the input unit 22 is a microphone, it is information obtained by voice input by the user. Alternatively, if the input unit 22 is a camera, it is information obtained by image processing of captured data.
[0115] Input information D inis information about the structure of the organic compound whose physical properties are to be predicted. If a structural formula, structural image, substance name, or other information other than a fingerprint notation is input, it is input to the prediction means in the calculation unit 24 after passing through an appropriate conversion means. The prediction means predicts the physical properties of the input organic compound based on the results learned in advance by the learning means.
[0116] The result of the prediction is output via the output unit 26 .
[0117] [Data Server 28] The data server 28 has a function of storing the program executed by the calculation unit 24. The data server 28 may also have a function of storing data generated by the calculation unit 24 (e.g., calculation results, analysis results, inference results), data input to the input unit 22, etc.
[0118] The data server 28 may have a database. Furthermore, the design system 20 may have a database separate from the data server 28. The design system 20 may have a function to retrieve data from a database that exists outside the data server 28 or outside the design system 20. Furthermore, the design system 20 may have a function to retrieve data from both its own database and an external database.
[0119] A file server may be used instead of the database. For example, when using files stored in a file server, it is preferable that the database has paths to files stored in the file server.
[0120] The data server 28 includes at least one of a volatile memory and a non-volatile memory. Examples of the volatile memory include a dynamic random access memory (DRAM) and a static random access memory (SRAM). Examples of the non-volatile memory include a resistive random access memory (ReRAM), a phase change random access memory (PRAM), a ferroelectric random access memory (FeRAM), a magnetoresistive random access memory (MRAM), and a flash memory. The data server 28 may also have a recording media drive, such as a hard disk drive (HDD) or a solid state drive (SSD).
[0121] [Calculation Unit 24] The calculation unit 24 may include, for example, an arithmetic circuit, a central processing unit (CPU), or a graphics processing unit (GPU).
[0122] The calculation unit 24 may have a microprocessor such as a DSP (Digital Signal Processor). The microprocessor may be implemented by a PLD (Programmable Logic Device) such as an FPGA (Field Programmable Gate Array) or an FPAA (Field Programmable Analog Array). The calculation unit 24 may also have a quantum processor. The calculation unit 24 can perform various data processing and program control by interpreting and executing instructions from various programs using the processor. Programs that can be executed by the processor are stored in at least one of the memory area of the processor and the data server 28.
[0123] When the arithmetic unit includes a neural network circuit, the neural network circuit preferably includes a product-sum operation circuit capable of performing product-sum operation. The product-sum operation circuit preferably includes a memory circuit for storing weight data. The memory element included in the memory circuit includes a transistor and a capacitor. The transistor is preferably a transistor including an oxide semiconductor in a semiconductor layer having a channel formation region (hereinafter referred to as an OS transistor). The OS transistor has an extremely small leakage current when it is off. Therefore, data can be stored by utilizing the property of being able to retain charge by turning the OS transistor off.
[0124] Another aspect of the present invention is a recording medium on which a control program and control software are recorded that can generate fingerprints that are concatenated or parallelized using these multiple fingerprint types, perform machine learning, and predict physical properties.
[0125] Embodiment 3 In this embodiment, a molecular orientation parameter will be described as a physical property value that can be used in a design system for an organic compound, which is one embodiment of the present invention.
[0126] Examples of organic compounds or organometallic complexes that constitute light-emitting devices with good luminous efficiency include substances in which the dot product of vector X, which connects the two most distant atoms in an excited state, and vector D, which is the transition dipole moment, is 2.5 or more, more preferably 4.0 or more. The direction of vector X is determined so that the angle θ between vector X and vector D is 90° or less. The unit of vector D is debye, and the unit of vector X is nm to align the magnitude of the vectors. The excited state is the lowest excited state, and the transition dipole moment is large in the lowest excited state, so that it affects the shape of the emission spectrum.
[0127] The direction of vector X connecting the two most distant atoms in the excited state is parallel to the direction of the longest side of the molecule (also called the long side direction). A small angle θ between this vector X and the transition dipole moment vector D is advantageous for molecular orientation. A large transition dipole moment is also advantageous for improving the luminescence quantum yield.
[0128] Therefore, an organic compound or organometallic complex having these characteristics has a large inner product of the vector X connecting the two most distant atoms in the excited state and the transition dipole moment vector D. When this value is 2.5 or greater, preferably 4.0 or greater, in the organic compound or organometallic complex, a light-emitting device or light-emitting device material containing the organic compound or organometallic complex can have good luminous efficiency and high reliability.
[0129] As mentioned above, it is known that light emission from organic compounds or organometallic complexes occurs in a direction perpendicular to the transition dipole of the molecule. In a molecule, the direction of the vector connecting the two most distant atoms (long side direction) is more likely to be aligned horizontally to the film surface when the molecule is formed into a film than other directions. Therefore, it is preferable that the angle θ between vector D, which is the transition dipole moment vector, and vector X is small.
[0130] Although the vector connecting the two most distant atoms in the ground state and the vector X connecting the two most distant atoms in the excited state are different vectors, a significant difference in direction is unlikely to occur, and therefore, the vector X can be used as an index of the ease of horizontal orientation with respect to the film formation surface.
[0131] Furthermore, the transition dipole moment indicates the ease of transition between two electronic states, and the larger the value, the easier the transition is. Therefore, it is advantageous for improving the luminescence quantum yield, and it is preferable that the value of the vector D is large.
[0132] Therefore, it is preferable that the dot product of vector X and vector D is large, and by using a material for a light-emitting device or a material for a light-emitting apparatus containing an organic compound or an organometallic complex whose dot product is 2.5 or more, more preferably 4.0 or more, light can be extracted more efficiently.
[0133] In addition, when there are multiple vectors connecting the two most distant atoms in the excited state in the same organic compound or organometallic complex, the vector that forms the smaller angle θ with vector D is regarded as vector X.
[0134] Furthermore, the organic compound or organometallic complex contained in the material for a light-emitting device or a material for a light-emitting device is intended to have the function of emitting light in the light-emitting device or the light-emitting device. Examples of materials that have the function of emitting light include a light-emitting center material and a color conversion material. The luminescence quantum yield of the organic compound or organometallic complex is preferably 0.60 or more, more preferably 0.70 or more. The luminescence quantum yield is preferably measured by forming a film on a quartz substrate using a drop-cast method, dispersing each material in PMMA (poly(methyl methacrylate)) at an appropriate concentration (e.g., 4.8 wt%) in deoxygenated dichloromethane as a solvent, and then drying the film under a nitrogen stream in a glove box (e.g., at room temperature for 30 minutes), followed by measuring the resulting PMMA (poly(methyl methacrylate)) film.
[0135] Furthermore, the material for a light-emitting device or a light-emitting apparatus may be composed solely of the organic compound or organometallic complex, or may contain other substances.
[0136] 12A is a diagram illustrating a light-emitting device of one embodiment of the present invention. The light-emitting device of one embodiment of the present invention includes a first electrode 1101, a second electrode 1102, and an EL layer 1103 over an insulating layer 1000. The EL layer 1103 includes a light-emitting layer 1113. Note that the EL layer 1103 may include other functional layers such as a hole-injection layer 1111, a hole-transport layer 1112, an electron-transport layer 1114, and an electron-injection layer 1115.
[0137] The light-emitting device of one embodiment of the present invention includes the above-described material for a light-emitting device or a material for a light-emitting device in the light-emitting layer 1113. Preferably, the light-emitting layer 1113 further includes a host material, and the material for a light-emitting device or a material for a light-emitting device is dispersed in the host material. Note that the host material may be composed of a plurality of organic compounds. Furthermore, an organic compound that functions as a host material may be included in the material for a light-emitting device or a material for a light-emitting device.
[0138] The light-emitting device of one embodiment of the present invention having the above-described structure contains, as a luminescent center substance, an organic compound or an organometallic complex in which the dot product of a vector X connecting the two most distant atoms in an excited state and a vector D of a transition dipole moment is 2.5 or more, preferably 4.0 or more, and therefore can be a light-emitting device with high luminous efficiency and high reliability.
[0139] Furthermore, when the molecular orientation parameter a of the light emitted by the light-emitting device is 0.25 or less, preferably 0.23 or less, the light has good orientation characteristics, making it easy to extract light and enabling a light-emitting device with good efficiency. That is, a light-emitting device is more preferred in which the light-emitting layer contains, as a luminescent center substance, an organic compound or an organometallic complex in which the dot product of vector X and vector D is 2.5 or more, more preferably 4.0 or more, and the molecular orientation parameter a of the light emitted by the light-emitting device is 0.25 or less, preferably 0.23 or less.
[0140] Furthermore, when used in a light-emitting layer, a light-emitting device can be provided in which the molecular orientation parameter a of light is 0.25 or less, preferably 0.23 or less, and the dot product of vector X and vector D is 2.5 or more, more preferably 4.0 or more, and the material for a light-emitting apparatus or light-emitting device contains an organic compound or an organometallic complex.
[0141] The molecular orientation parameter a is a value that estimates the molecular orientation from the light-emitting state of the device. The radiation angle dependence (spatial emission pattern) of the light-emitting device's emission intensity reflects the spatial distribution of the transition dipole of the luminescent center substance. The orientation state of the light-emitting device can be investigated by analyzing this spatial distribution. This method observes and analyzes the light emission itself of the light-emitting device, so as long as the luminescent center substance is emitting light, it is possible to investigate the orientation state of the luminescent center substance in the light-emitting layer in terms of the relationship between the light-emitting surface and the transition dipole moment, even if the luminescent center substance is dispersed in a host material and its concentration is low.
[0142] Therefore, a light emitting device using the above material for a light emitting apparatus or a material for a light emitting device in the light emitting layer and having a molecular orientation parameter a of 0.25 or less, preferably 0.23 or less, can be a light emitting device with good efficiency.
[0143] <Method of Determining the Scalar Product of Vector X and Vector D> Using a platinum complex as an example of an organic compound or organometallic complex contained in a material for a light-emitting device or a material for a light-emitting apparatus, a method of determining the scalar product of vector X connecting the two most distant atoms in the atomic configuration in an excited state and vector D of the transition dipole moment will be described.
[0144] Here, an example is shown in which the dot product of vector X and vector D is calculated for three platinum complexes, platinum complex 1, platinum complex 2, and platinum complex 3, which are represented by the following structural formulas.
[0145] It is known that a highly planar structure such as that of the organometallic complex represented by Platinum Complex 1 is advantageous for molecular orientation. However, Platinum Complex 2 and Platinum Complex 3, represented by the following structural formulae, have been found to realize light-emitting devices that exhibit better luminous efficiency than the platinum complex, despite not being highly planar.
[0146]
[0147] The structure for quantum chemical calculation was sampled using the Maestro GUI manufactured by Schrödinger, with conformational analysis performed using Macro Model. Using the quantum chemical calculation software Jaguar, the most stable structure in the singlet ground state was calculated using density functional theory (DFT), and the structure of the most stable conformation was determined. In this structure, the basis functions used were DYALL-2ZCVP_ZORA-J-PT-GEN++ for the Pt atom and LACVP** for the other atoms, and the functional was ωB97X-D (ω = 0.1). The triplet state was calculated as the excited state using time-dependent density functional theory (TD-DFT) with the spin-free ZORA relativistic Hamiltonian, and the most stable structure was obtained. For this structure, a single-point energy calculation was performed on the excited state using the spin-orbit ZORA relativistic Hamiltonian, and the transition dipole moment vector D was visualized. Furthermore, the vector X connecting the two most distant atoms in this structure was determined so that the angle between it and vector D was 90° or less, and the angle between it and vector D was calculated. As an example, a diagram showing the angles between each vector in platinum complex 2 is shown in Figure 12B. The results are also shown in Table 1.
[0148]
[0149] It has been found that platinum complexes 2 and 3 have a molecular orientation parameter a obtained through experiments described below, which is superior to platinum complex 1. This is thought to be correlated with the difference in the dot product between vector X and vector D, as shown in Table 1, and platinum complexes 2 and 3 have structures in which the transition dipole moment greatly contributes to improving the luminous efficiency. Therefore, a light-emitting device or light-emitting device using a material for a light-emitting device containing platinum complex 2 or platinum complex 3 can be a light-emitting device or light-emitting device with good luminous efficiency.
[0150] Platinum complex 1 has a small inner product of vector X and vector D, and it is believed that the transition dipole moment makes little contribution to improving the luminous efficiency.
[0151] In this way, a light emitting apparatus or a light emitting device using a material for a light emitting device having an organic compound or an organometallic complex in which the inner product of the vector X and the vector D is large can be a light emitting apparatus or a light emitting device with good efficiency.
[0152] The organic compound or organometallic complex contained in the material for a light-emitting device or a material for a light-emitting apparatus is preferably an organometallic complex, since it exhibits high phosphorescence efficiency due to a fast intersystem crossing process between the singlet state and the triplet state. Furthermore, the organometallic complex is preferably a cyclometallic complex, since it forms a strong carbon-metal bond and exhibits high phosphorescence efficiency due to metal-ligand charge transfer (MLCT) in the excited state.
[0153] In addition, it is preferable that the organometallic complex has a ring formed by the contained metal and some of the atoms contained in the ligand, since this increases the planarity of the organometallic complex. This is also preferable because the direction of the long side between the ground state and the excited state does not change significantly. The ring is preferably a six-membered or five-membered ring, which is stable. It is also preferable that the organometallic complex contains multiple such rings, since this increases the planarity and reduces the change in the direction of the long side between the ground state and the excited state.
[0154] Since the platinum complex is tetradentate, the molecular planarity is easily maintained, and the inner product of vector X and vector D tends to become large, which is preferable.
[0155] <Method for calculating molecular orientation parameter a> Next, a method for calculating the molecular orientation parameter a will be described. By comparing the angular dependence of the measured emission intensity of a light-emitting device with the calculated value of the angular dependence of the emission intensity calculated using a device simulator assuming a parameter a (see formula (3) below) that represents the orientation of the light-emitting molecules, it is possible to estimate a reasonable value for the molecular orientation parameter a for the measured light-emitting device and investigate the orientation state of the luminescence center substance in the light-emitting device (see non-patent document 2).
[0156] The inventors also focused on the shape of the emission spectrum obtained from the device simulator, and compared the measured and calculated values for the emission spectrum shape and the change in the shape of the emission spectrum depending on the angle, and performed a match. Furthermore, the emission intensity in the actual measurements and simulations was the area intensity of the emission spectrum, rather than the emission intensity of a specific wavelength. These newly applied techniques by the inventors enable highly accurate estimation of the parameter a, unlike Non-Patent Document 2.
[0157] 13 shows the relationship between the observation direction of the measuring instrument in measuring the spatial distribution of luminescence intensity and each vector component of the transition dipole moment on the substrate. Because the transition dipole moment is a vector, it can be combined and decomposed, and the average transition dipole moment of the luminescent center substance in the luminescent layer can be decomposed into the mutually orthogonal components in the x-axis direction (TEh component), the y-axis direction (TMh component), and the z-axis direction (TMv component).
[0158] Here, as mentioned above, it is known that light emission from molecules is emitted in a direction perpendicular to the transition dipole moment (any direction within a perpendicular plane). Of the components divided in the above three directions, the TEh component and the TMh component (x-axis direction and y-axis direction) have transition dipole moments parallel to the substrate surface, so their emission direction is perpendicular to the substrate, and they can be said to be components that exhibit light emission that is easy to extract. On the other hand, the TMv component (z-axis direction) has a transition dipole moment perpendicular to the substrate surface, so its emission direction is parallel to the substrate, and they are components that exhibit light emission that is difficult to extract.
[0159] In FIG. 13, the arrows pointing out from the center of the vectors of each component indicate the direction of the detector relative to the front of the substrate (θ D = 0 degrees) to almost horizontal with the substrate (detector angle θ D = 90 degrees), the linear distance from the center is proportional to the intensity.
[0160] For the TEh component, the detector is located in the direction of light emission, so even if the angle of the substrate is changed, the detected light intensity (i.e., the linear distance from the center of the arrow in the figure) is constant, and the figure emanating from the center of the arrow in the figure shows a neat fan shape. On the other hand, for the TMh and TMv components, the figures emanating from the center of the arrow in the figure are distorted, and the angle θ of the detector relative to the substrate D As shown in the figure, the TMh component changes significantly depending on the angle θ of the detector. D The intensity is strong in the observation in the small area (close to the front of the substrate), and the TMv component is D The intensity is stronger when observed in a region where the angle θ of the detector is large (direction angled relative to the substrate). D Emission intensity with respect to wavelength λ at λ (θ D , λ) can be expressed as equation (3).
[0161]
[0162] In the formula I TMv , I TMh , I TEh represents the spatial intensity distribution of light emitted from the transition dipoles arranged as shown in Figure 13, where a represents the proportion of transition dipoles arranged perpendicular to the film surface (TMv components). On the other hand, 1-a represents the proportion of transition dipoles arranged horizontally (TMh components, TEh components). In other words, a can be considered as a parameter representing the orientation of the transition dipoles of the luminescent molecules.
[0163] In the formula, if the transition dipoles are arranged only in a direction completely horizontal to the substrate, the TMv component will disappear, and a = 0. On the other hand, if the transition dipoles are arranged only in a direction perpendicular to the substrate, a = 1. Furthermore, if the transition dipoles are oriented randomly, the transition dipoles are considered to be oriented isotropically with a ratio of 1:1:1 relative to the x-axis, y-axis, and z-axis. Therefore, the ratio of the components perpendicular to the substrate (TMv component) to the components parallel to the substrate (TMh component and TEh component) is 1:2, and a = 1 / 3 (approximately 0.33).
[0164] Here, as mentioned above, ITEh The intensity of is constant regardless of the angle, but I TMv , I TMh is the angle of the substrate relative to the measuring instrument (detector angle θ D ) changes its size depending on the detector angle θ D By measuring the emission intensity while changing the angle θ D The value of a can be found from the change with respect to
[0165] In addition, the intensity does not change depending on the angle. TEh However, the amplitude direction of the electric field of the emitted light is the same as the direction of the transition dipole moment, so I TEh is S wave, I TMv , I TMh Since is a P wave, it is possible to measure it by excluding the TEh component by inserting a linear polarizer in the direction perpendicular to the substrate surface.
[0166] The structure of this embodiment mode can be used in appropriate combination with other structures.
[0167] F100 flowchart, S101: step, S102: step, S103: step, S104: step, S105: step, S106: step, S107: step, S108: step, S109: step, D101: step, 20: design system, 21: information terminal, 22: input unit, 24: calculation unit, 26: output unit, 28: data server, 50: learning data set, 51_1: learning data, 51_m: learning data, 52_1: input data, 52_m: input data, 53_1: teacher data, 53_m: teacher data, 1000: insulating layer, 1101: first electrode, 1102: second electrode, 1103: EL layer, 1113: light-emitting layer
Claims
setting a first group of objective variables from a first group of explanatory variables for a first group of organic compounds; A step of learning the molecular structure of the first organic compound group using an autoencoder to obtain a first latent variable set corresponding to the molecular structure of the first organic compound group; A step of training a regression model on the correlation between the first latent variable group and the first objective variable group; generating a second set of latent variables using random numbers; obtaining molecular structures of a second set of organic compounds from the second set of latent variables using the autoencoder; obtaining predicted values of a second set of objective variables from the second set of latent variables using the regression model; selecting a third group of organic compounds from the second group of organic compounds using the predicted values of the second group of dependent variables; and calculating a second group of explanatory variables for the third group of organic compounds; A method for designing an organic compound, wherein the first group of explanatory variables and the second group of explanatory variables have information regarding molecular orientation.
2. The method for designing an organic compound according to claim 1 , wherein the first group of explanatory variables and the second group of explanatory variables respectively include a length of a major axis of a molecule in the organic compound, a magnitude of a transition dipole moment in the organic compound, and an angle between the major axis and the transition dipole moment.
3. The method for designing an organic compound according to claim 1, wherein the first group of organic compounds to the third group of organic compounds are organometallic complexes having a central metal with a coordination number of tetradentate. An input unit, a calculation unit, and an output unit, the input unit has a function of inputting a molecular structure of a first group of organic compounds and a first group of explanatory variables; The calculation unit is A function of setting a first group of objective variables from the first group of explanatory variables in the first group of organic compounds; A function of learning the molecular structure of the first organic compound group using an autoencoder to obtain a first latent variable set corresponding to the molecular structure of the first organic compound group; A function of training a regression model on the correlation between the first latent variable group and the first objective variable group; A function for generating a second set of latent variables using random numbers; a function of acquiring a molecular structure of a second group of organic compounds from the second group of latent variables using the autoencoder; A function of obtaining predicted values of a second target variable group from the second latent variable group using the regression model; a function of selecting a third group of organic compounds from the second group of organic compounds by using the predicted values of the second group of objective variables; and a function of calculating a second group of explanatory variables for the third group of organic compounds; the output unit has a function of outputting a physical property value related to a molecular structure of the second organic compound or a molecular orientation of the second organic compound group; A system for designing an organic compound, wherein the first group of explanatory variables and the second group of explanatory variables have information regarding molecular orientation.
5. The system for designing an organic compound according to claim 4 , wherein the first group of explanatory variables and the second group of explanatory variables respectively include a length of a major axis of a molecule in the organic compound, a magnitude of a transition dipole moment in the organic compound, and an angle between the major axis and the transition dipole moment.
6. The system for designing organic compounds according to claim 4 or claim 5, wherein the autoencoder is JT-VAE.
6. The system for designing an organic compound according to claim 4 or 5, wherein the organic compound is an organometallic complex having a central metal with a coordination number of four.
Citation Information
Patent Citations
Chemical compound generation device, chemical compound generation method, learning device, learning method, and program
JP2021068410A
Computer-implemented method, system, and computer program product for training molecule generative model, and computer-implemented method and system for generating molecules and number of substructures of molecules (interpretable molecular generative models)
JP2022094334A
Method for simultaneous characterization and expansion of reference libraries for small molecule identification
US20200176087A1
Graphic user interface assisted chemical structure generation
US20200272702A1
Computer implemented method and system for small molecule drug discovery
US20230317202A1