Material design optimization method, system, and medium
Patent Information
- Application Number
- CN202410149280.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-01
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2044-02-01
AI Technical Summary
[0005]本发明的主要目地在于提供一种材料设计优化方法、装置、系统及介质,旨在解决常规的筛选方案筛选得到的目标材料对结构和机理复杂、分子量庞大的材料进行设计优化的效率较低的技术问题
[0036]This invention proposes a material design optimization method, system, and medium. By sampling experimental materials according to a preset sampling method, multiple specific substructures are obtained, enabling direct analysis based on the structure with the highest mechanistic relevance. This reduces interference from non-critical structures present in materials with high molecular weight and strong mechanistic structures, improving the accuracy of prediction results. The multiple specific substructures are input into a relationship exploration model to obtain target data, and the target material is designed and optimized based on this target data. The steps include: inputting multiple specific substructures into the relationship exploration model to obtain predicted property values of the experimental material, the importance of each specific substructure to the predicted property values, the interaction relationships between the specific substructures, and the attention distribution characteristics of each specific substructure to the predicted property values; determining the target specific substructure from the multiple specific substructures based on the importance, interaction relationships, and attention distribution characteristics, and then designing and optimizing the target material based on the target specific substructure. By acquiring crucial information for adjusting experiments and mechanistic explorations—namely, the importance of each specific substructure and their interactions—further analysis based on importance and interactions reveals connections between intermediate variables related to higher-level mechanisms. Additionally, data directly applicable to materials exploration and research, such as attention distribution characteristics, is obtained. Analysis of specific substructures based on these attention distribution characteristics enhances model interpretability. This, combined with crucial information for adjusting experiments and mechanistic explorations and data directly applicable to materials exploration and research, improves the efficiency of materials design optimization.
Smart Images

Figure CN117831685B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of materials design optimization technology, and in particular to a materials design optimization method, system and medium. Background Technology
[0002] Virtual screening based on molecular structure has been widely used in new materials exploration and drug discovery. Its main purpose is to screen high-potential target materials from a large pool of candidate materials that meet specific properties. The screening process involves predicting target properties using mathematical models based on ab initio computation or machine learning / deep learning methods. Molecules with the best properties are then experimentally validated to obtain the desired target materials.
[0003] However, the aforementioned screening scheme is only applicable to target materials with simple structures and mechanisms and small molecular weights. For target materials with complex structures and mechanisms and large molecular weights, the models may struggle to capture the key structures related to the material properties due to the complexity of the mechanism and the large molecular weight, making them unsuitable for direct use in material exploration and research. Furthermore, the accuracy of prediction results may decrease. Moreover, because the machine learning and deep learning methods currently used for material property prediction are black-box mathematical models, they cannot obtain crucial information for adjustments to experiments and mechanism exploration.
[0004] Therefore, conventional screening methods are less efficient at designing and optimizing target materials with complex structures and mechanisms and large molecular weights. Summary of the Invention
[0005] The main objective of this invention is to provide a material design optimization method, apparatus, system, and medium, which aims to solve the technical problem that conventional screening schemes have low efficiency in designing and optimizing target materials with complex structures and mechanisms and large molecular weights.
[0006] To achieve the above objectives, the present invention provides a material design optimization method, which includes the following steps:
[0007] The experimental materials were sampled according to a preset sampling method to obtain multiple specific substructures;
[0008] The target data is obtained by inputting the multiple specific substructures into the relationship exploration model, and the target material is designed and optimized based on the target data;
[0009] The step of inputting the multiple specific substructures into a relationship exploration model to obtain target data, and then optimizing the design of the target material based on the target data, includes:
[0010] By inputting the multiple specific substructures into the relationship exploration model, the predicted values of the experimental material's properties, the importance of each specific substructure to the predicted values of the properties, the interaction relationships between each specific substructure, and the attention distribution characteristics of each specific substructure to the predicted values of the properties are obtained.
[0011] Based on the importance, the interaction relationship, and the attention distribution characteristics, after determining the target specific substructure from the plurality of specific substructures, the target material is designed and optimized according to the target specific substructure.
[0012] Optionally, the step of sampling the experimental material according to a preset sampling method to obtain multiple specific substructures includes:
[0013] The key physicochemical elements of the experimental material are obtained, and the positioning information corresponding to the key physicochemical elements is obtained based on the key physicochemical elements. The key physicochemical elements are located according to the positioning information to obtain the central unit of the key physicochemical elements, and then the preset iterative sampling process is entered.
[0014] During the preset iterative sampling process, after selecting adjacent units of the key physicochemical elements whose distance from the central unit is less than or equal to the expected distance as physicochemical units, based on the physicochemical units and the central unit, new physicochemical units whose distance from the key physicochemical elements is less than or equal to the expected distance from the physicochemical units or the central unit are selected, until the preset iterative sampling process ends;
[0015] The specific substructure is formed based on the physicochemical relationship between the selected physicochemical unit and the central unit during the preset iterative sampling process.
[0016] Optionally, the step of inputting the plurality of specific substructures into the relationship exploration model to obtain the property prediction values of the experimental material includes:
[0017] After obtaining the initial codes corresponding to each specific substructure based on the preset computation theory, the efficient representations corresponding to each initial code are determined according to the preset coding model.
[0018] The multiple efficient representations are input into the structure-attribute model in the relationship exploration model, and the predicted attribute values are output.
[0019] Optionally, when selecting the self-attention model in the relationship exploration model to output the attribute prediction value, the step of inputting the multiple efficient representations into the relationship exploration model and outputting the attribute prediction value includes:
[0020] The multiple efficient representations are concatenated and then input into the self-attention model to obtain the weight coefficients corresponding to the multiple specific substructures respectively;
[0021] After normalizing the weight coefficients using a preset function to obtain the attention weights corresponding to the multiple specific substructures, the attention weights are fused and merged to obtain the output vector.
[0022] The output vector is input into the self-attention model to obtain the predicted attribute value;
[0023] The attention weight is used as the degree of importance of the corresponding specific substructure to the predicted value of the attribute.
[0024] Optionally, when selecting the converter model in the relationship exploration model to output the attribute prediction value, the step of inputting the multiple efficient representations into the relationship exploration model and outputting the attribute prediction value includes:
[0025] The multiple efficient representations are input into the converter model, and the embedding vectors corresponding to each specific substructure are output. The vector features of the embedding vectors are obtained through the converter encoder in the converter model.
[0026] The vector features are input into the fully connected layer of the converter model to obtain the predicted attribute values.
[0027] Optionally, after the step of obtaining the embedding vectors corresponding to each of the specific substructures, the material design optimization method further includes:
[0028] The embedding vectors corresponding to each specific substructure are input into the structure-attribute model to calculate the interaction relationships between each specific substructure.
[0029] Optionally, after obtaining the attention weights corresponding to the plurality of specific substructures, the material design optimization method further includes:
[0030] The attention weights are calculated to obtain the attention distribution characteristics corresponding to each specific substructure.
[0031] Optionally, the step of determining the target specific substructure from the plurality of specific substructures and then optimizing the design of the target material based on the target specific substructure includes:
[0032] Based on the importance level, a first specific substructure that has an impact on the predicted value of the attribute exceeding a preset impact level is determined from the plurality of specific substructures, and a second specific substructure whose interaction with other specific substructures exceeds a preset effect level is determined from the specific substructures based on the interaction relationship, and the specific substructures that are repeated in the first specific substructure and the second specific substructure are determined as the first target specific substructure.
[0033] After determining the second target-specific substructure from the plurality of specific substructures based on the attention distribution characteristics, the target material is designed and optimized based on the first target-specific substructure and the second target-specific substructure.
[0034] Furthermore, to achieve the above objectives, the present invention provides a material design optimization system, which includes a memory, a processor, and a computer processing program stored in the memory and executable on the processor. When the computer processing program is executed by the processor, it implements the steps of the material design optimization method as described above.
[0035] In addition, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a computer processing program, which, when executed by a processor, implements the steps of the material design optimization method as described above.
[0036] This invention proposes a material design optimization method, system, and medium. By sampling experimental materials according to a preset sampling method, multiple specific substructures are obtained, enabling direct analysis based on the structure with the highest mechanistic relevance. This reduces interference from non-critical structures present in materials with high molecular weight and strong mechanistic structures, improving the accuracy of prediction results. The multiple specific substructures are input into a relationship exploration model to obtain target data, and the target material is designed and optimized based on this target data. The steps include: inputting multiple specific substructures into the relationship exploration model to obtain predicted property values of the experimental material, the importance of each specific substructure to the predicted property values, the interaction relationships between the specific substructures, and the attention distribution characteristics of each specific substructure to the predicted property values; determining the target specific substructure from the multiple specific substructures based on the importance, interaction relationships, and attention distribution characteristics, and then designing and optimizing the target material based on the target specific substructure. By acquiring crucial information for adjusting experiments and mechanistic explorations—namely, the importance of each specific substructure and their interactions—further analysis based on importance and interactions reveals connections between intermediate variables related to higher-level mechanisms. Additionally, data directly applicable to materials exploration and research, such as attention distribution characteristics, is obtained. Analysis of specific substructures based on these attention distribution characteristics enhances model interpretability. This, combined with crucial information for adjusting experiments and mechanistic explorations and data directly applicable to materials exploration and research, improves the efficiency of materials design optimization. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the terminal structure of the hardware operating environment involved in the embodiments of the present invention;
[0038] Figure 2 This is a flowchart illustrating the first embodiment of the material design optimization method of the present invention;
[0039] Figure 3 This is a schematic diagram illustrating the process of handling experimental materials using a relationship exploration model according to the present invention;
[0040] Figure 4 This is a flowchart illustrating the second embodiment of the material design optimization method of the present invention;
[0041] Figure 5 This is a schematic diagram of the connection path between adjacent units and the central unit of the present invention;
[0042] Figure 5 (1) is a schematic diagram of the present invention where there is only one connection path between adjacent units and the central unit;
[0043] Figure 5(2) is a schematic diagram showing that there are two connection paths between the adjacent units and the central unit of the present invention;
[0044] Figure 6 This is a flowchart illustrating the third embodiment of the material design optimization method of the present invention;
[0045] Figure 7 This is a schematic diagram illustrating the process of processing a specific substructure using an encoding model according to the present invention;
[0046] Figure 8 This is a schematic diagram illustrating the process of attribute prediction numerical output using the self-attention model in the relation exploration model of this invention;
[0047] Figure 9 This is a schematic diagram of the structure of the interpretable model of the physicochemical properties of TADF materials;
[0048] Figure 10 The diagram shows the results obtained for the HOMO, LUMO, HOMO-1, and LUMO+1 orbitals based on the self-attention model and the converter model, respectively.
[0049] Figure 11 This is a schematic diagram showing the relationship between the attention weight variance of specific substructures of the molecular orbitals of TADF molecules containing three receptor structures in the experimental material and the predicted values of PLQY attributes.
[0050] The realization of the objective of this invention, its functional features and advantages will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0051] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0052] like Figure 1 As shown, Figure 1 This is a schematic diagram of the terminal structure of the hardware operating environment involved in the embodiments of the present invention.
[0053] The terminal implementation of this invention is a material design optimization system, such as... Figure 1As shown, the material design optimization system may include: a processor 1001, such as a CPU; a network interface 1004; a user interface 1003; a memory 1005; and a communication bus 1002. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface). The memory 1005 may be high-speed RAM or non-volatile memory, such as a disk drive. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0054] Optionally, the material design optimization system may also include RF (Radio Frequency) circuits, sensors, WiFi modules, etc. Sensors such as light sensors, motion sensors, and other sensors will not be elaborated upon here.
[0055] Those skilled in the art will understand that Figure 1 The material design optimization system structure shown does not constitute a limitation on the material design optimization system. It may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0056] like Figure 1 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a computer processing program.
[0057] exist Figure 1 In the material design optimization system shown, network interface 1004 is mainly used to connect to the backend server and communicate data with it; user interface 1003 is mainly used to connect to the client (user terminal) and communicate data with it; while processor 1001 can be used to call the computer processing program stored in memory 1005 and perform the following operations:
[0058] The experimental materials were sampled according to a preset sampling method to obtain multiple specific substructures;
[0059] The target data is obtained by inputting the multiple specific substructures into the relationship exploration model, and the target material is designed and optimized based on the target data;
[0060] The step of inputting the multiple specific substructures into a relationship exploration model to obtain target data, and then optimizing the design of the target material based on the target data, includes:
[0061] By inputting the multiple specific substructures into the relationship exploration model, the predicted values of the experimental material's properties, the importance of each specific substructure to the predicted values of the properties, the interaction relationships between each specific substructure, and the attention distribution characteristics of each specific substructure to the predicted values of the properties are obtained.
[0062] Based on the importance, the interaction relationship, and the attention distribution characteristics, after determining the target specific substructure from the plurality of specific substructures, the target material is designed and optimized according to the target specific substructure.
[0063] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0064] The steps of sampling experimental materials according to a preset sampling method to obtain multiple specific substructures include:
[0065] The key physicochemical elements of the experimental material are obtained, and the positioning information corresponding to the key physicochemical elements is obtained based on the key physicochemical elements. The key physicochemical elements are located according to the positioning information to obtain the central unit of the key physicochemical elements, and then the preset iterative sampling process is entered.
[0066] During the preset iterative sampling process, after selecting adjacent units of the key physicochemical elements whose distance from the central unit is less than or equal to the expected distance as physicochemical units, based on the physicochemical units and the central unit, new physicochemical units whose distance from the key physicochemical elements is less than or equal to the expected distance from the physicochemical units or the central unit are selected, until the preset iterative sampling process ends;
[0067] The specific substructure is formed based on the physicochemical relationship between the selected physicochemical unit and the central unit during the preset iterative sampling process.
[0068] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0069] The step of inputting the multiple specific substructures into the relationship exploration model to obtain the predicted property values of the experimental material includes:
[0070] After obtaining the initial codes corresponding to each specific substructure based on the preset computation theory, the efficient representations corresponding to each initial code are determined according to the preset coding model.
[0071] The multiple efficient representations are input into the structure-attribute model in the relationship exploration model, and the predicted attribute values are output.
[0072] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0073] The steps of inputting the multiple efficient representations into the relationship exploration model and outputting the predicted attribute values include:
[0074] The multiple efficient representations are concatenated and then input into the self-attention model to obtain the weight coefficients corresponding to the multiple specific substructures respectively;
[0075] After normalizing the weight coefficients using a preset function to obtain the attention weights corresponding to the multiple specific substructures, the attention weights are fused and merged to obtain the output vector.
[0076] The output vector is input into the self-attention model to obtain the predicted attribute value;
[0077] The attention weight is used as the degree of importance of the corresponding specific substructure to the predicted value of the attribute.
[0078] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0079] The steps of inputting the multiple efficient representations into the relationship exploration model and outputting the predicted attribute values include:
[0080] The multiple efficient representations are input into the converter model, and the embedding vectors corresponding to each specific substructure are output. The vector features of the embedding vectors are obtained through the converter encoder in the converter model.
[0081] The vector features are input into the fully connected layer of the converter model to obtain the predicted attribute values.
[0082] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0083] After the step of outputting the embedding vectors corresponding to each of the specific substructures, the material design optimization method further includes:
[0084] The embedding vectors corresponding to each specific substructure are input into the structure-attribute model to calculate the interaction relationships between each specific substructure.
[0085] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0086] After obtaining the attention weights corresponding to the multiple specific substructures, the material design optimization method further includes:
[0087] The attention weights are calculated to obtain the attention distribution characteristics corresponding to each specific substructure.
[0088] Furthermore, the processor 1001 can call the computer processing program stored in the memory 1005 and also perform the following operations:
[0089] After determining the target specific substructure from the plurality of specific substructures, the step of designing and optimizing the target material based on the target specific substructure includes:
[0090] Based on the importance level, a first specific substructure that has an impact on the predicted value of the attribute exceeding a preset impact level is determined from the plurality of specific substructures, and a second specific substructure whose interaction with other specific substructures exceeds a preset effect level is determined from the specific substructures based on the interaction relationship, and the specific substructures that are repeated in the first specific substructure and the second specific substructure are determined as the first target specific substructure.
[0091] After determining the second target-specific substructure from the plurality of specific substructures based on the attention distribution characteristics, the target material is designed and optimized based on the first target-specific substructure and the second target-specific substructure.
[0092] Reference Figure 2 In the first embodiment of the present invention, the material design optimization method includes:
[0093] Step S10: Sample the experimental materials according to the preset sampling method to obtain multiple specific substructures.
[0094] Step S20: Input the multiple specific substructures into the relationship exploration model to obtain target data, and optimize the design of the target material based on the target data.
[0095] The main method of this embodiment is to use a preset sampling method to split the structure of the input experimental material into multiple specific substructures that are highly related to the layers and mechanisms, and then use a deep learning model, namely a relationship exploration model, to process the multiple specific substructures to obtain the target data between mechanisms.
[0096] Specifically, refer to Figure 3As shown, to increase the accuracy of the relation exploration model prediction, before inputting the experimental materials into the relation exploration model for prediction, the key structures of the experimental materials, i.e., specific substructures, are extracted to eliminate the influence of non-key structures on the prediction results. In this example, multiple physicochemical elements 200, 201, ..., n (e.g., molecular structure) of the experimental material 100 are first obtained, and the central unit of each physicochemical element (i.e., ...) is calculated. Figure 3 After (300, 301, ..., n) in the model, sampling iteration is performed on the corresponding physicochemical elements based on the central unit. The specific substructure (i.e., ..., n) is formed according to the sampling iteration results. Figure 3 The values 400, 401, ..., n) are input into the critical exploration model U, enabling the critical exploration model to make high-precision result predictions based on the critical structure and output target data.
[0097] Optionally, the step of inputting the multiple specific substructures into the relationship exploration model to obtain target data in step S20, and then optimizing the design of the target material based on the target data, includes:
[0098] Step S201: Input the multiple specific substructures into the relationship exploration model to obtain the predicted value of the experimental material's properties, the importance of each specific substructure to the predicted value of the properties, the interaction relationship between each specific substructure, and the attention distribution characteristics of each specific substructure to the predicted value of the properties.
[0099] Step S202: Based on the importance, the interaction relationship, and the attention distribution characteristics, after determining the target specific substructure from the multiple specific substructures, the target material is designed and optimized according to the target specific substructure.
[0100] In this example, the relationship exploration model can not only predict the property prediction values of experimental materials based on the structure of key physicochemical property relationships extracted from the experimental materials, i.e., specific substructures, but also explore the importance of each specific substructure relative to the property prediction values, the interaction relationships between each specific substructure, and the attention distribution characteristics of each specific substructure to the property prediction values. By analyzing the importance of each specific substructure relative to the property prediction values and the interaction relationships between each specific substructure, the model can obtain the connections between intermediate variables of high-level mechanism relationships in the experimental materials. This facilitates the understanding of the prediction results by technicians, improves the accuracy of target material design optimization, and helps to verify and adjust the prediction results output by the black-box model in the early stage.
[0101] In this embodiment, the experimental material is sampled according to a preset sampling method to obtain multiple specific substructures. This allows for analysis based directly on the structure with the highest mechanistic relevance, reducing interference from non-critical structures present in materials with high molecular weight and responsible mechanistic structures, thus improving the accuracy of the prediction results. The multiple specific substructures are then input into a relationship exploration model to obtain target data, and the target material is designed and optimized based on this target data. The steps of inputting multiple specific substructures into a relationship exploration model to obtain target data and then designing and optimizing the target material include: inputting multiple specific substructures into a relationship exploration model to obtain the predicted property values of the experimental material, the importance of each specific substructure to the predicted property values, the interaction relationships between each specific substructure, and the attention distribution characteristics of each specific substructure to the predicted property values; based on the importance, interaction relationships, and attention distribution characteristics, a target specific substructure is determined from the multiple specific substructures, and the target material is designed and optimized based on the target specific substructure. By acquiring crucial information for adjusting experiments and mechanistic explorations—namely, the importance of each specific substructure and their interactions—further analysis based on importance and interactions reveals connections between intermediate variables related to higher-level mechanisms. Additionally, data directly applicable to materials exploration and research, such as attention distribution characteristics, is obtained. Analysis of specific substructures based on these attention distribution characteristics enhances model interpretability. This, combined with crucial information for adjusting experiments and mechanistic explorations and data directly applicable to materials exploration and research, improves the efficiency of materials design optimization.
[0102] Furthermore, based on the first embodiment of the material design optimization method of the present invention described above, a second embodiment of the material design optimization method of the present invention is proposed.
[0103] Reference Figure 4 In the second embodiment of the material design optimization method of the present invention, the step S10 above, which involves sampling the experimental material according to a preset sampling method to obtain multiple specific substructures, includes:
[0104] Step S101: Obtain the key physicochemical elements of the experimental material, and obtain the positioning information corresponding to the key physicochemical elements based on the key physicochemical elements. Position the key physicochemical elements according to the positioning information to obtain the central unit of the key physicochemical elements, and enter the preset iterative sampling process.
[0105] To ensure the physicochemical interpretability of each specific substructure, this example designs a pre-defined sampling method based on the integrity of physicochemical properties. Specifically, it involves obtaining the key physicochemical elements affecting the material properties of the experimental material through theoretical derivation, such as the key constituent elements, molecular orbitals, vibrational modes, etc. These key factors typically correspond to the initial substructure of the material, such as specific groups in organic molecules, pharmacophores in drug molecules, or chain or ring structures in polymer materials.
[0106] After listing these key physicochemical elements, the location information of each key physicochemical element is obtained by calculating based on DFT (Density Functional Theory). The location information is used to locate the corresponding key physicochemical element information. After obtaining the central unit corresponding to each key physicochemical element, the preset iterative sampling process is entered. Based on the central unit, non-key units in the key physicochemical elements, that is, units that do not have mechanistic physicochemical properties, are eliminated.
[0107] Step S102: In the preset iterative sampling process, after selecting the adjacent units of the key physicochemical elements whose distance from the central unit is less than or equal to the expected distance as physicochemical units, based on the physicochemical units and the central unit, new physicochemical units whose distance from the key physicochemical elements is less than or equal to the expected distance from the physicochemical units or the central unit are selected, until the preset iterative sampling process ends.
[0108] In the preset iterative sampling process, the number of iterative sampling times is set according to the computing power of the iterative sampling system. In the first iterative sampling, the expected distance of the central cell is calculated according to the DFT (the expected distance indicates that the adjacent cells within the expected distance of the central cell are physicochemical cells with mechanistic physicochemical properties). Adjacent cells with a distance less than or equal to the expected distance from the central cell are selected from the key physicochemical elements where the central cell is located, and these adjacent cells are identified as physicochemical cells. Then, the process proceeds to the second iterative sampling. In the second iterative sampling, the expected distance of the physicochemical cells obtained in the first iterative sampling is calculated according to the DFT. Adjacent cells with a distance less than or equal to the expected distance from both the physicochemical cells and the central cell are selected from the key physicochemical elements where they are located, and these adjacent cells are identified as new physicochemical cells.
[0109] Based on the above selection process of physicochemical units, the remaining iteration sampling times are performed sequentially until the last iteration sampling time is completed, at which point the preset iteration sampling process is determined to be over.
[0110] It should be noted that in each iteration of sampling, after acquiring the corresponding physicochemical units, an iterative sampling judgment process is required. This judgment process includes three judgment conditions. The first judgment condition is: determining whether adjacent units in the key physicochemical elements are directly or indirectly connected to the central unit. If the judgment is yes, then it is determined whether the initial iteration sampling number is the last iteration sampling number. If not, then it is determined whether adjacent units that are directly or indirectly connected to the central unit have a physicochemical relationship with the central unit. In this example, the physicochemical relationship is set to have at least two connection channels with the central unit, as shown in the reference. Figure 5 As shown, Figure 5 In (1), there is only one path between the adjacent unit and the central unit, for example, only... Figure 5 If path a is used in (1), it indicates that the adjacent unit and the central unit do not have a physical or chemical relationship. Figure 5 (2) in the example means that the adjacent unit and the central unit have two connection paths, for example, there are Figure 5 (2) If the path is a and the path is b1→b2→b3, it indicates that the adjacent unit has a physical-chemical relationship with the central unit. If the number of iterations is determined to be the last number of iterations, then after determining that the physical-chemical unit has a physical-chemical relationship with the central unit, the next iteration of sampling or the preset iteration sampling process is performed.
[0111] In addition, after each iteration, any recurring physicochemical relationships (i.e. structural connections) will be removed.
[0112] Step S103: Based on the physicochemical relationship between the physicochemical unit and the central unit selected in the preset iterative sampling process, the specific substructure is formed.
[0113] After the preset iterative sampling process is completed, the physicochemical units collected during the preset iterative sampling process will form specific substructures based on the physicochemical relationships between the physicochemical units and between the physicochemical units and the center.
[0114] In this embodiment, key physicochemical elements of the experimental materials are obtained, and positioning information corresponding to the key physicochemical elements is obtained based on the key physicochemical elements. The key physicochemical elements are located according to the positioning information to obtain the central unit of the key physicochemical elements. The process then enters a preset iterative sampling process. During the preset iterative sampling process, adjacent units of the key physicochemical elements whose distance from the central unit is less than or equal to the expected distance are selected as physicochemical units. Based on the physicochemical units and the central unit, new physicochemical units whose distance from the key physicochemical elements or the central unit is less than or equal to the expected distance are selected until the preset iterative sampling process ends. Based on the physicochemical relationship between the physicochemical units and the central unit selected during the preset iterative sampling process, a specific substructure is formed to ensure the physicochemical interpretability of each specific substructure.
[0115] Furthermore, based on the first embodiment of the material design optimization method of the present invention described above, a third embodiment of the material design optimization method of the present invention is proposed.
[0116] Reference Figure 6 In the third embodiment of the material design optimization method of the present invention, the step S201 above, which involves inputting the plurality of specific substructures into the relationship exploration model to obtain the property prediction values of the experimental material, includes:
[0117] Step A1: After obtaining the initial codes corresponding to each specific substructure based on the preset computation theory, determine the efficient representations corresponding to each initial code according to the preset coding model.
[0118] Step A2: Input the multiple efficient representations into the structure-attribute model in the relationship exploration model, and output the predicted attribute values.
[0119] It should be noted that before inputting the efficient representations corresponding to each specific substructure into the relationship exploration model to obtain the attribute prediction values, it is necessary to first construct a relationship model from physicochemical property to structure to attribute, and then train the relationship model from structure to attribute to obtain a relationship model that can output the corresponding attribute based on the efficient representation of the input material.
[0120] Specifically, the coded representation of the physicochemical mechanism correlation structure is used as the input to the relational model. First, an initial code for the specific self-structure is obtained by decomposing a sampling method based on DFT calculation and the integrity of physicochemical properties. Figure 7 N specific substructures in (i.e. Figure 7 The 400, 401, ..., n in the text correspond to N initial codes (i.e., ... Figure 7 (600, 601, ..., n in the original text). For different specific substructures, a matching characterization method is selected. This matching characterization method requires using key mechanistic elements or DFT calculation results based on the physicochemical properties of the specific substructure, such as characteristic atoms, spatial coordinates, spatial angles, and charge vibrations, as the initial encoding of the specific substructure. For organic molecules, molecular fingerprints can be directly used as part of the initial encoding. For the topological structure of other non-organic molecules, the spatial relationships of structural units and their key elements can be used as the initial encoding. After determining the initial encoding, a reasonable encoding model (i.e., ...) is selected. Figure 7 The information in the initial encoding (U700, U701, ..., UN) is further extracted to obtain an efficient representation. Finally, the efficient representation is input into the trained relational model (i.e., Figure 7In the U), the relational model is able to predict the attributes of the experimental materials based on the efficient representation of the input, and then output the predicted attribute values.
[0121] Optionally, the step A2, which involves inputting the multiple efficient representations into the relationship exploration model and outputting the predicted attribute value, includes:
[0122] Step A201: After concatenating the multiple efficient representations, input them into the self-attention model to obtain the weight coefficients corresponding to the multiple specific substructures respectively;
[0123] Step A202: Normalize the weight coefficients using a preset function to obtain the attention weights corresponding to the multiple specific substructures, and then merge the attention weights to obtain the output vector.
[0124] Step A203: Input the output vector into the self-attention model to obtain the predicted attribute value;
[0125] The attention weight is used as the degree of importance of the corresponding specific substructure to the predicted value of the attribute.
[0126] In this embodiment, there are two prediction methods for the attribute prediction value.
[0127] The first approach is to select a self-attention model from the relation exploration model to output attribute prediction values. Specifically, to explore the relationship between a specific substructure and its predicted value, self-attention is used to analyze N specific substructures to obtain information on the importance of those substructures. For example... Figure 8 As shown, the self-attention model in this example is implemented by efficiently representing the N specific substructures (i.e., Figure 8 After concatenating the 800, 801, ..., n values, the input is fed into the fully connected network layer (i.e., ..., n) in the self-attention model. Figure 8 In the FC (Functional Control Unit), N corresponding weight coefficients are generated (i.e., ... Figure 8 The specific process for obtaining the weight coefficients (K1, K2, ..., KN) is shown in Formula 1. Then, the softmax function (i.e., the preset function) is used to normalize the weight coefficients, resulting in the attention weights (i.e., the weights corresponding to the N specific substructures) respectively. Figure 8In the case of a1, a2, ..., aN), a two-step operation is performed based on the acquired attention weights. The first step is to merge the attention weights to obtain the output vector of the self-attention physicochemical relationship part of the experimental material (i.e., all specific substructures in the experimental material). The specific process of obtaining the output vector is shown in Formula 2. The output line is then fed back into the fully connected network layer in the self-attention model to output the attribute prediction data of the experimental material. The second step is to directly determine the obtained attention weights as the importance of each specific substructure to the attribute prediction value.
[0128] Formula 1 is:
[0129]
[0130] Among them, a i Let represent the weight coefficient corresponding to the i-th specific substructure, e and k are constants, N represents the total number of specific substructures, and n represents the number of specific substructures currently being calculated among all specific substructures.
[0131] Formula 2 is:
[0132] out=Σ i a i h i ————Formula 2
[0133] Where out represents the output vector, a i h represents the weight coefficient corresponding to the i-th specific substructure. i This represents the efficient representation corresponding to the i-th specific substructure.
[0134] Optionally, after obtaining the attention weights corresponding to the plurality of specific substructures in step A202, the material design optimization method further includes:
[0135] Step A204: Calculate the attention weights for each of the aforementioned attention weights to obtain the attention distribution characteristics corresponding to each of the specific substructures.
[0136] When obtaining the influence of a specific substructure on the predicted value of the target material's properties by statistically analyzing the average of the weight coefficients of N specific substructures, the attention weights of specific substructures with a greater influence are extracted. When there are specific substructures of the same type, the distribution characteristics of the attention weights of specific substructures of the same type can be calculated, and the distribution characteristics of the attention weights can be used as an indicator for designing and optimizing the target material.
[0137] Optionally, the step A2, which involves inputting the multiple efficient representations into the relationship exploration model and outputting the predicted attribute value, includes:
[0138] Step A205: Input the multiple efficient representations into the converter model, output the embedding vectors corresponding to each specific substructure, and obtain the vector features of the embedding vectors through the converter encoder in the converter model;
[0139] Step A206: Input the vector features into the fully connected layer of the converter model to obtain the attribute prediction value.
[0140] The second approach is to select the converter model in the relation exploration model to output the attribute prediction values.
[0141] Specifically, to explore the interaction relationships between specific substructures, a single-layer transformer model is used to analyze N specific substructures. This single-layer transformer model needs to include residual connections.
[0142] The efficient representations corresponding to each specific substructure are input as tokens into the converter model. Based on the converter model, embedding vectors corresponding to each specific substructure are generated, namely Query, Key and Value. According to the embedding vectors corresponding to each specific substructure, the vector features of each embedding vector are obtained through the converter encoder in the converter model. The vector features are then input into the fully connected layer of the converter model. Based on the output of the fully connected layer, the predicted attribute values corresponding to the experimental materials are obtained.
[0143] Optionally, after the step of outputting the embedding vectors corresponding to each of the specific substructures in step A205, the material design optimization method further includes:
[0144] Step A207: Input the embedding vectors corresponding to each specific substructure into the structure-attribute model to calculate the interaction relationship between each specific substructure.
[0145] In this example, the relational model trained by the model is interpretable and can output the interaction relationships between specific substructures based on the embedding vectors input to each specific substructure. Specifically, the structure-attribute model in the relational model is used to pad the Query and Key of each specific substructure to obtain numerical values, and the interaction relationships between the specific substructures are calculated based on the numerical values of each specific substructure. Specific Implementation
[0147] The step of designing and optimizing the target material based on the target specific substructure after determining the target specific substructure from the plurality of specific substructures in steps S201 and S202 includes:
[0148] Step S208: Based on the importance level, determine the first specific substructure from the plurality of specific substructures whose influence on the predicted value of the attribute exceeds a preset influence level; and based on the interaction relationship, determine the second specific substructure from the specific substructure whose interaction with other specific substructures exceeds a preset interaction level; and determine the specific substructures that are repeated in the first specific substructure and the second specific substructure as the first target specific substructure.
[0149] Step S209: After determining the second target specific substructure from the plurality of specific substructures based on the attention distribution characteristics, the target material is designed and optimized based on the first target specific substructure and the second target specific substructure.
[0150] To make the solution more comprehensive, this embodiment uses OLED display materials as an example for explanation.
[0151] OLES (Optical Orbital Electron Sequences) have attracted much attention as a new generation of display materials, especially thermally activated delayed fluorescence (TADF) materials, which have great potential for development due to their composition of only light elements and ability to achieve 100% luminescence quantum efficiency. The main mechanism involves electrons in the triplet state crossing back to the singlet state via antisystem crossings, and then transitioning back to the ground state to emit fluorescence. When studying TADF materials, the primary focus is on the energy level states of the electrons, given their luminescence principle. Currently, the most important mechanistic elements are HOMO and LUMO. Based on the luminescence mechanism of TADF materials, this case study designs a physicochemical mechanism interpretability model based on HOMO, LUMO, HOMO-1, and LUMO+1 molecular orbitals to optimize the molecular design of novel TADF materials.
[0152] The interpretability model of the physical and chemical properties of the designed TADF material is as follows: Figure 9 As shown, the input TADF molecule (i.e., experimental material) is split according to four molecular orbitals (i.e., key physicochemical elements) that have a significant impact on the predicted target property, fluorescence quantum yield (PLQY). The four sets of specific substructures after splitting are encoded to obtain the corresponding efficient characterization input relationship exploration model. The model is fully trained using TADF molecule-PLQY property data. Finally, the model is used to obtain the importance of each molecular orbital to the predicted PLQY property value and the interaction relationship between each specific substructure, and to analyze the hidden mechanism, thereby guiding the design and experimentation of new TADF molecules (i.e., target materials).
[0153] It should be noted that this example utilizes the non-linear relationship of attention variance as a characteristic of attention distribution.
[0154] Specifically: ① First, based on DFT calculations, the central units corresponding to the HOMO, LUMO, HOMO-1, and LUMO+1 orbitals are located. From the DFT calculations, the percentage contribution of each atom in the TADF material molecule to the orbitals can be obtained, and the atom with the highest contribution is selected as the central unit corresponding to each molecular orbital (i.e., Figure 9 (302, 303, 304, and 305 in the series).
[0155] ② The physicochemical units of TADF material molecules are sampled using a pre-defined iterative sampling method based on the integrity of physicochemical properties. During sampling, the m atoms with the highest contribution from the DFT calculation results for each molecular orbital are selected as the central units for the first iteration. After a pre-defined number of iterations, the sampling results that retain the integrity of the physicochemical properties are finally obtained. Observation reveals that these sampled structures are complete ring or chain structures containing central units.
[0156] ③ Select appropriate encoding as the initial encoding for the specific substructures corresponding to each molecular orbital. In this case, the ECFP fingerprint spectrum of the specific substructures of each molecular orbital is used as the first part of the initial encoding. Secondly, the contribution of the central unit to the molecular orbital and its spatial structure are selected as the second part of the primary encoding. Thirdly, a 3D spatial encoding of the associated structure generated using a graphical model is selected as the third part of the primary encoding. A self-attention model is used to obtain an efficient representation of the specific substructures of each molecular orbital from the primary encoding (i.e.,...). Figure 9 The input relationship exploration model includes 1100101...10, 0110100...11, 1110110...01 and 1010101...10.
[0157] ④ Design a self-attention model and a converter model that are dimensionally matched to the specific substructures of each input molecular orbital to extract the importance and interaction relationships. Figure 10 The results for HOMO, LUMO, HOMO-1, and LUMO+1 orbitals, obtained using both self-attention and converter-based models, are presented. These results reflect the importance and interaction relationships of each molecular orbital in the predicted PLQY properties. Figure 10 (1) It can be found that the predicted values of specific substructures corresponding to the HOMO, LUMO, and LUMO+1 orbits for the PLQY attribute are evenly distributed in the 0-0.8 range, while the predicted values of specific substructures corresponding to the HOMO-1 orbit are less evenly distributed in the 0-0.8 range. Figure 10As shown in the interaction distribution diagram of the specific substructures of each molecular orbital and the specific substructures of the HOMO orbital in (2), the interaction distribution between the specific substructures of the LUMO+1 orbital and the specific substructures corresponding to the HOMO orbital is more uniform in the 0-1 interval. Therefore, when designing and optimizing new TADF material molecules, the LUMO+1 orbital can be given priority as guiding data for designing and optimizing new TADF material molecules.
[0158] ⑤ When using the relationship exploration model of the physicochemical properties of TADF materials to guide the design and optimization of new molecules, the main methods used are the property prediction values of the model and the attention weights obtained from the relationship model. Analyzing the attention weights of specific substructures corresponding to the HOMO, LUMO, HOMO-1, and LUMO+1 orbitals calculated based on the self-attention model, it is easy to find that their attention distribution characteristics are negatively correlated with the property prediction values. For TADF molecules containing similar acceptor structures, the smaller the distribution characteristics of the attention weights of specific substructures corresponding to each molecular orbital, the higher the property prediction value of the molecule, and the better the luminous efficiency of the TADF device using the molecule. Figure 11 The paper illustrates the relationship between the attention weight distribution characteristics of specific substructures of molecular orbitals in TADF molecules with three receptor structures in hypothetical experimental materials and the predicted values of PLQY properties. Experimental data for TADF molecules using triazine as the receptor structure are abundant, showing a significant correlation. Furthermore, specific substructures with more similar molecular structures are located closer together in the regions shown in the figure. Figure 11 In (1), the only difference in the specific substructure of the TADF molecular orbitals represented by categories 1 and 2 is the modified R group. It can be found that although the regions are different, they maintain the same correlation trend, and the trend is relatively flat. Figure 11 The correlation trends of specific substructures of TADF molecular orbitals represented in (2) and (3) are significantly different from those of specific substructures of other TADF molecular orbitals, and the trends are noticeably steeper. Using the methods described above, when designing new TADF material molecules, the molecular orbitals corresponding to specific substructures of categories 1 and 2 can be given priority as guiding data for designing and optimizing new TADF material molecules.
[0159] Furthermore, the present invention also proposes a material design optimization system, which includes a memory, a processor, and a computer processing program stored in the memory and executable on the processor. When the computer processing program is executed by the processor, it implements the steps of the material design optimization method as described above.
[0160] Furthermore, the present invention also proposes a computer-readable storage medium storing a computer processing program, which, when executed by a processor, implements the steps of the material design optimization method described above.
[0161] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.
[0162] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0163] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0164] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.
Claims
1. A material design optimization method, characterized in that, The material design optimization method includes the following steps: The experimental materials were sampled according to a preset sampling method to obtain multiple specific substructures; The step of sampling experimental materials according to a preset sampling method to obtain multiple specific substructures includes: The key physicochemical elements of the experimental material are obtained, and the positioning information corresponding to the key physicochemical elements is obtained based on the key physicochemical elements. The key physicochemical elements are located according to the positioning information to obtain the central unit of the key physicochemical elements, and then the preset iterative sampling process is entered. During the preset iterative sampling process, after selecting adjacent units of the key physicochemical elements whose distance from the central unit is less than or equal to the expected distance as physicochemical units, based on the physicochemical units and the central unit, new physicochemical units whose distance from the key physicochemical elements is less than or equal to the expected distance from the physicochemical units or the central unit are selected, until the preset iterative sampling process ends; Based on the physicochemical relationship between the selected physicochemical unit and the central unit during the preset iterative sampling process, the specific substructure is formed; The target data is obtained by inputting the multiple specific substructures into the relationship exploration model, and the target material is designed and optimized based on the target data; The step of inputting the multiple specific substructures into a relationship exploration model to obtain target data, and then optimizing the design of the target material based on the target data, includes: By inputting the multiple specific substructures into the relationship exploration model, the predicted values of the experimental material's properties, the importance of each specific substructure to the predicted values of the properties, the interaction relationships between each specific substructure, and the attention distribution characteristics of each specific substructure to the predicted values of the properties are obtained. Based on the importance, the interaction relationship, and the attention distribution characteristics, after determining the target specific substructure from the plurality of specific substructures, the target material is designed and optimized according to the target specific substructure.
2. The material design optimization method as described in claim 1, characterized in that, The step of inputting the multiple specific substructures into the relationship exploration model to obtain the predicted property values of the experimental material includes: After obtaining the initial codes corresponding to each specific substructure based on the preset computation theory, the efficient representations corresponding to each initial code are determined according to the preset coding model. The multiple efficient representations are input into the structure-attribute model in the relationship exploration model, and the predicted attribute values are output.
3. The material design optimization method as described in claim 2, characterized in that, When selecting the self-attention model in the relationship exploration model to output the attribute prediction value, the step of inputting the multiple efficient representations into the relationship exploration model and outputting the attribute prediction value includes: The multiple efficient representations are concatenated and then input into the self-attention model to obtain the weight coefficients corresponding to the multiple specific substructures respectively; After normalizing the weight coefficients using a preset function to obtain the attention weights corresponding to the multiple specific substructures, the attention weights are merged to obtain the output vector. The output vector is input into the self-attention model to obtain the predicted attribute value; The attention weight is used as the degree of importance of the corresponding specific substructure to the predicted value of the attribute.
4. The material design optimization method as described in claim 3, characterized in that, When selecting the converter model in the relationship exploration model to output the attribute prediction value, the step of inputting the multiple efficient representations into the relationship exploration model and outputting the attribute prediction value includes: The multiple efficient representations are input into the converter model, and the embedding vectors corresponding to each specific substructure are output. The vector features of the embedding vectors are obtained through the converter encoder in the converter model. The vector features are input into the fully connected layer of the converter model to obtain the predicted attribute values.
5. The material design optimization method as described in claim 4, characterized in that, After the step of obtaining the embedding vectors corresponding to each of the specific substructures, the material design optimization method further includes: The embedding vectors corresponding to each specific substructure are input into the structure-attribute model to calculate the interaction relationships between each specific substructure.
6. The material design optimization method as described in claim 3, characterized in that, After obtaining the attention weights corresponding to the plurality of specific substructures, the material design optimization method further includes: The attention weights are calculated to obtain the attention distribution characteristics corresponding to each specific substructure.
7. The material design optimization method as described in claim 1, characterized in that, The step of determining the target specific substructure from the plurality of specific substructures and then designing and optimizing the target material based on the target specific substructure includes: Based on the importance level, a first specific substructure that has an impact on the predicted value of the attribute exceeding a preset impact level is determined from the plurality of specific substructures, and a second specific substructure whose interaction with other specific substructures exceeds a preset effect level is determined from the specific substructures based on the interaction relationship, and the specific substructures that are repeated in the first specific substructure and the second specific substructure are determined as the first target specific substructure. After determining the second target-specific substructure from the plurality of specific substructures based on the attention distribution characteristics, the target material is designed and optimized based on the first target-specific substructure and the second target-specific substructure.
8. A material design optimization system, characterized in that, The material design optimization system includes a memory, a processor, and a computer processing program stored in the memory and executable on the processor. When the computer processing program is executed by the processor, it implements the steps of the material design optimization method as described in any one of claims 1-7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer processing program, which, when executed by a processor, implements the steps of the material design optimization method according to any one of claims 1-7.
Citation Information
Patent Citations
Optimal design method for internal structure of composite material machine tool bed
CN108416158A
Training method and device of molecular attribute prediction model, equipment and storage medium
CN116959611A