Method for calculating high-entropy material structure descriptor by using large language model

By constructing a knowledge base of high-entropy material structure descriptors and utilizing large language models for automated computation, the problem of selecting high-entropy material structure descriptors is solved, achieving efficient and accurate feature extraction and supporting machine learning design of high-entropy materials.

CN122050658APending Publication Date: 2026-05-15UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to systematically and quickly select and calculate structural descriptors for high-entropy materials, resulting in inefficient high-entropy material design and a cumbersome and error-prone calculation process.

Method used

A knowledge base of high-entropy material structure descriptors is constructed using a large language model. By retrieving research literature through keywords, the large language model is trained to recognize and extract relevant information from structure descriptors. The calculation method is optimized by combining the frequency of literature use and material type, and the structure descriptors of high-entropy materials are calculated automatically.

Benefits of technology

It significantly improves the efficiency and accuracy of structure descriptor selection, increases the calculation speed by 5-10 times, and achieves an accuracy of over 95%, supporting the construction of efficient high-entropy material machine learning models and the research and development of new materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050658A_ABST
    Figure CN122050658A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of material informatics and artificial intelligence, in particular to a method for training and calculating a high-entropy material structure descriptor based on a large language model, and the method is used for a machine learning task of high-entropy material design and performance prediction. According to the 3D chemical structure model of the high-entropy material or the element information in the chemical structural formula, the corresponding information is extracted through the large language model, and the structure descriptor for the high-entropy material machine learning task is automatically calculated according to the element information with the highest utilization rate in the literature and the knowledge base, so that the method has the characteristics of simple operation, reliable data, high speed and the like; and the structure descriptor with the highest literature recognition degree can be obtained through calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of materials informatics and artificial intelligence, specifically to a method for training and computing high-entropy material structure descriptors based on a large language model, for machine learning tasks of high-entropy material design and performance prediction. Background Technology

[0002] High-entropy materials, as a novel class of materials, including high-entropy alloys, high-entropy ceramics, and high-entropy composites, have attracted widespread attention due to their unique compositional design and excellent properties. High-entropy materials typically consist of five or more main elements in equimolar or near-equimolar ratios, forming a material system with a simple solid solution structure. Due to their complex compositional space and microstructure, traditional material design methods struggle to efficiently predict their performance. Machine learning-based methods, however, offer a new approach for the design and optimization of high-entropy materials.

[0003] The performance of machine learning models largely depends on the quality of the input features, i.e., the choice of structure descriptors. Structure descriptors are feature values ​​derived from information such as the chemical composition and crystal structure of a material. For high-entropy materials, commonly used structure descriptors include, but are not limited to, atomic size difference parameters δ and mixing enthalpy ΔH. mix Mixed entropy ΔS mix Factors such as valence electron concentration (VEC) and electronegativity differences exist. However, existing technologies face the following problems:

[0004] The selection and construction of structural descriptors for S1 high-entropy materials often rely on researchers' experience and understanding of the literature, lacking systematicity and objectivity. Different research teams may choose different structural descriptors for similar materials.

[0005] As research on high-entropy materials deepens, new structural descriptors are constantly emerging, making it difficult for researchers to fully grasp all possible descriptors and their applicable conditions. The process of literature retrieval and knowledge integration is time-consuming and prone to overlooking crucial information.

[0006] Existing methods for calculating S3 structure descriptors are mostly manual, which is cumbersome and prone to errors. For large-scale screening of high-entropy materials, this method is inefficient and cannot meet the needs of high-throughput computing.

[0007] Although existing research has attempted to establish materials databases and automatic feature extraction methods, for the emerging field of high-entropy materials, there is still no systematic method that can combine literature knowledge and artificial intelligence technology to quickly and accurately calculate the most suitable structural descriptors. Therefore, a new method is needed to solve the aforementioned technical problems. Summary of the Invention

[0008] This invention provides a method for calculating high-entropy material structure descriptors using a large language model, comprising:

[0009] Step S1: Retrieve research literature related to high-entropy materials based on keywords and construct a knowledge base of high-entropy material structure descriptors;

[0010] Step S2 uses a large language model to train the knowledge base of high-entropy material structure descriptors in step (1), and identifies and extracts the associated information related to the high-entropy material structure descriptors;

[0011] Step S3 uses the trained large language model to identify the most frequently used high-entropy material structure descriptors and train the corresponding calculation methods.

[0012] Step S4: Input the 3D chemical structure model or chemical structure formula of the high-entropy material to be calculated, and extract the element types and proportion information used to calculate the high-entropy material structure descriptor.

[0013] Step S5 calculates the corresponding structure descriptor for the high-entropy material based on the identified and trained high-entropy material structure descriptor calculation method, combined with the element type and proportion information.

[0014] In some embodiments, in step S1, relevant research literature is searched based on keywords, and custom settings can be made as needed.

[0015] In some embodiments, step S1, constructing a knowledge base for high-entropy material structure descriptors, includes: collecting research literature in the field of high-entropy materials based on keywords; extracting the material element composition, structure descriptors, calculation methods, and applicable conditions used in the research literature; and calling the corresponding component information from an open-source database.

[0016] In some embodiments, the structure descriptor includes an atomic size difference parameter δ and a mixing enthalpy ΔH. mix Mixed entropy ΔS mix One or more of the following: valence electron concentration (VEC), electronegativity difference, atomic radius ratio, crystal structure parameters, topological descriptor, and electronic structure descriptor.

[0017] In some embodiments, step S2, training the large language model includes: inputting structured and unstructured high-entropy material literature data; setting the training objective as recognizing the association between material composition and applicable structural descriptors; and optimizing the model's recognition accuracy through manual verification and feedback mechanisms.

[0018] In some embodiments, step S3, identifying the most frequently used structural descriptor calculation method, includes: using a large language model to extract relevant structural descriptors from the literature; assigning weights to different structural descriptor calculation methods based on the frequency of use in the literature; adjusting the selection priority of structural descriptor calculation methods considering the target material type and expected performance; and outputting the structural descriptor calculation method with the highest comprehensive score and its parameters.

[0019] In some embodiments, in step S4, the 3D chemical structure model of the high-entropy material to be calculated includes common 3D atomic structure model format files, which can automatically identify and extract the chemical structural formula corresponding to the 3D chemical structure model.

[0020] In some embodiments, in step S4, the information required for calculating the structure descriptor is automatically extracted from the corresponding chemical structural formula automatically identified from the 3D chemical structure model of the high-entropy material to be calculated or from the directly provided chemical structural formula.

[0021] In some embodiments, the knowledge base constructed by collecting research literature based on keywords in step S1 can be automatically updated periodically according to settings.

[0022] In some embodiments, in step S1, the large language model covers different types of high-entropy materials, and the training process employs multiple rounds of manual verification and feedback optimization.

[0023] In some embodiments, the steps need to start from step S1, that is, from the start of building or updating the structure descriptor knowledge base.

[0024] In some embodiments, the steps can begin directly from step S4, where there is already a relevant structure descriptor knowledge base and no updates are required; the process can start directly from inputting the chemical structure model or chemical formula of the high-entropy material to be settled.

[0025] The beneficial effects of the method for calculating high-entropy material structure descriptors using a large language model provided in this application include, but are not limited to, the following:

[0026] This application addresses the technical problems of current high-entropy material structure descriptor acquisition, such as reliance on manual methods, long processing times, and poor objectivity. It proposes to utilize Large Language Modeling (LLM) to deeply mine literature data and automatically generate structure descriptors with high literature acceptance. It has the advantages of simple operation, reliable data, and high speed. Compared with traditional manual calculation methods, the calculation speed is increased by 5-10 times, and the accuracy of structure descriptor selection reaches over 95%, which greatly accelerates the construction of high-entropy material machine learning models and the research and development process of new materials. Attached Figure Description

[0027] The following accompanying drawings describe in detail the exemplary embodiments disclosed in this application. The same reference numerals denote similar structures in several views of the drawings. Those skilled in the art will understand that these embodiments are non-limiting and exemplary, and the drawings are for illustrative purposes only and are not intended to limit the scope of this application. Other embodiments may similarly fulfill the inventive intent of this application. It should be understood that the drawings are not drawn to scale. Wherein:

[0028] Figure 1 Logical framework diagram for; Detailed Implementation

[0029] The following description provides specific application scenarios and requirements for this application, intended to enable those skilled in the art to make and use the content of this application. Various partial modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments and applications without departing from the spirit and scope of this application. Therefore, this application is not limited to the embodiments shown, but rather to the widest scope consistent with the claims.

[0030] The core idea of ​​this application is to leverage the powerful natural language processing and information extraction capabilities of large language models to integrate the high-entropy materials and their structural descriptor calculation rules scattered in a large number of research papers, thereby achieving automated and highly reliable calculation of structural descriptors.

[0031] High-entropy materials are a new class of materials, typically containing five or more main elements. These materials exhibit excellent mechanical properties, thermal stability, and corrosion resistance due to their unique "high-entropy effect," and have broad application prospects in aerospace, energy, and biomedicine. However, the complex compositional space of high-entropy materials makes traditional trial-and-error methods inefficient for developing new materials, while machine learning-based methods offer new approaches to the design of high-entropy materials.

[0032] Machine learning models need to transform material information into numerical features, i.e., structural descriptors, in order to perform calculations and predictions. Traditional methods for constructing structural descriptors rely on manual selection based on expert experience, making it difficult to systematically perform high-throughput calculations. In the rapidly developing field of high-entropy materials, new structural descriptors are constantly emerging, and how to quickly and accurately select appropriate descriptors has become a challenge.

[0033] This invention addresses the aforementioned technical problems by proposing a method that combines a large language model with literature knowledge to automatically identify and calculate the most suitable structural descriptors for specific high-entropy materials. The core idea behind this invention is that structural descriptors frequently used in scientific literature often possess high reliability and applicability. By automatically extracting and analyzing this knowledge through a large language model, the efficiency and accuracy of structural descriptor selection can be significantly improved.

[0034] The following will describe in detail, with reference to the accompanying drawings and embodiments, a method for calculating high-entropy material structure descriptors using a large language model provided by the present invention.

[0035] Figure 1 The present invention provides a method for calculating high-entropy material structure descriptors using a large language model, specifically including:

[0036] Step S1: Input keywords in the field of materials. The system searches and downloads research literature related to high-entropy materials through Web of Science and Google Scholar to build a knowledge base of original literature on high-entropy materials. Based on the knowledge base of original literature, the system uses a large language model to extract the material element composition, structural descriptors, calculation methods, and applicable conditions. According to the extracted element types, the system extracts the corresponding constituent element information through open source databases. The above steps constitute or update the knowledge base of high-entropy material structure descriptors.

[0037] Step S2: Based on the knowledge base, using structured and unstructured high-entropy material literature data, with the corresponding material element composition as the target, a large language model is used for training; using the extracted high-entropy material element types, composition, and proportion information as basic input, with structural descriptors as the target, a large language model is used for training; the recognition accuracy of the model is optimized through manual verification and feedback mechanisms.

[0038] Step S3: Using the trained large language model, assign weights to different structural descriptor calculation methods based on their frequency of use in the literature; consider the target material type and expected performance, and adjust the selection priority of structural descriptor calculation methods; output the structural descriptor calculation method with the highest comprehensive score and its parameters.

[0039] Step S4 (Starting step of Method 2): Input the 3D chemical structure model or chemical formula of the high-entropy material to be calculated. The 3D chemical structure model of the high-entropy material to be calculated includes common 3D atomic structure model format files such as mol, cif, and xsd. The system automatically identifies and extracts the chemical formula corresponding to the chemical structure model, and then extracts information such as the types, composition, and proportion of elements used in the calculation structure descriptor.

[0040] Step S5: Calculate the corresponding structural descriptor for the high-entropy material based on the identified and trained structural descriptor calculation method, combined with the element type and proportion information.

[0041] Example 1

[0042] A method for calculating the structure descriptor of high-entropy materials using a large language model, taking the calculation of the structure descriptor for a specific high-entropy alloy material (CoCrFe2MnNi) as an example, specifically includes the following steps:

[0043] Step S1: Input the keywords "High-entropy alloy", "descriptors", and "machine learning". The system will automatically search for research literature and download the original texts through Web of Science and Google Scholar. The system will then use a large language model to extract information from the downloaded research literature, including the elemental composition, structural descriptors, calculation methods, and applicable conditions of high-entropy alloy materials. Based on the extracted high-entropy alloy element types, the system will extract basic elemental information from open-source databases. Finally, the system will build or update a knowledge base of high-entropy alloy structural descriptors.

[0044] Step S2: Based on the knowledge base, using structured and unstructured high-entropy material literature data, with the corresponding material element composition as the target, a large language model is used for training; using the extracted high-entropy material element types, composition, and proportion information as basic input, with structural descriptors as the target, a large language model is used for training; the recognition accuracy of the model is optimized through manual verification and feedback mechanisms.

[0045] Step S3: Using the trained large language model, assign weights to different structural descriptor calculation methods based on their frequency of use in the literature; consider the target material type and expected performance, and adjust the selection priority of structural descriptor calculation methods; output the structural descriptor calculation method with the highest comprehensive score and its parameters.

[0046] Step S4: Input the chemical formula CoCrFe2MnNi of the high-entropy alloy material to be calculated. The system extracts the element types, composition, and proportion information used in the calculation structure descriptor, which are Co, Cr, Fe, Mn, and Ni elements, with proportions of 0.167, 0.167, 0.332, 0.167, and 0.167, respectively.

[0047] Step S5: Based on the identified and trained structure descriptor calculation method, and combined with the aforementioned element types and proportion information, calculate the structure descriptor of this high-entropy alloy material, which are the atomic size difference parameter δ=0.0121 and the mixing enthalpy ΔH. mix =-3.81 kJ / mol, mixing entropy ΔS mix =1.56R, valence electron concentration VEC=8.0, electronegativity difference=0.13.

[0048] Example 2

[0049] A method for calculating the structure descriptor of high-entropy materials using a large language model, taking the calculation of the structure descriptor (method 2) for a specific high-entropy oxide material ((NbCeGdSnTi)O2) as an example, specifically includes the following steps:

[0050] Step S1: Input the keywords "High-entropy alloy", "descriptors", and "machine learning". The system will automatically search for research literature and download the original texts through Web of Science and Google Scholar. The system will then use a large language model to extract information from the downloaded research literature, including the elemental composition, structural descriptors, calculation methods, and applicable conditions of high-entropy alloy materials. Based on the extracted high-entropy alloy element types, the system will extract basic elemental information from open-source databases. Finally, the system will build or update a knowledge base of high-entropy alloy structural descriptors.

[0051] Step S2: Based on the knowledge base, using structured and unstructured high-entropy material literature data, with the corresponding material element composition as the target, a large language model is used for training; using the extracted high-entropy material element types, composition, and proportion information as basic input, with structural descriptors as the target, a large language model is used for training; the recognition accuracy of the model is optimized through manual verification and feedback mechanisms.

[0052] Step S3: Using the trained large language model, assign weights to different structural descriptor calculation methods based on their frequency of use in the literature; consider the target material type and expected performance, and adjust the selection priority of structural descriptor calculation methods; output the structural descriptor calculation method with the highest comprehensive score and its parameters.

[0053] Step S4: Directly upload the 3D chemical structure model (cif format) of the high-entropy oxide material to be calculated. The system extracts its chemical formula (NbCeGdSnTi)O2 from the content of the cif file and extracts the main element types, composition and proportion information used in the calculation structure descriptor, which are Nb, Ce, Gd, Sn and Ti, with a proportion of 0.2 for each.

[0054] Step S5: Based on the extracted element types and proportions, the structure descriptor for the high-entropy oxide material is calculated using the structure descriptor calculation method stored in the system's large model. These descriptors are the atomic size difference parameter δ=0.148 and the mixing enthalpy ΔH. mix =-4.56 kJ / mol, mixing entropy ΔS mix =1.61R, Valence electron concentration VEC=7.8, Electronegativity difference=0.302, Atomic radius ratio=2.329.

[0055] In summary, the method for calculating high-entropy material structure descriptors using a large language model provided by this invention has advantages such as simple operation, reliable data, and high speed. This invention solves the technical problems of difficult selection and cumbersome calculation of high-entropy material structure descriptors, and provides an efficient and accurate feature extraction method for machine learning-aided design of high-entropy materials, which has significant scientific significance and application value.

[0056] It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects may be any one or a combination of the above, or any other possible beneficial effects.

[0057] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is merely an example and does not constitute a limitation of this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements and corrections to this application. Such modifications, improvements and corrections are suggested in this specification, and therefore such modifications, improvements and corrections still fall within the spirit and scope of the exemplary embodiments of this application.

[0058] Furthermore, when terms such as "S1," "S2," and "S3" are used in this application specification to describe various features, these terms are only used to distinguish these features and should not be construed as indicating or implying the correlation or relative importance between features, or implicitly specifying the number of features indicated. In fact, the features of the embodiments are fewer than all the features of the single embodiments disclosed above.

[0059] Finally, it should be understood that the embodiments described in this application are merely illustrative of the principles of the embodiments of this application. Other modifications may also fall within the scope of this application. Therefore, alternative configurations of the embodiments of this application are considered as examples and not limitations, and are regarded as consistent with the teachings of this application. Accordingly, the embodiments of this application are not limited to the embodiments explicitly described and illustrated in this application.

Claims

1. A method for calculating high-entropy material structure descriptors using a large language model, characterized in that, include: Step S1: Search for research literature related to high-entropy materials based on keywords, and construct a knowledge base of high-entropy material structure descriptors; Step S2: Use a large language model to train the knowledge base of high-entropy material structure descriptors in step (1), and identify and extract the related information of high-entropy material structure descriptors; Step S3: Using the trained large language model, identify the most frequently used high-entropy material structure descriptors and train the corresponding calculation methods. Step S4: Input the 3D chemical structure model or chemical structure formula of the high-entropy material to be calculated, and extract the element types and proportion information used to calculate the high-entropy material structure descriptor; Step S5: Calculate the corresponding structure descriptor for the high-entropy material based on the identified and trained high-entropy material structure descriptor calculation method, combined with the element type and proportion information.

2. The method for calculating high-entropy material structure descriptors using a large language model according to claim 1, characterized in that, In step S1, relevant research literature is retrieved based on keywords, and custom settings can be made as needed.

3. The method for calculating high-entropy material structure descriptors using a large language model according to claim 2, characterized in that, In step S1, constructing a knowledge base for high-entropy material structure descriptors includes: collecting research literature in the field of high-entropy materials based on keywords; extracting the material element composition, structure descriptors, calculation methods, and applicable conditions used in the research literature; and calling the corresponding component information from open-source databases.

4. The method for calculating high-entropy material structure descriptors using a large language model according to claim 3, characterized in that, The structure descriptor includes atomic size difference parameters. δ Mixed enthalpy ΔH mix Mixed entropy ΔS mix Valence electron concentration VEC One or more of the following: electronegativity difference, atomic radius ratio, crystal structure parameters, topological descriptor, and electronic structure descriptor.

5. The method for calculating high-entropy material structure descriptors using a large language model according to claim 4, characterized in that... In step S2, training the large language model includes: inputting structured and unstructured high-entropy material literature data; setting the training objective as recognizing the association between material composition and applicable structural descriptors; and optimizing the model's recognition accuracy through manual verification and feedback mechanisms.

6. The method for calculating high-entropy material structure descriptors using a large language model according to claim 5, characterized in that, In step S3, identifying the most frequently used structural descriptor calculation method includes: using a large language model to extract relevant structural descriptors from the literature; assigning weights to different structural descriptor calculation methods based on the frequency of use in the literature; adjusting the selection priority of structural descriptor calculation methods considering the target material type and expected performance; and outputting the structural descriptor calculation method with the highest comprehensive score and its parameters.

7. The method for calculating high-entropy material structure descriptors using a large language model according to claim 1, characterized in that, In step S4, the 3D chemical structure model of the high-entropy material to be calculated includes common 3D atomic structure model format files, which can automatically identify and extract the chemical structural formula corresponding to the 3D chemical structure model.

8. The method for calculating high-entropy material structure descriptors using a large language model according to claim 7, characterized in that, In step S4, the information required for calculating the structure descriptor is automatically extracted from the corresponding chemical structural formula automatically identified from the 3D chemical structural model of the high-entropy material to be calculated or from the directly provided chemical structural formula.

9. The method for calculating high-entropy material structure descriptors using a large language model according to claim 1, characterized in that, A knowledge base built from research literature collected based on keywords can be automatically updated periodically according to settings.

10. The method for calculating high-entropy material structure descriptors using a large language model according to claim 1, characterized in that, The large language model covers different types of high-entropy materials, and the training process employs multiple rounds of manual verification and feedback optimization.