Chemical material data processing method and system based on large language model
By employing a chemical materials data processing method based on a large language model, and utilizing a two-layer data processing structure and quality assessment mechanism, textual and numerical data are standardized, solving the problem of low data processing accuracy in existing technologies and achieving unified and automated data processing.
Patent Information
- Application Number
- CN202511056705.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-11
AI Technical Summary
In existing technologies, standardized processing methods rely on expert rules or template matching, which can easily lead to information omissions or incorrect matching when dealing with unstructured text, resulting in low accuracy in data processing.
A chemical materials data processing method based on a large language model is adopted. The text and numerical data are standardized through a two-layer data processing structure. The terminology and units are unified by using a chemical terminology standard database and a unit conversion relation database. A material quality assessment mechanism is introduced to generate a data quality index.
It achieves semantic consistency of heterogeneous data, improves the flexibility and automation of data processing, and ensures the accuracy and consistency of data quality.
Smart Images

Figure CN120930651A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic processing technology, specifically to a method and system for processing chemical material data based on a large language model. Background Technology
[0002] In the field of chemical materials, heterogeneity in experimental data and literature is prevalent, including inconsistent data formats, differing terminology, inconsistencies in units, and a lack of naming conventions. This severely impacts the efficiency of automated data processing, analysis, and reuse, hindering the in-depth development of intelligent materials research, knowledge graph construction, and high-throughput computing. Currently, data processing methods largely rely on expert rules or template matching, which perform well with structured data but are prone to information omissions or mismatches with unstructured text. Because they depend on predefined rules, when new terms or expressions exist in the data that are not covered by the rules, these rules fail to capture the information, leading to omissions. Furthermore, since rules are often manually formulated and adjusted, end-to-end automation is difficult to achieve, and they lack self-learning and evolutionary capabilities. When faced with new data formats, terminology, or units, mismatches easily occur, affecting the accuracy of data processing.
[0003] In summary, existing technologies suffer from the problem that standardized processing methods rely on expert rules or template matching, which can easily lead to information omissions or incorrect matching when dealing with unstructured text, resulting in low accuracy in data processing. Summary of the Invention
[0004] This application provides a chemical materials data processing method and system based on a large language model, which addresses the technical problem in the prior art where standardized processing methods rely on expert rules or template matching, which easily leads to information omissions or incorrect matching when dealing with unstructured text, resulting in low accuracy of data processing.
[0005] In view of the above problems, this application provides a method and system for chemical material data processing based on a large language model.
[0006] Firstly, this application provides a chemical material data processing method based on a large language model. This method is implemented through a chemical material data processing system based on a large language model. The method includes: acquiring chemical material data, which comprises text data and numerical data; activating a first channel layer in a two-layer data processing structure to standardize the text data, obtaining standardized text; activating a second channel layer in the two-layer data processing structure to standardize the numerical data, obtaining standardized values; and evaluating and analyzing the standardized text and standardized values according to a material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
[0007] Optionally, a first chemical entity is extracted from the text data; the first chemical entity is standardized by combining it with the chemical terminology standard database embedded in the first channel layer to obtain a first target entity; and the standardized text is formed based on the first target entity.
[0008] Optionally, the first chemical entity is traversed in the chemical terminology standard database to obtain a first traversal result; if the first traversal result does not conform to the predetermined traversal constraints, an interaction instruction is issued; a first supplementary question about the first chemical entity is generated based on the interaction instruction, and a first supplementary answer to the first supplementary question is obtained; the first chemical entity is then subjected to equivalent replacement processing in combination with the first supplementary answer.
[0009] Optionally, the unit conversion relation database embedded in the second channel layer is obtained; the unit conversion relation database is visualized using units as nodes and conversion relations as edges to obtain a directed graph of unit conversion; a first chemical value is extracted from the numerical data, and the first chemical value corresponds to a first chemical unit; the first chemical value is standardized by coordinating the directed graph of unit conversion and the first chemical unit to obtain a first target value; the standardized value is formed based on the first target value.
[0010] Optionally, a first node corresponding to the first chemical unit is matched in the unit conversion directed graph; a first standard unit corresponding to the first chemical unit is obtained, and a second node corresponding to the first standard unit is matched in the unit conversion directed graph; a conversion path optimization is performed with the first node as the initial node and the second node as the target node to obtain a first optimal conversion path; the first chemical value is standardized according to the first optimal conversion path to obtain the first target value.
[0011] Optionally, a first set of adjacent nodes of the initial node is constructed, wherein the first set of adjacent nodes includes a third node; the path from the initial node to the third node is calculated to obtain the exact path length; the path between the third node and the target node is predicted to obtain the predicted path length; the exact path length and the predicted path length are summed to obtain the third converted path of the third node; the third node is sorted in ascending order based on the third converted path to obtain an ascending sequence list; the first node in the ascending sequence list is taken to construct the first optimal converted path.
[0012] Optionally, a first set of relevant index parameters for the first chemical value is constructed; a first parameter corresponding to the first relevant index in the first set of relevant index parameters is matched in the numerical data; a first predetermined correlation coefficient corresponding to the first relevant index is retrieved and combined with the first parameter to obtain a first verification value; the first chemical value is verified using the first verification value to obtain a first effective support; if the first effective support does not reach the predetermined support limit, a first warning instruction is issued; and the first chemical value is subjected to anomaly warning and repair processing based on the first warning instruction.
[0013] Optionally, the physical quantity mapping relationship database embedded in the second channel layer is obtained; a second chemical value is extracted from the numerical data, and the second chemical value corresponds to a second chemical unit; according to the physical quantity mapping relationship database, it is determined whether the first chemical unit and the second chemical unit have a mapping relationship; if they do, a third chemical value is derived by combining the first chemical value and the second chemical value; the numerical data is verified using the third chemical value to obtain a verification result; if the verification result does not meet the predetermined conditions, a second warning instruction is issued; and the numerical data is subjected to abnormal warning and repair processing based on the second warning instruction.
[0014] Optionally, the frequency of issuing the interaction command, the first warning command, and the second warning command is statistically analyzed according to the material quality assessment mechanism; the frequency of issuing the command is normalized to obtain the data quality index.
[0015] Secondly, this application also provides a chemical material data processing system based on a large language model, used to execute the chemical material data processing method based on a large language model as described in the first aspect, wherein the chemical material data processing system based on a large language model includes: a data acquisition module for acquiring chemical material data, wherein the chemical material data includes text data and numerical data; a text standardization processing module for activating a first channel layer in a two-layer data processing structure to standardize the text data to obtain standardized text; a numerical standardization processing module for activating a second channel layer in the two-layer data processing structure to standardize the numerical data to obtain standardized numerical values; and a data quality assessment module for evaluating and analyzing the standardized text and the standardized numerical values according to a material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
[0016] One or more technical solutions provided in this application have at least the following beneficial effects: By acquiring chemical material data, which includes text data and numerical data; activating the first channel layer in a two-layer data processing structure to standardize the text data, obtaining standardized text; activating the second channel layer in the two-layer data processing structure to standardize the numerical data, obtaining standardized values; and evaluating and analyzing the standardized text and standardized values according to a material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data. In other words, by standardizing chemical material data through a two-layer data processing structure, all data is unified under the same knowledge framework, ensuring semantic consistency of heterogeneous data from different sources and in different formats. Introducing a material quality assessment mechanism to evaluate standardized values and obtain a data quality index improves the flexibility and automation level of data processing. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating the chemical materials data processing method based on a large language model, as described in this application.
[0019] Figure 2 This is a schematic diagram of the chemical materials data processing system based on a large language model, as described in this application.
[0020] Figure labeling: Data acquisition module 11, text standardization processing module 12, numerical standardization processing module 13, data quality assessment module 14. Detailed Implementation
[0021] This application provides a chemical materials data processing method and system based on a large language model to address the technical problem of low accuracy in existing technologies where standardization methods rely on expert rules or template matching, leading to information omissions or mismatches when dealing with unstructured text. The method employs a two-layer data processing structure to standardize chemical materials data, unifying all data within the same knowledge framework. This ensures semantic consistency across heterogeneous data from different sources and in different formats. A material quality assessment mechanism is introduced to evaluate the standardized values, yielding a data quality index, thereby improving the flexibility and automation of data processing.
[0022] Example 1, as Figure 1 As shown, this application provides a chemical material data processing method based on a large language model. The method is applied to a chemical material data processing system based on a large language model, and specifically includes the following steps:
[0023] Acquire chemical material data, wherein the chemical material data includes text data and numerical data.
[0024] Specifically, raw data related to chemical materials is collected from various sources (such as laboratory measurements, public databases, and literature). This includes information describing the properties, structure, performance, reactions, and synthesis methods of chemical materials. Chemical material data can be broadly categorized into textual data and numerical data. Textual data is descriptive information presented in written form, typically including the properties, reaction processes, experimental methods, and research results of chemical materials. It can be unstructured textual data such as articles, experimental reports, papers, and patents. Numerical data is quantifiable chemical material property data, usually expressed in numerical form, such as density, melting point, electrical conductivity, and the coefficient of thermal expansion of materials. It is obtained through experimental measurements or calculations.
[0025] The first channel layer in the two-layer data processing structure is activated to perform standardization processing on the text data, resulting in standardized text.
[0026] Furthermore, this application also includes the following steps: extracting a first chemical entity from the text data; standardizing the first chemical entity by combining it with the chemical terminology standard database embedded in the first channel layer to obtain a first target entity; and forming the standardized text based on the first target entity.
[0027] Furthermore, this application also includes the following steps: traversing the first chemical entity in the chemical terminology standard database to obtain a first traversal result; if the first traversal result does not conform to a predetermined traversal constraint, issuing an interaction instruction; generating a first supplementary question about the first chemical entity based on the interaction instruction, and obtaining a first supplementary answer to the first supplementary question; and performing an equivalent replacement process on the first chemical entity in conjunction with the first supplementary answer.
[0028] Specifically, the Chemical Terminology Standardization Database is a knowledge base dedicated to storing information on chemical terminology, naming rules, unit conversions, and physical quantity mappings. It includes standardized expressions for common chemical names, material names, physical quantities, and various chemical reactions encountered in chemical materials and experiments. The database encompasses common chemical terms and material names (such as chemical substance names, abbreviations, and molecular formulas), physical quantities and performance indicators (such as density, melting point, and electrical conductivity, with standardized naming rules), chemical reactions, and other important terms (standardizing different chemical reaction formulas and experimental procedures to ensure uniformity and accuracy). The database is constructed using existing standards (such as IUPAC naming rules and SI unit systems), mainstream databases (such as the Materials Project, PubChem, and NIST), and manually compiled glossaries.
[0029] The two-layer data processing structure divides the data processing process into two main processing channels, each responsible for standardizing different types of data. The first channel layer specifically processes text data, while the second channel layer handles numerical data. Optimization is performed for different data types to improve processing efficiency and quality. The first channel layer is dedicated to text data processing, including terminology standardization, vocabulary unification, and naming standardization. It ensures that chemical terms in the text data conform to existing standards or domain specifications by calling a chemical terminology standardization database. This database serves as the core module of the first channel layer, providing standardization support for chemical terms. When text data enters the first channel layer, the chemical terminology standardization database is automatically invoked to retrieve and standardize the chemical terms, units, and names appearing in the text. For example, when encountering polyethylene, PE, or Polyethylene, it will be automatically identified and standardized as polyethylene.
[0030] The first chemical entity is randomly extracted from the text data, such as the name of a chemical substance, its chemical formula, or the name of a material. This includes the name of the chemical substance (e.g., polyvinyl alcohol, titanium dioxide), the symbol of a chemical element (e.g., Fe, Cu), the chemical formula (e.g., H₂O), material performance indicators (e.g., dielectric constant), and chemical processes or operations (e.g., polymerization reactions). The first chemical entity is a chemical entity in the text data that serves as the first target to be identified.
[0031] The first chemical entity is traversed through the chemical terminology standard database, yielding the first traversal result. This result may be an explicit match (a corresponding standardized entry was found), a fuzzy match (multiple possible entries were found), or no match at all. Predefined traversal constraints are pre-set standards or criteria used to determine the acceptability of the traversal result, such as requiring the finding of a unique and explicit standardized entry; allowing multiple possible entries for further confirmation; and disallowing ambiguity.
[0032] When the first traversal result does not meet the predetermined traversal constraints, it indicates that the standardized expression corresponding to the first chemical entity has not been found in the chemical terminology standard database. An interactive instruction is automatically generated to trigger subsequent interactive behaviors, instructing the system to ask a question to the large language model to obtain more information. Based on the interactive instruction, a first supplementary question for the first chemical entity is generated, aiming to confirm more contextual information about the first chemical entity. A corresponding first supplementary answer is obtained based on the first supplementary question, and the first chemical entity is replaced with a more accurate expression after clarification. For example, if the first chemical entity is BTO, and no match is found in the chemical terminology standard database, the large language model initially judges it to be barium titanium oxide, but it could also refer to other substances (such as bis-perovskite oxide). Due to the lack of clear contextual clues (such as no mention of monomers or specific applications), the certainty of its judgment is low. Therefore, an interactive instruction is issued to generate a first supplementary question for the first chemical entity, such as requesting confirmation of the chemical formula or crystal structure type of BTO, and this is sent to the large language model; the large language model returns a first supplementary answer, such as the chemical formula BaTiO, which belongs to the perovskite structure. This process of asking and answering supplementary questions may not be confirmed in one go and needs to be repeated multiple times until the uncertainty is eliminated or the preset limit for the number of interaction rounds is reached. Each round of interaction may provide new information, guiding the large language model to gradually focus on the most likely interpretation. After multiple rounds of interaction, the accuracy and reliability of the final identification, interpretation, or standardization results of the chemical entity are significantly improved compared to the initial uncertain results. Based on the specific meaning determined by the first supplementary answer, the first chemical entity is replaced with an equivalent one, ensuring that the representation of the entity in the text is consistent, standardized, and correct.
[0033] By combining a standardized chemical terminology database, the first chemical entity is standardized. This involves converting it into a unified and standardized form according to the standardized expressions specified in the database, including standardized nomenclature of chemical substances and terms for the same properties. This eliminates heterogeneity in the original text and ensures consistent information expression. The first target entity is the standardized chemical entity obtained after standardization, representing the chemical concept pointed to by the first chemical entity in the original text, but existing in a unique standard form. For example, PVA, Polyvinyl Alcohol, and Polyvinyl Alcohol all exist in the form of Polyvinyl Alcohol (CAS: 9002-89-5). Standardized text is generated based on the first target entity. The aforementioned steps are repeated for all chemical entities in the text data, replacing them with unified standard terms to ensure the consistency and accuracy of the text data.
[0034] In other words, word embedding technology converts text into vector form. Combined with a standardized chemical terminology database, it identifies and standardizes terms such as material names and performance descriptions. Simultaneously, it normalizes units and numerical expressions in the text, unifying and standardizing the names of material synthesis processes with different expressions. Word embedding technology is a technique that converts words in text into numerical vectors, ensuring that words with similar meanings are close together in the vector space. Techniques like Word2Vec, GloVe, and BERT capture the context and semantic relationships of words. Input text data (such as material names and performance descriptions) is converted into word vectors. Through word embedding techniques (such as Word2Vec and BERT), each word is converted into a high-dimensional vector, capturing its semantic information and ensuring that similar or related words are located close together in the vector space. For example, polyethylene is a common plastic material with good electrical conductivity. Polyethylene and other plastic materials have a small distance in the vector space, indicating a similar semantic relationship. The converted vectors are then identified, standardized, and normalized using a standardized chemical terminology database. To enable large language models to accurately understand and process input text, dynamic prompts are generated based on the domain characteristics of the text and the processing task, guiding the large language model to focus on specific knowledge domains. For example, when processing text on organic chemistry materials, prompts are automatically generated, including those related to IUPAC nomenclature rules; when processing material performance test reports, prompts emphasizing unit conversions and numerical standardization are generated to ensure numerical consistency during processing. Through the dynamic prompt generation module, the large language model is guided to accurately understand text data from specific domains.
[0035] The second channel layer in the dual-layer data processing structure is activated to standardize the numerical data, resulting in standardized values.
[0036] Furthermore, this application also includes the following steps: obtaining the unit conversion relation database embedded in the second channel layer; visualizing the unit conversion relation database with units as nodes and conversion relations as edges to obtain a directed unit conversion graph; extracting a first chemical value from the numerical data, wherein the first chemical value corresponds to a first chemical unit; standardizing the first chemical value by coordinating the directed unit conversion graph and the first chemical unit to obtain a first target value; and forming the standardized value based on the first target value.
[0037] Specifically, in the two-layer data processing structure, the second layer is dedicated to processing numerical data, performing unit unification, numerical range standardization, and unit conversion. This second layer ensures that numerical data from different sources and represented in different units can be compared and analyzed under the same standard. The unit conversion database is a database containing conversion relationships between different units, defining conversion rules from one unit to another. For example, it defines g / cm³. 3 and kg / m 3 It can automatically convert between units, such as the relationship between temperature units Celsius and Kelvin, to ensure the consistency and standardization of numerical data.
[0038] Information on unit conversions is retrieved from the embedded unit conversion database, including conversion rules and relationships between different units, such as 1 g / cm³. 3 =1000kg / m 3 This method transforms the units and conversion relationships in the database into a directed graph, with units as nodes and conversion relationships as edges. Each edge has a direction, representing the conversion process from one unit to another, providing a visual representation of the conversion path between units. This is particularly useful for complex unit conversion scenarios, such as kg / (m·s). 2 This system automatically parses complex units, breaks them down into combinations of basic units, generates the optimal conversion path using a unit conversion database, and completes the unit conversion step by step. During numerical processing, it provides confidence intervals based on the calculated values, representing the error range of those values. This is particularly helpful for data obtained through indirect measurements, calculating a confidence interval to assess the reliability of the data. For example, a density calculated as 1.2 g / cm³... 3 The output confidence interval is 1.2 ± 0.1 g / cm³. 3 This indicates that the numerical value has a certain degree of uncertainty. A unit conversion directed graph is a mathematical structure in which each edge has a direction, representing the conversion path from one unit to another.
[0039] The first chemical value is randomly extracted from the numerical data. The first chemical value corresponds to the first chemical unit, such as a density of 1.2 g / cm³. 3 The first chemical value was 1.2, and the first chemical unit was g / cm³. 3 Based on the first chemical unit, find the first node corresponding to the first chemical unit in the unit conversion directed graph. At the same time, determine the first standard unit corresponding to the first chemical unit, that is, the international standard unit used to uniformly represent the physical quantity, and find the second node corresponding to the first standard unit in the unit conversion directed graph.
[0040] Using the first node as the initial node and the second node as the target node, a conversion path optimization is performed. That is, the path from the current unit to the standard unit is searched in the directed graph of unit conversion, and the optimal conversion path (the path using the fewest steps) is selected as the first optimal conversion path. Based on the first optimal conversion path, the first chemical value is standardized, that is, it is converted to a value in the standard unit to obtain the first target value. For example, 1.2 g / cm³... 3 Converted to 1200kg / m 3 .
[0041] For other chemical values in the numerical data, repeat the above steps to form standardized values. Standardized values are the result of uniformly standardizing numerical data, typically including numerical conversion and unit unification to conform to standard units. Through standardization, all numerical data are ensured to conform to a unified standard unit, avoiding comparison difficulties and errors caused by different units.
[0042] Furthermore, this application also includes the following steps: matching a first node corresponding to the first chemical unit in the unit conversion directed graph; obtaining a first standard unit corresponding to the first chemical unit, and matching a second node corresponding to the first standard unit in the unit conversion directed graph; performing conversion path optimization with the first node as the initial node and the second node as the target node to obtain a first optimal conversion path; and standardizing the first chemical value according to the first optimal conversion path to obtain the first target value.
[0043] Furthermore, this application also includes the following steps: constructing a first set of adjacent nodes of the initial node, wherein the first set of adjacent nodes includes a third node; calculating the path from the initial node to the third node to obtain the exact path length; predicting the path between the third node and the target node to obtain the predicted path length; summing the exact path length and the predicted path length to obtain the third converted path of the third node; sorting the third node in ascending order based on the third converted path to obtain an ascending sequence list; taking the first node in the ascending sequence list to construct the first optimal converted path.
[0044] Specifically, in the unit conversion directed graph, the first node corresponding to the first chemical unit is matched; that is, the node representing the first chemical unit. The first standard unit corresponding to the first chemical unit is then obtained, which is the standard unit used to uniformly represent this physical quantity (such as the SI unit system). For example, the standard unit of density is kg / m³. 3 Instead of g / cm 3 Or other units. Similarly, in the directed graph of unit conversion, match the second node corresponding to the first standard unit, that is, the node representing the first standard unit.
[0045] The first node is the starting unit of the conversion process, and the second node is the target unit. The conversion path leads from the first node to the second node. The nodes directly connected to the first node are examined, forming the first set of adjacent nodes. In a graph structure, an adjacent node is a node directly connected to the current node. The first set of adjacent nodes represents all nodes directly connected to the first node; these nodes can be considered as the next step in the conversion path. The first set of adjacent nodes includes the third node. The path from the initial node to the third node is calculated to obtain the exact path length, i.e., the actual conversion path length from the first node to the third node. For example, if the conversion is from g / cm... 3 up to kg / m 3 After a direct conversion step, the exact path length is 1. Using a prediction algorithm, the path length is predicted based on the conversion relationship between the current node (third node) and the target node (second node), determining the path complexity required to reach the target node from the third node. The predicted path length is an estimate or prediction of the path length required to reach the target node (representing standard units) from the third node, predicted through a heuristic method or based on the statistical regularity of the average distance between nodes in the graph. For example, if the first node (initial node) is g / cm... 3 The second node (target node) is kg / m 3 The third node is g / m 3 If the predicted path length between the third node and the target node is 1 step; if the first node is g / cm 3 The second node is kg / m 3 The third node is g / dm 3 Therefore, the predicted path length between the three nodes and the target node may be 2 steps.
[0046] The exact path length is added to the predicted path length to obtain the third converted path for the third node. For other nodes in the first set of adjacent nodes, the above steps are repeated to obtain multiple converted paths. All possible converted paths are sorted in ascending order, that is, ordered by path length (number of conversion steps) from smallest to largest, and the shortest path is selected as the final converted path. In simple terms, the shortest path is found in the set of adjacent nodes of each node to obtain the first optimal converted path, which serves as the optimal converted path from the first node to the target node, representing the most efficient and shortest conversion method from one unit to another.
[0047] The first chemical value is standardized according to the first optimal conversion path to obtain the first target value. For example, if the density of a certain novel polymer film in the experimental report is 1.2 g / cm³... 3 All density data must be expressed in kg / m³ 3 Stored in units, recognizing the original value 1.2 and the original unit g / cm³. 3 In the unit conversion directed graph, find the corresponding nodes, then find the optimal path based on the nodes, and calculate the conversion factor as 0.001 / (0.01). 3 =1000, multiply the original value 1.2 by the composite factor 1000 to obtain the standardized value 1200 kg / m³. 3 For complex units, the large language model generates the optimal conversion path based on the knowledge base, such as kg / (m·s). 2 First, break it down into N / m 2 The value is then converted to Pa. Accurate and automated conversion of composite unit values is achieved by utilizing a directed graph of unit conversion and an optimal path optimization algorithm. During numerical processing, the large language model provides uncertainty quantification indicators while outputting numerical values. For example, for parameters calculated through indirect measurement, a confidence interval is given to help subsequent analysis and evaluation of data quality. A context-aware anomaly detection and collaborative repair mechanism is provided. Anomalies in numerical values are identified by combining domain knowledge and textual context. For example, when the text describes a lightweight material but the density value is too high, an anomaly alarm is triggered. When an abnormal value is detected, a repair decision is made by integrating the outputs of multiple domain models. Unit dimension conversion (e.g., converting g / cm³) is implemented. 3 kg / m 3 Convert to SI units kg / m 3 This includes numerical normalization (such as adjusting the melting point under different test conditions to the value under standard atmospheric pressure) to solve the problems of inconsistent physical quantity units and numerical benchmarks, and to ensure the physical comparability of data within the same knowledge framework.
[0048] Furthermore, this application also includes the following steps: constructing a first set of relevant index parameters for the first chemical value; matching the first parameter corresponding to the first relevant index in the first set of relevant index parameters in the numerical data; retrieving the first predetermined correlation coefficient corresponding to the first relevant index and combining it with the first parameter to obtain a first verification value; verifying the first chemical value with the first verification value to obtain a first effective support; if the first effective support does not reach the predetermined support limit, issuing a first warning command; and performing abnormal warning and repair processing on the first chemical value based on the first warning command.
[0049] Obtain the physical quantity mapping relationship database embedded in the second channel layer; extract the second chemical value from the numerical data, and the second chemical value corresponds to the second chemical unit; determine whether the first chemical unit and the second chemical unit have a mapping relationship according to the physical quantity mapping relationship database; if they do, deduce the third chemical value by combining the first chemical value and the second chemical value; verify the numerical data with the third chemical value to obtain the verification result; if the verification result does not meet the predetermined conditions, issue a second warning command; perform abnormal warning and repair processing on the numerical data based on the second warning command.
[0050] Specifically, a set of first relevant index parameters related to the first chemical value is obtained, which includes other variables related to that chemical value. For example, indices related to density (first chemical value) include the material's melting point, glass transition temperature, tensile strength, Young's modulus, etc. A first relevant index is randomly selected from the set of first relevant index parameters, and the first parameter corresponding to the first relevant index is matched in the numerical data; that is, the actual measured or recorded value corresponding to the first relevant index (tensile strength). The first predetermined correlation coefficient is a preset coefficient relating to the relationship between the first relevant index and the first chemical value, representing the expected correlation strength or influence between the first relevant index and the first chemical value. For example, assuming a predetermined temperature coefficient of 0.02, it means that the influence coefficient of temperature on density is 0.02.
[0051] A first verification value is calculated by combining a first parameter and a first predetermined correlation coefficient. This value represents the theoretical value calculated based on the predetermined correlation coefficient and the actual parameter, and is used to determine the accuracy of the first chemical value. For example, if the first parameter (temperature) is 25°C and the first predetermined correlation coefficient is 0.02, the first verification value might be the theoretical value of density as a function of temperature: The theoretical density = current density + 0.02 * 25 = actual density + 0.5. The first valid support is calculated by comparing the difference between the first verification value and the first chemical value. If these two values are very close, the verification result is valid and the support is high. If the difference between them is large, the support is low. For example, if the actual measured density is 1.2 g / cm³... 3 The verification value was 1.7 g / cm³. 3 The support rate is 70.6%.
[0052] If the first effective support does not reach the predetermined support limit, a first warning instruction is issued, indicating that the data may have problems and requires further inspection or correction. The predetermined support limit is a pre-set threshold used to determine whether the first effective support is high enough. Only when the first effective support is higher than this limit is the first chemical value considered reasonable and reliable. The first warning instruction is an automatically triggered warning signal when the first effective support falls below the predetermined support limit.
[0053] The process involves acquiring the physical quantity mapping relationship database embedded in the second channel layer. This database contains conversion rules and mapping relationships between physical quantities, including conversion relationships between many physical quantities (such as density, temperature, and conductivity). It retrieves the rules and relationships related to the conversion between different physical quantities. A second chemical value is extracted from the numerical data; this is any chemical value different from the first chemical value and corresponding to a second chemical unit. The unit of the first chemical value (first chemical unit) is compared with the second chemical unit, and the physical quantity mapping relationship database is queried to determine if a usable mapping relationship exists between the physical quantities represented by these two units. For example, suppose the first chemical value is the bulk modulus K of a certain alloy, with a value of 120 GPa (first chemical unit), and the second chemical value is the Young's modulus E of the same alloy, with a value of 200 GPa (second chemical unit). Querying the physical quantity mapping relationship database reveals that for isotropic materials, there is a mapping relationship between the bulk modulus K and Young's modulus E through Poisson's ratio ν: E = 3K(1-2ν). Because this mapping relationship exists, it is determined that a mapping relationship exists.
[0054] By querying the physical quantity mapping relationship database, it is determined whether a valid mapping relationship exists between the first and second chemical units. If a mapping relationship exists, these two chemical values are combined to derive a third chemical value through the physical quantity mapping relationship. Using the mapping relationship formulas provided in the physical quantity mapping relationship database, and combining known or estimated intermediate parameters, a calculation is performed to obtain a third chemical value. The third chemical value is a new value derived by combining the first and second chemical values and applying the mapping relationship.
[0055] The numerical data is validated using a third chemical numerical method to obtain validation results. If the calculated value is within a reasonable range, the validation passes; if it deviates significantly from the expected value, the validation fails, indicating an anomaly. If the validation result does not meet the predetermined conditions (e.g., the derived value exceeds a reasonable range), a second warning instruction is issued, indicating data anomaly and requiring further repair. Upon receiving the second warning instruction, the anomaly warning and repair mechanism is activated, automatically repairing the data according to preset rules, or suggesting manual review by the user, including adjusting values, correcting unit conversion errors, or marking data as awaiting manual review. Multi-model collaborative decision-making (e.g., combining expert rules and language models to understand context) is invoked to generate a more intelligent repair plan, and this information is recorded for subsequent manual review or automatic correction. Through the validation and warning mechanism, the inherent consistency and physical rationality of the numerical data are ensured. Utilizing the interrelationship of physical quantities in materials science, it identifies anomalous combinations that violate physical laws and are difficult to detect when viewing individual data points in isolation. Standardized values and metadata (units, confidence intervals, anomaly markers) are output in CSV format for data analysis or storage.
[0056] The standardized text and standardized numerical values are evaluated and analyzed according to the material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
[0057] Furthermore, this application also includes the following steps: statistically analyzing the frequency of the issuance of the interaction command, the first warning command, and the second warning command according to the material quality assessment mechanism; and normalizing the frequency of the command issuance to obtain the data quality index.
[0058] Specifically, the materials quality assessment mechanism is a module used to evaluate the quality of chemical materials data. Based on preset rules, standards, and algorithms, it analyzes and evaluates the data to determine whether it meets quality standards. Standardized text is text data that has undergone numerical verification and formatting. All numerical values and related information have conformed to predetermined standard specifications, ensuring that data from different sources can be compared and analyzed. Standardized numerical values are numerical data that have undergone standardization processing, typically converted to the International System of Units (SI units). Standardized numerical values ensure the uniformity and consistency of numerical values, making different values comparable under the same standard. The materials quality assessment mechanism evaluates and analyzes standardized text and standardized numerical values, checking whether the text and values meet predetermined quality standards, such as whether the format is correct, whether the values are within a reasonable range, and whether the units are consistent.
[0059] The interactive command is issued when a standardized expression is not found in the chemical terminology standard database; the first warning command is issued when processing numerical data, based on the first set of relevant index parameters and correlation coefficients, and when the numerical relationship between the first chemical value and its relevant index is found to be inconsistent with expectations; the second warning command is issued when processing numerical data, based on the physical quantity mapping relationship database, and when the numerical relationship between the first chemical value and the second chemical value is found to be inconsistent with known physical laws or mapping relationships.
[0060] The frequency of each interactive command, first warning command, and second warning command is statistically analyzed. This frequency is then normalized to a uniform metric, facilitating comparison and analysis by adjusting the frequencies of different command types to a consistent standard. After normalization, a data quality index is generated to quantify the data quality level. A higher index indicates better data quality, while a lower index suggests more problems. For example, if the interactive command frequency is 2, the first warning command frequency is 1, and the second warning command frequency is 0, normalization yields an interactive command index of 0.2, a first warning command index of 0.1, and a second warning command index of 0. A weighted average of these values results in a data quality index of 0.85, indicating relatively good data quality. Generating a data quality index quantifies data quality, provides a clear measurement standard, identifies high-frequency anomalies, prioritizes the handling of potentially serious problems, and improves the efficiency of anomaly detection and remediation.
[0061] In summary, the chemical material data processing method based on a large language model provided in this application has the following beneficial effects: By acquiring chemical material data, which includes text data and numerical data; activating the first channel layer in the two-layer data processing structure to standardize the text data, obtaining standardized text; activating the second channel layer in the two-layer data processing structure to standardize the numerical data, obtaining standardized values; and evaluating and analyzing the standardized text and standardized values according to a material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data. In other words, by standardizing chemical material data through a two-layer data processing structure, all data is unified under the same knowledge framework, ensuring semantic consistency of heterogeneous data from different sources and in different formats. Introducing a material quality assessment mechanism to evaluate the standardized values and obtain a data quality index improves the flexibility and automation level of data processing.
[0062] Example 2: Based on the same inventive concept as the chemical material data processing method based on a large language model in Example 1, this application also provides a chemical material data processing system based on a large language model. Please refer to the appendix. Figure 2 The chemical materials data processing system based on a large language model includes:
[0063] The data acquisition module 11 is used to acquire chemical material data, wherein the chemical material data includes text data and numerical data; the text standardization processing module 12 is used to activate the first channel layer in the dual-layer data processing structure to standardize the text data to obtain standardized text; the numerical standardization processing module 13 is used to activate the second channel layer in the dual-layer data processing structure to standardize the numerical data to obtain standardized numerical values; the data quality assessment module 14 is used to evaluate and analyze the standardized text and the standardized numerical values according to the material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
[0064] Furthermore, the text standardization processing module 12 in the chemical materials data processing system based on a large language model is also used for:
[0065] Extract the first chemical entity from the text data; combine it with the chemical terminology standard database embedded in the first channel layer to standardize the first chemical entity and obtain the first target entity; form the standardized text based on the first target entity.
[0066] Furthermore, the text standardization processing module 12 in the chemical materials data processing system based on a large language model is also used for:
[0067] The first chemical entity is traversed in the chemical terminology standard database to obtain a first traversal result; if the first traversal result does not conform to the predetermined traversal constraints, an interaction instruction is issued; a first supplementary question about the first chemical entity is generated based on the interaction instruction, and a first supplementary answer to the first supplementary question is obtained; the first chemical entity is then subjected to equivalent replacement processing in combination with the first supplementary answer.
[0068] Furthermore, the numerical standardization processing module 13 in the chemical materials data processing system based on a large language model is also used for:
[0069] Obtain the unit conversion relationship database embedded in the second channel layer; visualize the unit conversion relationship database using units as nodes and conversion relationships as edges to obtain a directed unit conversion graph; extract the first chemical value from the numerical data, where the first chemical value corresponds to the first chemical unit; standardize the first chemical value by combining the directed unit conversion graph with the first chemical unit to obtain a first target value; form the standardized value based on the first target value.
[0070] Furthermore, the numerical standardization processing module 13 in the chemical materials data processing system based on a large language model is also used for:
[0071] In the directed unit conversion graph, a first node corresponding to the first chemical unit is matched; a first standard unit corresponding to the first chemical unit is obtained, and a second node corresponding to the first standard unit is matched in the directed unit conversion graph; a conversion path optimization is performed with the first node as the initial node and the second node as the target node to obtain a first optimal conversion path; the first chemical value is standardized according to the first optimal conversion path to obtain the first target value.
[0072] Furthermore, the numerical standardization processing module 13 in the chemical materials data processing system based on a large language model is also used for:
[0073] A first set of adjacent nodes of the initial node is constructed, wherein the first set of adjacent nodes includes a third node; the path from the initial node to the third node is calculated to obtain the exact path length; the path between the third node and the target node is predicted to obtain the predicted path length; the exact path length and the predicted path length are summed to obtain the third converted path of the third node; the third node is sorted in ascending order based on the third converted path to obtain an ascending sequence list; the first node in the ascending sequence list is taken to construct the first optimal converted path.
[0074] Furthermore, the numerical standardization processing module 13 in the chemical materials data processing system based on a large language model is also used for:
[0075] A first set of relevant index parameters is constructed for the first chemical value; a first parameter corresponding to the first relevant index in the first set of relevant index parameters is matched in the numerical data; a first predetermined correlation coefficient corresponding to the first relevant index is retrieved and combined with the first parameter to obtain a first verification value; the first chemical value is verified using the first verification value to obtain a first effective support; if the first effective support does not reach the predetermined support limit, a first warning instruction is issued; and the first chemical value is subjected to anomaly warning and repair processing based on the first warning instruction.
[0076] Furthermore, the numerical standardization processing module 13 in the chemical materials data processing system based on a large language model is also used for:
[0077] Obtain the physical quantity mapping relationship database embedded in the second channel layer; extract the second chemical value from the numerical data, and the second chemical value corresponds to the second chemical unit; determine whether the first chemical unit and the second chemical unit have a mapping relationship according to the physical quantity mapping relationship database; if they do, deduce the third chemical value by combining the first chemical value and the second chemical value; verify the numerical data with the third chemical value to obtain the verification result; if the verification result does not meet the predetermined conditions, issue a second warning command; perform abnormal warning and repair processing on the numerical data based on the second warning command.
[0078] Furthermore, the data quality assessment module 14 in the chemical materials data processing system based on a large language model is also used for:
[0079] The frequency of the interactive command, the first warning command, and the second warning command is statistically analyzed based on the material quality assessment mechanism; the data quality index is obtained by normalizing the command issuance frequency.
[0080] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Figure 1 The chemical material data processing method and specific examples based on large language models in Example 1 are also applicable to the chemical material data processing system based on large language models in this example. Through the foregoing detailed description of the chemical material data processing method based on large language models, those skilled in the art can clearly understand the chemical material data processing system based on large language models in this example. Therefore, for the sake of brevity, it will not be described in detail here.
[0081] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0082] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of this application and its equivalents, this application also intends to include such modifications and variations.
Claims
1. A chemical materials data processing method based on a large language model, characterized in that, include: Acquire chemical material data, wherein the chemical material data includes text data and numerical data; The first channel layer in the two-layer data processing structure is activated to perform standardization processing on the text data, resulting in standardized text. The second channel layer in the dual-layer data processing structure is activated to perform standardization processing on the numerical data to obtain standardized values. The standardized text and standardized numerical values are evaluated and analyzed according to the material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
2. The chemical material data processing method based on a large language model according to claim 1, characterized in that, Activating the first channel layer in the two-layer data processing structure to perform standardization processing on the text data, resulting in standardized text, including: Extract the first chemical entity from the text data; By combining the chemical terminology standard database embedded in the first channel layer, the first chemical entity is standardized to obtain the first target entity; The standardized text is formed based on the first target entity.
3. The chemical material data processing method based on a large language model according to claim 2, characterized in that, Before standardizing the first chemical entity by combining it with the chemical terminology standard database embedded in the first channel layer to obtain the first target entity, the process further includes: The first chemical entity is traversed through the chemical terminology standard database to obtain the first traversal result; If the first traversal result does not meet the predetermined traversal constraints, an interactive command is issued; Based on the interactive instructions, a first supplementary question about the first chemical entity is generated, and a first supplementary answer to the first supplementary question is obtained; The first chemical entity is replaced by the same method based on the first supplementary answer.
4. The chemical material data processing method based on a large language model according to claim 3, characterized in that, Activating the second channel layer in the dual-layer data processing structure to perform standardization processing on the numerical data, obtaining standardized values, including: Obtain the unit conversion relation database embedded in the second channel layer; Using units as nodes and conversion relationships as edges, the unit conversion relationship database is visualized to obtain a directed graph of unit conversion. Extract the first chemical value from the numerical data, and the first chemical value corresponds to the first chemical unit; By coordinating the unit conversion directed graph with the first chemical unit, the first chemical value is standardized to obtain the first target value; The standardized value is formed based on the first target value.
5. The chemical material data processing method based on a large language model according to claim 4, characterized in that, By coordinating the unit conversion directed graph with the first chemical unit, and standardizing the first chemical value to obtain the first target value, the following steps are taken: In the unit conversion directed graph, match the first node corresponding to the first chemical unit; Obtain the first standard unit corresponding to the first chemical unit, and match the second node corresponding to the first standard unit in the unit conversion directed graph; Using the first node as the initial node and the second node as the target node, the conversion path is optimized to obtain the first optimal conversion path; The first chemical value is standardized according to the first optimal conversion path to obtain the first target value.
6. The chemical material data processing method based on a large language model according to claim 5, characterized in that, Using the first node as the initial node and the second node as the target node, a conversion path optimization is performed to obtain the first optimal conversion path, including: Construct a first set of adjacent nodes for the initial node, wherein the first set of adjacent nodes includes a third node; Calculate the path from the initial node to the third node to obtain the exact path length; The path between the third node and the target node is predicted to obtain the predicted path length; The exact path length and the predicted path length are summed to obtain the third converted path of the third node; The third node is sorted in ascending order based on the third conversion path to obtain an ascending sequence list; Take the first node in the ascending sequence list and construct the first optimal conversion path.
7. The chemical material data processing method based on a large language model according to claim 4, characterized in that, After forming the standardized value based on the first target value, the method further includes: Construct a first set of relevant index parameters for the first chemical value; Match the first parameter corresponding to the first relevant indicator in the first relevant indicator parameter set in the numerical data; Retrieve the first predetermined correlation coefficient corresponding to the first relevant indicator, and combine it with the first parameter to obtain the first verification value; The first chemical value is verified using the first verification value to obtain the first effective support. If the first effective support does not reach the predetermined support limit, a first warning instruction will be issued; Based on the first warning instruction, abnormal warnings and repair processing are performed on the first chemical value.
8. The chemical material data processing method based on a large language model according to claim 7, characterized in that, After performing abnormal warning and repair processing on the first chemical value based on the first warning instruction, the process further includes: Obtain the physical quantity mapping relationship database embedded in the second channel layer; Extract the second chemical value from the numerical data, and the second chemical value corresponds to the second chemical unit; Based on the physical quantity mapping relationship database, determine whether the first chemical unit and the second chemical unit have a mapping relationship; If present, the third chemical value is derived by combining the first chemical value and the second chemical value. The numerical data is verified using the third chemical value to obtain the verification result; If the verification result does not meet the predetermined conditions, a second warning instruction will be issued; Based on the second warning instruction, abnormal warnings and repair processes are performed on the numerical data.
9. The chemical material data processing method based on a large language model according to claim 8, characterized in that, The standardized text and standardized values are evaluated and analyzed according to the material quality assessment mechanism to obtain a data quality index, including: The frequency of issuing the interactive command, the first warning command, and the second warning command is statistically analyzed based on the material quality assessment mechanism. The data quality index is obtained by normalizing the frequency of the instructions issued.
10. A chemical materials data processing system based on a large language model, characterized in that, The steps for implementing the chemical material data processing method based on a large language model according to any one of claims 1 to 9, wherein the chemical material data processing system based on a large language model comprises: A data acquisition module is used to acquire chemical material data, wherein the chemical material data includes text data and numerical data; The text standardization processing module is used to activate the first channel layer in the two-layer data processing structure to standardize the text data and obtain standardized text. The numerical standardization processing module is used to activate the second channel layer in the two-layer data processing structure to standardize the numerical data and obtain standardized values. The data quality assessment module is used to evaluate and analyze the standardized text and the standardized numerical values according to the material quality assessment mechanism to obtain a data quality index, wherein the data quality index is used to quantitatively characterize the quality of the chemical material data.
Citation Information
Patent Citations
Quality evaluation method and device for operator data, server and medium
CN113064890A
Clinical information text standardization method and device, equipment and medium
CN117422074A
Medical named entity recognition and clinical term standardization method and device
CN117993391A
Transformer substation system dynamics safety assessment method considering mutual verification
CN119180004A
Smart city multi-modal data acquisition and fusion method and system
CN120197130A
Cited By
High-quality metal material process data set construction method based on large language model
CN121789818A