Text mining method and system for identifying chemical formula of material
The automatic chemical formula extraction algorithm, which involves text preprocessing, chemical element symbol marking, index shifting, and k-means clustering, solves the problems of irregular chemical formula expression and chaotic typesetting, and achieves efficient and accurate chemical formula extraction.
Patent Information
- Application Number
- CN202510941677.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-23
AI Technical Summary
Existing text analysis technologies are inefficient and inaccurate when processing complex chemical formula expressions. Chemical formulas in literature are often expressed in irregular and chaotic formats, making them difficult to identify and parse.
A text-based automatic chemical formula extraction algorithm is adopted, including text preprocessing, chemical element symbol marking, index offset processing, differential processing, k-means clustering and other steps. A systematic solution for chemical formula mining is constructed. Through the marking-transformation-clustering-extraction process, the chemical element positions are accurately marked, the chemical formula components are distinguished and the chemical formula is extracted.
It significantly improves the accuracy of chemical formula extraction in complex chemical formula text scenarios, solves the recognition problems caused by non-standardization and chaotic typesetting, and improves extraction efficiency and accuracy.
Smart Images

Figure CN120688496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a text mining method and system for identifying chemical formulas of materials. Background Art
[0002] With the rapid development of materials science, the processing and utilization of chemical information has become increasingly important in materials research. Materials science research typically involves the analysis and processing of a large number of chemical reactions and chemical formulas. Accurate and rapid acquisition of chemical formula information is crucial for developing new materials and optimizing the performance of existing materials, especially in the synthesis, modification, characterization, and application of materials. However, with the rapid increase in the amount of scientific literature, experimental reports, and patent texts, how to effectively extract key information such as chemical formulas from these massive data sources has become a major challenge facing materials researchers.
[0003] Existing text analysis technologies mostly rely on traditional text retrieval or rule-based recognition methods, which are often inefficient and inaccurate when dealing with complex chemical formulas. Chemical formulas often appear in literature in various forms, and their expressions in text can be non-standard and typographically confusing. They may include unknown numbers (such as 0.25Pb(In1 / 2Nb1 / 2)O3-(0.75-x)PbZrO3-xPbTiO3) or unexpected spaces (such as BaTi0.9Fe0.1O3), making recognition and parsing by existing algorithms extremely difficult. Summary of the Invention
[0004] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a text mining method and system for identifying material chemical formulas.
[0005] The purpose of the present invention can be achieved by the following technical solutions:
[0006] According to one aspect of the present invention, a text mining method for identifying chemical formulas of materials is provided, the method steps comprising:
[0007] S1. Obtain a set of texts to be processed, preprocess them, and output a target text set;
[0008] S2. Mark the position of each target text in the target text set according to the chemical element symbol to obtain an original vector; and delete the incorrectly marked positions in the original vector to output a marked vector;
[0009] S3, perform exponential offset processing on the marker vector in sequence and output an offset vector;
[0010] S4. Perform differential processing on the offset vector and output a differential vector;
[0011] S5. Calculate the maximum ascending subsequence length of the outlier sequence in the differential vector;
[0012] S6. Setting the number of clusters based on the maximum ascending subsequence length, performing k-means clustering on the offset vectors according to the number of clusters, and outputting a list of offset position-category pairs;
[0013] S7. Based on the offset position-category pair list, perform a binding operation on the tag vector and output a tag position-category pair list;
[0014] S8. Extract chemical formulas from the text to be processed based on the list of tag position-category pairs to complete text mining.
[0015] As a preferred technical solution, the original vector refers to a vector in which the numerical value of each dimension corresponds to the position of a character marked with a chemical element symbol; the exponential offset refers to the operation of increasing the qualified numerical value in the marked vector by an exponentially growing numerical value; the "offset position-category" pair refers to a tuple with the position of the character marked as the chemical element symbol after the exponential offset processing and the category into which this character is divided by k-means clustering as elements; the "marked position-category" tuple refers to a tuple with the position of the character marked as the chemical element symbol in the marked vector and the category into which this character is divided by k-means clustering as elements.
[0016] As a preferred technical solution, the preprocessing process in S1 includes a segmentation process and a retention process. The segmentation process divides long text segments into sentences to obtain a set of sentences to be processed. The retention process retains a subset of sentences in the set of sentences to be processed that contain chemical elements. This subset of sentences is the target text set. Specifically, the retention process uses character strings of all chemical elements to match tokens in all sentences in the set of sentences to be processed. If a token exists whose length exceeds a preset proportion of the token's length, the sentence corresponding to the token is added to the subset of sentences containing chemical elements.
[0017] As a preferred technical solution, S2 marks the position of each target text in the target text set according to the chemical element symbols, and the specific process of obtaining the original vector includes: using a preset common metal element symbol table, a preset common acid anion ion symbol table, a preset common metal cation ion symbol table and a preset common material chemical formula abbreviation table to mark the corresponding text, thereby obtaining the original vector composed of the position index of the marked characters in the character string.
[0018] As a preferred technical solution, the specific steps of deleting the position of the erroneous mark and outputting the mark vector in S2 include:
[0019] S21. Decompose the text into a number of tokens according to the space character, and delete the punctuation marks and numbers in the tokens to obtain a first token;
[0020] S22. If a first token is included in the preset banned word list, clear all tags in the text corresponding to the token;
[0021] S23. Calculate the ratio of the number of characters marked by chemical elements in the first token to the number of characters in the entire token;
[0022] S24, calculating the first token length;
[0023] S25. If the length of the first token is less than a preset token length threshold, and the first token is not in the preset ion symbol list of common metal cations, clear all marks in the text corresponding to the token;
[0024] S26. If the ratio of the number of characters marked by chemical elements in the first token to the total number of characters in the entire token is lower than a preset character ratio threshold, clear all marks in the text corresponding to the token;
[0025] S27. After traversing all tokens, obtain the tag vector.
[0026] As a preferred technical solution, in S3, the specific steps of sequentially performing exponential offset processing on the marker vectors and outputting the offset vectors include:
[0027] S31, decomposing the text into a plurality of tokens according to space characters and removing punctuation marks and numbers from the tokens to obtain second tokens;
[0028] S32. If the second token is included in the preset banned word list, clear all tags in the text corresponding to the token;
[0029] S33, calculating the ratio of the number of characters marked by chemical elements in the second token to the number of characters in the entire token;
[0030] S34, calculating the second token length;
[0031] S35. If the length of the second token is less than a preset token length threshold, and the second token is not in the preset ion symbol list of common metal cations, clear all marks in the text corresponding to the token;
[0032] S36. If the ratio of the number of characters marked by chemical elements in the second token to the total number of characters in the token is lower than a preset character ratio threshold, clear all marks in the text corresponding to the token;
[0033] S37. If the length of the token after the clearing mark is 0, and this token is the first token after the previous token whose length after the clearing mark is not 0, then add a preset index offset value to the position index in the original vector after the position index of the character corresponding to this token, and update the index offset value to twice the original value;
[0034] S38. After traversing all tokens, obtain the offset vector.
[0035] As a preferred technical solution, the offset vector in S3 has the same dimension as the differential variable in S4; and the first dimension of the differential vector is equal to 1; when i ≥ 2, the i-th dimension of the differential vector is equal to the i-th dimension of the offset vector minus the i-1-th dimension. The specific formula is:
[0036]
[0037] Among them, A[i] represents the i-th dimension of the difference vector; B[i] represents the i-th dimension of the offset vector; and i is the number of dimensions.
[0038] As a preferred technical solution, the specific process of calculating the maximum ascending subsequence length of the outlier sequence in the differential vector in S5 includes:
[0039] S51, copy the differential vector to obtain a copy vector;
[0040] S52, calculating the average of all elements in the replica vector;
[0041] S53, deleting elements in the replica vector that are greater than a preset multiple of the average calculated in S53;
[0042] S54, determine whether there is an element to be deleted in S53, if yes, return to step S52; if not, go to step S55;
[0043] S55. Retain the average of all elements in the current replica vector as the retained average;
[0044] S56. Initialize a peak list, where the peak list is: a list consisting of outliers in the difference vector;
[0045] S57, traverse all elements in the difference vector, and if there is a value greater than a preset multiple of the retained average, put it into the peak list;
[0046] S58. Calculate the maximum rising subsequence length of the peak list that is greater than a preset threshold, that is, the maximum rising subsequence length of the outlier sequence in the differential vector.
[0047] As a preferred technical solution, the preset multiples in S53 and S57 are set between 1.5 and 2.5; and the preset threshold in S58 is set within 1 times of the exponential offset value δ.
[0048] As a preferred technical solution, the specific process of S7 includes:
[0049] S71, initializing the mark position-category pair list to an empty list;
[0050] S72. Traverse the offset position-category pair list and the tag vector by position; and combine the position index in the tag vector with the category in the offset position-category pair to form a tag position-category pair;
[0051] S73. Put all the marker position-category pairs into the empty list initialized in S71, and output the marker position-category pair list.
[0052] As a preferred technical solution, the specific process of extracting the chemical formula from the text to be processed according to the list of marker position-category pairs in S8 includes:
[0053] Classify all characters between the minimum and maximum mark positions of the same category in the list of mark position-category pairs into the current category;
[0054] If different categories exist in the same token or a token is incompletely labeled, the categories of all characters in the token are uniformly modified to the category with the largest number of characters;
[0055] Characters in the same category are extracted as the same chemical formula.
[0056] According to another aspect of the present invention, a text mining system for identifying chemical formulas of materials is provided, the system comprising a text preprocessing module, an element labeling and error correction module, a vector transformation module, and a clustering and chemical formula extraction module;
[0057] The text preprocessing module is used to obtain a set of texts to be processed, preprocess them, and output a target text set;
[0058] The element marking and error correction module marks the position of each target text in the target text set according to the chemical element symbol to obtain the original vector; and deletes the position of the incorrect mark in the original vector to output the marked vector;
[0059] The vector transformation module performs exponential offset processing on the marker vector in sequence and outputs an offset vector; and performs differential processing on the offset vector and outputs a differential vector;
[0060] The clustering and chemical formula extraction module is used to calculate the maximum ascending subsequence length of the outlier sequence in the differential vector, then set the number of clusters based on the maximum ascending subsequence length, and perform k-means clustering on the offset vector according to the number of clusters, outputting a list of offset position-category pairs. Based on the offset position-category pair list, the tag vector is bound and a list of tag position-category pairs is output. The chemical formula in the text to be processed is extracted based on the tag position-category pair list to complete text mining.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] 1. In the present invention, firstly, a text set is obtained and pre-processed to obtain a target text set; then the chemical element positions are marked and error corrected to obtain a mark vector, which is then subjected to exponential offset and differential processing to obtain a differential vector; then the maximum ascending subsequence length of the outlier sequence is calculated, and the number of clusters is set based on this, and the offset vector is subjected to k-means clustering, and finally the text is marked according to the result of k-means clustering. Through the complete process of "marking-transformation-clustering-extraction" in the present invention, a systematic solution for chemical formula mining is constructed. First, the chemical element positions are accurately marked, feature associations are mined through vector transformation, and then different chemical formula components are distinguished by clustering, and finally, extraction is performed. Each link is closely linked to each other, and the chemical formula information can be effectively stripped out from the text, solving the recognition problem of chemical formulas in the text due to non-standardization, chaotic typesetting, etc., and significantly improving the accuracy of chemical formula extraction in complex chemical formula text scenarios.
[0063] 2. The processing process in this invention includes a segmentation step and a retention step. This not only reduces text processing complexity by segmenting the text, but also uses a "token matching ratio" to screen for sentences containing chemical elements. Compared to directly processing the entire text, this reduces irrelevant text interference and focuses on content containing chemical elements. Using a preset ratio, it can flexibly identify sentences containing element fragments, avoiding missed detections and improving the efficiency of subsequent processing for extracting chemical formulas from text.
[0064] 3. In the present invention, a preset ion symbol list of common metal cations is utilized when deleting the incorrectly marked position; a preset ion symbol list of common metal cations is utilized during index offset processing; a preset common metal element symbol table, a preset ion symbol table of common acid anions, a preset ion symbol table of common metal cations and a preset common material chemical formula abbreviation table are used to mark the corresponding text, thereby obtaining the original vector composed of the position index of the marked characters in the character string, thereby improving the effectiveness and efficiency of extracting material chemical formulas from the text.
[0065] 4. In the present invention, the tag vectors are sequentially subjected to exponential offset processing, and the offset vector is output. The technique of exponential offset plus differential vector is used, which can effectively calculate the number of chemical formulas in the text. By dynamically adjusting the position index (offset during tag clearing), the continuous and discontinuous element tagging scenarios in the text can be distinguished, so that the offset vector is more in line with the actual distribution relationship of the elements in the chemical formula. Through the exponential offset technology, the probability of characters in the same chemical formula being classified as the same chemical formula during the k-means clustering process is greatly improved.
[0066] 5. In this invention, chemical formula extraction takes into account "category position range, category conflicts within tokens, and incomplete tags." Characters are aggregated by category position range to address cross-token chemical formula extraction. Unified rules for categories within tokens (selecting the category with the largest number of characters) address ambiguous or incomplete tags, ensuring the completeness and accuracy of the extracted chemical formula at the text level. This results in more standardized mining results that better align with actual chemical formula expressions. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 A schematic diagram of the steps of a text mining method for identifying chemical formulas of materials in the present invention;
[0068] Figure 2 This is a flowchart of the implementation of preprocessing a sentence set in the embodiment;
[0069] Figure 3 This is a flowchart for extracting chemical formulas from text in the examples;
[0070] Figure 4 This is a physical structure diagram of a text mining electronic device for identifying chemical formulas of materials in an embodiment. DETAILED DESCRIPTION
[0071] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0072] Existing text analysis technologies mostly rely on traditional text retrieval or rule-based recognition methods, which are often inefficient and inaccurate when dealing with complex chemical formulas. Chemical formulas often appear in literature in various forms, and their expressions in text can be non-standard, with confusing layouts, including unknown numbers or unexpected spaces. This makes recognition and parsing by existing algorithms extremely difficult.
[0073] To meet this need, this application proposes a text-based automatic chemical formula extraction algorithm that can effectively handle the complex and diverse expressions of chemical formulas, helping researchers to more conveniently obtain and utilize relevant information, and accelerate research and innovation in the field of materials science.
[0074] Example 1
[0075] In this embodiment, a text mining method for identifying chemical formulas of materials is used, and the method steps are as follows: Figure 1 As shown, specifically including:
[0076] S1. Obtain a set of texts to be processed, preprocess them, and output a target text set;
[0077] S2. Mark the position of each target text in the target text set according to the chemical element symbol to obtain an original vector; and delete the incorrectly marked positions in the original vector to output a marked vector;
[0078] S3, perform exponential offset processing on the marker vector in sequence and output an offset vector;
[0079] S4. Perform differential processing on the offset vector and output a differential vector;
[0080] S5. Calculate the maximum ascending subsequence length of the outlier sequence in the differential vector;
[0081] S6. Setting the number of clusters based on the maximum ascending subsequence length, performing k-means clustering on the offset vectors according to the number of clusters, and outputting a list of offset position-category pairs;
[0082] S7. Based on the offset position-category pair list, perform a binding operation on the tag vector and output a tag position-category pair list;
[0083] S8. Extract chemical formulas from the text to be processed based on the list of tag position-category pairs to complete text mining.
[0084] In this embodiment, this solution is specifically implemented by taking the extraction of ferroelectric material chemical formulas from scientific literature in the field of ferroelectric material research as an example.
[0085] Implementation steps include:
[0086] S1, obtain the text set to be extracted;
[0087] S2, extract chemical formula using preset tools.
[0088] Figure 2 This is the implementation flow chart for preprocessing sentence sets, such as Figure 2 As shown, the method includes:
[0089] S101, obtaining a set of texts to be extracted;
[0090] S102, dividing the long text into sentences to obtain a set of sentences to be processed;
[0091] S103, screening the sentence subset containing chemical elements;
[0092] S104, matching all tokens in the sentence.
[0093] Figure 3 The implementation flow chart of extracting text chemical formula for text chemical formula extraction tool is as follows: Figure 3 As shown, the process includes:
[0094] S201, obtaining the original vector according to the position of the chemical element symbol mark in the text;
[0095] S202, deleting the incorrectly marked positions to obtain a marking vector;
[0096] S203, performing exponential offset processing on the marker vector in sequence to obtain an offset vector;
[0097] S204, performing differential processing on the offset vector to obtain a differential vector;
[0098] S205, calculating the maximum ascending subsequence length of the outlier sequence in the differential vector;
[0099] S206, setting the number of clusters to the maximum ascending subsequence length of the outlier sequence in the differential vector plus one;
[0100] S207, performing k-means clustering on the offset vector to obtain a list of "offset position-category" pairs;
[0101] S208, using the "offset position-category" pair list to perform a binding operation on the tag vector to obtain a "tag position-category" pair list;
[0102] S209: extracting chemical formulas from the text according to the “mark position-category” pair list.
[0103] In this example, this method is used to mark all chemical formulas in the following sentences:
[0104] "Kutty X-ray photoelectron spectroscopic investigations on cubicBaTiO 3,BaTi0.9Fe 0.1O 3and Ba 0.9Nd 0.1TiO 3systems Appl.";
[0105] Specifically, it is carried out through the following methods:
[0106] First, label the above sentence according to the list of chemical elements:
[0107] " K utty X-ray photoelectron spectroscopic investigations on cubic BaTiO 3, BaTi 0.9 Fe 0.1 O 3and Ba 0.9 Nd 0.1 TiO 3systems Appl."
[0108] Filter out error tags:
[0109] "Kutty X-ray photoelectron spectroscopic investigations on cubic BaTiO 3, BaTi 0.9 Fe 0.1 O 3and Ba 0.9 Nd 0.1 TiO 3systems Appl."
[0110] Generate a token vector:
[0111] The positions of the marked characters in the string are: 64, 65, 66, 67, 68, 74, 75, 76, 77, 83, 84, 90, 98, 99, 105, 106, 112, 113, 114;
[0112] So the generated token vector is:
[0113] [64,65,66,67,68,74,75,76,77,83,84,90,98,99,105,106,112,113,114]
[0114] Exponential offset generates an offset vector:
[0115] Initialize the index offset value to 20;
[0116] In the sequence "BaTiO 3,BaTi", "," is not a special word, so it causes a shift of 20 to the following token, and the index shift value becomes 40;
[0117] The offset vector becomes:
[0118] [64,65,66,67,68,94,95,96,97,103,104,110,118,119,125,126,132,133,134];
[0119] In the sequence "O 3and Ba", "and" is not a special word, so it causes a shift of 40 on the following tokens;
[0120] The offset vector becomes:
[0121] [64,65,66,67,68,94,95,96,97,103,104,110,158,159,165,166,172,173,174];
[0122] At the end of the traversal, the offset vector is:
[0123] [64,65,66,67,68,94,95,96,97,103,104,110,158,159,165,166,172,173,174];
[0124] Use the offset vector to calculate the difference vector. Let the offset vector be A, A[i] represents the i-th dimension element of A, and the difference vector be B, B[i] represents the i-th dimension element of B.
[0125] When i=1, B[i]=1;
[0126] When i=2, B[i]=A[i]-A[i-1]=65-64=1;
[0127] When i=3, B[i]=A[i]-A[i-1]=66-65=1;
[0128] When i=19, we get B=[1,1,1,1,1,26,1,1,1,6,1,6,48,1,6,1,6,1,1];
[0129] The copy vector will be [1,1,1,1,1,26,1,1,1,6,1,6,48,1,6,1,6,1,1];
[0130] Loop through the copy vectors to calculate the average and delete the elements in the copy vector that are greater than 2 times the average.
[0131] In the first loop, the average is 5.84210. The elements 26 and 48 are twice the average, so the replica vector deletes them.
[0132] In the second loop, the average is 2.17647. There are four 6s that are twice the average, so the copy vector deletes these four 6s.
[0133] In the third loop, the average is 1, and there is no element greater than twice the average, so the loop exits;
[0134] The final average is recorded as 1;
[0135] Initialize the spike list to an empty list;
[0136] The elements in the offset vector that are greater than twice the average are 26, 6, 6, 48, 6, 6, and are added to the peak list in sequence. The peak list is [26, 6, 6, 48, 6, 6].
[0137] The largest ascending subsequence of elements greater than 20 in the peak list is [26,48], and its length is 2;
[0138] Set the number of clusters to be 3, the length of the largest ascending subsequence of elements greater than 20 in the spike list plus one, and perform k-means clustering on the marker vector to obtain a list of "marker position-category" pairs:
[0139] [[64,1],[65,1],[66,1],[67,1],[68,1],[94,2],[95,2],[96,2],[97,2],[103,2],[104,2],[110,2],[158,3],[159,3],[165,3],[166,3],[172,3],[173,3],[174,3]]
[0140] Return the marker vector and perform a binding operation to obtain a list of "marker position-category" pairs:
[0141] [[64,1],[65,1],[66,1],[67,1],[68,1],[94,2],[95,2],[96,2],[97,2],[103,2],[104,2],[110,2],[118,3],[119,3],[125,3],[126,3],[132,3],[133,3],[134,3]];
[0142] All characters between the minimum mark position and the maximum mark position of the same category are classified into the current category;
[0143] When the same token has different categories or a token is incompletely labeled, the categories of all characters in the token are uniformly modified to the category with the largest number of characters;
[0144] Characters in the same category are considered to be the same chemical formula;
[0145] Get the final list of "mark position-category" pairs:
[0146] [[64,1],[65,1],[66,1],[67,1],[68,1],[94,2],[95,2],[96,2],[97,2],[98,2],[99,2],[100,2],[101,2],[102,2],[103,2],[104,2],[105,2],[106,2],[107,2],[108,2],[ 109,2],[110,2],[118,3],[119,3],[120,3],[121,3],[122,3],[123,3],[124,3],[125,3],[126,3],[127,3],[128,3],[129,3],[130,3],[131,3],[132,3],[133,3],[134,3]]
[0147] Return to the original sentence and mark the chemical formula in the original sentence:
[0148] "Kutty X-ray photoelectron spectroscopic investigations on cubic BaTiO 3, BaTi0.9Fe0.1O 3and Ba0.9Nd0.1TiO 3systems Appl."
[0149] That completes the text mining process.
[0150] In summary, compared with existing methods, the accuracy (F1-score) of this method in extracting chemical formulas from text is 82.1%, which is higher than existing methods. By constructing a complete process of "marking-transformation-clustering-extraction", a systematic solution for chemical formula mining was constructed. First, the positions of chemical elements are accurately marked, and feature associations are mined through vector transformation. Then, clustering is used to distinguish different chemical formula components, and finally extraction is carried out. Each link is closely linked to each other, which can effectively extract chemical formula information from the text, solve the recognition problems caused by non-standard chemical formulas in the text and chaotic typesetting, and significantly improve the accuracy of chemical formula extraction in complex chemical formula text scenarios.
[0151] Example 2
[0152] In this embodiment, a physical structure of an electronic device is used, such as Figure 4As shown, the electronic device includes: a bus S301, a processor S302, a memory S303, and an input / output device S304. The processor S302, the memory S303, and the input / output device S304 communicate with each other via the bus S301. The input / output device S304 can be used to transmit information to the electronic device. The processor S302 can call logic instructions in the memory S303 to execute the following method:
[0153] Get the text to be extracted;
[0154] Extract chemical formulas from text using the preset text chemical formula extraction tool.
[0155] Before obtaining the text to be extracted, the method for extracting chemical formulas from text further includes: obtaining a set of texts to be extracted; and preprocessing the set of texts to be extracted to obtain a target text set.
[0156] Before obtaining the target text set, the preprocessing of the text set to be extracted includes: segmenting the long text into sentences to obtain a set of sentences to be processed; and retaining a subset of the sentence set to be processed that contains chemical elements.
[0157] A subset of sentences containing chemical elements in the set of sentences to be processed is retained, including:
[0158] Use the strings of all chemical elements to match all tokens in the sentence. If there is a token whose matched length is more than a certain proportion of the token itself, then add the sentence to the subset of sentences containing chemical elements.
[0159] Extract chemical formulas from text using the preset text chemical formula extraction tool, including:
[0160] According to the position of the chemical element symbol in the text, an original vector is obtained. In the present invention, an original vector refers to a vector in which the value of each dimension corresponds to the position of a character marked by a chemical element symbol;
[0161] Delete the incorrectly marked positions in the original vector to obtain the marked vector;
[0162] The tag vectors are sequentially subjected to exponential shift processing to obtain offset vectors. In the present invention, exponential shift refers to the operation of increasing the qualified value in the tag vector by an exponentially growing value.
[0163] Perform differential processing on the offset vector to obtain a differential vector;
[0164] Calculate the maximum rising subsequence length of the outlier sequence in the difference vector;
[0165] Set the number of clusters to the maximum ascending subsequence length of the outlier sequence in the vector plus one, perform k-means clustering on the offset vector, and obtain a list of "offset position-category" pairs. In the present invention, an "offset position-category" pair refers to a pair consisting of the position of a character marked with a chemical element symbol after exponential offset processing and the category into which this character is classified by k-means clustering.
[0166] Using the "offset position-category" two-tuple list to bind the tag vector, a "tag position-category" pair list is obtained. In the present invention, the "tag position-category" two-tuple refers to a two-tuple with the position of the character marked by the chemical element symbol in the tag vector and the category of the character classified by k-means clustering as elements;
[0167] Extract chemical formulas from text based on a list of token position-category pairs.
[0168] According to the position of the chemical element symbol in the text, the original vector is obtained, including:
[0169] Use common metal element symbols, common acid anion symbols, common metal cation symbols, and common material chemical formula abbreviations to mark the text, and obtain a vector consisting of the position index of the marked characters in the string.
[0170] Delete the incorrectly marked positions and get the marking vector, including:
[0171] The first step is to decompose the text into several tokens according to the space character and remove the punctuation and numbers in the tokens;
[0172] In the second step, if the processed token is in the banned word list, all tags in the text corresponding to the token are cleared;
[0173] The third step is to calculate the ratio of the number of characters marked by chemical elements in the processed token to the number of characters in the entire token;
[0174] The fourth step is to calculate the length of the processed token;
[0175] Step 5: If the length of the processed token is less than a preset token length threshold and the processed token is not in the ion symbol list of common metal cations, then clear all tags in the text corresponding to the token, wherein the preset token length threshold is greater than or equal to 2 and less than or equal to 4;
[0176] Step 6: If the ratio of the number of characters marked by chemical elements in the processed token to the total number of characters in the token is lower than a preset character ratio threshold, then all marks in the text corresponding to the token are cleared, wherein the preset character ratio threshold is between 1 / 2 and 3 / 4;
[0177] Step 7: After traversing all tokens, we get the token vector.
[0178] The original vector is exponentially offset in sequence to obtain the offset vector, which includes the first six steps of obtaining the marked vector, including:
[0179] The first step is to decompose the text into several tokens according to the space character and remove the punctuation and numbers in the tokens;
[0180] In the second step, if the processed token is in the banned word list, all tags in the text corresponding to the token are cleared;
[0181] The third step is to calculate the ratio of the number of characters marked by chemical elements in the processed token to the number of characters in the entire token;
[0182] The fourth step is to calculate the length of the processed token;
[0183] Step 5: If the length of the processed token is less than a preset token length threshold and the processed token is not in the ion symbol list of common metal cations, then clear all tags in the text corresponding to the token, wherein the preset token length threshold is greater than or equal to 2 and less than or equal to 4;
[0184] Step 6: If the ratio of the number of characters marked by chemical elements in the processed token to the total number of characters in the token is lower than a preset character ratio threshold, then all marks in the text corresponding to the token are cleared, wherein the preset character ratio threshold is between 1 / 2 and 3 / 4;
[0185] Step 7: If the length of the token after the cleared mark is 0, and this token after the cleared mark is the first token after the previous token whose length after the cleared mark is not 0, then the index of the position in the original vector after the position index of the character corresponding to this token is increased by the index offset value, and then the index offset value becomes twice the original value;
[0186] Formally expressed as:
[0187] If the length of the token after the clear marker is 0, and this token is the first clear marker token after the previous clear marker whose length is not 0, then update the offset vector:
[0188] That is, add the exponential offset value δ to all positions after the position index in the offset vector:
[0189]
[0190] Where j is the position index of the current token in the original vector; V is the offset vector; j is the position index in the offset vector; δ is the exponential offset value
[0191] And update the exponential offset value δ to twice the original value:
[0192] δ←2·δ
[0193] After traversing all tokens, we get the offset vector.
[0194] Perform differential processing on the offset vector to obtain the differential vector, including:
[0195] The dimension of the difference vector is the same as the offset vector dimension;
[0196] In particular, the first dimension of the difference vector is equal to 1;
[0197] When i ≥ 2, the i-th dimension of the difference vector is equal to the i-th dimension of the offset vector minus the i-1-th dimension;
[0198] The formal expression is:
[0199] Let A be the difference vector, A[i] represents the i-th dimension of the difference vector;
[0200] Let B be the offset vector and B[i] represent the i-th dimension of the offset vector.
[0201] Then we have:
[0202]
[0203] Calculate the maximum ascending subsequence length of the outlier sequence in the difference vector, including:
[0204] Copy the difference vector to obtain a copy vector;
[0205] Perform the following operations in a loop:
[0206] Calculate the average of all elements in the replica vector;
[0207] Delete the elements in the replica vector that are a multiple of the average, where this multiple should be guaranteed to be between 1.5 and 2.5;
[0208] If the deletion operation of the elements in the copy vector that are greater than a multiple of the average does not delete an element, the loop is exited.
[0209] Keep the average of all elements in the last calculated replica vector;
[0210] Initialize a peak list, where the peak list in the present invention refers to a list of outliers in the differential vector;
[0211] Traverse all elements in the difference vector. If its value is greater than a certain multiple of the average of all elements in the replica vector calculated last time, put it into the peak list, where this multiple should be guaranteed to be between 1.5 and 2.5;
[0212] Calculate the maximum rising subsequence length of the peak list that is greater than a certain threshold, that is, the maximum rising subsequence length of the outlier sequence in the difference vector, where this threshold should be guaranteed to be within 1 times the exponential offset value δ.
[0213] The "offset position-category" pair is a two-tuple, where the first element is an element of the offset vector, that is, a position index, and the second element is the category into which this element is classified after k-means clustering;
[0214] The "mark position-category" pair is a two-tuple, where the first element is an element of the mark vector, that is, a position index, and the second element is the category assigned to this element after the binding operation;
[0215] Use the "offset position-category" pair list to bind the original vector to obtain the "mark position-category" pair list, including:
[0216] Initialize the "mark position-category" pair list to an empty list;
[0217] Traverse the "offset position-category" pair list and the original vector by position, take the position index in the original vector and the category in the "offset position-category" pair to form a "marked position-category" pair, and then put this "marked position-category" pair into the "marked position-category" pair list;
[0218] The formal expression is:
[0219] Let A be a list of "offset position-category" pairs, A[i] represents the i-th element of A, and A[i][j] represents the j-th element of the i-th "offset position-category" pair of A;
[0220] B is a label vector, B[i] represents the i-th element of B;
[0221] C is a list of "mark position-category" pairs, C[i] represents the i-th element of C, and C[i][j] represents the j-th element of the i-th "mark position-category" pair of C;
[0222] Then we have:
[0223]
[0224] Extract chemical formulas from text based on a list of "token position-category" pairs, including:
[0225] All characters between the minimum mark position and the maximum mark position of the same category are classified into the current category;
[0226] When the same token has different categories or a token is incompletely labeled, the categories of all characters in the token are uniformly modified to the category with the largest number of characters;
[0227] Characters in the same category are considered to belong to the same chemical formula.
[0228] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A text mining method for identifying chemical formulas of materials, characterized in that: The method steps include: S1. Obtain a set of texts to be processed, preprocess them, and output a target text set; S2. Mark the position of each target text in the target text set according to the chemical element symbol to obtain an original vector; and delete the incorrectly marked positions in the original vector to output a marked vector; S3, perform exponential offset processing on the marker vector in sequence and output an offset vector; S4. Perform differential processing on the offset vector and output a differential vector; S5. Calculate the maximum ascending subsequence length of the outlier sequence in the differential vector; S6. Setting the number of clusters based on the maximum ascending subsequence length, performing k-means clustering on the offset vectors according to the number of clusters, and outputting a list of offset position-category pairs; S7. Based on the offset position-category pair list, perform a binding operation on the tag vector and output a tag position-category pair list; S8. Extract chemical formulas from the text to be processed based on the list of tag position-category pairs to complete text mining.
2. A text mining method for identifying material chemical formulas according to claim 1, characterized in that: The preprocessing process in S1 includes a segmentation process and a retention process; the segmentation process is to segment a long paragraph of text into sentences to obtain a set of sentences to be processed; the retention process is to retain a subset of sentences containing chemical elements in the set of sentences to be processed, and this subset of sentences is the target text set. The retention process is specifically: using character strings of all chemical elements to match tokens of all sentences in the set of sentences to be processed, if there is a token whose matched length is greater than a preset proportion of the length of the token itself, then the sentence corresponding to the token is added to the subset of sentences containing chemical elements.
3. The text mining method for identifying chemical formulas of materials according to claim 1, characterized in that: In the S2, the position of each target text in the target text set is marked in the text according to the chemical element symbol, and the specific process of obtaining the original vector includes: using a preset common metal element symbol table, a preset common acid anion ion symbol table, a preset common metal cation ion symbol table and a preset common material chemical formula abbreviation table to mark the corresponding text, thereby obtaining the original vector composed of the position index of the marked characters in the character string.
4. A text mining method for identifying chemical formulas of materials according to claim 3, characterized in that: The specific steps of deleting the position of the error mark in S2 and outputting the mark vector include: S21. Decompose the text into a number of tokens according to the space character, and delete the punctuation marks and numbers in the tokens to obtain a first token; S22. If a first token is included in the preset banned word list, clear all tags in the text corresponding to the token; S23. Calculate the ratio of the number of characters marked by chemical elements in the first token to the number of characters in the entire token; S24, calculating the first token length; S25. If the length of the first token is less than a preset token length threshold, and the first token is not in the preset ion symbol list of common metal cations, clear all marks in the text corresponding to the token; S26. If the ratio of the number of characters marked with chemical elements in the first token to the total number of characters in the token is lower than a preset character ratio threshold, clear all marks in the text corresponding to the token; S27. After traversing all tokens, obtain the tag vector.
5. The text mining method for identifying chemical formulas of materials according to claim 4, characterized in that: In the aforementioned S3, the specific steps of sequentially performing exponential offset processing on the marker vectors and outputting the offset vectors include: S31, decomposing the text into a plurality of tokens according to space characters and removing punctuation marks and numbers from the tokens to obtain second tokens; S32. If the second token is included in the preset banned word list, clear all tags in the text corresponding to the token; S33, calculating the ratio of the number of characters marked by chemical elements in the second token to the number of characters in the entire token; S34, calculating the second token length; S35. If the length of the second token is less than a preset token length threshold, and the second token is not in the preset ion symbol list of common metal cations, clear all marks in the text corresponding to the token; S36. If the ratio of the number of characters marked with chemical elements in the second token to the total number of characters in the token is lower than a preset character ratio threshold, clear all marks in the text corresponding to the token; S37. If the length of the token after the cleared mark is 0, and this token is the first token after the previous token whose length after the cleared mark is not 0, then add a preset index offset value to the position index in the original mark after the position index of the character corresponding to this token, and update the index offset value to twice the original value; S38. After traversing all tokens, obtain the offset vector.
6. A text mining method for identifying chemical formulas of materials according to claim 5, characterized in that: The offset vector in S3 has the same dimension as the differential variable in S4; and the first dimension of the differential vector is equal to 1; when i ≥ 2, the i-th dimension of the differential vector is equal to the i-th dimension of the offset vector minus the i-1-th dimension. The specific formula is: Among them, A[i] represents the i-th dimension of the difference vector; B[i] represents the i-th dimension of the offset vector; and i is the number of dimensions.
7. The text mining method for identifying chemical formulas of materials according to claim 1, characterized in that: The specific process of calculating the maximum ascending subsequence length of the outlier sequence in the differential vector in S5 includes: S51, copy the differential vector to obtain a copy vector; S52, calculating the average of all elements in the replica vector; S53, deleting elements in the replica vector that are greater than a preset multiple of the average calculated in S53; S54, determine whether there is an element to be deleted in S53, if yes, return to step S52; if not, go to step S55; S55. Retain the average of all elements in the current replica vector as the retained average; S56. Initialize a peak list, wherein the peak list is a list consisting of outliers in the differential vector; S57, traverse all elements in the difference vector, and if there is a value greater than a preset multiple of the retained average, put it into the peak list; S58 , calculating the maximum ascending subsequence length of the peak list that is greater than a preset threshold, that is, the maximum ascending subsequence length of the outlier sequence in the differential vector.
8. The text mining method for identifying chemical formulas of materials according to claim 1, characterized in that: The specific process of S7 includes: S71, initializing the mark position-category pair list to an empty list; S72, traversing the offset position-category pair list and the tag vector by position; and combining the position index in the tag vector and the category in the offset position-category pair to form a tag position-category pair; S73. Put all the marker position-category pairs into the empty list initialized in S71, and output the marker position-category pair list.
9. The text mining method for identifying chemical formulas of materials according to claim 1, characterized in that: The specific process of extracting the chemical formula from the text to be processed according to the tag position-category pair list in S8 includes: Classify all characters between the minimum and maximum mark positions of the same category in the list of mark position-category pairs into the current category; If different categories exist in the same token or a token is incompletely labeled, the categories of all characters in the token are uniformly modified to the category with the largest number of characters; Characters in the same category are extracted as the same chemical formula.
10. A text mining system for identifying chemical formulas of materials, characterized in that: The system is applied to a text mining method for identifying chemical formulas of materials as described in any one of claims 1 to 9, and the system includes a text preprocessing module, an element labeling and error correction module, a vector transformation module, and a clustering and chemical formula extraction module; The text preprocessing module is used to obtain a set of texts to be processed, preprocess them, and output a target text set; The element marking and error correction module marks the position of each target text in the target text set according to the chemical element symbol to obtain an original vector; and deletes the position of the incorrect mark in the original vector to output a marked vector; The vector transformation module sequentially performs exponential offset processing on the marker vector and outputs an offset vector; and performs differential processing on the offset vector and outputs a differential vector; The clustering and chemical formula extraction module is used to calculate the maximum ascending subsequence length of the outlier sequence in the differential vector, then set the number of clusters based on the maximum ascending subsequence length, and perform k-means clustering on the offset vector according to the number of clusters, outputting a list of offset position-category pairs, and based on the offset position-category pair list, perform a binding operation on the tag vector and output a list of tag position-category pairs; and extract the chemical formula in the text to be processed based on the tag position-category pair list to complete text mining.