Material science text entity category intelligent labeling method and system based on cluster analysis

CN122817464APending Publication Date: 2026-09-25ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611025909.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]直接判定成本高;现有方法需对每条文本独立执行完整的实体识别与分类流程,面对千万级数据时计算消耗巨大,处理效率低,难以满足对大规模材料文本实体与属性的类别的快速标注需求

Benefits of technology

[0069](1)提升了类别判定的准确性;采用分层关键词匹配与三层语义模板匹配双重判定机制,关键词按组分结构、工艺方法、材料性能等类别细分并加权融合,语义模板动态适应文本表述变化,更准确地识别专业术语语义。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817464A_ABST
    Figure CN122817464A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a material text entity category intelligent labeling method and system based on clustering analysis; the method comprises the following steps: obtaining text data and semantic embedding vectors, performing dimension reduction processing on the semantic embedding vectors to obtain a low-dimensional vector dataset; adopting multiple clustering algorithms to cluster the low-dimensional vector dataset, aggregating text entities with similar semantics into multiple clustering clusters, and calculating the silhouette coefficients of each clustering algorithm; adopting a double judgment mechanism combining rule matching and semantic template matching, determining the material entity attribution category and cluster level confidence of each clustering cluster based on an intelligent boundary category judgment strategy provided with double condition verification; taking the product of the silhouette coefficient and the cluster level confidence of each clustering algorithm as the voting weight of the clustering algorithm; selecting a final classification category based on the voting weight, and obtaining a final confidence based on the confidence and the silhouette coefficient of each algorithm; and outputting the final classification category and the final confidence of each material text entity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of information processing technology, particularly the field of materials science information processing technology, and specifically relates to a method and system for intelligent labeling of material science text entity categories based on cluster analysis. Background Technology

[0002] The field of materials science has accumulated massive amounts of textual data, including academic papers, patent documents, technical reports, and experimental records. This unstructured textual data contains a wealth of fine-grained knowledge about material entity categories. Structured extraction and annotation of this knowledge, followed by writing it back into a knowledge graph, is a crucial prerequisite for supporting intelligent question-answering retrieval and design-aided systems for new materials. The main technical challenges in entity category annotation for materials science texts are as follows:

[0003] Direct determination is costly; existing methods require a complete entity recognition and classification process to be performed independently for each text, which consumes a huge amount of computation when dealing with tens of millions of data points, resulting in low processing efficiency and making it difficult to meet the need for rapid labeling of entity and attribute categories in large-scale material texts.

[0004] The category determination is inaccurate; relying on fixed keyword matching is difficult to deal with the common problem of semantic similarity but different expression in text entities. The understanding ability is limited, which easily leads to a large number of missed and wrong judgments, seriously weakening the recall and accuracy of the annotation.

[0005] Multi-model result fusion is difficult; a single model can only capture some features of the data. When multiple heterogeneous models are introduced, the inconsistent granularity of the output labels and the lack of a unified and stable fusion mechanism lead to conflicting results, poor consistency, and an inability to effectively leverage complementary advantages.

[0006] Improper handling of boundary categories: For ambiguous or entirely new text entities that cannot be clearly classified into preset categories, existing methods generally lack reasonable fallback strategies, often forcibly classifying them into a similar category, introducing structural noise into the knowledge graph, and reducing the purity and usability of the knowledge base.

[0007] The lack of confidence assessment means that classification results typically only provide definitive labels without quantitative confidence information. This makes it difficult for downstream validation and applications to effectively distinguish between reliable annotations and low-quality inferences. Unreliable information can easily pollute the knowledge graph and hinder the automated construction of high-quality knowledge bases. Summary of the Invention

[0008] In view of this, on the one hand, some embodiments disclose a method for intelligent annotation of material science text entity categories based on cluster analysis, including the following steps:

[0009] S1. Text Data Acquisition and Preprocessing: Extract material science text entities and their associated semantic embedding vectors in batches from the knowledge graph database, and perform preprocessing.

[0010] S2. Dimensionality reduction: The preprocessed semantic embedding vectors are reduced in dimensionality to obtain a low-dimensional vector dataset.

[0011] S3. Multi-model clustering analysis and quality assessment: Multiple clustering algorithms are used to cluster the low-dimensional vector dataset, and semantically similar material science text entities are aggregated into multiple clusters. The silhouette coefficient of each clustering algorithm is calculated as the clustering quality assessment value.

[0012] S4. Cluster-level dual determination: A dual determination mechanism combining rule matching and semantic template matching is adopted. Based on the intelligent boundary category determination strategy with dual condition verification, the material entity belonging category and cluster-level confidence of each cluster are determined.

[0013] S5. Multi-model weighted voting fusion: The product of the silhouette coefficient and the cluster-level confidence of each clustering algorithm is used as the voting weight of the clustering algorithm; the voting weights of all decision algorithms for each candidate category under the cluster are accumulated, and the candidate category with the highest accumulated voting weight score is used as the final classification category. The final confidence is obtained by weighted averaging of the confidence of each algorithm under the final classification category using the silhouette coefficient.

[0014] S6. Output Results: Output the final classification category and final confidence score for each material text entity, completing the intelligent labeling of material science text entity categories.

[0015] Furthermore, in some embodiments of the intelligent annotation method for material science text entity categories based on cluster analysis, the preprocessing in step S1 includes:

[0016] Data cleaning, filtering out invalid samples with vector values ​​of "None";

[0017] Data transformation: Convert the data list into a NumPy array;

[0018] Standardize the data using StandardScaler or MinMaxScaler.

[0019] Some embodiments of the intelligent labeling method for material science text entity categories based on cluster analysis include the following steps in step S2: dimensionality reduction of the semantic embedding vectors: first, principal component analysis is used to reduce the dimensionality, retaining 90-95% of the variance information; then, a second dimensionality reduction is performed through uniform manifold approximation and projection to a preset low-dimensional target dimension.

[0020] Some embodiments disclose a method for intelligent annotation of material science text entity categories based on cluster analysis, wherein step S3 includes:

[0021] S301. Determine the set of multi-model clustering algorithms, including K-means clustering, DBSCAN clustering, hierarchical clustering, and Gaussian mixture model (GMM);

[0022] S302. Automatic parameter optimization is performed on each clustering algorithm, including:

[0023] K-means clustering: Automatically searches for the optimal combination of the number of clusters and two initialization methods, and uses the elbow rule and silhouette coefficient to determine the optimal K value;

[0024] DBSCAN clustering: Automatically searches for the optimal combination of neighborhood radius and minimum number of samples. It first filters out combinations with fewer than three clusters or more than half of the clusters being noisy, and then selects the one with the highest silhouette coefficient. If none of these conditions are met, it downgrades to a weighted score of silhouette coefficient multiplied by the noise ratio to select the best one.

[0025] Hierarchical clustering: Automatically searches for the optimal combination of cluster size, linking method, and distance metric;

[0026] Gaussian Mixture Model (GMM): Automatically searches for the optimal combination of component number and covariance type;

[0027] S303. Structured storage of results: The clustering results of each clustering algorithm and the corresponding silhouette coefficient evaluation index are stored in a structured manner for subsequent category determination.

[0028] Some embodiments disclose a smart annotation method for material science text entity categories based on cluster analysis. In step S4, the material entity's category includes component structure category, process method category, material property category, and "other" category; the dual determination of cluster-level categories specifically includes the following steps:

[0029] S401. Extract the high-frequency keywords of all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary of each category, and obtain the rule matching score.

[0030] S402. For cases where keywords do not match exactly in words but are semantically similar, use the embedding model to calculate the cosine similarity of these keywords to the vectors of the general template layer, the specific template layer, and the Chinese template layer, respectively. Take the maximum value and multiply it by the preset weight of the template to obtain the weighted maximum similarity of the keyword. Then, calculate the semantic matching score according to the adaptive threshold weighting.

[0031] S403. The rule matching score and semantic matching score are weighted and fused to obtain the fusion score of each category. The category with the highest fusion score is selected as the preliminary judgment category of the cluster. The fusion score is also used as the confidence level of the judgment.

[0032] S404. For the initial classification, perform a boundary category determination with dual conditions. If the conditions are met, the category of the cluster is reclassified as the boundary category "Other".

[0033] In some embodiments of the intelligent annotation method for material science text entity categories based on cluster analysis, step S401, the method for calculating the rule matching score includes:

[0034] Each keyword is assigned a weight w; if no weight is specified, the default value is 1; the score for rule matching is calculated based on the keyword weights: the sum of matching weights W match It equals the sum of the weights of all successfully matched keywords; the normalized score, coverage α, is defined as the sum of the matching weights and the sum of the weights of all keywords, W. total The ratio is calculated using the following formula:

[0035] ;

[0036] Calculate the average weight of all keywords To match the total weight W match Divide by the average weight to obtain the absolute importance:

[0037] ;

[0038] Then, the absolute importance is normalized to the interval [0,1) using the following formula:

[0039] ;

[0040] Weighting factors are:

[0041] ;

[0042] Weighted normalization rule matching score:

[0043] .

[0044] In some embodiments of the intelligent annotation method for material science text entity categories based on cluster analysis, step S402, the semantic matching score calculation method includes:

[0045] For keyword embedding vectors First, calculate its relationship with all template embedding vectors. The cosine similarity between them is taken as the maximum value and then multiplied by the preset weight of the template. The weighted maximum similarity of the keyword is obtained. This forms a similarity sequence for all keywords;

[0046] ;

[0047] Calculate the adaptive similarity threshold: take the average of the similarity sequences. Multiply by a preset coefficient to obtain a baseline threshold, and take the larger of the baseline threshold and the preset minimum threshold as the adaptive threshold:

[0048] ;

[0049] Then, the number of keywords in the similarity sequence that exceed the adaptive threshold is counted, and divided by the total number of keywords to obtain the proportion of highly similar keywords, p.

[0050] Finally, the average similarity of all keywords is multiplied by the proportion of highly similar keywords (p), and the resulting product is the semantic matching score for that category.

[0051] .

[0052] In some embodiments of the intelligent annotation method for material science text entity categories based on cluster analysis, the formula for calculating the fusion score in step S403 is as follows:

[0053] ;

[0054] Among them, R score To weight the normalized rule matching score, α is its preset weight; S score The semantic matching score is represented by β, which is its preset weight.

[0055] On the other hand, some embodiments disclose a materials science text entity category intelligent annotation system based on cluster analysis, used to implement a materials science text entity category intelligent annotation method based on cluster analysis, including:

[0056] The text data acquisition and preprocessing module is configured to extract materials science text entities and their associated semantic embedding vectors in batches from the knowledge graph database, and then perform preprocessing.

[0057] The dimensionality reduction module is configured to perform dimensionality reduction processing on the preprocessed semantic embedding vectors to obtain a low-dimensional vector dataset.

[0058] The multi-model clustering analysis and quality assessment module is configured to use multiple clustering algorithms to cluster the low-dimensional vector dataset, aggregate semantically similar material science text entities into multiple clusters, and calculate the silhouette coefficient of each clustering algorithm as the clustering quality assessment value.

[0059] The cluster-level category dual determination module is configured to use a dual determination mechanism that combines rule matching and semantic template matching. Based on an intelligent boundary category determination strategy with dual condition verification, it determines the category to which the material entity belongs and the cluster-level confidence of each cluster.

[0060] The multi-model weighted voting fusion module is configured to use the product of the silhouette coefficient and the cluster-level confidence of each clustering algorithm as the voting weight of that clustering algorithm; accumulate the voting weights of all decision algorithms for each candidate category under the cluster, and take the candidate category with the highest accumulated voting weight score as the final classification category; the final confidence is obtained by weighting the confidence of each algorithm under the final classification category with the silhouette coefficient.

[0061] The results output module is configured to output the final classification category and final confidence score of each material text entity, thus completing the intelligent labeling of material science text entity categories.

[0062] Furthermore, some embodiments of the intelligent annotation system for material science text entity categories based on cluster analysis disclose a cluster-level category dual determination module including:

[0063] The rule matching submodule is configured to extract high-frequency keywords from all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary for each category, and obtain the rule matching score.

[0064] The semantic template matching submodule is configured to calculate the cosine similarity of these keywords to the general template layer, specific template layer and Chinese template layer vectors when the keywords are not exactly matched in words but are semantically similar. The maximum value is then multiplied by the preset weight of the template to obtain the weighted maximum similarity of the keyword. Finally, the semantic matching score is calculated by weighting according to the adaptive threshold.

[0065] The dual-judgment fusion submodule is configured to perform weighted fusion of rule matching score and semantic matching score to obtain fusion score for each category, select the category with the highest fusion score as the initial judgment category for that cluster, and the fusion score is also used as the confidence level of the judgment.

[0066] The intelligent boundary category determination submodule is configured to perform a boundary category determination under two conditions for the initial category determination. If the conditions are met, the category of the cluster is re-determined as the boundary category "Other".

[0067] The present invention discloses an intelligent labeling method and system for material science text entity categories based on cluster analysis. The method first extracts and preprocesses text data in the field of materials science through a knowledge graph, then performs dimensionality reduction using PCA and UMAP, then executes multiple clustering algorithms with different principles in parallel to generate clusters, and performs dual category determination on each cluster. Finally, the clusters are classified based on a multi-model voting fusion strategy, thereby achieving intelligent classification of text entity attributes in the field of materials science.

[0068] It has at least the following beneficial technical effects:

[0069] (1) Improved the accuracy of category determination; adopted a dual determination mechanism of hierarchical keyword matching and three-layer semantic template matching. Keywords are subdivided and weighted according to categories such as component structure, process method, and material properties. The semantic template dynamically adapts to changes in text expression, and more accurately identifies the semantics of professional terms.

[0070] (2) Enhanced the stability of classification results; based on the product of the silhouette coefficient of each clustering model and its confidence in the category of each entity's cluster, the model's comprehensive voting weight is used, so that models with good clustering quality and high confidence in the judgment have greater influence in weighted voting; the random fluctuation of a single algorithm is suppressed by dynamic weights, reducing the difference in results between different rounds.

[0071] (3) Effective control of the proportion of boundary categories; only when all major categories cannot be effectively matched, or when the scores of multiple major categories are close and difficult to determine clearly, is it judged as a boundary category.

[0072] (4) It provides quantifiable classification confidence; the weighted voting fusion process outputs a normalized confidence of 0 to 1 for each text, reflecting the consistency of the algorithm's judgment, which can provide direct quantitative basis for downstream quality control and manual review and sorting.

[0073] (5) The method of this invention is a weakly supervised or unsupervised learning method that does not require manual labeling of training data, thus reducing deployment and manpower costs and making it suitable for professional fields where labeled data is scarce.

[0074] (6) It has good scalability; the keywords adopt a modular hierarchical dictionary, the semantic template adopts a pluggable multi-layer architecture, and the clustering algorithm set can be flexibly expanded; when migrating to a new field, only the keyword dictionary and semantic template content need to be replaced, and when introducing a new clustering algorithm, the implementation can be added in the clustering module, and the core process does not need to be changed. Attached Figure Description

[0075] Figure 1 , one The flowchart of the intelligent annotation method for material science text entity categories based on cluster analysis disclosed in these embodiments is shown.

[0076] Figure 2 , one The flowchart of the cluster-level category dual determination method disclosed in these embodiments.

[0077] Figure 3 Pie chart showing the classification results of the intelligent annotation method for material science text entity categories based on cluster analysis disclosed in Example 1. Detailed Implementation

[0078] The term "embodiment" used herein, as an example, is not necessarily to be construed as superior to or better than other embodiments. Performance testing in these embodiments of the invention, unless otherwise specified, employs conventional testing methods in the art. It should be understood that the terminology used in these embodiments is merely for describing particular implementations and is not intended to limit the scope of the disclosure of these embodiments.

[0079] Unless otherwise stated, the technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the embodiments of this invention pertain; other experimental methods and technical means not specifically noted in the embodiments of this invention refer to experimental methods and technical means commonly used by one of ordinary skill in the art.

[0080] The terms “basic” and “approximately” as used herein are used to describe small fluctuations. For example, they can mean less than or equal to ±5%, such as less than or equal to ±2%, such as less than or equal to ±1%, such as less than or equal to ±0.5%, such as less than or equal to ±0.2%, such as less than or equal to ±0.1%, such as less than or equal to ±0.05%. Numerical data presented or expressed in range format herein are used for convenience and brevity only, and should therefore be interpreted flexibly to include not only the explicitly listed values ​​that define the range, but also all independent values ​​or subranges contained within that range. For example, a numerical range of “1–5%” should be interpreted to include not only the explicitly listed values ​​from 1% to 5%, but also the independent values ​​and subranges within the indicated range. Thus, this numerical range includes independent values ​​such as 2%, 3.5%, and 4%, and subranges such as 1%–3%, 2%–4%, and 3%–5%, etc. This principle also applies to ranges that list only one value. Furthermore, this interpretation applies regardless of the width of the range or the characteristics described.

[0081] In this document, including in the claims, conjunctions such as "comprising," "including," "with," "having," "containing," "involving," and "accommodating" are understood to be open-ended, meaning "including but not limited to." Only the conjunctions "consisting of" and "composed of" are closed conjunctions.

[0082] To better illustrate the content of this invention, numerous specific details are provided in the following detailed embodiments. Those skilled in the art should understand that the invention can be practiced even without certain specific details. In the embodiments, some methods, means, instruments, and devices well-known to those skilled in the art are not described in detail, in order to highlight the main points of the invention.

[0083] Without conflict, the technical features disclosed in the embodiments of the present invention can be combined arbitrarily, and the resulting technical solution belongs to the content disclosed in the embodiments of the present invention.

[0084] In some implementations, such as Figure 1 As shown, the intelligent annotation method for material science text entity categories based on cluster analysis includes the following steps:

[0085] S1. Text Data Acquisition and Preprocessing: Extract material science text entities and their associated semantic embedding vectors in batches from the knowledge graph database, and perform preprocessing.

[0086] In some embodiments, the material text entities and their semantic embedding vectors are directly extracted from the nodes of the Neo4j graph database, and the embedding vectors have been generated and written by the embedding model during the knowledge graph construction stage.

[0087] In some embodiments, preprocessing includes: data cleaning, filtering out invalid samples with vector values ​​of "None"; data transformation, converting the data list into a NumPy array; and standardization, using StandardScaler or MinMaxScaler to standardize the data.

[0088] S2. Dimensionality Reduction: The preprocessed semantic embedding vectors are dimensionality reduced to obtain a low-dimensional vector dataset. Typically, dimensionality reduction of semantic embedding vectors includes: first, using Principal Component Analysis (PCA) to reduce dimensionality while retaining 90-95% of the variance information; then, using Uniform Manifold Approximation and Projection UMAP to further reduce dimensionality to a preset low-dimensional target dimension, such as 50 dimensions.

[0089] S3. Multi-model clustering analysis and quality assessment: Multiple clustering algorithms are used to cluster the low-dimensional vector dataset, aggregating semantically similar material science text entities into multiple clusters, and calculating the silhouette coefficient of each clustering algorithm as the clustering quality assessment value. Typically, multiple clustering algorithms with different principles are executed in parallel on the dimensionality-reduced low-dimensional vector dataset, and hyperparameter optimization is performed on each clustering algorithm. The results of each clustering algorithm are saved. In some embodiments, multi-model clustering analysis and quality assessment specifically include:

[0090] S301. Determine the set of multi-model clustering algorithms, including K-means clustering, DBSCAN clustering, hierarchical clustering, and Gaussian mixture model (GMM);

[0091] S302. Automatic parameter optimization is performed on each clustering algorithm, including:

[0092] K-means clustering: Automatically searches for the optimal combination of the number of clusters and two initialization methods, and uses the elbow rule and silhouette coefficient to determine the optimal K value;

[0093] DBSCAN clustering: Automatically searches for the optimal combination of neighborhood radius and minimum number of samples. It first filters out combinations with fewer than three clusters or more than half of the clusters being noisy, and then selects the one with the highest silhouette coefficient. If none of these conditions are met, it downgrades to a weighted score of silhouette coefficient multiplied by the noise ratio to select the best one.

[0094] Hierarchical clustering: Automatically searches for the optimal combination of cluster size, linking method, and distance metric;

[0095] Gaussian Mixture Model (GMM): Automatically searches for the optimal combination of component number and covariance type;

[0096] S303. Structured storage of results: The clustering results of each clustering algorithm and the corresponding silhouette coefficient evaluation index are stored in a structured manner for subsequent category determination.

[0097] S4. Cluster-level dual determination: Taking each cluster as a unit, a dual determination mechanism combining rule matching and semantic template matching is adopted. Based on the intelligent boundary category determination strategy with dual condition verification, the material entity belonging category and cluster-level confidence of each cluster are determined. Typically, the belonging category of material science text entities includes component structure category, process method category, material performance category and "other" category.

[0098] In some embodiments, such as Figure 2 As shown, the cluster-level category dual determination specifically includes the following steps:

[0099] S401, Rule Matching Determination: Extract high-frequency keywords from all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary for each category, and obtain the rule matching score; typically, the TF-IDF algorithm is used to extract high-frequency keywords from all texts within a cluster. After segmenting all texts within the cluster and removing stop words, the word frequency of each word is counted, and the top 10 to 20 words in terms of word frequency are selected as cluster keywords;

[0100] In some embodiments, the rule matching score calculation method includes:

[0101] Convert all keywords in the cluster to lowercase and assign a weight w to each keyword; if no weight is specified, the default value is 1; the rule matching score is calculated based on the keyword weights: the sum of the matching weights W. match It equals the sum of the weights of all successfully matched keywords; the normalized score, coverage α, is defined as the sum of the matching weights and the sum of the weights of all keywords, W. total The ratio is calculated using the following formula:

[0102] ;

[0103] Calculate the average weight of all keywords To match the total weight W match Divide by the average weight to obtain the absolute importance:

[0104] ;

[0105] Then, the absolute importance is normalized to the interval [0,1) using the following formula:

[0106] ;

[0107] Weighting factors are:

[0108] ;

[0109] Weighted normalization rule matching score:

[0110] .

[0111] S402, Semantic Template Matching Determination; used to capture cases where keywords do not match perfectly literally but are semantically similar; the cosine similarity of these keywords to the vectors of the general template layer, the specific template layer and the Chinese template layer is calculated using the embedding model, the maximum value is taken and then multiplied by the preset weight of the template to obtain the weighted maximum cosine similarity of the keyword, and then the semantic matching score is calculated according to the adaptive threshold weighting.

[0112] In some embodiments, the embedding model includes third-party provided embedding models, such as, but not limited to, Alibaba Cloud DashScope and some locally running embedding models.

[0113] In some embodiments, the weight coefficient of the general template layer is 1.0, the weight coefficient of the specific template layer is 1.5, and the weight coefficient of the Chinese template layer is 1.2; the matching results of each template layer are weighted and fused according to the weight coefficient to obtain the weighted maximum similarity of the keyword.

[0114] In some embodiments, the semantic matching score calculation method includes:

[0115] For keyword embedding vectors First, calculate its relationship with all template embedding vectors. The cosine similarity between them is taken as the maximum value and then multiplied by the preset weight of the template. The weighted maximum similarity of the keyword is obtained. This forms a similarity sequence for all keywords;

[0116] ;

[0117] Calculate the adaptive similarity threshold: take the average of the similarity sequences. Multiply by a preset coefficient to obtain a baseline threshold, and take the larger of the baseline threshold and the preset minimum threshold as the adaptive threshold:

[0118] ;

[0119] Then, the number of keywords in the similarity sequence that exceed the adaptive threshold is counted, and divided by the total number of keywords to obtain the proportion of highly similar keywords, p.

[0120] Finally, the average similarity of all keywords is multiplied by the proportion of highly similar keywords (p), and the resulting product is the semantic matching score for that category.

[0121] .

[0122] S403. The rule matching score and semantic matching score are weighted and fused to obtain the fusion score of each category. The category with the highest fusion score is selected as the preliminary judgment category of the cluster. The fusion score is also used as the confidence level of the judgment.

[0123] In some embodiments, the fusion score is calculated using the following formula:

[0124] ;

[0125] Among them, R score To weight the normalized rule matching score, α is its preset weight; S score The semantic matching score is represented by β, which is its preset weight.

[0126] In some embodiments, α and β are each preset to a weight of 0.5. The category with the highest fusion score is taken as the initial classification category for that cluster, and this score is used as the confidence level for the classification.

[0127] S404, Intelligent Boundary Category Determination; For the initial category determination, perform a boundary category determination with dual conditions. If the conditions are met, the category of the cluster is re-determined as the boundary category "Other".

[0128] In some embodiments, the dual conditions include condition A and condition B; condition A is: when the highest score among the three main categories is lower than a preset low confidence threshold, and the score of the "other" category is significantly higher than the highest score, typically more than 1.5 times; condition B is: the second highest score is not lower than 20% of the highest score, and the score of the "other" category is higher than the highest score; when conditions A and / or B are met, the cluster is determined to be a boundary category; when neither condition A nor B is met, the cluster is determined to be a belonging category, such as component structure, process method, or material properties.

[0129] In some embodiments, the preset low confidence threshold, multiple, and proportion threshold in step S404 are all configurable hyperparameters, which can be optimized through classification performance on the validation set or set by the user. A 20% threshold is chosen to ensure that the "process method" category can be correctly identified. Since text related to process methods often also involves material composition, there may be cases where both component structure and process method scores are high. A lower threshold ensures that such mixed text can be processed correctly.

[0130] S5. Multi-model weighted voting fusion: The product of the silhouette coefficient and the cluster-level confidence of each clustering algorithm is used as the voting weight of the clustering algorithm; the voting weights of all decision algorithms for each candidate category under the cluster are accumulated, and the candidate category with the highest accumulated voting weight score is used as the final classification category. The final confidence is obtained by weighted averaging of the confidence of each algorithm under the final classification category using the silhouette coefficient.

[0131] In some embodiments, the specific calculation method for weighted voting fusion is as follows: for each clustering algorithm i, its voting weight formula is calculated:

[0132] ;

[0133] Where m i Let c be the silhouette coefficient of this clustering algorithm. i The confidence score of each cluster obtained by matching rules and semantics is given; this product serves as the overall voting weight for the model; for each candidate category, the voting weights W of all clustering algorithms that classify it as belonging to that category are summed. i The weighted voting score for that category is obtained; the category with the highest weighted voting score is selected as the final classification category, and the final confidence score is the average of the confidence scores of all models under that category, weighted by the silhouette coefficient.

[0134] S6. Output Results: Output the final classification category and final confidence score for each material text entity, completing the intelligent labeling of material science text entity categories. Optionally, the final classification category and final confidence score for each material text entity can be output and optionally written back to the knowledge graph database.

[0135] Typically, the intelligent annotation method disclosed in this invention does not require manually annotated training samples, but only relies on a keyword dictionary and semantic template predefined by domain experts. It is a weakly supervised or unsupervised automated classification method. Furthermore, the cluster-level dual-determination strategy reduces the computational complexity of category determination from O(N) to O(M), where N is the number of descriptions of the material text entity and M is the number of clusters.

[0136] Some embodiments disclose a materials science text entity category intelligent annotation system based on cluster analysis, used to implement a materials science text entity category intelligent annotation method based on cluster analysis, including:

[0137] The text data acquisition and preprocessing module is configured to extract materials science text entities and their associated semantic embedding vectors in batches from a knowledge graph database, and then perform preprocessing. In some embodiments, the text entities to be classified are extracted from the knowledge graph database, the text is cleaned, segmented, and stop word removed, and each text is converted into a fixed-dimensional semantic vector using a pre-trained language model.

[0138] The dimensionality reduction module is configured to perform dimensionality reduction on the preprocessed semantic embedding vectors to obtain a low-dimensional vector dataset. In some embodiments, a two-stage dimensionality reduction strategy combining principal component analysis with uniform manifold approximation and projection algorithms is employed to compress high-dimensional semantic vectors to a lower dimension suitable for clustering.

[0139] The multi-model clustering analysis and quality assessment module is configured to use multiple clustering algorithms to cluster the low-dimensional vector dataset, aggregating semantically similar materials science text entities into multiple clusters, and calculating the silhouette coefficient of each clustering algorithm as the clustering quality assessment value. In some embodiments, the multi-model clustering analysis and quality assessment module executes multiple clustering algorithms with different principles in parallel and calculates the clustering quality assessment index of each algorithm.

[0140] The cluster-level category dual-determination module is configured to employ a dual-determination mechanism combining rule matching and semantic template matching. Based on an intelligent boundary category determination strategy with dual-condition verification, it determines the category to which the material entity belongs and the cluster-level confidence level for each cluster. In some embodiments, the clusters generated by clustering are used as the smallest determination unit, replacing the traditional text-by-text determination strategy. The category to which each cluster belongs is determined through a dual mechanism of keyword matching and multi-layer semantic template matching.

[0141] In some embodiments, the cluster-level category dual determination module includes:

[0142] The rule matching submodule is configured to extract high-frequency keywords from all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary for each category, and obtain the rule matching score.

[0143] The semantic template matching submodule is configured to calculate the cosine similarity of these keywords to the general template layer, specific template layer and Chinese template layer vectors when the keywords are not exactly matched in words but are semantically similar. The maximum value is then multiplied by the preset weight of the template to obtain the weighted maximum similarity of the keyword. Finally, the semantic matching score is calculated by weighting according to the adaptive threshold.

[0144] The dual-judgment fusion submodule is configured to perform weighted fusion of rule matching score and semantic matching score to obtain fusion score for each category, select the category with the highest fusion score as the initial judgment category for that cluster, and the fusion score is also used as the confidence level of the judgment.

[0145] The intelligent boundary category determination submodule is configured to perform a dual-condition boundary category determination for the initially determined category. If the conditions are met, the category of the cluster is reclassified as the boundary category "Other". In some embodiments, the intelligent boundary category determination submodule uses an adaptive threshold to determine the fuzzy sample whose classification cannot be determined as the boundary category.

[0146] The multi-model weighted voting fusion module is configured to use the product of the silhouette coefficient and the cluster-level confidence score of each clustering algorithm as the voting weight for that algorithm. It accumulates the voting weights of all decision algorithms for each candidate category within a cluster, and selects the candidate category with the highest accumulated score as the final classification category. The final confidence score is obtained by weighted averaging of the confidence scores of each algorithm under the final classification category using the silhouette coefficient. Typically, the weighted voting fusion module performs weighted voting fusion on the classification results output by each model based on the quality evaluation metrics of each clustering model, obtaining a stable final classification result with quantifiable confidence.

[0147] The results output module is configured to output the final classification category and final confidence score for each material text entity, completing the intelligent labeling of material science text entity categories. In some embodiments, the results output module writes the classification results back to the knowledge graph database to update the category of the corresponding node.

[0148] The technical details are further illustrated below with reference to the embodiments.

[0149] Example 1

[0150] Example 1 provides an intelligent annotation method for material science text entity categories based on cluster analysis. This method classifies and annotates the attributes of 784 material text entities, specifically including the following steps:

[0151] S1. Text Data Acquisition and Vectorization

[0152] The system extracts 784 material science text entities and their 1356-dimensional semantic embedding vectors from the graph database, performs standard normalization preprocessing, and saves them as structured binary data files.

[0153] S2, Dimensionality Reduction

[0154] First, principal component analysis is used to reduce the dimensionality of the semantic embedding vectors while retaining 90% of the variance information. Then, unified manifold approximation and projection are used to further reduce the dimensionality to 50.

[0155] S3, Multi-model Parallel Clustering and Quality Assessment

[0156] After dimensionality reduction, the system ran four clustering algorithms in parallel, with each algorithm undergoing automatic parameter optimization. After optimization, the optimal number of clusters for the K-means algorithm was 8, initialized with K-means++; the optimal neighborhood radius for the density-based clustering algorithm was 0.4, with a minimum sample size of 3, yielding 10 valid clusters and 42 noisy points; the optimal number of clusters for the hierarchical clustering algorithm was 8; and the optimal number of components for the Gaussian mixture model was 8. The silhouette coefficient of each model was calculated as the clustering quality evaluation value for that model on the current dataset.

[0157] S4, Cluster-level Dual Classification Determination

[0158] Using clusters as units, the attribute categories of each cluster can be determined in batches. Taking the 8 clusters obtained by K-means clustering as an example, 10 representative keywords are first extracted for each cluster, and then its category is identified through a dual-judgment mechanism.

[0159] The first step is rule-based matching: comparing representative keywords with keyword sets of various predefined categories and calculating keyword matching scores.

[0160] For example, given 4 keywords in a certain category, the total weight of all keywords is 26, and the average weight is 26 / 4 = 6.5. The sum of the weights of the matched keywords is 15. First, the normalized score (coverage α) is calculated as 15 / 26 ≈ 0.5769. Then, the absolute importance is 15 / 6.5 ≈ 2.3077, the normalized importance is 2.3077 / (1.0+2.3077) ≈ 0.6978, and the weight factor is 1.0+0.6978=1.6978. Finally, the keyword matching score is obtained.

[0161] .

[0162] The second layer is semantic template matching: The system pre-sets multi-level semantic templates for each category. For example, templates for the component structure category include "material composition elements, stoichiometry, crystal structure type," etc., and are stored separately in three categories: general, specific, and Chinese. The system automatically selects the appropriate template based on the language of the input keywords. During calculation, the top 15 representative keywords of each cluster and up to 10 template texts under each template are taken and encoded into fixed-dimensional vectors using an embedding model. For each keyword vector, its cosine similarity with all template vectors is calculated, and the maximum value is multiplied by the template weight to obtain the weighted similarity of that keyword. Based on this, the mean of the weighted similarities of all keywords is calculated (e.g., 0.596). Then, an adaptive threshold is applied.

[0163] ;

[0164] Then, examining the 15 keywords above, 11 have a weighted similarity score not lower than the threshold. Therefore, the percentage of high-quality matching keywords is approximately 11 / 15 ≈ 0.733. The final semantic matching score is:

[0165] ;

[0166] Finally, the two scores are weighted and combined:

[0167] ;

[0168] The category with the highest fusion score is selected as the attribute category of the cluster, and this fusion score is used as the confidence level for the decision.

[0169] S5, Multi-model Weighted Voting Fusion

[0170] For each text entity category, the results of four clustering models are combined, and a weighted fusion is performed using the product of the silhouette coefficient and the corresponding cluster confidence as the voting weight. Taking the category of a certain text entity as an example, the judgments and weighted scores of each model are as follows:

[0171] The DBSCAN clustering model classifies it as "component structure", with a cluster confidence of 0.85, silhouette coefficient of 0.62, and weighted score of 0.527.

[0172] The K-means clustering model identified it as having a "component structure," with a cluster confidence score of 0.72, a silhouette coefficient of 0.48, and a weighted score of 0.346.

[0173] The hierarchical clustering model was classified as "process method", with a cluster confidence of 0.60, a silhouette coefficient of 0.51, and a score of 0.306.

[0174] The Gaussian mixture clustering model was classified as "component structure" with a cluster confidence of 0.55, a silhouette coefficient of 0.44, and a score of 0.242.

[0175] The scores are accumulated by category. "Component Structure" scores 1.115 and "Process Method" scores 0.306, so the final category is "Component Structure". The final confidence level is calculated only by the model that hits this category: multiply the confidence level by the profile coefficient and sum them to get 1.115. Then divide by the sum of the profile coefficients of these models to get 1.54, and the final confidence level is about 0.724.

[0176] S6. Result Output and Database Update

[0177] The system writes the final classification results of 784 text entities back to the graph database, updating the category labels of the corresponding nodes. The classification results are as follows: Figure 3 As shown.

[0178] The technical solutions and technical details disclosed in the embodiments of this invention are merely illustrative of the inventive concept of this invention and do not constitute a limitation on the technical solutions of the embodiments of this invention. Any conventional changes, substitutions, or combinations made to the technical details disclosed in the embodiments of this invention have the same inventive concept as this invention and are within the protection scope of the claims of this invention.

Claims

1. A method for intelligent annotation of entity categories in materials science texts based on cluster analysis, characterized in that: Including the following steps: S1. Text Data Acquisition and Preprocessing: Extract material science text entities and their associated semantic embedding vectors in batches from the knowledge graph database, and perform preprocessing. S2. Dimensionality reduction: The preprocessed semantic embedding vectors are reduced in dimensionality to obtain a low-dimensional vector dataset. S3. Multi-model clustering analysis and quality assessment: Multiple clustering algorithms are used to cluster the low-dimensional vector dataset, and semantically similar material science text entities are aggregated into multiple clusters. The silhouette coefficient of each clustering algorithm is calculated as the clustering quality assessment value. S4. Cluster-level dual determination: A dual determination mechanism combining rule matching and semantic template matching is adopted. Based on the intelligent boundary category determination strategy with dual condition verification, the material entity belonging category and cluster-level confidence of each cluster are determined. S5. Multi-model weighted voting fusion: The product of the silhouette coefficient and the cluster confidence of each clustering algorithm is used as the voting weight of that clustering algorithm; The voting weights of all decision algorithms for each candidate category under the cluster are accumulated, and the candidate category with the highest accumulated voting weight score is taken as the final classification category. The final confidence score is obtained by weighting the confidence scores of each algorithm under the final classification category with the silhouette coefficient. S6. Output Results: Output the final classification category and final confidence score for each material text entity, completing the intelligent labeling of material science text entity categories.

2. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 1, characterized in that, In step S1, the preprocessing includes: Data cleaning, filtering out invalid samples with vector values ​​of "None"; Data transformation: Convert the data list into a NumPy array; Standardize the data using StandardScaler or MinMaxScaler.

3. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 1, characterized in that, Step S2, the dimensionality reduction processing of the semantic embedding vector includes: First, use principal component analysis to reduce dimensionality, retaining 90-95% of the variance information; The dimensionality is reduced to a pre-defined low-dimensional target dimension through uniform manifold approximation and projection.

4. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 1, characterized in that, Step S3 includes the following steps: S301. Determine the set of multi-model clustering algorithms, including K-means clustering, DBSCAN clustering, hierarchical clustering, and Gaussian mixture model (GMM); S302. Automatic parameter optimization is performed on each clustering algorithm, including: K-means clustering: Automatically searches for the optimal combination of the number of clusters and two initialization methods, and uses the elbow rule and silhouette coefficient to determine the optimal K value; DBSCAN clustering: Automatically searches for the optimal combination of neighborhood radius and minimum number of samples. It first filters out combinations with fewer than three clusters or more than half of the clusters being noisy, and then selects the one with the highest silhouette coefficient. If none of these conditions are met, it downgrades to a weighted score of silhouette coefficient multiplied by the noise ratio to select the best one. Hierarchical clustering: Automatically searches for the optimal combination of cluster size, linking method, and distance metric; Gaussian Mixture Model (GMM): Automatically searches for the optimal combination of component number and covariance type; S303. Structured storage of results: The clustering results of each clustering algorithm and the corresponding silhouette coefficient evaluation index are stored in a structured manner for subsequent category determination.

5. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 1, characterized in that, In step S4, the material entity classification includes component structure category, process method category, material property category, and "other" category; The specific steps involved in the dual determination of cluster-level categories are as follows: S401. Extract the high-frequency keywords of all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary of each category, and obtain the rule matching score. S402. For cases where keywords do not match exactly in words but are semantically similar, use the embedding model to calculate the cosine similarity of these keywords to the vectors of the general template layer, the specific template layer, and the Chinese template layer, respectively. Take the maximum value and multiply it by the preset weight of the template to obtain the weighted maximum similarity of the keyword. Then, calculate the semantic matching score according to the adaptive threshold weighting. S403. The rule matching score and semantic matching score are weighted and fused to obtain the fusion score of each category. The category with the highest fusion score is selected as the preliminary judgment category of the cluster. The fusion score is also used as the confidence level of the judgment. S404. For the initial classification, perform a boundary category determination with dual conditions. If the conditions are met, the category of the cluster is reclassified as the boundary category "Other".

6. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 5, characterized in that, In step S401, the method for calculating the rule matching score includes: Each keyword is assigned a weight w; if no weight is specified, the default value is 1; the score for rule matching is calculated based on the keyword weights: the sum of matching weights W match It equals the sum of the weights of all successfully matched keywords; the normalized score, coverage α, is defined as the sum of the matching weights and the sum of the weights of all keywords, W. total The ratio is calculated using the following formula: ; Calculate the average weight of all keywords To match the total weight W match Divide by the average weight to obtain the absolute importance: ; Then, the absolute importance is normalized to the interval [0,1) using the following formula: ; Weighting factors are: ; Weighted normalization rule matching score: 。 7. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 5, characterized in that, In step S402, the semantic matching score calculation method includes: For keyword embedding vectors First, calculate its relationship with all template embedding vectors. The cosine similarity between them is taken as the maximum value and then multiplied by the preset weight of the template. The weighted maximum similarity of the keyword is obtained. This forms a similarity sequence for all keywords; ; Calculate the adaptive similarity threshold: take the average of the similarity sequences. Multiply by a preset coefficient to obtain a baseline threshold, and take the larger of the baseline threshold and the preset minimum threshold as the adaptive threshold: ; Then, the number of keywords in the similarity sequence that exceed the adaptive threshold is counted, and divided by the total number of keywords to obtain the proportion of highly similar keywords, p. Finally, the average similarity of all keywords is multiplied by the proportion of highly similar keywords (p), and the resulting product is the semantic matching score for that category. 。 8. The intelligent annotation method for material science text entity categories based on cluster analysis according to claim 5, characterized in that, In step S403, the formula for calculating the fusion score is: ; Among them, R score To weight the normalized rule matching score, α is its preset weight; S score The semantic matching score is represented by β, which is its preset weight.

9. A materials science text entity category intelligent annotation system based on cluster analysis, used to implement the method described in any one of claims 1 to 8, characterized in that, include: The text data acquisition and preprocessing module is configured to extract materials science text entities and their associated semantic embedding vectors in batches from the knowledge graph database, and then perform preprocessing. The dimensionality reduction module is configured to perform dimensionality reduction processing on the preprocessed semantic embedding vectors to obtain a low-dimensional vector dataset. The multi-model clustering analysis and quality assessment module is configured to use multiple clustering algorithms to cluster the low-dimensional vector dataset, aggregate semantically similar material science text entities into multiple clusters, and calculate the silhouette coefficient of each clustering algorithm as the clustering quality assessment value. The cluster-level category dual determination module is configured to use a dual determination mechanism that combines rule matching and semantic template matching. Based on an intelligent boundary category determination strategy with dual condition verification, it determines the category to which the material entity belongs and the cluster-level confidence of each cluster. The multi-model weighted voting fusion module is configured to use the product of the silhouette coefficient and the cluster-level confidence of each clustering algorithm as the voting weight of that clustering algorithm; The voting weights of all decision algorithms for each candidate category under the cluster are accumulated, and the candidate category with the highest accumulated voting weight score is taken as the final classification category. The final confidence score is obtained by weighting the confidence scores of each algorithm under the final classification category with the silhouette coefficient. The results output module is configured to output the final classification category and final confidence score of each material text entity, thus completing the intelligent labeling of material science text entity categories.

10. The intelligent annotation system for material science text entity categories based on cluster analysis according to claim 9, characterized in that, The cluster-level category dual determination module includes: The rule matching submodule is configured to extract high-frequency keywords from all texts within a single cluster to form a cluster keyword set, calculate the matching degree between the cluster keyword set and the predefined keyword dictionary for each category, and obtain the rule matching score. The semantic template matching submodule is configured to calculate the cosine similarity of these keywords to the general template layer, specific template layer and Chinese template layer vectors when the keywords are not exactly matched in words but are semantically similar. The maximum value is then multiplied by the preset weight of the template to obtain the weighted maximum similarity of the keyword. Finally, the semantic matching score is calculated by weighting according to the adaptive threshold. The dual-judgment fusion submodule is configured to perform weighted fusion of rule matching score and semantic matching score to obtain fusion score for each category, select the category with the highest fusion score as the initial judgment category for that cluster, and the fusion score is also used as the confidence level of the judgment. The intelligent boundary category determination submodule is configured to perform a boundary category determination under two conditions for the initial category determination. If the conditions are met, the category of the cluster is re-determined as the boundary category "Other".