Material data cleaning method and system fusing large language model and knowledge graph
By integrating large language models and knowledge graphs, the problems of semantic ambiguity and difficulty in parsing ambiguous meanings in material data cleaning are solved, and high-precision and explainable material data cleaning is achieved, which adapts to changes in corporate business and provides continuous learning capabilities.
Patent Information
- Application Number
- CN202511254237.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-04
AI Technical Summary
Existing technologies in material master data cleaning have problems such as semantic ambiguity and difficulty in resolving ambiguous meanings, difficulty in capturing implicit associated knowledge, and poor interpretability of the cleaning process, which lead to cleaning errors or omissions, and the single AI model lacks industry standard basis.
Integrate the large language model with the knowledge graph, extract candidate entities through the large language model and retrieve similar nodes in combination with the knowledge graph, construct knowledge-enhanced prompt words for parsing, and perform secondary verification and attribute completion in combination with the knowledge graph to ensure that the output complies with industry standards.
It improves the accuracy and credibility of material data cleaning, realizes highly reliable and traceable intelligent cleaning, can adapt to changes in enterprise context and continuously learn, and improves the accuracy and transparency of cleaning results.
Smart Images

Figure CN120744327A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data governance technology, and specifically to a material data cleaning method and system that integrates a large language model and a knowledge graph. Background Art
[0002] The statements in this section merely provide background art related to the present invention and do not necessarily constitute prior art.
[0003] In the digital transformation of manufacturing, material master data serves as the core foundation for production, procurement, and inventory management, and its quality directly impacts operational efficiency. Currently, the industry relies primarily on three approaches for master data cleansing: manual verification, where professionals check feature items and values, which is inefficient and susceptible to subjective factors; rule engines, which identify errors based on preset rules (such as regular expressions), but struggle to cover diverse, non-standard descriptions; and the application of single AI technologies, such as using only large models to process text (which can easily produce output that doesn't meet industry standards) or relying solely on knowledge graphs (which struggle to handle fuzzy semantics).
[0004] The above-mentioned master data cleaning method also has the following problems: (1) Semantic ambiguity and ambiguity resolution are difficult. Traditional rule engines cannot effectively resolve terminology conflicts (such as "cold-rolled plate" and "cold-hardened plate"), fuzzy expressions (such as "about 5cm steel pipe") and context-dependent semantics in material descriptions, resulting in cleaning errors or omissions; (2) Implicit associated knowledge is difficult to capture and apply. The complex matching, substitution, and component relationships between materials (such as the matching of seals and bearings) are difficult to convert into explicit cleaning rules, resulting in loss of data correlation and affecting the integrity of master data; (3) The cleaning process and results are poorly interpretable. The cleaning decision-making process of a single AI model (such as a pure large language model) is not transparent and lacks traceability based on industry standards or corporate specifications, which affects trust and manual review efficiency. Summary of the Invention
[0005] In order to address the shortcomings of the existing technology, the present invention provides a material data cleaning method and system that integrates a large language model and a knowledge graph. The knowledge graph provides the large language model with precise structured knowledge anchors and standardized constraints, effectively suppressing the "hallucinations" of the large language model and ensuring that the output content complies with industry standards; the large language model injects powerful natural language understanding, fuzzy semantic parsing and unstructured information extraction capabilities into the knowledge graph, making up for the lack of flexibility of the knowledge graph. The two are deeply integrated to jointly solve the key problems of "understanding the semantics of the text" and "determining whether the semantics conform to the standards", thereby improving the accuracy of material data cleaning.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a material data cleaning method that integrates a large language model and a knowledge graph.
[0007] A material data cleaning method that integrates a large language model and a knowledge graph includes the following steps: Extract candidate entities from the language description input by the user based on the large language model; Vectorize the candidate entities and retrieve multiple similar nodes in the preset knowledge graph based on the vectorized candidate entities; Based on the node information of each similar node, knowledge enhancement prompt words are constructed to enable the large language model to generate language description recognition results, which include material entity link relationships; A secondary verification is performed based on the material entity link relationship through the knowledge graph. Once the secondary verification passes, a large language model is used to generate material classification prediction results based on the material entity link relationship. Based on the material classification prediction results, attribute verification is performed in combination with the knowledge graph; The large language model completes the missing attributes based on the verification results and the context, and performs standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after completing the missing attributes are combined with the preset knowledge graph template to obtain a standardized description; Calculate the comprehensive similarity score between the standardized description and the verified entity standardized description. When the comprehensive score is greater than the set threshold, it is determined to be a duplicate entity; otherwise, the standardized description is stored as new standardized entity data.
[0008] In an implementation of the first aspect of the present invention, the comprehensive similarity score is a weighted sum of the large language model similarity and the knowledge graph relationship similarity.
[0009] In one implementation of the first aspect of the present invention, when the output confidence of the large language model is less than a set threshold, or the link relationship violates the preset constraints in the knowledge graph, the secondary verification fails, and the secondary verification result is sent to the client for manual review; If the standardization check fails, the completed attribute value will be sent to the client for manual review.
[0010] In an implementation of the first aspect of the present invention, the knowledge graph includes: a standard layer, an external data source layer, and an instance layer; The standard layer is used to store standard texts converted into structured rules; The external data source layer is used to connect to external dynamic data sources to dynamically update industry terminology, new materials, and new specifications; The instance layer is used to store standardized descriptions of entities that have been verified in the enterprise history.
[0011] In an implementation of the first aspect of the present invention, the knowledge enhancement prompt word includes: a prompt word instruction input by a user, a knowledge base generated according to node information of similar nodes, and a task output format.
[0012] In one implementation of the first aspect of the present invention, a large language model is used to generate a material classification prediction result based on a material entity link relationship, including: The prediction enhancement prompt words are generated according to the material entity link relationship. The large language model selects the main category from the knowledge graph as the material classification prediction result based on the prediction enhancement prompt words.
[0013] In one implementation of the first aspect of the present invention, attribute verification is performed in conjunction with the knowledge graph, including: required attribute verification, hierarchical rationality verification, and attribute conflict verification; The completed attribute values are standardized and verified based on the knowledge graph, including: value range verification, unit conversion, and attribute consistency verification.
[0014] In one implementation of the first aspect of the present invention, the large language model is fine-tuned based on manual review results; When the frequency of a new term is greater than a first set threshold, or the manual confirmation rate is greater than a second set threshold, the new term is added to the knowledge graph; When the number of conflicts for the same attribute exceeds a third set threshold, the conflict pattern is extracted and candidate constraint rules are generated for manual confirmation. After manual confirmation, the candidate constraint rules are added to the knowledge graph; A fine-tuned large language model is used to parse manual annotations and output triples, which are then added to the knowledge graph after confidence filtering.
[0015] In a second aspect, the present invention provides a material data cleaning system that integrates a large language model and a knowledge graph.
[0016] A material data cleaning system that integrates a large language model and knowledge graph, including: An entity extraction unit is configured to: extract candidate entities from the language description input by the user based on the large language model; The node retrieval unit is configured to: vectorize the candidate entity, and retrieve multiple similar nodes in a preset knowledge graph based on the vectorized candidate entity; A link relationship recognition unit is configured to: construct knowledge enhancement prompt words based on the node information of each similar node so that the large language model generates a language description recognition result, and the language description recognition result includes the material entity link relationship; The verification unit is configured to: perform secondary verification based on the material entity link relationship through the knowledge graph; after the secondary verification passes, generate a material classification prediction result based on the material entity link relationship using the large language model; and perform attribute verification based on the material classification prediction result in combination with the knowledge graph; The attribute completion unit is configured to: use the large language model to complete the missing attributes based on the verification results and context, and perform standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after completing the missing attributes are combined with the preset knowledge graph template to obtain a standardized description; The data cleaning unit is configured to: calculate the comprehensive similarity score between the standardized description and the verified entity standardized description; when the comprehensive score is greater than a set threshold, it is determined to be a duplicate entity; otherwise, the standardized description is stored as new standardized entity data.
[0017] In a third aspect, the present invention provides a computer device comprising: a processor and a computer-readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by the processor, implements the material data cleaning method for integrating a large language model with a knowledge graph as described in the first aspect of the present invention.
[0018] Compared with the prior art, the present invention has the following beneficial effects: This invention innovatively proposes a material data cleaning method that integrates a large language model and a knowledge graph. The knowledge graph provides the large language model with precise structured knowledge anchors and standardized constraints, effectively suppressing the "hallucinations" of the large language model and ensuring that the output content complies with industry standards. The large language model injects powerful natural language understanding, fuzzy semantic parsing, and unstructured information extraction capabilities into the knowledge graph to make up for the lack of flexibility of the knowledge graph. The two are deeply integrated to jointly solve the key problems of "understanding the semantics of the text" and "determining whether the semantics conform to the standards", thereby improving the accuracy of material data cleaning and providing a highly reliable and traceable intelligent solution for complex material data governance.
[0019] This invention innovatively proposes a material data cleaning method that integrates a large language model and a knowledge graph. The closed-loop mechanism synchronously feeds back the manual review results to the model fine-tuning and knowledge graph update, optimizes the large language model parameters by manually annotating data, and enhances adaptability to the enterprise context; at the same time, the confirmed new terms and rules are injected into the knowledge graph in the form of triples to form a dynamically expanding knowledge network, so that the system has continuous learning capabilities and can automatically adapt to business changes (such as new product lines, terminology adjustments), avoiding frequent reconstruction of traditional rule bases. It has significant advantages in cleaning accuracy, cross-domain adaptability and long-term maintenance costs, and provides an intelligent solution for complex material data governance in manufacturing and other industries.
[0020] This invention innovatively proposes a material data cleaning method that integrates a large language model with a knowledge graph. The knowledge graph stores core knowledge such as the enterprise's material classification system, term definitions, and quality rules in the form of structured triples, providing clear standard basis for cleaning operations (such as material coding specifications, attribute value ranges, and association verification rules), ensuring that each cleaning decision can be traced back to a specific knowledge node, fundamentally solving the problem of unexplainable results caused by the "black box" cleaning method used in traditional methods. At the same time, the large language model uses natural language reasoning capabilities to convert the rule constraints of the knowledge graph into highly readable text descriptions, making the entire cleaning process, from rule matching to result output, transparent. This not only significantly improves the credibility of the cleaning results, but also enables the system to continuously adapt to changes in the enterprise's terminology system and business rules through the dynamic update mechanism of the knowledge graph.
[0021] Advantages of additional aspects of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0023] Figure 1 A three-layer knowledge graph architecture diagram provided for an exemplary embodiment of the present invention; Figure 2 A schematic diagram of a flow chart of a material data cleaning method that integrates a large language model and a knowledge graph, provided as an exemplary embodiment of the present invention; Figure 3 A closed-loop flow chart of feedback learning provided for an exemplary embodiment of the present invention; Figure 4 A schematic diagram of a material data cleaning system that integrates a large language model and a knowledge graph, provided as an exemplary embodiment of the present invention; Figure 5A schematic diagram of a computer device is provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0025] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0026] This implementation proposes a material data cleaning method that integrates a large language model and a knowledge graph, and constructs a cleaning framework that is collaboratively driven by two engines and has dynamic and autonomous evolution capabilities.
[0027] In this implementation, a domain knowledge graph (KG) for material master data governance is constructed as the structured knowledge base and standardized constraint source for the cleansing process; the LLM (Large Language Model) and KG collaborative cleansing module receives the original material description text and, through deep interaction between LLM and KG, completes intelligent parsing, linking, verification, standardization, and matching, outputting high-quality, standardized master data records.
[0028] In this implementation, a domain knowledge graph for material master data governance is constructed, such as Figure 1 As shown, Figure 1 The system in this context refers to the client-side system that supports user operations. The core goal is to build a domain KG with dynamic mapping capabilities to provide structured knowledge anchors and real-time constraints for cleaning. Specifically, it includes: Step 1.1: Material body definition.
[0029] (1) Cleaning-oriented ontology attribute design, specifically including: class MaterialOntology: mandatory_attrs=[classification, material] mandatory attributes (supports verification in step 2); constraint_rules={ Cold rolled sheet: Thickness: {unit: mm; min: 0.3; max: 3.0}; Range constraints (for standardization in step 3); Related standards: GB / T 708 standard reference (used for description generation in the subsequent step 4); } } (2) Relationship system design, specifically, as shown in Table 1.
[0030] Table 1: Relationship system
[0031] Step 1.2: Initialize the three-layer knowledge architecture.
[0032] The standard layer integrates authoritative specifications such as national standards (GB / T) and international standards (ISO, IEC), providing the most fundamental standard basis.
[0033] Convert natural language standards into structured rules (to support strong verification in step 3). Specifically, this includes: def convert_standard(text): Example: Analyze GB / T 20878-Stainless steel grade 06Cr19Ni10 corresponding to 304 return { source:GB / T 20878; rule: IF material = 304 stainless steel THEN standardized to = 06Cr19Ni10 Drools executable rule } The external data source layer connects to external dynamic data sources such as product databases and industry standard libraries, and continuously updates industry terminology, new materials, new specifications, etc.
[0034] The instance layer stores the company's historically verified, high-quality master data, which is used to supplement rules and provide statistical information. Specifically, it includes: { Entity: 6205-2RS bearing; Attributes: {Inner diameter: 25mm; Outer diameter: 52mm}; Statistics: Width frequency: {120mm: 82%; 125mm: 18%} / / Supports LLM attribute completion in step 3 } } The following is an introduction taking automobile bearing KG as an example: Ontology definition, specifically, includes: Core Concepts concepts=[deep groove ball bearing; seal ring; roller] Attribute constraints (support cleaning and verification) attrs={ Inner diameter: {unit: mm; precision: 0.01}; Accuracy level: {enum: [P0; P6; P5]} in accordance with GB / T 307.1 } Key relationships relations=[(deep groove ball bearing; hasPart; sealing ring)] is used to complete implicit knowledge.
[0035] The three-layer initialization is shown in Table 2.
[0036] Table 2: Three-layer initialization results
[0037] Dynamic mapping converts external terms into standard layer mappings, specifically including: Automatically analyze the correspondence between SKF data fields and GB standards map_skf_to_gb(skf_data): if skf_data[type] == contact_seal: return {Standard seal type: RS; Based on: GB / T 276 Article 4.2} Output the standard code available for cleaning.
[0038] The example execution, specifically, includes: Enter SKF data: {Model: 6205-2RS; Seal description: Double-sided rubber contact seal}; Output mapping: {Seal type: RS; Standard source: GB / T 276}.
[0039] like Figure 2 As shown, this implementation provides a collaborative cleaning method between LLM and KG. Through deep, bidirectional interaction between LLM and KG, accurate semantic parsing, dynamic verification, and standardization are achieved. The LLM engine performs natural language understanding and reasoning, subject to real-time KG constraints; the KG engine provides structured knowledge (standard nodes, attribute constraints, and classification trees). Its core collaborative steps include entity recognition and linking, material classification prediction and verification, attribute value standardization and completion, and description standardization and duplicate detection.
[0040] Example: The user inputs 304 stainless steel cold-rolled plate with a thickness of 0.5±0.05mm.
[0041] Step 1: Entity recognition and linking.
[0042] (1) LLM initially extracts candidate entities in the description (e.g., material = 304 stainless steel; category = cold-rolled plate; thickness = 0.5 ± 0.05 mm); (2) Vectorize the candidate entities and retrieve the top-3 similar nodes in the KG (based on embedding cosine similarity); (3) Return node information: {ID: KG_MAT_304; Standard name: ASTM_A276-304, Synonym: [304 stainless steel], Constraint: Thickness unit = mm, Value range [0.3, 3.0]}; (4) Construct knowledge enhancement prompt words, integrate KG retrieval results and constraints, and guide LLM decisions.
[0043] Specifically, the examples are as follows: [System Instructions] You are a material data cleaning expert. Please parse the text strictly according to the knowledge base constraints: [Input text] 304 stainless steel cold-rolled plate, about 0.5cm thick [Knowledge Base] 1. KG_MAT_304: {Standard Name: ASTM_A276-304; Synonym: [304 Stainless Steel]; Unit Constraint: Thickness Unit = mm}; 2. KG_PLATE: {Classification: Cold-rolled Plate; Required Attributes: [Material, Thickness, Width]} [Task] Output JSON: {Material: ASTM_A276-304; Thickness: {value: 5.0; Unit: mm}; Missing attribute: [width]}.
[0044] (5) KG secondary verification: If the LLM output confidence is <0.7, or the link result violates the KG constraint (e.g., thickness = 0.5 mm but the cold-rolled plate node requires the thickness unit to be cm), it is marked as a low-confidence entity and transferred to the manual review queue.
[0045] Step 2: Material classification prediction and verification.
[0046] (1) LLM preliminary prediction, input enhanced prompt words (including linked entities).
[0047] Specifically, the examples are as follows: Known entity: {Material=ASTM_A276-304, Category=GB / T_708-ColdRolledSheet}; Select the main category from the KG category tree: [Steel / Hot-rolled Plate, Steel / Cold-rolled Sheet…]; Output: {primary_class: steel / cold-rolled sheet, confidence: 0.95}. (2) KG verification includes mandatory attribute check, hierarchical rationality, and attribute conflict detection. For example, mandatory attribute check: under the cold-rolled plate classification, material, thickness, and width are required, which triggers the width attribute completion, as shown in Table 3.
[0048] Table 3: KG calibration types
[0049] Step 3: Attribute completion and standardization. This step uses LLM semantic parsing and the KG rule engine to achieve accurate standardization of attribute values. LLM attribute value extraction and completion: For missing attributes, LLM dynamically completes them based on the context.
[0050] Specifically, the examples are as follows: Known: Classification = Steel / Cold-rolled Sheet, Thickness = 0.5±0.05mm; Material = 304 Stainless Steel; Generate the width value according to the KG constraint: {cold-rolled sheet width range [1000, 1500mm]}; Output: {attribute name: width; value: 1200mm; basis: default width of similar materials in KG}.
[0051] The KG strong constraint verification engine executes the check logic based on the Drools rule engine. Specifically, it includes: Range check: If width = 2000mm (outside the KG range), reject immediately; Intelligent unit conversion: automatically convert 5cm to 50mm (according to KG conversion factor); Property consistency check: For example, if Surface Treatment = Galvanized and Material = Stainless Steel → conflict (according to GB / T 13911).
[0052] Step 4: Standardized description generation and duplicate detection.
[0053] LLM fills in the KG template, specifically including: Template: {category} | Standard: {standard number} | Material: {material} | Thickness: {thickness}; Output: Cold-rolled steel plate | Standard: GB / T 708 | Material: 06Cr19Ni10 | Thickness: 0.50±0.05mm.
[0054] Repeat testing, specifically including: Fusion of LLM semantic similarity + KG graph relationship similarity (such as supplier and BOM association); If the weighted score is greater than the set threshold, it is determined to be a duplicate material; otherwise, it is considered unique standardized data (i.e., a new material).
[0055] This implementation also provides a feedback learning and knowledge evolution solution, which converts the uncertain results and manual review feedback in the cleaning process into system knowledge, drives the continuous optimization of LLM and KG, and realizes dynamic self-evolution.
[0056] Example: Manual review found that a "nylon washer" was mistakenly cleaned as a "rubber washer"; Step (1): Manual intervention, marking the correct result: material = nylon 66 (PA66), temperature range: -40℃~120℃; Step (2): LLM fine-tuning, adding new training data: the correct label of the original description "white high-temperature resistant gasket" is PA66, strengthening the association between "color + temperature resistance" and material; Step (3): KG dynamic update, specifically including: Rule 1: Term addition (automation).
[0057] If the frequency of new terms exceeds the set threshold AND the manual confirmation rate exceeds 95%, the following terms are added: KG.add_node(name=PA66; type=material; props={melting point: 265℃}); Rule 2: Constraint generation (semi-automatic).
[0058] If the number of conflicts for the same attribute is greater than 5, the conflict pattern is automatically extracted and candidate rules are generated. After manual confirmation, the rule is generated: if material = PA66 AND operating temperature > 120°C, an error is output.
[0059] Rule 3: Relation completion (LLM assistance).
[0060] Use the fine-tuned LLM to parse manual annotations (e.g., this gasket is used for bearing seal): output triples [gasket; used for; bearing seal], which are filtered by confidence and then added to the KG.
[0061] More specifically, Figure 3 As shown in the figure, starting with manual review, the review data is converted into structured knowledge graph triples, and then the update type is determined through intelligent diversion: terms / attributes are directly injected into the knowledge graph for real-time expansion, and rules / constraints are used for constraint perception training of the large language model. After the dual engines complete the cleaning in collaboration, the system automatically identifies low-confidence results and feeds back manual review, forming a closed-loop optimization mechanism of "review-diversion-update-cleaning-feedback", which not only ensures rule compliance through the knowledge graph, but also improves semantic understanding ability with the help of the large language model, and ultimately achieves continuous iterative improvement of cleaning accuracy and standard adaptability.
[0062] Figure 4 A material data cleaning system that integrates a large language model and a knowledge graph is shown, including: The entity extraction unit 401 is configured to: extract candidate entities from the language description input by the user according to the large language model; The node retrieval unit 402 is configured to: vectorize the candidate entity, and retrieve multiple similar nodes in a preset knowledge graph based on the vectorized candidate entity; The link relationship identification unit 403 is configured to: construct knowledge enhancement prompt words based on the node information of each similar node so that the large language model generates a language description recognition result, and the language description recognition result includes the material entity link relationship; Verification unit 404 is configured to: perform secondary verification based on the material entity link relationship through the knowledge graph; after the secondary verification passes, generate a material classification prediction result based on the material entity link relationship using the large language model; and perform attribute verification based on the material classification prediction result in combination with the knowledge graph; The attribute completion unit 405 is configured to: complete the missing attributes based on the verification results and the context of the large language model, and perform standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after the missing attributes are completed are combined with the preset knowledge graph template to obtain a standardized description; The data cleaning unit 406 is configured to calculate a comprehensive similarity score between the standardized description and the verified entity standardized description, and determine the entity as a duplicate when the comprehensive score is greater than a set threshold; otherwise, store the standardized description as new standardized entity data.
[0063] It is understandable that each of the above-mentioned units can be separately or completely combined into one or several other units to form a unit, or one (or some) of the units can be further divided into multiple functionally smaller units to form a unit, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the system may also include other units. In actual applications, these functions can also be implemented with the assistance of other units and can be implemented by the collaboration of multiple units.
[0064] According to another embodiment of the present application, the system described in this embodiment can be constructed by running a computer program (including program code) capable of executing the steps involved in the corresponding method of the present invention on a general-purpose computing device such as a computer, which includes processing elements and storage elements such as a central processing unit (CPU), random access memory (RAM), and read-only memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.
[0065] Figure 5A computer device is shown, which includes a processor 501, a communication interface 502, and a computer-readable storage medium 503. The processor 501, the communication interface 502, and the computer-readable storage medium 503 may be connected via a bus or other means.
[0066] Among them, the communication interface 502 is used to receive and send data, the computer-readable storage medium 503 can be stored in the memory of the electronic device, the computer-readable storage medium 503 is used to store computer programs, the computer programs include program instructions, and the processor 501 is used to execute the program instructions stored in the computer-readable storage medium 503.
[0067] The processor 501 is the computing core and control core of the electronic device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement corresponding method processes or corresponding functions.
[0068] The processor 501 is configured to perform the following process: Extract candidate entities from the language description input by the user based on the large language model; Vectorize the candidate entities and retrieve multiple similar nodes in the preset knowledge graph based on the vectorized candidate entities; Based on the node information of each similar node, knowledge enhancement prompt words are constructed to enable the large language model to generate language description recognition results, which include material entity link relationships; A secondary verification is performed based on the material entity link relationship through the knowledge graph. Once the secondary verification passes, a large language model is used to generate material classification prediction results based on the material entity link relationship. Based on the material classification prediction results, attribute verification is performed in combination with the knowledge graph; The large language model completes the missing attributes based on the verification results and the context, and performs standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after completing the missing attributes are combined with the preset knowledge graph template to obtain a standardized description; Calculate the comprehensive similarity score between the standardized description and the verified entity standardized description. When the comprehensive score is greater than the set threshold, it is determined to be a duplicate entity; otherwise, the standardized description is stored as new standardized entity data.
[0069] Those skilled in the art will appreciate that the units and algorithmic steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technical personnel may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0070] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. A computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data processing device such as a server or data center that integrates one or more available media. Available media can include magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives).
[0071] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A material data cleaning method integrating a large language model and a knowledge graph, characterized in that: The following processes are included: Extract candidate entities from the language description input by the user based on the large language model; Vectorize the candidate entities and retrieve multiple similar nodes in the preset knowledge graph based on the vectorized candidate entities; Based on the node information of each similar node, knowledge enhancement prompt words are constructed to enable the large language model to generate language description recognition results, which include material entity link relationships; A secondary verification is performed based on the material entity link relationship through the knowledge graph. Once the secondary verification passes, a large language model is used to generate material classification prediction results based on the material entity link relationship. Based on the material classification prediction results, attribute verification is performed in combination with the knowledge graph; The large language model completes the missing attributes based on the verification results and the context, and performs standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after completing the missing attributes are combined with the preset knowledge graph template to obtain a standardized description; Calculate the comprehensive similarity score between the standardized description and the verified entity standardized description. When the comprehensive score is greater than the set threshold, it is determined to be a duplicate entity; otherwise, the standardized description is stored as new standardized entity data.
2. The material data cleaning method integrating a large language model and a knowledge graph according to claim 1 is characterized in that: The comprehensive similarity score is the weighted sum of the large language model similarity and the knowledge graph relationship similarity.
3. The material data cleaning method integrating a large language model and a knowledge graph as claimed in claim 1 is characterized in that: When the output confidence of the large language model is less than the set threshold, or the link relationship violates the preset constraints in the knowledge graph, the secondary verification fails and the secondary verification result is sent to the client for manual review; If the standardization check fails, the completed attribute value will be sent to the client for manual review.
4. The material data cleaning method integrating a large language model and a knowledge graph according to claim 1 is characterized in that: Knowledge graph, including: standard layer, external data source layer and instance layer; The standard layer is used to store standard texts converted into structured rules; The external data source layer is used to connect to external dynamic data sources to dynamically update industry terminology, new materials, and new specifications; The instance layer is used to store standardized descriptions of entities that have been verified in the enterprise history.
5. The material data cleaning method integrating a large language model and a knowledge graph according to claim 1 is characterized in that: The knowledge enhancement prompt word includes: a prompt word instruction input by the user, a knowledge base generated according to the node information of similar nodes, and a task output format.
6. The material data cleaning method integrating a large language model and a knowledge graph according to claim 1 is characterized in that: Based on the material entity link relationship, a large language model is used to generate material classification prediction results, including: The prediction enhancement prompt words are generated according to the material entity link relationship. The large language model selects the main category from the knowledge graph as the material classification prediction result based on the prediction enhancement prompt words.
7. The material data cleaning method integrating a large language model and a knowledge graph according to claim 1 is characterized in that: Combined with the knowledge graph to perform attribute verification, including: required attribute verification, hierarchical rationality verification, and attribute conflict verification; The completed attribute values are standardized and verified based on the knowledge graph, including: value range verification, unit conversion, and attribute consistency verification.
8. The material data cleaning method integrating a large language model and a knowledge graph according to any one of claims 1 to 7, characterized in that: Fine-tune the large language model based on manual review results; When the frequency of a new term is greater than a first set threshold, or the manual confirmation rate is greater than a second set threshold, the new term is added to the knowledge graph; When the number of conflicts for the same attribute exceeds a third set threshold, the conflict pattern is extracted and candidate constraint rules are generated for manual confirmation. After manual confirmation, the candidate constraint rules are added to the knowledge graph; A fine-tuned large language model is used to parse manual annotations and output triples, which are then added to the knowledge graph after confidence filtering.
9. A material data cleaning system integrating a large language model and a knowledge graph, characterized in that: include: An entity extraction unit is configured to: extract candidate entities from the language description input by the user based on the large language model; The node retrieval unit is configured to: vectorize the candidate entity, and retrieve multiple similar nodes in a preset knowledge graph based on the vectorized candidate entity; A link relationship recognition unit is configured to: construct knowledge enhancement prompt words based on the node information of each similar node so that the large language model generates a language description recognition result, and the language description recognition result includes the material entity link relationship; The verification unit is configured to: perform secondary verification based on the material entity link relationship through the knowledge graph; after the secondary verification passes, generate a material classification prediction result based on the material entity link relationship using the large language model; and perform attribute verification based on the material classification prediction result in combination with the knowledge graph; The attribute completion unit is configured to: use the large language model to complete the missing attributes based on the verification results and context, and perform standardization verification on the completed attribute values based on the knowledge graph. After the standardization verification passes, the classification prediction results after completing the missing attributes are combined with the preset knowledge graph template to obtain a standardized description; The data cleaning unit is configured to: calculate the comprehensive similarity score between the standardized description and the verified entity standardized description; when the comprehensive score is greater than a set threshold, it is determined to be a duplicate entity; otherwise, the standardized description is stored as new standardized entity data.
10. A computer device, characterized in that: include: a processor and a computer-readable storage medium; a processor adapted to execute a computer program; A computer-readable storage medium having a computer program stored therein, wherein the computer program, when executed by the processor, implements the material data cleaning method for integrating a large language model with a knowledge graph as described in any one of claims 1 to 8.
Citation Information
Patent Citations
DCS intelligent decision-making method and system fusing large language model and knowledge graph
CN118820778A
Traditional Chinese medicine question-answering system construction method based on large language model and knowledge graph
CN118838996A
Knowledge graph fraud-related subject analysis and completion method based on large language model
CN119066132A
Knowledge graph construction method based on fine-tuning large language model
CN119808917A
Intelligent maintenance reasoning method based on knowledge graph and large language model
CN119886334A
Cited By
Industrial multi-modal data alignment method and device for carbon emission management and medium
CN121598284A
Large model corpus knowledge graph management method and system
CN122287822A