Data management agent system and implementation method
By using a data governance intelligent agent system to achieve unified data parsing and quality optimization, the system solves the problem of integrating structured and unstructured data in enterprise data governance, improves data governance efficiency and quality, and is suitable for professional data governance in the manufacturing and financial industries.
Patent Information
- Application Number
- CN202511000682.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-18
AI Technical Summary
In enterprise data governance, the integration of structured and unstructured data is difficult, resulting in serious data silos. Traditional governance methods rely on manual parsing, which is inefficient and lacks effective monitoring of the accuracy of data parsing, failing to meet the specialized data governance needs of specific industries.
A data governance intelligent agent system is adopted, including a data parsing module, a data storage module, a user response module, a quality monitoring and feedback module, and a domain-specific large language model module, to achieve unified data parsing, automated knowledge construction, and quality optimization. The data parsing module uses multimodal parsing and knowledge graph construction, storing the data in a graph database and a vector database. The user response module supports natural language interaction, the quality monitoring module optimizes data quality through closed-loop feedback, and the domain-specific large language model adapts to the needs of vertical domains.
Breaking down data barriers, reducing reliance on manual labor, and improving the efficiency and quality of data governance are applicable to the specialized data governance needs of industries such as manufacturing and finance.
Smart Images

Figure CN120973780A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing, and more specifically, to a data governance intelligent agent system and its implementation method. Background Technology
[0002] Currently, enterprise data governance commonly faces the challenge of integrating structured and unstructured data, leading to severe data silos. Traditional governance methods rely on manual parsing, entity extraction, and relation construction, heavily depending on the expertise of data engineers and domain specialists. This not only incurs high labor costs but also results in long governance cycles and low efficiency. Furthermore, the lack of effective monitoring mechanisms for data parsing accuracy and the difficulty in closing the loop on user feedback lead to poor data quality consistency, hindering the support of intelligent enterprise decision-making. In addition, general-purpose large language models lack sufficient accuracy in entity recognition and relation extraction within vertical domains, failing to meet the specialized data governance needs of specific industries. Therefore, there is an urgent need for an intelligent data governance solution that can achieve unified data parsing, automated knowledge construction, dynamic quality optimization, and domain adaptation. Summary of the Invention
[0003] The purpose of this invention is to provide a data governance intelligent agent system and its implementation method.
[0004] In a first aspect, embodiments of the present invention provide a data governance intelligent agent system, comprising:
[0005] The system includes a data parsing module, a data storage module, a user response module, a quality monitoring and feedback module, and a domain-specific large language model module; among them,
[0006] The data parsing module is configured to perform unified parsing and structured modeling of structured and unstructured data within the enterprise.
[0007] The data storage module is configured to store structured knowledge data processed by the data parsing module, including a graph database and a vector database;
[0008] The user response module is configured to support users to interact with the system through natural language in order to achieve intelligent question answering of the structured knowledge data.
[0009] The quality monitoring and feedback module is configured to optimize data governance quality through a closed-loop feedback mechanism, including monitoring data parsing results, collecting user feedback, and iteratively optimizing the model.
[0010] The domain-specific large language model module is configured to build and optimize large language models for vertical domains to adapt to the data parsing, user interaction, and quality optimization needs of specific industries or scenarios.
[0011] In one possible implementation, the data parsing module includes an unstructured data processing unit, which is configured to:
[0012] Text extraction, knowledge graph construction, and knowledge vectorization processing of unstructured data within enterprises;
[0013] The text extraction involves using a multimodal large model to parse the scanned document to extract the text content, and combining engineering techniques to extract the structured text content from the non-scanned document.
[0014] The knowledge graph is constructed by using a domain-specific large language model to identify entities, entity attributes, and relationships between entities from the extracted text, and storing the extracted entities and their relationships in the graph database to form a preliminary knowledge graph.
[0015] The knowledge vectorization involves using engineering methods to perform semantic segmentation on the extracted text, and then using an embedding model to generate semantic vectors from the text segments, which are then stored in the vector database.
[0016] In one possible implementation, the data parsing module further includes a structured data processing unit, which is configured to:
[0017] For internal SQL databases, the system automatically identifies table structure, field types, and relationships between fields based on the database schema information, and constructs a data structure ontology to define entity categories, attributes, and relationships.
[0018] The data fields are abstracted and categorized using engineering methods, and the categorization results are organized into entities and their relationships based on the data structure ontology, thereby enriching the content of the knowledge graph.
[0019] In one possible implementation, the data storage module includes: the graph database is Neo4j, configured to store entities and relationships in a graph structure, supporting complex relationship retrieval and reasoning; the vector database is Qdrant, configured to store semantic vectors of text fragments, supporting semantic retrieval based on similarity calculation.
[0020] In one possible implementation, the user response module includes a user intent parsing subunit, a data query subunit, and an answer generation subunit;
[0021] The user intent parsing subunit is configured to parse the user's natural language query intent based on the domain big language model, and identify key entities, events, time, location and category information from the original query input;
[0022] The data query subunit is configured to construct a graph database query statement based on the recognition result of the user intent parsing subunit to retrieve relevant entities and their relationships in the graph database using the domain big language model.
[0023] Simultaneously, user queries are converted into semantic vectors, and relevant document fragments or knowledge pieces are retrieved from the vector database.
[0024] The answer generation subunit is configured to combine the retrieval results obtained from the graph database and vector database by the data query subunit, and generate a structured or natural language answer using the domain-specific large language model.
[0025] In one possible implementation, the quality monitoring and feedback module includes a model supervision subunit, a user feedback collection subunit, and a model fine-tuning subunit.
[0026] The model supervision subunit is configured to use the domain big language model to perform legality verification on the entities and relations extracted from the SQL database by the data parsing module, and automatically mark the entities or relations that do not conform to the preset verification rules for manual review and confirmation.
[0027] The user feedback collection subunit is configured to collect user feedback on the system output results through the user interaction interface, including illegal entities or relationships marked by the user, and omitted entities or relationships.
[0028] The model fine-tuning subunit is configured to automatically include the erroneous labeled samples into the training dataset when the cumulative number of erroneous labeled samples from the model supervision subunit and the user feedback collection subunit exceeds a preset threshold, and to fine-tune the domain-specific large language model based on the training dataset to generate an optimized version.
[0029] In one possible implementation, the domain-wide language model module includes a base model selection subunit, a model fine-tuning subunit, and a continuous optimization subunit.
[0030] The base model selection subunit is configured to select a large language model with a preset parameter scale and support for vertical domain knowledge transfer as the initial base model. The preset parameter scale is set according to the data complexity and knowledge reasoning requirements of the target domain.
[0031] The model fine-tuning subunit is configured to continuously fine-tune the initial base model based on the error-labeled samples and user feedback data collected by the quality monitoring and feedback module, so as to enhance its entity recognition, relation extraction and question answering capabilities in a specific domain.
[0032] The continuous optimization subunit is configured to improve the accuracy, robustness, and knowledge understanding ability of the domain-specific large language model through multiple rounds of iterative fine-tuning and effect evaluation.
[0033] In one possible implementation, the model supervision subunit's verification of the legality of the data parsing results includes entity legality verification and relation legality verification;
[0034] The entity legality verification includes determining whether the entity belongs to a preset domain entity category;
[0035] The relationship validity verification includes determining whether the relationship between two entities conforms to preset domain relationship rules.
[0036] In one possible implementation, the quality monitoring and feedback module further includes a manual review subunit;
[0037] The manual review subunit is configured to receive entities or relationships marked by the model supervision subunit that do not conform to the preset verification rules, and send the erroneous entities or relationships confirmed by manual review to the user feedback collection subunit for aggregation.
[0038] In a second aspect, embodiments of the present invention provide a method for implementing a data governance intelligent agent, applied to the data governance intelligent agent system described in the first aspect, comprising:
[0039] The data parsing module performs unified parsing and structured modeling of structured and unstructured data within the enterprise to obtain structured knowledge data.
[0040] The structured knowledge data processed through data parsing and modeling steps is stored in the graph database and vector database of the data storage module;
[0041] The user response module receives natural language queries from users, parses the user's intent, retrieves relevant knowledge data from graph databases and vector databases, and generates structured or natural language answers to feed back to the user.
[0042] The quality monitoring and feedback module verifies the legality of the data parsing results, collects erroneous entities, relationships and omissions reported by users, accumulates error samples, and fine-tunes the domain-wide language model when the error exceeds a preset threshold.
[0043] The initial base model is selected through the domain-specific large language model module. Based on the fine-tuning data from the quality monitoring and feedback optimization steps, the model is continuously iterated and optimized to adapt to the data governance needs of specific industries or scenarios.
[0044] Compared to existing technologies, the beneficial effects provided by this invention include: The data governance intelligent agent system and implementation method disclosed in this invention pertain to the field of data processing. This system includes a data parsing module, a data storage module, a user response module, a quality monitoring and feedback module, and a domain-specific large language model module. The data parsing module provides unified parsing and modeling for both structured and unstructured data; the data storage module uses graph databases and vector databases to store knowledge data; the user response module supports natural language interaction and intelligent question answering; the quality monitoring and feedback module optimizes data quality through closed-loop feedback; and the domain-specific large language model module constructs optimized models adapted to vertical domains. This invention, through multimodal parsing, knowledge graph construction, dual-database collaborative storage, and dynamic model optimization, breaks down data barriers, reduces reliance on manual intervention, and improves the efficiency, quality, and intelligence level of data governance. It is suitable for the specialized data governance needs of industries such as manufacturing and finance. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 A schematic block diagram of the data governance intelligent agent system provided in an embodiment of the present invention;
[0047] Figure 2 This is a flowchart illustrating the steps of a data governance intelligent agent implementation method provided in an embodiment of the present invention;
[0048] Figure 3 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0050] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0051] In order to solve the technical problems mentioned in the background art Figure 1This is a schematic block diagram of the data governance intelligent agent system provided in the embodiments of this disclosure. The data governance intelligent agent system will be described in detail below.
[0052] This invention provides a data governance intelligent agent system 110, comprising:
[0053] The module comprises a data parsing module 1101, a data storage module 1102, a user response module 1103, a quality monitoring and feedback module 1104, and a domain-specific large language model module 1105; among which,
[0054] The data parsing module 1101 is configured to perform unified parsing and structured modeling of structured and unstructured data within the enterprise.
[0055] The data storage module 1102 is configured to store structured knowledge data processed by the data parsing module 1101, including a graph database and a vector database;
[0056] The user response module 1103 is configured to support users to interact with the system through natural language in order to achieve intelligent question answering of the structured knowledge data.
[0057] The quality monitoring and feedback module 1104 is configured to optimize data governance quality through a closed-loop feedback mechanism, including monitoring data parsing results, collecting user feedback, and iteratively optimizing the model.
[0058] The domain-specific large language model module 1105 is configured to build and optimize large language models for vertical domains to adapt to the data parsing, user interaction, and quality optimization needs of specific industries or scenarios.
[0059] In this embodiment of the invention, exemplarily targeting a data governance scenario in a manufacturing enterprise, the specific implementation process of the data governance intelligent agent system 110 described in this invention is explained in detail, with the server as the execution entity. This manufacturing enterprise needs to govern data across its entire internal procurement, production, and sales process, including structured SQL business databases (such as customer information tables and supplier contract tables) and unstructured document data (such as scanned purchase contracts, email correspondence, and product manual PDFs). This system enables unified data parsing, intelligent retrieval, and dynamic optimization, supporting core business scenarios such as enterprise procurement decisions and supplier management.
[0060] The server initiates a unified parsing and structured modeling process for internal enterprise data through the data parsing module 1101. This module includes unstructured data processing units and structured data processing units, which perform specialized processing for different types of data.
[0061] For unstructured data, the server first calls the unstructured data processing unit to classify the non-scanned documents (such as Word format procurement contract templates and Excel quotation forms) and scanned documents (such as scanned copies of paper contracts and scanned copies of handwritten approval forms) stored by the enterprise. For scanned documents, the server loads a multi-modal large model (such as the InternVL3-78B model adapted for document parsing), and extracts the text content by combining OCR technology with multi-modal semantic understanding. For example, for a scanned copy of a "Procurement Contract with Supplier A" in October 20xx, the server first uses the model to identify the seal and signature areas in the document to locate the key information segments, and then performs text extraction on the body area to obtain structured text such as "Purchaser: XX Manufacturing Co., Ltd.; Supplier: Supplier A; Procurement Item: Precision Gears; Quantity: 500 pieces; Unit Price: 200 yuan / piece; Delivery Date: December 31, 20xx". For non-scanned documents (such as PDF product manuals), the server directly extracts the text using engineering technical means (such as the PyPDF2 tool combined with regular expressions), and distinguishes areas such as titles, bodies, and tables through layout analysis algorithms to ensure the integrity of the text structure.
[0062] After text extraction is completed, the server starts the knowledge graph construction process: calls the current optimized version model in the domain large language model module 1105 (a large language model fine-tuned based on the manufacturing domain) to identify entities, entity attributes, and relationships in the extracted text. Taking the text of the "Procurement Contract with Supplier A" as an example, the model automatically identifies the entities "XX Manufacturing Co., Ltd." (type: purchaser), "Supplier A" (type: supplier), "Precision Gears" (type: product), "500 pieces" (type: quantity), "200 yuan / piece" (type: unit price), "December 31, 20xx" (type: delivery date); the entity attributes include "Specification Model: M10×20" and "Material: Steel No. 45" of "Precision Gears"; the relationships between entities include "XX Manufacturing Co., Ltd. - Procurement - Precision Gears", "Supplier A - Supply - Precision Gears", "Precision Gears - Quantity - 500 pieces", and "Precision Gears - Unit Price - 200 yuan / piece". The server stores the above entities, attributes, and relationships in the graph database of the data storage module 1102 in the form of triples (such as <XX Manufacturing Co., Ltd., Procurement, Precision Gears>) to form a preliminary procurement domain knowledge graph.
[0063] Meanwhile, the server performs semantic segmentation on the extracted text using knowledge vectorization subunits: using engineering methods (such as sentence boundary detection algorithms), the full contract text is semantically divided into segments such as "Contract Main Clauses," "Subject Description Clauses," "Price Clauses," and "Delivery Clauses," with each segment's length controlled within 512 tokens; then, an embedding model (such as paraphrase-multilingual-mpnet-base-v2) is called to generate a 768-dimensional semantic vector for each segment. For example, the vector generated from the "Subject Description Clauses" segment "The subject matter of the purchase is precision gears, specification model M10×20, material 45 steel, and execution of GB / T10095.1-2021 standard" will be stored in the vector database for subsequent semantic similarity retrieval.
[0064] For structured data, the server connects to the enterprise's internal SQL database (such as the "Customer Information Table" and "Supplier Contract Table" in a MySQL database) through a structured data processing unit. Taking the "Supplier Contract Table" as an example, the server first reads the database schema information and automatically identifies the table structure: table name "supplier_contract", fields include "contract_id" (primary key, string type), "supplier_id" (foreign key, related to the "id" field of the "supplier" table), "product_id" (foreign key, related to the "id" field of the "product" table), "quantity" (integer type), "unit_price" (floating-point type), and "delivery_date" (date type). Based on the above schema, the server constructs a data structure ontology, defining entity categories as "contract", "supplier", and "product". Entity attributes include "contract_id" and "delivery_date" for "contract", "supplier_id" and "name" for "supplier", and "product_id" and "specification" for "product". Relationships between entities include "contract-related-supplier" (based on the "supplier_id" foreign key) and "contract-containing-product" (based on the "product_id" foreign key). Subsequently, the server abstracts and categorizes data fields using engineering methods (such as SQL queries combined with field mapping rules). For example, the value "500" in the "quantity" field is categorized as the "quantity" attribute, and the value "200.00" in the "unit_price" field is categorized as the "unit price" attribute. Based on the data structure ontology, these categorization results are organized into entity relation triples (such as <Contract C001, Contains, Product P001>, <Contract C001, Unit Price, 200.00 yuan / piece>), which are then added to the knowledge graph of the graph database to achieve knowledge fusion between structured and unstructured data.
[0065] After the server completes data parsing, it stores the structured knowledge data into the graph database and the vector database respectively through the data storage module 1102, forming a dual storage architecture of "relational knowledge + semantic knowledge".
[0066] The graph database uses Neo4j, and the server uses the Cypher query language to store entities, attributes, and relationships generated during the knowledge graph construction process in a graph structure. For example, for knowledge related to "Supplier A's purchase contract", the server executes the Cypher statement "CREATE(c:Contract{contract_id:C001,delivery_date:20xx-12-31})-[:Purchase]->(p:Product{product_id:P001,name:Precision Gear,specification:M10×20})", creating the "Contract C001" node, the "Product P001" node, and the "Purchase" relationship edge between them in the graph database, and adding attribute information to the nodes. This graph database supports complex retrieval based on relational paths. For example, the server can quickly obtain the supplied products and corresponding contract numbers of all suppliers since 20xx by querying “MATCH(s: Supplier)-[: Supply]->(p: Product)<-[: Procurement]-(c: Contract)WHERE c.delivery_date>20xx-01-01RETURNs.name,p.name,c.contract_id”, which meets the needs of enterprises to sort out supply chain relationships.
[0067] Vector database uses Q d The rant database uses an HTTP API to store the semantic vectors of text segments generated during knowledge vectorization, along with corresponding metadata (such as the document ID to which the segment belongs and paragraph titles), into the "contract_segments" collection of the vector database. Each vector record contains an "id" (a unique identifier for the segment), a "vector" (a 768-dimensional semantic vector), and a "payload" (metadata, such as "doc_id": "C001", "segment_title": "subject description clause"). The vector database supports approximate nearest neighbor retrieval based on cosine similarity. When a user query involves specific clause details, the server can quickly locate relevant text segments by calculating the similarity between the query vector and vectors in the database. For example, for a query "What are the execution standards for precision gears?", the vector database can return the segment with the highest similarity, "subject description clause", ensuring the semantic relevance of the search results.
[0068] When a purchasing employee enters a natural language query through the system's front-end interface, asking "Please query the total quantity and corresponding contract number of precision gears in all purchase contracts signed with Supplier A in 20xx," the server executes an intelligent question-and-answer process through user response module 1103. The specific steps are as follows:
[0069] First, the server invokes the user intent parsing subunit, loading a domain-specific large language model to identify the user's intent and extract key information from the query. The model uses Named Entity Recognition (NER) technology to identify the key entities "20xx year" (time), "Supplier A" (supplier), and "precision gears" (product) from the query. Through an intent classification model, it determines the user's intent as "data statistics query" (target: total purchase quantity; additional requirement: contract number). Simultaneously, the model normalizes the query, generating structured query conditions: "Time range: 20xx year; Supplier: Supplier A; Product: Precision gears; Statistical dimension: total quantity; Return field: contract number."
[0070] Subsequently, the server initiated the data query subunit, performing a dual-database collaborative retrieval based on the aforementioned structured query conditions. For graph database retrieval, the server converted the query conditions into a Cypher query statement using a domain-specific large language model: "MATCH(s:Supplier{name:Supplier A})-[:Supply]->(p:Product{name:Precision Gears})<-[:Purchase]-(c:Contract)WHEREc.delivery_date>=20xx-01-01ANDc.delivery_date<=20xx-12-31RETURNc.contract_id,c.quantity". After executing this query, the graph database returned the results: Contract No. C001 (Quantity 500 pieces), Contract No. C003 (Quantity 300 pieces). For vector database retrieval, the server converts the user query "total purchase quantity of precision gears" into a semantic vector, performs a similarity search in the "contract_segments" set, and returns text fragments related to "purchase quantity" (such as the "quantity clause" of contract C001: "This purchase involves 500 pieces of precision gears, to be delivered in two batches") as supplementary evidence to the graph database retrieval results.
[0071] Finally, the server calls the answer generation subunit and combines the search results from both databases to generate a natural language answer. The domain-specific large language model summarizes and calculates the contract numbers and quantities returned by the graph database (500 + 300 = 800 items), and integrates the clause details returned by the vector database to generate a structured answer: "Two precision gear purchase contracts were signed with supplier A in 20xx, with contract numbers C001 and C003, and a total purchase quantity of 800 items (500 items in contract C001 and 300 items in contract C003). The relevant contract terms show that contract C001 stipulates delivery in two batches, while contract C003 stipulates delivery in one batch." The server then feeds this answer back to the user through the front-end interface, completing the intelligent question-and-answer interaction.
[0072] The server uses the quality monitoring and feedback module 1104 to perform closed-loop quality optimization of the entire data governance process, ensuring the accuracy of the knowledge graph and the reliability of the system response. The specific process is as follows:
[0073] During the data parsing phase, the server activates the model supervision subunit to perform validity checks on the entities and relationships extracted from the SQL database by the structured data processing unit. Pre-defined validation rules include: entity validity rules (e.g., the "type" attribute of the "supplier" entity can only be "enterprise" or "individual business owner," not "individual") and relationship validity rules (e.g., in the "contract-include-product" relationship, the "product" entity must be associated with a valid "specification" attribute). For example, when the server parses the "supplier B contract," if it finds that the "supplier_type" field value in the SQL database is "individual," the model supervision subunit determines that the entity does not conform to the preset rules for the "supplier" category, automatically marks the entity as an "entity item that does not conform to the preset validation rules," and adds it to the manual review queue.
[0074] The server pushes the tagged item to the enterprise data administrator's review interface through the manual review sub-unit. The interface displays the content to be reviewed: "Entity: Supplier B; Attribute: type = Individual; Validation result: Does not comply with preset rules (allowed value: Enterprise / Individual Business)". After verification, the data administrator confirms that the tag is correct (Supplier B is actually an individual business, and the database field was entered incorrectly), and submits the review result ("Error confirmed, correct attribute value: Individual Business") to the system through the interface operation. After receiving the manual review result, the server's user feedback collection sub-unit stores the error sample (original entity attribute, correct attribute, error type) into the "Error Tag Sample Library".
[0075] During the user interaction phase, if a user disagrees with the system's returned answer (e.g., "The query result omitted Contract C005, signed in November 20xx"), they can submit feedback via the "Feedback" button on the front-end interface. The user feedback collection subunit records this feedback information: "Query intent: 20xx Supplier A precision gear contract; Omitted entity: Contract C005 (number C005, quantity 200 pieces)", and adds it to the "Error Labeling Sample Library" as an error sample of the "Omitted Entity" type.
[0076] When the number of erroneous samples accumulated in the "erroneous labeling sample library" reaches a preset threshold (e.g., 100), the server triggers a model fine-tuning subunit: 80% of the samples are extracted from the sample library as the training set and 20% as the validation set to fine-tune the domain-specific language model. The fine-tuning process employs LoRA (Low-Rank Adaptation) technology, freezing most model parameters and optimizing only low-rank matrix parameters related to entity recognition and relation extraction. For example, for erroneous samples of the "supplier type" attribute, the model adjusts the weights of the entity attribute classifier by learning from correctly labeled data ("individual" → "sole proprietorship"); for "missing contract" samples, the model optimizes the sensitivity of the relation extraction module to the "contract-purchase-product" relationship to ensure that similar contract data is correctly identified in subsequent parsing. After fine-tuning, the server continuously optimizes the subunit to evaluate the new model's performance (e.g., entity recognition accuracy, relation extraction F1 score). If the evaluation metrics are better than the old model (e.g., accuracy improved by 5%), the new model is deployed as the currently active model, achieving iterative optimization of system performance.
[0077] The server uses the Domain Large Language Model Module 1105 to build and continuously optimize a vertical domain large language model adapted to manufacturing enterprises, supporting the core needs of data parsing, user interaction, and quality optimization. The specific process is as follows:
[0078] In the initial stage of model building, the server selects sub-units from the base model to evaluate the characteristics of enterprise data: Manufacturing domain data contains a large number of professional terms (such as "tolerance grade IT7" and "surface roughness Ra1.6"), complex product structure relationships (such as "assembly-component-part" hierarchy), and has high requirements for entity recognition accuracy (such as distinguishing between "supplier" and "manufacturer"). Based on this, the server selects a large language model with a parameter scale of more than 200 bytes and support for domain knowledge transfer as the initial base model (such as a general large language model with a parameter scale adapted to complex knowledge reasoning). This model has already learned basic language understanding capabilities through pre-training, but needs to be specifically optimized for the manufacturing domain.
[0079] During the model fine-tuning phase, the server constructs a fine-tuning dataset based on the accumulated erroneous labeled samples (such as entity attribute errors, relationship omissions, and intention comprehension deviations in user interactions) from the quality monitoring and feedback module 1104, combined with data from the enterprise's internal "manufacturing domain terminology library" and "procurement process specification documents." For example, for the entity attribute extraction task of "product specifications and models," the server selects typical samples such as "M10×20 (module × number of teeth)" and "Φ50H7 (diameter × tolerance zone)" from the terminology library, mixes them with erroneous labeled samples, and then fine-tunes the base model. During the fine-tuning process, the server adopts a "multi-round iterative fine-tuning" strategy: the first round of fine-tuning focuses on entity recognition and relationship extraction capabilities (optimizing the requirements of the data parsing module 1101), the second round of fine-tuning enhances intention comprehension and answer generation capabilities (optimizing the requirements of the user response module 1103), and after each round of fine-tuning, the model performance is evaluated using a test set (containing 1000 manufacturing domain labeled data). If the entity recognition accuracy does not reach 90%, 500 special samples are added to start the next round of fine-tuning.
[0080] During the continuous optimization phase, the server periodically (e.g., monthly) tracks the model's performance through continuous optimization sub-units, collecting metrics such as the entity extraction error rate (target ≤3%) of the data parsing module 1101, the answer satisfaction score of the user response module 1103 (target ≥4.5 / 5 points), and the manual review pass rate of the quality monitoring module (target ≥95%). If a metric fails to meet the target for two consecutive weeks (e.g., the entity extraction error rate rises to 5%), the server automatically triggers incremental fine-tuning: it selects key samples from the latest accumulated error samples (e.g., the high-frequency error type "product specification confusion") and updates the model's parameters locally, combining this with newly added domain data (e.g., the company's newly released "20XX version of the procurement terminology standard"), avoiding the resource consumption caused by full-scale fine-tuning. Through this mechanism, the domain-specific large language model gradually adapts to the professional data characteristics of manufacturing enterprises, continuously improving entity recognition accuracy, relation extraction completeness, and user intent understanding precision, ultimately supporting the data governance intelligent agent system 110 in achieving efficient and accurate enterprise data governance goals.
[0081] In this embodiment of the invention, the data parsing module 1101 includes an unstructured data processing unit, which is configured as follows:
[0082] Text extraction, knowledge graph construction, and knowledge vectorization processing of unstructured data within enterprises;
[0083] The text extraction involves using a multimodal large model to parse the scanned document to extract the text content, and combining engineering techniques to extract the structured text content from the non-scanned document.
[0084] The knowledge graph is constructed by using a domain-specific large language model to identify entities, entity attributes, and relationships between entities from the extracted text, and storing the extracted entities and their relationships in the graph database to form a preliminary knowledge graph.
[0085] The knowledge vectorization involves using engineering methods to perform semantic segmentation on the extracted text, and then using an embedding model to generate semantic vectors from the text segments, which are then stored in the vector database.
[0086] In an embodiment of the present invention, for example, the server takes the scenario of a manufacturing company’s “20XX Supplier B CNC Lathe Procurement Contract” as an example. It processes the scanned contract (including paper scans with handwritten signatures) and electronic PDF attachments through an unstructured data processing unit to achieve the structured transformation of unstructured information.
[0087] The server first identifies the type of file to be processed: scanned document "B Supplier Purchase Contract Scan.pdf" (including seal and handwritten signature) and non-scanned document "CK6150 Lathe Technical Parameters.pdf" (electronically generated PDF, including tables).
[0088] For scanned documents, the server calls a multimodal large model (such as InternVL3-78B) to perform parsing. The model uses visual-text fusion technology to locate the following areas on the page: title area (“Purchase Contract”), body text area (printed text), signature area (handwritten “Wang Wu”), and seal area (“XX Manufacturing Contract Seal”). When performing OCR extraction on the body text area, the model combines semantic correction to correct recognition errors, correcting “Delivery Date: 20XX.03.31” to “Delivery Date: March 31, 20XX”, and finally extracting the structured text: “Purchaser: XX Manufacturing Co., Ltd. (Unified Social Credit Code: 91310XXX); Supplier: Supplier B (Contact Person: Zhao Liu, Tel: 139XXXX8901); Purchased Item: CNC Lathe (Model CK6150); Quantity: 2 units; Unit Price: 185,000 yuan / unit; Total Amount: 370,000 yuan; Quality Guarantee Period: 18 months.”
[0089] For non-scanned technical attachments, the server uses the PyPDF2 tool to directly read the text stream, and combines regular expressions to identify table boundaries, extracting data from the "CK6150 Technical Parameter Table": "Spindle speed: 30-2000r / min; Machining diameter ≤500mm; X-axis travel 280mm; Z-axis travel 1000mm; Positioning accuracy: X-axis ≤0.015mm, Z-axis ≤0.02mm"; At the same time, the main text clause is extracted through layout analysis: "The equipment must comply with GB / T17421.1-2016 'General Rules for Machine Tool Inspection', and includes 1 set of three-jaw chuck and 4 sets of anchor bolts."
[0090] The server calls a large language model in the field (fine-tuned in the manufacturing field) to perform entity, attribute, and relationship recognition on the extracted text.
[0091] In the entity recognition stage, the model recognizes core entities based on preset categories: "XX Manufacturing Co., Ltd." (purchaser), "Supplier B" (supplier), "CNC lathe CK6150" (equipment), "2 units" (quantity), "185,000 yuan / unit" (unit price), "GB / T17421.1-2016" (standard number).
[0092] In the attribute recognition stage, attributes are matched to entities: "XX Manufacturing Co., Ltd." → "Unified social credit code: 91310XXX"; "CNC lathe CK6150" → "Model: CK6150", "Positioning accuracy: X-axis ≤ 0.015mm", "Quality assurance period: 18 months".
[0093] In the relationship recognition stage, associations are constructed based on domain rules: "XX Manufacturing Co., Ltd. - Purchase - CNC lathe CK6150", "Supplier B - Supply - CNC lathe CK6150", "CNC lathe CK6150 - Quantity - 2 units", "CNC lathe CK6150 - Conforms to standard - GB / T17421.1-2016".
[0094] The server writes the above triples (such as <XX Manufacturing Co., Ltd., Purchase, CNC lathe CK6150>) into the Neo4j graph database through Cypher statements, forming a preliminary knowledge graph containing procurement entities and equipment parameters.
[0095] The server performs semantic sharding on the extracted text: The scanned document is split into 4 segments of "contract entity", "procurement subject matter", "price terms", and "quality assurance terms" (300 - 500 characters per segment) according to chapters and semantic logic, and the technical attachments are split into 3 segments of "parameter table description", "performance indicators", and "standard requirements". For example, the content of the "procurement subject matter clause": "Procurement subject matter: CNC lathe (model CK6150); Quantity: 2 units; Quality assurance period: 18 months, starting from the date of acceptance."
[0096] Subsequently, an embedding model (paraphrase-multilingual-mpnet-base-v2) is called to generate 768-dimensional semantic vectors for each segment, and the vectors of the "performance indicators" segment (including spindle speed, machining diameter, etc.) represent the core meaning by encoding keyword semantics. The server writes the vectors and metadata (document ID "B_20XX01", segment title "procurement subject matter clause") into the "lathe_contract_segments" collection of the Qdrant vector database through the API, completing the vectorized storage of unstructured data.
[0097] In this embodiment of the invention, the data parsing module 1101 further includes a structured data processing unit, which is configured to:
[0098] For internal SQL databases, the system automatically identifies table structure, field types, and relationships between fields based on the database schema information, and constructs a data structure ontology to define entity categories, attributes, and relationships.
[0099] The data fields are abstracted and categorized using engineering methods, and the categorization results are organized into entities and their relationships based on the data structure ontology, thereby enriching the content of the knowledge graph.
[0100] In this embodiment of the invention, for example, the server uses the SQL database (MySQL 8.0) of the manufacturing enterprise's procurement management system as the processing object. This database contains three core business tables: suppliers, products, and contracts, which are used to store cooperation data between the enterprise and its suppliers. The server parses these tables through a structured data processing unit, transforming the structured data into entities and relationships in a knowledge graph, thereby achieving the integration with unstructured data knowledge.
[0101] The server first connects to the SQL database via JDBC driver and queries the INFORMATION_SCHEMA system table to obtain the schema information of the target table. For the contracts table, the server identifies the table structure details: table name contracts, primary key contract_id (string type, length 20, not null), foreign key supplier_id (related to the id field of the suppliers table, integer type), foreign key product_id (related to the id field of the products table, integer type), and ordinary fields quantity (integer type, describing the purchase quantity), unit_price (DECIMAL type, precision 10.2, describing the unit price), delivery_date (DATE type, describing the delivery date), and sign_date (DATE type, describing the contract signing date).
[0102] Based on the above schema, the server constructs a data structure ontology: defining entity categories as "Contract", "Supplier", and "Product"; entity attributes include "Contract" with contract_id (unique identifier), delivery_date (delivery date), and sign_date (signing date); "Supplier" with id (supplier number), name (supplier name), and contact_phone (contact number); and "Product" with id (product number), name (product name), and specification (specification). Relationships between entities are defined based on foreign key associations: "Contract-Association-Supplier" (using the supplier_id foreign key to describe the supplier to which the contract belongs) and "Contract-Includes-Product" (using the product_id foreign key to describe the product involved in the contract).
[0103] The server uses an engineering approach (such as SQL queries combined with a field mapping rule base) to abstract and categorize the common fields of the contracts table. The field mapping rule base pre-defines common attribute categories in the manufacturing field, such as "quantity," "unit price," "total amount," and "date," and defines the mapping logic: the quantity field (purchase quantity) maps to the "quantity" attribute, the unit_price field (unit price) maps to the "unit price" attribute, and delivery_date and sign_date map to the "date" attribute (further subdivided into "delivery date" and "contract signing date").
[0104] The server executes the SQL query `SELECT contract_id, supplier_id, product_id, quantity, unit_price, delivery_date FROM contracts WHERE sign_date>=20XX-01-01` to extract contract data since 20XX. For a single record in the returned results (contract_id=HT20XX001, supplier_id=102, product_id=503, quantity=300, unit_price=150.50, delivery_date=20XX-04-15), the server abstracts and categorizes it according to a rule base: quantity=300 is categorized as the "quantity" attribute, unit_price=150.50 as the "unit price" attribute, and delivery_date=20XX-04-15 as the "delivery date" attribute.
[0105] Based on the data structure ontology, the server organizes the abstracted and categorized results into entity and relation triples, and adds them to the knowledge graph of the graph database.
[0106] First, associate the entities corresponding to the foreign key: query the suppliers table using supplier_id=102 to get the supplier name name=C Precision Components Factory, and determine that the "supplier" entity is "C Precision Components Factory"; query the products table using product_id=503 to get the product name name=bearing and specification=6205-ZZ, and determine that the "product" entity is "bearing (specification 6205-ZZ)".
[0107] Then, generate entity relationship triples: based on the "contract-association-supplier" relationship, generate <Contract HT20XX001, Associated, C Precision Components Factory>; based on the "contract-inclusion-product" relationship, generate <Contract HT20XX001, Inclusion, Bearing (Specification 6205-ZZ)>; based on the abstract classification attributes, generate <Contract HT20XX001, Quantity, 300 pieces>, <Contract HT20XX001, Unit Price, 150.50 yuan / piece>, and <Contract HT20XX001, Delivery Date, April 15, 20XX>.
[0108] Finally, the server uses Cypher statements to write the aforementioned triples into the Neo4j graph database. For example, executing `MATCH(c:Contract{contract_id:HT20XX001}),(s:Supplier{name:C Precision Components Factory})CREATE(c)-[:Relationship]->(s)` completes the transformation of structured data into a knowledge graph. At this point, the graph database contains both the "scanned copy of the purchase contract" entity extracted from unstructured data and the "Contract HT20XX001" entity transformed from the SQL database, achieving cross-data source knowledge fusion and providing complete data support for subsequent user queries.
[0109] In this embodiment of the invention, the data storage module 1102 includes: the graph database is Neo4j, configured to store entities and relationships in a graph structure, supporting complex relationship retrieval and reasoning; the vector database is Qdrant, configured to store semantic vectors of text fragments, supporting semantic retrieval based on similarity calculation.
[0110] In this embodiment of the invention, for example, the server uses knowledge data from the manufacturing enterprise's procurement domain as the storage object. It utilizes the Neo4j graph database and the Qdrant vector database to achieve structured storage of entity relationships and vectorized storage of textual semantics, respectively, supporting subsequent complex relationship reasoning and semantic retrieval needs. The following section details the storage process and retrieval capabilities using a scenario involving the enterprise's procurement data governance from January to March 20XX.
[0111] The server stores the entities, attributes, and relationships generated by the data parsing module 1101 into the Neo4j database in a graph structure. Specifically, each entity in the database corresponds to a node, which contains a unique identifier (such as contract_id) and attribute key-value pairs (such as delivery_date:"20XX-03-31"); the association between entities corresponds to directed relationship edges, which contain relationship type (such as "purchase" or "supply") and attribute (such as sign_date:"20XX-01-15").
[0112] Taking the "CNC lathe purchase contract from supplier B" and the "bearing contract from precision component manufacturer C" as examples, the server completes the storage using Cypher statements:
[0113] Write the contract node: CREATE(c:contract{contract_id:HT20XX001,sign_date:20XX-01-15,delivery_date:20XX-03-31});
[0114] Write to the supplier node: CREATE(s:supplier{id:102,name:C Precision Components Factory,contact:138XXXX5678});
[0115] Write to the product node: CREATE(p:product{id:503,name:bearing,specification:6205-ZZ,material:GCr15});
[0116] Establish relationship edges: MATCH(c:contract{contract_id:HT20XX001}),(s:supplier{id:102})CREATE(c)-[:association{type:primary supplier}]->(s), MATCH(c:contract{contract_id:HT20XX001}),(p:product{id:503})CREATE(c)-[:contains{quantity:300,unit_price:150.50}]->(p).
[0117] After storage, Neo4j supports complex relationship retrieval and reasoning. For example, when the purchasing department needs to query "whether there are any historical quality problem records for the products involved in the bearing contract signed with C Precision Components Factory in 20XX", the server performs a multi-hop relationship query:
[0118] MATCH(s: Supplier {name: C Precision Components Factory}) <- [: Association] - (c: Contract) - [: Contains] -> (p: Product {name: Bearing}) - [: Quality Issue Exists] -> (q: Quality Record) WHERE c.sign_date >= 20XX-01-01 RETURN c.contract_id, p.specification, q.problem_desc. Database return result: contract_id: HT20XX001, specification: 6205-ZZ, problem_desc: Radial clearance out of tolerance in December 20xx batch, helping companies quickly locate potential supply chain risks.
[0119] The server stores the text segment semantic vectors and metadata generated by knowledge vectorization into the Qd rant database. Each vector record contains three parts: a 768-dimensional semantic vector (representing the semantics of the text), a unique ID (e.g., seg_20XX01_005), and metadata (e.g., doc_id:tech_attach_20XX01, entity:bearing6205-ZZ, segment_title:technical parameters and quality standards).
[0120] Taking the text fragment of "Bearing 6205-ZZ Technical Parameter Attachment" as an example, the server writes a vector through Qdrant's HTTP API:
[0121] The request body contains: {"points":[{"id":1001,"vector":[0.0xx,-0.156,...,0.089](768-dimensional vector),"payload":{"doc_id":"tech_attach_20XX01","entity":"Bearing 6205-ZZ","segment_title":"Technical Parameters and Quality Standards","content":"Bearing 6205-ZZ complies with GB / T307.1-2017 standard, radial clearance range 20-40μm, material is high carbon chromium bearing steel GCr15, hardness HRC58-62"}}]};
[0122] The server receives the response {"status":"ok"}, confirming that vector storage is complete.
[0123] Qdrant supports semantic retrieval based on cosine similarity. When the technical department queries "What is the radial clearance standard for bearing 6205-ZZ?", the server generates a semantic vector from the query text "Radial clearance standard for bearing 6205-ZZ" using an embedding model and calls Q... d rant search API:
[0124] POST / collections / tech_segments / search{"vector":[0.031,-0.142,...,0.095],"limit":1,"filter":{"must":[{"key":"entity","match":{"value":"Bearing 6205-ZZ"}}]}}.
[0125] The database returns the fragment with the highest similarity (0.92), whose payload contains "radial clearance range 20-40μm". The server extracts this information and feeds it back to the user to achieve accurate semantic matching.
[0126] Through the collaborative storage of Neo4j and Qdrant, the server achieves both structured organization of entity relationships and fine-grained representation of text semantics, providing dual-engine support for intelligent question answering and decision support of data governance agents.
[0127] In this embodiment of the invention, the user response module 1103 includes a user intent parsing subunit, a data query subunit, and an answer generation subunit;
[0128] The user intent parsing subunit is configured to parse the user's natural language query intent based on the domain big language model, and identify key entities, events, time, location and category information from the original query input;
[0129] The data query subunit is configured to construct a graph database query statement based on the recognition result of the user intent parsing subunit to retrieve relevant entities and their relationships in the graph database using the domain big language model.
[0130] Simultaneously, user queries are converted into semantic vectors, and relevant document fragments or knowledge pieces are retrieved from the vector database.
[0131] The answer generation subunit is configured to combine the retrieval results obtained from the graph database and vector database by the data query subunit, and generate a structured or natural language answer using the domain-specific large language model.
[0132] In this embodiment of the invention, for example, the server takes a natural language query from a purchasing department employee of a manufacturing company as an example. Through the collaborative work of the three sub-units of the user response module 1103, the entire process from intent parsing to answer generation is completed. The user inputs the query through the system front end: "What is the total purchase quantity of bearings of model 6205-ZZ in the first quarter of 20XX purchased from C Precision Components Factory? Please list the corresponding contract number and the quantity of each purchase."
[0133] The server invokes the user intent parsing subunit, loading a domain-specific large language model (fine-tuned for the manufacturing and procurement domain) to parse the query text. The model first uses an intent classification algorithm to determine the user intent as a "statistical query," specifically requiring "total purchase quantity calculation" and "contract detail listing." Subsequently, the model performs Named Entity Recognition (NER) and key information extraction.
[0134] Entity recognition: Extract the core entities "C Precision Components Factory" (supplier category), "Bearing" (product category), and "6205-ZZ" (product model);
[0135] Time information: "First quarter of 20XX" is interpreted as the time range "from 20XX-01-01 to 20XX-03-31";
[0136] Attribute requirements: Identify the attributes that users need to obtain: "Total purchase volume", "Contract number", and "Quantity per purchase".
[0137] The model structures the above information and outputs it as query conditions: {Intent:"Statistical Query", Supplier:"C Precision Components Factory", Product:"Bearing", Model:"6205-ZZ", Time Range:["20XX-01-01","20XX-03-31"], Return Fields:["Total Purchase Quantity","Contract Number","Quantity per Purchase"]}, which serve as the input for the data query sub-unit.
[0138] The server initiates collaborative retrieval between the graph database and the vector database based on structured query conditions and data query subunits.
[0139] Graph Database Retrieval: The domain-specific large language model transforms query conditions into Neo4j Cypher query statements, aiming to extract contract and product association data that meet the conditions from the knowledge graph. The generated Cypher statement is: MATCH(s:Supplier{name:C Precision Components Factory})<-[:Association]-(c:Contract)-[:Contains]->(p:Product{name:Bearing,specification:6205-ZZ})WHERE c.sign_date>=20XX-01-01ANDc.sign_date<=20XX-03-31RETURNc.contract_idAS Contract Number,c.quantityAS Purchase Quantity.
[0140] After the server executes the statement, Neo4j returns the following result set: [{Contract No.:HT20XX001, Purchase Quantity:300},{Contract No.:HT20XX015, Purchase Quantity:200}].
[0141] Vector Database Retrieval: To verify the accuracy of the purchase quantity, the server generates a 768-dimensional semantic vector from the user query "purchase quantity of bearing model 6205-ZZ" using an embedding model (paraphrase-multilingual-mpnet-base-v2), and then calls the Qdrant vector database retrieval interface. The retrieval criteria are: {vector: [generated query vector], filter criteria: {entity: bearing 6205-ZZ, time_range: 20XXQ1}, similarity threshold: 0.85, number of returned values: 2}. Qdrant returns two highly similar text fragments:
[0142] Segment 1 (from Attachment to Contract HT20XX001): "Purchase Item: Bearing 6205-ZZ, Quantity 300 pieces, Unit Price 150.50 RMB / piece, Delivery Batch: February 10, 20XX (200 pieces), March 5, 20XX (100 pieces)";
[0143] Section 2 (from Attachment to Contract HT20XX015): "Quantity of 200 6205-ZZ bearings to be purchased, technical requirements conforming to GB / T307.1-2017, acceptance criteria refer to Attachment 3."
[0144] The server invokes the answer generation subunit to fuse the structured results from the graph database with the text fragments from the vector database. First, the domain-specific large language model summarizes and calculates the purchase quantity returned by the graph database: 300 items + 200 items = 500 items, confirming the total purchase quantity. Then, the model integrates the text details from the vector database, supplementing each contract with delivery batch or technical standard information to ensure the accuracy and richness of the answer.
[0145] Finally, the model generates a natural language answer: "Two contracts were signed in the first quarter of 20XX for the purchase of 6205-ZZ model bearings from C Precision Components Factory, with a total purchase quantity of 500 pieces. The specific contract details are as follows: 1. Contract No. HT20XX001, purchase quantity 300 pieces (delivered in two batches: 200 pieces on February 10, 20XX, and 100 pieces on March 5, 20XX); 2. Contract No. HT20XX015, purchase quantity 200 pieces (technical standard conforms to GB / T307.1-2017)." The server then feeds this answer back to the user through the front-end interface, completing the intelligent question-and-answer interaction.
[0146] Throughout the process, the user response module 1103 achieves efficient response to complex business queries through precise intent parsing, dual-database collaborative retrieval, and deep result fusion, supporting rapid decision-making on enterprise procurement data.
[0147] In this embodiment of the invention, the quality monitoring and feedback module 1104 includes a model supervision subunit, a user feedback collection subunit, and a model fine-tuning subunit;
[0148] The model supervision subunit is configured to use the domain big language model to perform legality verification on the entities and relations extracted from the SQL database by the data parsing module 1101, and automatically mark the entities or relations that do not conform to the preset verification rules for manual review and confirmation.
[0149] The user feedback collection subunit is configured to collect user feedback on the system output results through the user interaction interface, including illegal entities or relationships marked by the user, and omitted entities or relationships.
[0150] The model fine-tuning subunit is configured to automatically include the erroneous labeled samples into the training dataset when the cumulative number of erroneous labeled samples from the model supervision subunit and the user feedback collection subunit exceeds a preset threshold, and to fine-tune the domain-specific large language model based on the training dataset to generate an optimized version.
[0151] In this embodiment of the invention, taking a manufacturing enterprise procurement data governance scenario as an example, the server forms a closed loop through the three sub-units of the quality monitoring and feedback module 1104 to continuously optimize the accuracy of data parsing and model performance. The following describes the execution process in detail with a case study of actual data processing from March to April 20XX.
[0152] While the server is processing the SQL database in data parsing module 1101, it simultaneously starts the model supervision subunit. This subunit uses a domain-specific large language model (a version fine-tuned for the procurement domain) to perform validity checks on the entities and relationships extracted by the structured data processing unit. The preset validation rule base includes general rules from the manufacturing domain.
[0153] Entity validity rules: The "type" attribute of the "supplier" entity is only allowed to have two values: "enterprise" and "individual business owner"; the "specification" attribute (specification model) of the "product" entity cannot be empty.
[0154] Relationship validity rules: In the "contract-include-product" relationship, the "product" entity must be associated with the "material" attribute; in the "contract-association-supplier" relationship, the "sign_date" must be earlier than the "delivery_date".
[0155] When the server parses a new record (id=108, name=D Hardware Store, type=Individual, contact=135XXXX1xx4) in the "suppliers" table of the SQL database, the structured data processing unit extracts it as a "supplier" entity (name: D Hardware Store, type: Individual). The model supervision subunit calls the domain large language model to validate this entity and finds that "type=Individual" does not conform to the preset category rule for the "supplier" entity (allowed values: enterprise / individual business). It automatically marks the entity as an "entity item that does not conform to the preset validation rule" and generates a validation report: {Entity type: Supplier, Entity name: D Hardware Store, Error attribute: type=Individual, Violation rule: Supplier type must be enterprise or individual business, Suggested correction value: individual business}. Subsequently, this marked item is pushed to the manual review interface of the enterprise data administrator.
[0156] During user interaction, if a user finds an error in the system output, they can submit feedback through the "Feedback" function on the front-end interface. For example, when a purchasing department employee queries "all bearing contracts signed with C Precision Components Factory in March 20XX", the system returns contract numbers HT20XX001 (300 pieces) and HT20XX015 (200 pieces), but the user is aware that a third contract, HT20XX022 (150 pieces), has been omitted. The user clicks the "Feedback on Omission Information" button below the answer, enters the following in the pop-up window: {Query Question: Bearing contracts with C Precision Components Factory in March 20XX, Omission Entity Type: Contract, Omission Entity Details: Contract Number HT20XX022, Purchase Quantity 150 pieces, Signing Date 20XX-03-28}, and submits feedback.
[0157] The server receives this information through the user feedback collection subunit, automatically classifies it as an "omitted entity" error sample, adds it to the error annotation sample library, and records the feedback time, user ID, and the original query log associated with the issue to ensure traceability. Furthermore, if a user finds that the system identifies an entity attribute error (e.g., "the material of bearing 6205-ZZ is marked as 45 steel, but should actually be GCr15"), they can mark the invalid attribute through the "attribute correction" function, and the server will also add it to the sample library.
[0158] The server continuously monitors the number of errors in the error-annotated sample library, with a preset threshold of 100. As of April 15, 20XX, the model supervision subunit had marked 65 entity relationship errors such as "incorrect supplier type" and "missing product specifications," and 42 user-reported errors such as "missing contract" and "incorrect attribute value," totaling 107, exceeding the preset threshold. The server automatically triggers the model fine-tuning subunit, initiating the iterative optimization process of the domain-wide language model.
[0159] First, the server divides the sample library into a training set (86 entries) and a validation set (21 entries) in an 8:2 ratio. The training set includes typical error types such as "incorrect supplier type" (e.g., "individual → sole proprietorship"), "missing relationship attributes" (e.g., "product material not identified"), and "omitted entities" (e.g., "contract HT20XX022 not associated"). Then, LoRA (Low-Rank Adaptation) technology is used to fine-tune the domain-specific language model: 90% of the model's pre-training parameters are frozen, and only low-rank matrix parameters related to entity recognition and relation extraction are optimized (e.g., output layer weights of the entity classifier and attention mechanism parameters for relation extraction). During fine-tuning, the server sets a learning rate of 2e-5 and five training epochs. After each epoch, the model performance (entity recognition accuracy and relation extraction F1 score) is evaluated using the validation set.
[0160] After fine-tuning, the new model's entity recognition accuracy on the validation set improved from 88% to 95%, and its relation extraction F1 score improved from 85% to 92%. It successfully corrected high-frequency errors such as "misidentification of supplier types" and "omission of contract entities." The server deployed the new model as the current active version and cleared the error-labeled sample library count, awaiting the next round of accumulation and optimization, forming a closed-loop quality improvement mechanism of "monitoring-feedback-fine-tuning."
[0161] Through the above process, the quality monitoring and feedback module 1104 continuously ensures the accuracy of data governance, enabling the domain-wide language model to gradually adapt to the actual business data characteristics of enterprises, and improving the overall reliability and intelligence level of the system.
[0162] In this embodiment of the invention, the domain-wide language model module 1105 includes a base model selection subunit, a model fine-tuning subunit, and a continuous optimization subunit;
[0163] The base model selection subunit is configured to select a large language model with a preset parameter scale and support for vertical domain knowledge transfer as the initial base model. The preset parameter scale is set according to the data complexity and knowledge reasoning requirements of the target domain.
[0164] The model fine-tuning subunit is configured to continuously fine-tune the initial base model based on the error-labeled samples and user feedback data collected by the quality monitoring and feedback module 1104, so as to enhance its entity recognition, relation extraction and question answering capabilities in a specific domain.
[0165] The continuous optimization subunit is configured to improve the accuracy, robustness, and knowledge understanding ability of the domain-specific large language model through multiple rounds of iterative fine-tuning and effect evaluation.
[0166] In an embodiment of the present invention, for example, the server takes the construction and optimization of a large language model in the procurement field of manufacturing enterprises as the scenario, and creates a vertical domain model that is adapted to the professional needs of enterprises through three sub-units: base model selection, continuous fine-tuning and iterative optimization.
[0167] The server initiates the initial model selection by selecting sub-units from the base model, first analyzing the core requirements of the manufacturing procurement field:
[0168] Data complexity: The domain contains 500+ technical terms (such as “tolerance grade IT7”, “surface roughness Ra1.6”, “radial clearance C3 group”), 20+ entity relationship types (such as “contract-purchase-product”, “product-compliance-national standard”, “supplier-related-quality problem”), and unstructured data (such as technical attachments and inspection reports) accounts for 60%, requiring the model to have strong terminology understanding capabilities.
[0169] Knowledge reasoning requirements: Frequent user queries involve multi-hop relationship reasoning (such as "querying the quality problems and corresponding repair records of CNC lathes purchased from supplier B in the past 3 years") and attribute association reasoning (such as "automatically matching applicable testing standards based on product specification 6205-ZZ"), requiring the model to support the parsing of complex logic chains.
[0170] Based on the above characteristics, the server was designed with a preset parameter scale standard: it must support more than 200 bytes of parameters to handle complex semantic understanding and reasoning capabilities, while also possessing domain knowledge transfer interfaces (such as supporting LoRA fine-tuning and providing a dedicated output layer for entity relation extraction). After adaptability testing on mainstream large language models (including parameter scale, transfer capability, and inference speed), the server ultimately selected a general-purpose large language model with xx5 bytes of parameters and support for vertical domain knowledge transfer as the initial base model. This model has mastered basic language understanding capabilities through pre-training, but its accuracy is insufficient in specific tasks such as manufacturing terminology recognition and procurement relation extraction (entity recognition accuracy 82%, relation extraction F1 score 78%), requiring further optimization.
[0171] The server calls the model fine-tuning subunit, using the accumulated error-labeled samples and user feedback data from the quality monitoring and feedback module 1104 as the core, to perform specific fine-tuning on the initial base model.
[0172] Dataset Construction: The server selects typical error types from the error sample library and expands it into a fine-tuned dataset (5000 samples in total) by combining it with domain data.
[0173] Entity recognition error samples (1500): such as "supplier type misjudgment" (identifying "individual business owner" as "individual"), "product model confusion" (identifying "6205-ZZ" as "6204-ZZ");
[0174] Errors in relation extraction (2000 records): such as "Contract-Product Association Omission" (failed to identify the inclusion relationship between contract HT20XX022 and product 6205-ZZ), "Attribute Relationship Error" (associating "Delivery Date" with the "Contract Signing Date" field);
[0175] User feedback Q&A sample (1500 entries): such as "Answer omitted delivery batch" (the system only returned the total purchase quantity and did not mention the batch delivery information) and "Terminology interpretation error" (interpreting "radial clearance" as "axial clearance").
[0176] Fine-tuning execution: The server employs LoRA (Low-Rank Adaptation) lightweight fine-tuning technology, freezing 90% of the pre-trained parameters of the base model and updating only the low-rank matrices of the entity recognition head, relation extraction attention layer, and question-answering generation decoder. Specific configuration: learning rate 3e-5, 8 training epochs, batch size 16, with performance evaluated using a validation set (1000 domain-labeled data points) in each epoch. During fine-tuning, the model corrects the parameter weights corresponding to erroneous samples through gradient descent—for example, for "supplier type misjudgment" samples, the model adjusts the boundary weights between the "individual business owner" and "individual" categories in the entity classifier; for "relationship omission" samples, it enhances the attention weights for "contract-inclusion-product" relationship keywords (such as "purchase target" and "quantity").
[0177] The server establishes a long-term tracking and iteration mechanism for model performance by continuously optimizing sub-units, ensuring that the model adapts to dynamic changes in the domain.
[0178] Regular evaluation: The server performs monthly model performance evaluations, with key metrics including: entity recognition accuracy (target ≥95%), relation extraction F1 score (target ≥92%), and user question-and-answer satisfaction (target ≥4.5 / 5). The March 20XX evaluation showed that after the first round of fine-tuning, the model's entity recognition accuracy improved to 91%, but the F1 score for "product specification - standard matching" relation extraction was only 88% (below the target), and 12% of user feedback indicated "vague explanations of technical standards."
[0179] Incremental Fine-tuning: The server triggers the incremental fine-tuning process, adding training data including: 300 newly accumulated error samples in March 20XX (mainly "standard matching errors"), and the company's newly released "20XX Edition Procurement Technical Standards Compilation" (containing 200+ new terms and standard correspondences). Fine-tuning focuses on the "product-compliance-standard" relationship extraction layer, enhancing the model's learning of the "specification model → standard number" mapping by introducing standard clause text (e.g., "6205-ZZ → GB / T307.1-2017" "CK6150 → GB / T17421.1-2016").
[0180] Results solidified: An evaluation in April 20XX showed that after incremental fine-tuning, the F1 score for model relation extraction improved to 94%, user satisfaction reached 4.7 points, and the "standard matching error" issue was successfully fixed. The server deployed this version of the model as an active model in the production environment and saved a snapshot of the fine-tuned parameters as the baseline for the next iteration. Simultaneously, the continuous optimization sub-unit automatically recorded the performance change curves of each version of the model, providing a basis for decision-making regarding subsequent parameter scaling adjustments (such as upgrading to a 300B parameter model after the domain data volume increases).
[0181] Through the above process, the domain-specific large language model module 1105 gradually evolves from a general base model to a professional model adapted to the manufacturing and procurement field. Its entity recognition, relationship extraction, and question-answer generation capabilities continue to closely align with the actual business needs of enterprises, providing core AI support for the data governance intelligent agent system 110.
[0182] In this embodiment of the invention, the legality verification of the data parsing results by the model supervision subunit includes entity legality verification and relation legality verification;
[0183] The entity legality verification includes determining whether the entity belongs to a preset domain entity category;
[0184] The relationship validity verification includes determining whether the relationship between two entities conforms to preset domain relationship rules.
[0185] In this embodiment of the invention, for example, in the manufacturing enterprise procurement data parsing process, the server performs legality verification on the entities and relationships output by the data parsing module 1101 through the model supervision subunit to ensure the accuracy of the knowledge graph. The following describes in detail the execution process of entity legality verification and relationship legality verification, using a scenario from April 20XX involving the verification of newly accessed supplier and contract data.
[0186] The server first loads a pre-defined domain entity category rule library, which defines core entity categories and attribute constraints for the manufacturing and procurement domain, including:
[0187] Suppliers: Allowed categories are "Enterprise" and "Individual Business". They must include "name" and "type" attributes, and the "type" value can only be one of the two categories mentioned above.
[0188] Products: Allowed categories are "Raw Materials", "Components", and "Equipment". They must include the attributes "name" and "specification", and the "specification" attribute cannot be empty.
[0189] Contracts: Allowed categories are "Purchase Contract" and "Sales Contract". They must include the attributes "contract_id" (number) and "sign_date" (signing date), and the date format must conform to "YYYY-MM-DD".
[0190] When the data parsing module 1101 extracts a new supplier entity from the "suppliers" table in the SQL database (record: id=112, name=E Parts Store, type=Individual, contact=139XXXX7890), the server initiates the entity validity verification process. The domain-wide language model first identifies the candidate category of the entity as "supplier," and then calls the rule base to verify its attributes:
[0191] Check whether the "type" attribute value "individual" belongs to the "supplier" category, which includes "enterprise" and "sole business owner". The result is "no".
[0192] The required attribute "name" was checked and found to be "yes" (name = E Parts Store), but the core attribute "type" was found to be invalid.
[0193] The server automatically marks the entity as an anomaly that "does not conform to the preset domain entity category" and generates a verification report: {Entity ID: supplier_112, Entity Name: E Parts Store, Candidate Category: Supplier, Violation Attribute: type = Individual, Allowed Value: Enterprise / Individual Business}. The marked item is then pushed to the data administrator's manual review interface, waiting for manual confirmation of the correction direction (such as changing "Individual" to "Individual Business").
[0194] The server simultaneously loads a pre-defined domain relationship rule base, defining the allowed relationship types and constraints between entities. The core rules include:
[0195] Contract-Related-Supplier: The relationship type must be "Related", and the contract's "sign_date" (signing date) must be earlier than the "delivery_date" (delivery date);
[0196] Contract-Contains-Product: The relationship type must be "Contains" and must be associated with "quantity" and "unit_price" attributes, and the values must be positive.
[0197] Product-Compliant-Standard: The relationship type must be "Compliant", and the "Standard Number" must be in the current list of valid standards of the National Standardization Management Committee (e.g., "GB / T307.1-2017" is valid, while "GB / T307.1-2005" has been abolished).
[0198] When the data parsing module 1101 extracts the relationship from the associated data of the "contracts" and "products" tables (relationship triple: <Contract HT20XX030, Compliant, Bearing 6205-ZZ>, associated attribute: Standard Number = GB / T307.1-2005), the server initiates the relationship validity verification process. The domain-wide language model first identifies the subject entity "Contract HT20XX030", the relationship type "Compliant", and the object entity "Bearing 6205-ZZ", and then calls the rule base for verification.
[0199] Check whether the relationship type "compliant" is a permitted relationship type between the "contract" and "product" entities. In the rule base, "contract-product" only allows "inclusion" relationships. The "compliant" relationship should exist between the "product-standard" entities. The relationship type is deemed to be in violation.
[0200] Check if the associated attribute "Standard Number = GB / T307.1-2005" is valid. By calling the interface of the Standardization Administration of China, it was confirmed that the standard has been abolished (the current valid version is GB / T307.1-2017), and the attribute value was determined to be in violation.
[0201] The server automatically marks this relationship as an anomaly that "does not conform to the preset domain relationship rules," generating a verification report: {Relationship ID: rel_20XX030_01, Subject Entity: Contract HT20XX030, Relationship Type: Conforms, Object Entity: Bearing 6205-ZZ, Reason for Violation: 1. A conforming relationship is not allowed between Contract and Product (allowed relationship: containment); 2. Standard number GB / T307.1-2005 has been abolished (current version: GB / T307.1-2017)}. This flag is then pushed to the manual review interface for the administrator to confirm whether the relationship type (e.g., change to "Contract-Containment-Product") and standard number need to be corrected.
[0202] Through the aforementioned entity and relationship legitimacy checks, the model supervision subunit can accurately identify category errors and relationship violations during the data parsing process, providing clear guidance for manual review and ensuring that the entities and relationships entering the knowledge graph conform to the business logic of the manufacturing and procurement field, thus laying an accurate data foundation for subsequent user queries and decision support.
[0203] In this embodiment of the invention, the quality monitoring and feedback module 1104 further includes a manual review subunit;
[0204] The manual review subunit is configured to receive entities or relationships marked by the model supervision subunit that do not conform to the preset verification rules, and send the erroneous entities or relationships confirmed by manual review to the user feedback collection subunit for aggregation.
[0205] In an embodiment of the present invention, for example, in the manufacturing enterprise's procurement data governance process, the server connects model supervision and user feedback collection through a manual review sub-unit, and submits abnormal entities or relationship items marked by the model to manual confirmation to ensure the accuracy of erroneous samples.
[0206] After the model supervision subunit completes the legality verification of the newly parsed data, the server automatically aggregates the marked "entities or relationships that do not conform to the preset verification rules" into an audit task. On April 10, 20XX, the model supervision subunit marked 3 supplier records and 2 contract-product relationship records newly added to the SQL database as abnormal items. The server generated a task package containing 5 items to be audited through the manual audit subunit. Each item to be audited includes: a unique audit ID, entity / relationship type, original parsed data, reason for violation, suggested correction value from the model, and the source of related data (such as SQL table name and row number).
[0207] The server pushes the review task to the data administrator's "Data Governance Review Workbench" interface via the enterprise's internal office system API. After the administrator logs in, the system automatically loads the list of items to be reviewed. The first item to be reviewed is an entity exception: {Review ID: check_20XX0410_001, Type: Entity, Entity Name: E Parts Store, Entity Category: Supplier, Original Attributes: {type: Individual, name: E Parts Store, contact: 139XXXX7890}, Reason for Violation: The entity type "Individual" does not belong to the preset domain entity category (allowed: Enterprise / Individual Business), Suggested Correction Value: type = Individual Business, Data Source: suppliers table id = row 112}; The second item indicates a relationship anomaly: {Audit ID: check_20XX0410_002, Type: Relationship, Relationship Triplet: <Contract HT20XX030, Contains, Bearing 6205-ZZ>, Association Attribute: {quantity: -50, unit_price: 150.50}, Reason for violation: Relationship attribute quantity = -50 does not conform to the preset rule (quantity must be positive), Suggested correction value: quantity = 50, Data source: Association between row id = 205 of the contracts table and row id = 503 of the products table}.
[0208] The data administrator reviewed the first entity anomaly, "E Parts Store": Using the "View Original Data" button on the interface, they retrieved the complete record of row 112 in the suppliers table of the SQL database, confirming that the "type" field was indeed entered as "Individual"; then, they queried the supplier's business registration information through the enterprise CRM system, showing its type as "Individual Business Owner" (with the unified social credit code ending in "MAXXX", conforming to the coding rules for individual businesses). The administrator determined that the model's suggested correction value "type = Individual Business Owner" was correct, and clicked the "Confirm Error and Adopt Suggestion" button on the review interface. The system automatically updated the corrected attribute value to "type = Individual Business Owner" and recorded the review opinion: "After verification of business information, E Parts Store is indeed an individual business owner; the original entry was incorrect; the correction suggestion is adopted."
[0209] Regarding the second anomaly, "Contract HT20XX030-Includes-Bearing 6205-ZZ", the administrator reviewed the original scanned copy of the contract (accessed via the "Related Documents" link in the interface) and found that the contract text clearly stated "Purchase Quantity: 50 pieces". The value "-50" in the quantity field of the SQL database was an incorrect negative sign added during data entry. The administrator manually modified the suggested corrected value to "quantity=50" (consistent with the model suggestion) in the review interface and clicked the "Confirm Error and Correct" button. The system recorded the review comment: "The original contract quantity is 50 pieces; database entry error; correct quantity to 50."
[0210] For the other three items pending review (such as "Supplier F Trading Company is labeled as an enterprise but the business registration information shows that it has been deregistered" and "Contract HT20XX031-Includes-Gear M10 relationship omitting unit_price attribute"), the administrator completed the review by reviewing the original data, related documents, or external system information. Among them, two adopted the model's suggested correction values, and for one item, due to an error in the model's suggestion (such as suggesting "deregistered enterprise" to "individual business household" instead of actually being marked as "invalid supplier"), the administrator manually modified the correction direction and submitted it.
[0211] After receiving the review results submitted by the administrator, the server's manual review subunit verifies the legality of the review results (such as confirming whether the format of the correction value meets the requirements and whether the review comments are complete). After passing the verification, the entities or relations that are "confirmed as incorrect" are converted into standardized error samples.
[0212] Taking the "E Parts Store" entity exception as an example, the server generates the following error sample: {Sample ID: error_20XX0410_001, Error Type: Entity Attribute Error, Entity Type: Supplier, Original Data: {name: E Parts Store, type: Individual}, Error Field: type, Error Value: Individual, Correct Value: Sole Proprietor, Error Reason: Database Entry Error, Reviewer: admin_01, Review Time: 20XX-04-10 14:30:22}; Taking "Contract HT20" as an example... Taking the abnormal relationship "XX030-containing-bearing6205-ZZ" as an example, an error sample is generated: {Sample ID: error_20XX0410_002, Error type: Relationship attribute error, Relationship triple: <Contract HT20XX030, Contains, Bearing6205-ZZ>, Error attribute: quantity, Error value: -50, Correct value: 50, Error reason: Database entry error, Reviewer: admin_01, Review time: 20XX-04-10 14:35:10}.
[0213] Subsequently, the manual review subunit sends the aforementioned error samples to the user feedback collection subunit through an internal interface. The user feedback collection subunit merges these samples with feedback samples collected during user interaction (such as "omitted contract HT20XX022") and stores them in the error-labeled sample library. This provides standardized training data for the subsequent model fine-tuning subunit to accumulate error samples and initiate model iteration optimization.
[0214] By manually reviewing the connection of sub-units, the server ensures that the abnormal items marked by the model supervision sub-units are manually confirmed, avoiding the inclusion of model misjudgments into the error samples. At the same time, by standardizing the error sample format, it provides reliable data input for the closed-loop optimization of the quality monitoring and feedback module 1104.
[0215] To more clearly describe the solutions provided in the embodiments of the present invention, a more complete implementation method is provided below.
[0216] This invention proposes a data governance intelligent agent system 110 based on a large language model, multimodal model, and agent architecture. The system aims to achieve unified parsing, knowledge construction, and intelligent question answering of structured and unstructured data, thereby improving the intelligence and automation capabilities of data governance. The system includes the following five core modules:
[0217] 1. Data parsing module 1101.
[0218] This module is responsible for uniformly parsing and structurally modeling the structured and unstructured data within the enterprise, providing basic support for subsequent data governance and intelligent question answering.
[0219] 1.1 Unstructured Data Processing
[0220] Text extraction:
[0221] Scanned document processing: The scanned document is parsed using a multimodal large model (InternVL3-78B) to extract the text content;
[0222] General document processing: Combining engineering techniques to extract structured text content from non-scanned documents.
[0223] Knowledge graph construction:
[0224] Entity extraction: Using a domain-specific large language model, entities, entity attributes, and relationships between them are identified from the text extracted in the above steps;
[0225] Entity storage: The extracted entities and their relationships are stored in a graph database to form a preliminary knowledge graph structure.
[0226] Knowledge vectorization:
[0227] Text segmentation: Semantic segmentation of extracted text using engineering techniques;
[0228] Vector generation: Semantic vectors are generated from text segments using an embedding model (paraphrase-multilingual-mpnet-base-v2) and stored in a vector database for subsequent semantic retrieval and question answering.
[0229] 1.2 Structured Data Processing (SQL Database)
[0230] Data structure ontology construction: Based on the database schema information, the table structure, field types and relationships between fields are automatically identified to construct the data structure ontology and define entity categories, attributes and relationships;
[0231] Entity Relationship Extraction: Data fields are abstracted and categorized using engineering methods, and organized into entities and their relationships based on the constructed ontology, further enriching the knowledge graph content.
[0232] 2. Data storage module 1102.
[0233] This module is responsible for storing structured knowledge data and supports efficient and flexible queries.
[0234] Graph data storage: Entities and relationships are stored in a graph database in a graph structure, facilitating subsequent complex relationship retrieval and reasoning. Neo4j is the chosen graph database.
[0235] Vector data storage: The semantic vectors of text segments are stored in a vector database to support similarity calculation and semantic retrieval. Qdrant is selected as the vector database.
[0236] 3. User Response Module 1103.
[0237] This module allows users to interact with the system using natural language, enabling intelligent question answering of knowledge data.
[0238] User intent parsing: Based on a domain-specific large language model, user query intent is parsed to identify key entities, events, times, locations, and categories from the raw input;
[0239] Data Query:
[0240] Graph data query: Based on the intent recognition results, the large language model constructs the corresponding graph database query statement to retrieve relevant entities and their relationships;
[0241] Vector data query: Convert user queries into semantic vectors and retrieve relevant document fragments or knowledge pieces from the vector database.
[0242] Answer generation: Combining the search results from graph databases and vector databases, and leveraging a domain-specific large language model, structured or natural language answers are generated and provided to users.
[0243] 4. Quality monitoring and feedback module 1104.
[0244] This module ensures continuous optimization of data governance quality and drives system capability evolution through a closed-loop feedback mechanism.
[0245] Model supervision: Utilize a large language model to perform legality checks on entities and relations extracted from the SQL database, automatically marking items that are "potentially invalid" for manual review and confirmation;
[0246] User feedback collection: Collecting users' marked invalid entities / relationships, as well as omitted entities / relationships, for subsequent optimization;
[0247] Model fine-tuning: When the cumulative number of mislabeled samples exceeds a threshold (e.g., 100), they are automatically included in the training dataset to fine-tune the large language model and generate an optimized version.
[0248] 5. Domain-wide Large Language Model Module 1105.
[0249] This module focuses on building and optimizing large language models for vertical domains to adapt to the data governance needs of specific industries or scenarios.
[0250] Base model selection: Qwen3 xx5B was selected as the initial base for the large language model;
[0251] Model fine-tuning: Based on the data collected by the quality monitoring and feedback module 1104, the large language model is continuously fine-tuned to enhance its ability to recognize entities, extract relations, and generate questions in specific domains.
[0252] Continuous optimization: Through multiple rounds of iterative optimization, the model's accuracy, robustness, and knowledge understanding capabilities are continuously improved, enabling high-quality extraction and application of structured knowledge.
[0253] In summary, this invention addresses the common challenges faced by enterprises in data governance, including difficulties in integrating structured and unstructured data, high reliance on professional personnel, high governance costs and cycles, and difficulty in ensuring data quality. It proposes an intelligent data governance solution based on a large language model, intelligent agents, and a multimodal large model, offering the following significant advantages: First, by using a multimodal large model to extract content from unstructured data such as scanned documents, and combining this with the deep understanding of natural language text by the large language model, it achieves structured expression of unstructured information such as text and images, and maps it uniformly with structured database information to construct a complete knowledge graph. This effectively eliminates data silos and information fragmentation, breaking down the barriers between structured and unstructured data. Second, by leveraging the autonomous execution capabilities of intelligent agents in data parsing, entity extraction, vectorization, and query construction, it replaces traditional manual processes, significantly reducing reliance on professional personnel such as data engineers and database experts, thereby significantly lowering operational barriers and labor costs. Simultaneously… The system, through the collaboration of intelligent agents and a large language model, automates the entire process from data collection to knowledge generation, avoiding repetitive manual operations and significantly shortening the governance project cycle. Furthermore, the continuous learning and feedback optimization mechanism of the model reduces long-term maintenance and upgrade costs. In addition, by leveraging model supervision mechanisms and a closed-loop user feedback system, inconsistencies and illegal content in the data are automatically labeled and verified, continuously optimizing extraction accuracy and entity consistency, improving data quality and consistency assurance capabilities, and achieving the construction of high-quality, highly reliable data assets. Based on real business feedback data, the large language model is continuously fine-tuned and iterated to build a domain-specific language model for the enterprise, enabling the system to perform better in specific business contexts. This self-evolving domain model system supports flexible use and long-term evolution across different business departments. Finally, the system adopts a decoupled modular design, allowing each sub-module to be deployed independently and optimized in combination, facilitating implementation and future expansion, and possessing good technical portability and project delivery flexibility. In summary, this invention, driven by AI intelligent agents and integrating multimodal and large language model capabilities, significantly improves the intelligence and automation level of data governance, providing a feasible technical solution for enterprises to build a low-cost, high-efficiency, and high-quality data governance system.
[0254] Please refer to the following: Figure 2 , Figure 2This is a flowchart illustrating the steps of a data governance intelligent agent implementation method provided in an embodiment of the present invention, applied to the aforementioned data governance intelligent agent system 110, including:
[0255] Step S201: Through the data parsing module, the structured and unstructured data within the enterprise are uniformly parsed and structured modeled to obtain structured knowledge data;
[0256] Step S202: The structured knowledge data processed by the data parsing and modeling steps is stored in the graph database and vector database of the data storage module;
[0257] Step S203: Receive the user's natural language query through the user response module, parse the user's intent, retrieve relevant knowledge data from the graph database and vector database, and generate a structured or natural language answer to be fed back to the user.
[0258] Step S204: The data parsing results are validated through the quality monitoring and feedback module. Error entities, relationships and omissions reported by users are collected. Error samples are accumulated and the domain language model is fine-tuned when they exceed a preset threshold.
[0259] Step S205: Select an initial base model through the domain-specific large language model module, and continuously iterate and optimize the model based on the fine-tuning data from the quality monitoring and feedback optimization steps to adapt to the data governance needs of specific industries or scenarios.
[0260] It should be noted that the implementation principle of the aforementioned data governance intelligent agent implementation method can refer to the implementation principle of the aforementioned data governance intelligent agent system 110, and will not be repeated here.
[0261] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned data governance intelligent agent system 110. For example... Figure 3 As shown, Figure 3 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a data governance intelligent agent system 110, a memory 111, a processor 112, and a communication unit 113.
[0262] To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The data governance intelligent agent system 110 includes at least one software function module that can be stored in the memory 111 or embedded in the operating system (OS) of the computer device 100 in the form of software or firmware. The processor 112 is used to execute the data governance intelligent agent system 110 stored in the memory 111, such as the software function modules and computer programs included in the data governance intelligent agent system 110.
[0263] This invention provides a readable storage medium, which includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the aforementioned data governance intelligent agent system 110.
[0264] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.
Claims
1. A data governance intelligent agent system, characterized in that, include: The system includes a data parsing module, a data storage module, a user response module, a quality monitoring and feedback module, and a domain-specific large language model module; among them, The data parsing module is configured to perform unified parsing and structured modeling of structured and unstructured data within the enterprise. The data storage module is configured to store structured knowledge data processed by the data parsing module, including a graph database and a vector database; The user response module is configured to support users to interact with the system through natural language in order to achieve intelligent question answering of the structured knowledge data. The quality monitoring and feedback module is configured to optimize data governance quality through a closed-loop feedback mechanism, including monitoring data parsing results, collecting user feedback, and iteratively optimizing the model. The domain-specific large language model module is configured to build and optimize large language models for vertical domains to adapt to the data parsing, user interaction, and quality optimization needs of specific industries or scenarios.
2. The data governance intelligent agent system according to claim 1, characterized in that, The data parsing module includes an unstructured data processing unit, which is configured as follows: Text extraction, knowledge graph construction, and knowledge vectorization processing of unstructured data within enterprises; The text extraction involves using a multimodal large model to parse the scanned document to extract the text content, and combining engineering techniques to extract the structured text content from the non-scanned document. The knowledge graph is constructed by using a domain-specific large language model to identify entities, entity attributes, and relationships between entities from the extracted text, and storing the extracted entities and their relationships in the graph database to form a preliminary knowledge graph. The knowledge vectorization involves using engineering methods to perform semantic segmentation on the extracted text, and then using an embedding model to generate semantic vectors from the text segments, which are then stored in the vector database.
3. The data governance intelligent agent system according to claim 2, characterized in that, The data parsing module further includes a structured data processing unit, which is configured as follows: For internal SQL databases, the system automatically identifies table structure, field types, and relationships between fields based on the database schema information, and constructs a data structure ontology to define entity categories, attributes, and relationships. The data fields are abstracted and categorized using engineering methods, and the categorization results are organized into entities and their relationships based on the data structure ontology, thereby enriching the content of the knowledge graph.
4. The data governance intelligent agent system according to claim 1, characterized in that, In the data storage module: the graph database is Neo4j, which is configured to store entities and relationships in a graph structure and supports complex relationship retrieval and reasoning; the vector database is Qdrant, which is configured to store semantic vectors of text fragments and supports semantic retrieval based on similarity calculation.
5. The data governance intelligent agent system according to claim 1, characterized in that, The user response module includes a user intent parsing subunit, a data query subunit, and an answer generation subunit; The user intent parsing subunit is configured to parse the user's natural language query intent based on the domain big language model, and identify key entities, events, time, location and category information from the original query input; The data query subunit is configured to construct a graph database query statement based on the recognition result of the user intent parsing subunit to retrieve relevant entities and their relationships in the graph database using the domain big language model. Simultaneously, user queries are converted into semantic vectors, and relevant document fragments or knowledge pieces are retrieved from the vector database. The answer generation subunit is configured to combine the retrieval results obtained from the graph database and vector database by the data query subunit, and generate a structured or natural language answer using the domain-specific large language model.
6. The data governance intelligent agent system according to claim 1, characterized in that, The quality monitoring and feedback module includes a model supervision subunit, a user feedback collection subunit, and a model fine-tuning subunit. The model supervision subunit is configured to use the domain big language model to perform legality verification on the entities and relations extracted from the SQL database by the data parsing module, and automatically mark the entities or relations that do not conform to the preset verification rules for manual review and confirmation. The user feedback collection subunit is configured to collect user feedback on the system output results through the user interaction interface, including illegal entities or relationships marked by the user, and omitted entities or relationships. The model fine-tuning subunit is configured to automatically include the erroneous labeled samples into the training dataset when the cumulative number of erroneous labeled samples from the model supervision subunit and the user feedback collection subunit exceeds a preset threshold, and to fine-tune the domain-specific large language model based on the training dataset to generate an optimized version.
7. The data governance intelligent agent system according to claim 1, characterized in that, The domain-specific large language model module includes a base model selection subunit, a model fine-tuning subunit, and a continuous optimization subunit. The base model selection subunit is configured to select a large language model with a preset parameter scale and support for vertical domain knowledge transfer as the initial base model. The preset parameter scale is set according to the data complexity and knowledge reasoning requirements of the target domain. The model fine-tuning subunit is configured to continuously fine-tune the initial base model based on the error-labeled samples and user feedback data collected by the quality monitoring and feedback module, so as to enhance its entity recognition, relation extraction and question answering capabilities in a specific domain. The continuous optimization subunit is configured to improve the accuracy, robustness, and knowledge understanding ability of the domain-specific large language model through multiple rounds of iterative fine-tuning and effect evaluation.
8. The data governance intelligent agent system according to claim 6, characterized in that, The model supervision subunit performs validity checks on the data parsing results, including entity validity checks and relation validity checks. The entity legality verification includes determining whether the entity belongs to a preset domain entity category; The relationship validity verification includes determining whether the relationship between two entities conforms to preset domain relationship rules.
9. The data governance intelligent agent system according to claim 8, characterized in that, The quality monitoring and feedback module also includes a manual review subunit; The manual review subunit is configured to receive entities or relationships marked by the model supervision subunit that do not conform to the preset verification rules, and send the erroneous entities or relationships confirmed by manual review to the user feedback collection subunit for aggregation.
10. A method for implementing a data governance intelligent agent, characterized in that, The data governance intelligent agent system applied to any one of claims 1-9 includes: The data parsing module performs unified parsing and structured modeling of structured and unstructured data within the enterprise to obtain structured knowledge data. The structured knowledge data processed through data parsing and modeling steps is stored in the graph database and vector database of the data storage module; The user response module receives natural language queries from users, parses the user's intent, retrieves relevant knowledge data from graph databases and vector databases, and generates structured or natural language answers to feed back to the user. The quality monitoring and feedback module verifies the legality of the data parsing results, collects erroneous entities, relationships and omissions reported by users, accumulates error samples, and fine-tunes the domain-wide language model when the error exceeds a preset threshold. The initial base model is selected through the domain-specific large language model module. Based on the fine-tuning data from the quality monitoring and feedback optimization steps, the model is continuously iterated and optimized to adapt to the data governance needs of specific industries or scenarios.
Citation Information
Patent Citations
Large model-based medical text information governance method and system
CN118114718A
Knowledge question-answering system based on large language model
CN119396975A
Method and system for realizing intelligent number asking based on large model
CN119415538A
Multi-source hybrid question answering method and system thereof
KR101662450B1
Method and system for context-aware telecommunications, cellular, and radio based generative pre-trained transformer
US20250139140A1
Cited By
Construction project data management method and system based on multiple modes and large model
CN121210474A
Multi-agent typesetting generation system based on regular formal injection
CN121279294A
Person-social vertical domain large model optimization training method and system
CN121279465A
Intelligent interaction system based on large language model and knowledge graph
CN121597842A