Document-level Knowledge Extraction and Fusion Method and System Based on Large Language Model

By applying the document-level knowledge extraction and fusion method of large language models in the robot field, the problem of knowledge extraction in document-level unstructured data in the robot field is solved, efficient and accurate knowledge extraction and fusion is achieved, and data quality and system performance are improved.

CN119358546BActive Publication Date: 2025-06-17江淮前沿技术协同创新中心
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411561355.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-06-17
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

In the field of robotics, it is difficult for the prior art to extract and integrate knowledge from document-level unstructured data efficiently and accurately, resulting in a low quality of information extraction.

Method used

The document-level knowledge extraction and fusion method based on a large language model is adopted. By determining the scope of key information and establishing a keyword dictionary, the document is divided into paragraphs, and the knowledge extraction and fusion is extracted and fusion using the producer-consumer model integrated asynchronous architecture.

Benefits of technology

It realizes efficient and accurate knowledge extraction and fusion, improves data quality and extraction accuracy, reduces the difficulty of large models to process complex documents, and improves the system's concurrency processing capabilities and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119358546B_ABST
    Figure CN119358546B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for document-level knowledge extraction and fusion based on large language models, belonging to the field of industrial robots, including: determining the range of required key information and establishing a keyword dictionary; dividing the unstructured data at the document level into paragraphs according to the keyword dictionary to obtain the divided sub-documents; using the producer-consumer mode to integrate the asynchronous architecture of the large model to build a software system, and using the software system to sequentially perform knowledge extraction tasks on the divided sub-documents to extract key information from the unstructured data of the sub-documents; integrating and classifying all the key information extracted from the same sub-document to obtain regularized data, and then performing knowledge fusion processing on the regularized data; the degree of association between paragraphs is combined with the keyword dictionary to divide the document. After division, the content of the sub-documents is highly aggregated, reducing the difficulty of the large model in processing complex documents. The producer-consumer mode is integrated into the large model to avoid system blocking and improve the concurrent processing ability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial robots, and specifically relates to the knowledge extraction and integration of relevant documents in the field of robots. Background Art

[0002] In the field of robots, a detailed robot design scheme usually includes contents from parts to components and even electrical design and control design. These schemes cover a large amount of information such as drawings, BOM tables, and design specifications. However, due to the large amount of information and inefficient retrieval methods, it is difficult to identify and utilize general or reusable components, increasing labor costs. A knowledge graph is a good solution. It organizes a large number of design schemes in a graph structure and establishes the association relationships between the schemes, which can effectively solve the above-mentioned problems.

[0003] However, how to mine the association relationships between the schemes from the large amount of unstructured data in the design schemes has become the primary problem to be solved. Traditional natural language processing algorithms often rely on a large amount of training data, and at the same time, they have poor processing effects on document-level unstructured data and are difficult to efficiently and accurately mine the required content. This is a common problem in the development of special domain knowledge graph systems.

[0004] In the prior art, the Chinese patent application for invention "A Method for Extracting Knowledge from Unstructured Text Data Based on a Large Language Model" with the publication number CN118036734A includes: Step S1, performing named entity recognition on the obtained original document to obtain named entity candidate items; Step S2, classifying the named entity candidate items to determine the corresponding specific domain named entities; Step S3, performing permutation and combination processing on the specific domain named entities to obtain multiple named entity pair candidate items; Step S4, performing subject complementation processing on the original text; Step S5, based on the relationship reasoning of the named entity pair candidate items, using the text after subject complementation, determining the corresponding entity relationship triple candidate items. In Step S1, the original document is segmented, and based on each segment, in response to the given first prompt information, the entities to be confirmed are extracted from the current document. It can be seen that the division of the document-level unstructured data in this patent application is a simple division based on keywords. When extracting information from the text in the field of robots, due to the existence of a large number of special terms and complex sentence patterns in the field of robots, this method is difficult to accurately understand and extract key information and does not have strong domain understanding ability, and the knowledge extraction quality in the field of robots is low. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to achieve efficient and accurate knowledge extraction and integration in the document-level unstructured data in the field of robots.

[0006] The present invention solves the above technical problems through the following technical solutions: A method for document-level knowledge extraction and fusion based on a large language model, comprising the following steps:

[0007] Step 1: Determine the range of required key information and establish a keyword dictionary;

[0008] Step 2: Divide the unstructured data at the document level into paragraphs according to the keyword dictionary to obtain the divided sub-documents;

[0009] Step 3: Use the producer-consumer mode to integrate the asynchronous architecture of the large model to build a software system, and use the software system to sequentially perform knowledge extraction tasks on the divided sub-documents to extract key information from the unstructured data of the sub-documents;

[0010] Step 4: Integrate and classify all the key information extracted from the same sub-document to obtain regular data, and then perform knowledge fusion processing on the regular data.

[0011] Beneficial effects: The present invention takes paragraphs as the smallest granularity, divides the document according to the degree of association between paragraphs in cooperation with the keyword dictionary, integrates paragraphs with higher association degrees, that is, paragraphs with strong context relationships together, realizes the high aggregation of the content of the divided sub-documents, and the highly aggregated similar knowledge points can significantly improve the data quality. It does not only rely on simple keyword matching, but retains the structural relationship between paragraphs, does not destroy the context connection of the document, can ensure the coherence of information and clear structure of the divided sub-documents, and the fine-grained processing method can ensure the high precision of subsequent classification and knowledge extraction, thereby reducing the difficulty of the large model in processing complex documents. By integrating the producer-consumer mode in the large model, the decoupled design of the producer thread and the consumer thread can decompose the document processing task into independent knowledge extraction tasks, effectively avoid system blocking, and improve the system's concurrent processing ability and system performance.

[0012] Preferably, the step 1 includes:

[0013] 1.1. Collect key parameters and terms from product manuals, technical specifications, and industry standards to obtain keywords;

[0014] 1.2. Define synonyms and related terms for each keyword;

[0015] 1.3. Dynamically update and maintain the keyword dictionary;

[0016] 1.4. Perform configuration management on the keyword dictionary.

[0017] Beneficial effects: By defining synonyms and related terms for each keyword, classifying and summarizing paragraphs according to the keywords and synonyms, and achieving the aggregation of similar knowledge points, it not only simplifies the processing tasks of the large model but also improves the data quality and the accuracy of knowledge extraction. The configuration management of keywords and their synonyms enables the system to have flexible dynamic update capabilities, enabling the system to adapt to the continuously changing and expanding knowledge base, ensuring the accuracy and comprehensiveness of keyword matching.

[0018] Preferably, the second step includes:

[0019] 2.1. Load the unstructured data at the document level into the system;

[0020] 2.2. Use a document parsing tool to split the unstructured data at the document level so that each paragraph or sentence of the document becomes an independent processing unit;

[0021] 2.3. The system scans the text in each paragraph or sentence to find terms or phrases that match the entries in the keyword dictionary;

[0022] 2.4. Calculate the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs;

[0023] 2.5. Divide the document into several sub-documents according to the correlation degree between paragraphs;

[0024] 2.6. Store the divided sub-documents and generate a structured output.

[0025] Beneficial effects: The correlation degree not only depends on simple keyword matching but also conducts in-depth semantic understanding through context analysis. The system comprehensively calculates the correlation degree based on the keyword frequency, context semantics, and paragraph position in the paragraph, ensuring that the correlation degree calculation is not only based on surface vocabulary matching but can reflect the logical and thematic connections between paragraphs. In addition to keyword matching, the system also analyzes the logical structure between paragraphs, such as the position and front-back relationship of paragraphs in the document, and this information occupies a certain proportion in the correlation degree calculation. This ensures that the system can retain the overall structure of the document when dividing paragraphs without destroying its context connection, making the correlation degree calculation not only rely on surface keyword matching but also introduce in-depth semantic understanding and paragraph structure retention, thus greatly improving the accuracy of paragraph division and correlation degree calculation.

[0026] Preferably, calculating the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs includes the following steps:

[0027] 2.4.1. Let the keyword sets in paragraph P i and paragraph P j be K i and Kj , the weight of keyword k in paragraph P i and paragraph P j are W(k, P i ), W(k, P j ), respectively, and the calculation formula is:

[0028]

[0029] where f(k, P i ), f(k, P j ) respectively represent the occurrence frequency of keyword k in paragraph P i and paragraph P j , respectively represent the total occurrence frequency of all keywords in paragraph P i and paragraph P j ;

[0030] 2.4.2. Based on the importance of keywords in the field, the weights of keywords are corrected. The corrected weights of keyword k in paragraph P i and paragraph P j are W′(k, P i ) and W′(k, P j ), respectively:

[0031] W′(k, P i ) = W(k, P i ) · I(k), W′(k, P j ) = W(k, P j ) · I(k)

[0032] where I(k) is the domain importance weight of keyword k;

[0033] 2.4.3. Based on the corrected weights of keywords, calculate the similarity Sim(P i and paragraph P j ): i , P j ):

[0034]

[0035] 2.4.4. Calculate the structural relationship weight S(P i and paragraph P j ): i , P j ):

[0036]

[0037] where α is an adjustment parameter;

[0038] 2.4.5. Based on Paragraph P i , Paragraph P j 's similarity Sim(P i , P j ) and structural relationship weight S(P i , P j ), calculate the correlation degree R(P i and Paragraph P j : i , P j ):

[0039] R(P i , P j ) = β · Sim(P i , P j ) + (1 - β) · S(P i , P j )

[0040] where β is a weighting parameter.

[0041] Preferably, the third step includes:

[0042] 3.1. Deliver the divided sub - documents to the producer thread through the network, and the producer thread stores the feature information of the divided sub - documents into the Redis message queue in the generation order of the sub - documents;

[0043] 3.2. The consumer thread retrieves tasks from the Redis message queue and delivers the tasks to the large - model for knowledge extraction;

[0044] 3.3. Select a large - model based on specific application requirements and hardware conditions;

[0045] 3.4. Customize the relationship list of the model according to specific application requirements, and based on the relationship list, extract the relationships between entities in the divided sub - documents and output them in a structured format.

[0046] Beneficial effects: Through the integration of the producer - consumer pattern and the asynchronous architecture, the producer thread puts tasks into the message queue, and the consumer thread retrieves tasks from the queue and processes them, effectively solving the problems of computing power limitation and waiting time, significantly improving the concurrent processing ability and overall efficiency of the system, ensuring that when the system processes a large number of complex documents, while guaranteeing the quality of information extraction, reducing the computational load of the model, so as to better meet the actual application requirements. In addition, by dynamically adjusting the number of consumer threads, maintaining the hardware load at an ideal level, and improving system stability.

[0047] Preferably, the third step further includes a process of retraining the model using the proprietary text in the field of robotics and a process of optimizing the relation extraction ability of the retrained model. The process of retraining the model using the proprietary text in the field of robotics includes:

[0048] 3.5.1. Collect the proprietary text in the field of robotics, which covers relevant terms, technical parameters, and usage scenarios in the field of robotics;

[0049] 3.5.2. Clean and annotate the collected proprietary text of robotics;

[0050] 3.5.3. Use the annotated data to retrain the large model;

[0051] The process of optimizing the relation extraction ability of the retrained model includes:

[0052] 3.6.1. In the field of robotics, clarify the types of relationships between entities that need to be recognized;

[0053] 3.6.2. Annotate the entities and their relationships in the data, and define the head entity, relation, and tail entity;

[0054] 3.6.3. Input the annotated data into the retrained model for optimization training. During the optimization training process, continuously adjust the weights, learn how to recognize different relation triples, and output these relationships in a structured manner.

[0055] Beneficial effects: By retraining the large model using the proprietary text in the field of robotics, the large model can better adapt to the terms and expressions in the field of robotics, enhance the entity recognition ability. Through introducing specific data in the field of robotics for retraining, the large model can more accurately understand and extract key information related to robotics, improving the quality and accuracy of the data extraction results of the large model. Based on the well-retrained model, further training the model from the perspective of association relationships makes the model more intelligent and efficient when processing complex texts, not only improving the understanding and extraction ability of information in the proprietary field, but also enhancing the accuracy of information generation and analysis. After the model is retrained with the dedicated dataset and deeply trained with association relationships, it has a stronger domain understanding ability.

[0056] Preferably, the fourth step includes:

[0057] 4.1. Extract the attribute names and attribute values of each attribute triple;

[0058] 4.2. Use a similarity algorithm to remove duplicate information in the attribute triples from the existing knowledge graph and construct new attribute triples.

[0059] Beneficial effects: By processing various triple information from different data sources, analyzing, integrating, deduplicating, and disambiguating it with the existing data in the graph database, triples that can be stored in the database are finally obtained, effectively integrating the information from different data sources, constructing a more comprehensive, accurate, and valuable knowledge system, and providing a reliable basis for subsequent data analysis and applications.

[0060] The present invention also provides a document-level knowledge extraction and fusion system based on a large language model, including:

[0061] A keyword construction module for determining the range of required key information and establishing a keyword dictionary;

[0062] A document division module for dividing the unstructured data at the document level into paragraphs according to the keyword dictionary to obtain the divided sub-documents;

[0063] A knowledge extraction module for using the producer-consumer pattern to integrate the asynchronous architecture of the large model to build a software system, and using the software system to sequentially perform knowledge extraction tasks on the divided sub-documents to extract key information from the unstructured data of the sub-documents;

[0064] A knowledge fusion module for integrating and classifying all the key information extracted from the same sub-document to obtain regularized data, and then performing knowledge fusion processing on the regularized data.

[0065] Preferably, the document division module is further configured to:

[0066] 2.1. Load the unstructured data at the document level into the system;

[0067] 2.2. Use a document parsing tool to split the unstructured data at the document level so that each paragraph or sentence of the document becomes an independent processing unit;

[0068] 2.3. The system scans the text in each paragraph or sentence to find terms or phrases that match the entries in the keyword dictionary;

[0069] 2.4. Calculate the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs;

[0070] 2.5. Divide the document into several sub-documents according to the correlation degree between paragraphs;

[0071] 2.6. Store the divided sub-documents and generate a structured output.

[0072] Preferably, calculating the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs includes the following steps:

[0073] 2.4.1. Let paragraph P i and paragraph Pj The sets of keywords in i are K j , i keyword k has weights W(k, P j ) and W(k, P i ) in paragraphs P j respectively, and the calculation formula is:

[0074]

[0075] where f(k, P i ) and f(k, P j ) represent the occurrence frequencies of keyword k in paragraphs P i and P j respectively, represent the total occurrence frequencies of all keywords in paragraphs P i and P j respectively;

[0076] 2.4.2. Based on the importance of keywords in the domain, the weights of keywords are corrected. The corrected weights W′(k, P i ) and W′(k, P j ) of keyword k in paragraphs P i and P j are respectively:

[0077] W′(k, P i ) = W(k, P i ) · I(k), W′(k, P j ) = W(k, P j ) · I(k)

[0078] where I(k) is the domain importance weight of keyword k;

[0079] 2.4.3. Based on the corrected weights of keywords, calculate the similarity Sim(P i , P j ) between paragraphs P i and P j :

[0080]

[0081] 2.4.4. Calculate the structural relationship weight S(P i , P j ) between paragraphs P i and P j :

[0082]

[0083] Among them, α is an adjustment parameter;

[0084] 2.4.5. Based on paragraph P i , paragraph P j 's similarity Sim(P i , P j ) and structural relationship weight S(P i , P j ), calculate the correlation degree R(P i and paragraph P j : i j i ) between:

[0085] R(P i , P j ) = β · Sim(P i , P j ) + (1 - β) · S(P i , P j )

[0086] Among them, β is a weighting parameter.

[0087] Preferably, the knowledge extraction module is further configured to:

[0088] 3.1. Deliver the divided sub-documents to the producer thread through the network, and the producer thread stores the feature information of the divided sub-documents into the Redis message queue in the generation order of the sub-documents;

[0089] 3.2. The consumer thread obtains tasks from the Redis message queue and delivers the tasks to the large model for knowledge extraction;

[0090] 3.3. Select a large model based on specific application requirements and hardware conditions;

[0091] 3.4. Customize the relationship list of the model according to specific application requirements, and based on the relationship list, extract the relationships between entities in the divided sub-documents and output them in a structured format.

[0092] Preferably, the knowledge extraction module further includes a secondary training unit and a relationship extraction ability training unit. The secondary training unit is used to perform secondary training on the model using the proprietary text in the robot field, specifically including the following processes:

[0093] 3.5.1. Collect the proprietary text in the robot field, which covers relevant terms, technical parameters, and usage scenarios in the robot field;

[0094] 3.5.2. Clean and annotate the collected robot proprietary text;

[0095] 3.5.3. Use the labeled data to perform secondary training on the large model;

[0096] The relation extraction ability training unit is used to optimize the relation extraction ability of the model after secondary training, and specifically includes the following process:

[0097] 3.6.1. In the field of robotics, clarify the types of relationships between entities that need to be recognized;

[0098] 3.6.2. Label the entities and their relationships in the data, and define the head entity, relation, and tail entity;

[0099] 3.6.3. Input the labeled data into the model after secondary training for optimization training. During the optimization training process, continuously adjust the weights, learn how to recognize different relation triples, and output these relationships in a structured manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0100] Figure 1 It is a flowchart of the document-level knowledge extraction and fusion method based on the large language model provided by the embodiment of the present invention;

[0101] Figure 2 It is a flowchart of the working process of the knowledge extraction module in the document-level knowledge extraction and fusion system based on the large language model provided by the embodiment of the present invention;

[0102] Figure 3 It is a flowchart of the working process of the knowledge fusion module in the document-level knowledge extraction and fusion system based on the large language model provided by the embodiment of the present invention;

[0103] Figure 4 It is an effect diagram of the document-level knowledge extraction and fusion method based on the large language model provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following combines specific embodiments and refers to the accompanying drawings to clearly and completely describe the technical solutions of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0105] Open-source or commercial large language models usually provide general functions such as text generation, dialogue systems, content summarization, etc., and have a certain degree of information extraction ability. However, when facing the extraction tasks of texts in specific fields (such as the robotics field), their performance is not outstanding. The main reason is that there are a large number of professional terms and complex sentence patterns in these fields, making it difficult for general models to accurately understand and extract key information. Therefore, this invention proposes a document-level knowledge extraction and fusion method and system based on large language models.

[0106] Embodiment 1

[0107] As Figure 1 shown, this embodiment provides a document-level knowledge extraction and fusion method based on large language models, including the following steps:

[0108] Step 1: Determine the scope of required key information and establish a keyword dictionary.

[0109] In the robotics field, one usually pays more attention to the various design indicators, performance parameters, and component information of a certain product. These parameters intuitively reflect the various performances, power consumption, and composition of the product. Among them, design indicators usually include dimensions, weight, materials, and manufacturing processes, etc.; performance parameters include speed, accuracy, power consumption, and durability, etc.; component information includes the various components of the product and their interrelationships. Through design indicators, performance parameters, and component information, users can quickly evaluate whether the product meets their own needs, judge whether components can be reused, and thus make more reasonable decisions. Based on these needs, creating a dictionary containing keywords and synonyms becomes the starting point. The establishment of these keyword dictionaries can not only be used for document parsing and classification but also for subsequent data mining and knowledge extraction.

[0110] Step 1 specifically includes the following processes:

[0111] 1.1. Collect key parameters and terms: Collect common design indicators, performance parameters, and component terms from product manuals, technical specifications, and industry standards. These parameters and terms will form the core of the keyword dictionary.

[0112] Common key parameters in the robotics field include degrees of freedom, payload, load, workspace, maximum movement range, repeatability, positioning accuracy, resolution, power capacity, etc. By summarizing and classifying the common key parameters in the robotics field and brainstorming with the team to check for omissions, relatively complete key information can be obtained.

[0113] 1.2. Define synonyms and related terms: Define synonyms and related terms for each keyword to ensure the comprehensiveness and accuracy of the keyword dictionary and to ensure that the system can adapt to different usage habits and industry specifications. For example: "degree of freedom" and "number of degrees of freedom" are synonyms, and power consumption can include synonyms such as "energy consumption" and "power consumption".

[0114] 1.3. Dynamically update and maintain the keyword dictionary. As technology develops and new products are launched, the keyword dictionary needs to be updated and maintained regularly to add new terms and delete outdated vocabulary to maintain its timeliness and practicality.

[0115] 1.4. Configure and manage the keyword dictionary so that the keyword dictionary can be flexibly called and updated in the system, ensuring that the system can quickly adapt to different scenarios and improving the adaptability and flexibility of the system.

[0116] Step 2: Divide the unstructured data at the document level into paragraphs to obtain the divided sub-documents. The specific process includes the following:

[0117] 2.1. Load the unstructured data at the document level: Load the unstructured data document to be processed into the system. The document formats include but are not limited to docx, pdf, txt, etc.

[0118] 2.2. Parse the document content: Use a document parsing tool to convert the document content into a processable text format. The parsing tools can be selected as python-docx, PyPDF2, etc. The parsing tool divides the document into paragraphs or sentences, making each paragraph or sentence an independent processing unit. By dividing the document into the smallest units by paragraphs or sentences, the document content can be processed at a finer granularity, which helps to improve the accuracy of subsequent classification and extraction.

[0119] 2.3. Perform keyword matching on the parsed text content. The system scans the text in each paragraph or sentence to find terms or phrases that match the entries in the keyword dictionary. During the matching process, the system will consider the synonyms of the keywords to ensure comprehensive coverage.

[0120] 2.4. Calculate the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs: The system will dynamically assign weights to different paragraphs according to the frequency of keyword appearance in the paragraph, context semantics, paragraph position, and importance of keywords, and calculate the correlation degree between paragraphs in combination with the structural relationship between paragraphs.

[0121] The specific process of calculating the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs includes:

[0122] 2.4.1. Keyword matching calculation: The system first calculates the keyword matching in each paragraph and assigns weights to the keywords based on their frequencies of occurrence and their importance in the context.

[0123] Let the set of keywords in paragraph P i and paragraph P j be K i and K j respectively. The weight of keyword k in paragraph P i is W(k, P i ). The calculation formula for the weight W(k, P i ) is:

[0124]

[0125] where f(k, P i ) represents the frequency of occurrence of keyword k in paragraph P i , represents the total frequency of occurrence of all keywords in paragraph P i .

[0126] Similarly, the calculation formula for the weight W(k, P j ) of keyword k in paragraph P j is:

[0127]

[0128] where f(k, P j ) represents the frequency of occurrence of keyword k in paragraph P j , represents the total frequency of occurrence of all keywords in paragraph P j .

[0129] 2.4.2. Keyword weight correction: To better reflect the importance of keywords in a specific field, the weights of each keyword are corrected, and some core domain vocabulary will be assigned higher weights.

[0130] Let the domain importance weight of keyword k be I(k). Then the corrected weight W′(k, P i ) is:

[0131] W′(k, P i ) = W(k, P i ) · I(k)

[0132] The corrected weight W′(k, P j ) is:

[0133] W′(k, P j ) = W(k, P j)·I(k)

[0134] Among them, I(k) reflects the importance of keyword k in the field, and the system determines I(k) through pre-training or domain thesaurus.

[0135] 2.4.3. Similarity calculation between paragraphs: Based on the weights of keywords after correction, the system calculates the similarity between two paragraphs. The present invention calculates the similarity between paragraphs by means of weighted cosine similarity, and the calculation formula is:

[0136]

[0137] Among them, Sim(P i , P j ) represents the similarity between paragraph P i and paragraph P j , K i represents the set of keywords in paragraph P i , K j represents the set of keywords in paragraph P j , K i ∩ K j represents the intersection of the set of keywords K i in paragraph P i and the set of keywords K j in paragraph P j , W′(k, P i ), W′(k, P j ) respectively represent the weights of the corrected keyword k in paragraph P i , paragraph P j .

[0138] 2.4.4. Revision of the structural relationship between paragraphs: When calculating the similarity between two paragraphs, the system also needs to consider the structural relationship between paragraphs, such as the relative position and logical order of paragraphs in the document.

[0139] Let the distance between paragraph P i , paragraph P j in the document be d(P i , P j ), and the distance can be calculated by means of paragraph numbers, positions, etc. The calculation formula for the structural relationship weight S(P i , P j ) is:

[0140]

[0141] Among them, α is an adjustment parameter, reflecting the influence of the structural relationship on the similarity. The greater the distance, the smaller the structural relationship weight, and vice versa.

[0142] 2.4.5. Calculation of the Relevance Degree between Paragraphs: Considering both keyword similarity and structural relationship comprehensively, for paragraph P i and paragraph P j , the relevance degree R(P i , P j ) is as follows:

[0143] R(P i , P j ) = β · Sim(P i , P j ) + (1 - β) · S(P i , P j )

[0144] Among them, β is a weighting parameter used to adjust the influence of similarity and structural relationship on the relevance degree, and the optimal value is usually determined through experiments.

[0145] Paragraphs with a high relevance degree usually indicate that their contents are closely related logically and thematically, while paragraphs with a low relevance degree reflect a relatively large distance in terms of their themes or semantics. Dynamically assigning weights means that for the core domain vocabulary appearing in paragraphs, such as technical terms related to robots, the system will assign higher weights. Dynamically assigning weights according to the context in which keywords appear in paragraphs and the importance of keywords can improve the accuracy of calculating the relevance degree between paragraphs, enabling the system to better capture the hidden structure and logical relationship in the document.

[0146] Calculating the relevance degree between paragraphs in combination with the structural relationship between paragraphs means that by analyzing the logical structure between paragraphs, such as the position of paragraphs in the document, the front - back relationship, etc., to determine the proportion occupied by the structural relationship between paragraphs, which can ensure that the overall structure of the document is retained when the system performs paragraph division and the context connection of the document is not damaged.

[0147] When calculating the relevance degree between paragraphs, the present invention not only relies on surface keyword matching, but also introduces in - depth semantic understanding and paragraph structure preservation, which can greatly improve the accuracy of paragraph division and relevance degree calculation.

[0148] 2.5. Document Division: Divide the document into several sub - documents according to the relevance degree between paragraphs. The system integrates paragraphs with a relatively high relevance degree together to form sub - documents with highly aggregated content. Each sub - document contains paragraphs with strong context relationships and similar themes, ensuring that the information in the divided sub - documents is coherent and the structure is clear.

[0149] 2.6. Sub - document Generation: Store the divided sub - documents and generate a structured output. The system can generate output files in various formats, including structured database entries, marked document files, etc. These structured sub - documents will be used as the input for large - scale language model processing, which can improve the efficiency and accuracy of model processing.

[0150] Large language models are mainly applied in the field of natural language processing (NLP), including tasks such as text generation, text understanding, and text classification. However, large language models are usually limited by hardware performance and understanding ability, and it is difficult to efficiently process document-level unstructured data. The main reason is that the information covered by document-level data is too large, with more knowledge points, more advanced sentence grammars, more implicit correlation relationships, and a large number of context connections, making it difficult for the model to comprehensively understand and perform relationship triple extraction. In the present invention, the unstructured data at the document level is divided, with paragraphs as the smallest granularity, and different weights are assigned according to the degree of association between paragraphs in cooperation with a keyword dictionary to achieve the division of the document. Paragraphs with a higher degree of association, that is, paragraphs with strong context relationships, are integrated together to achieve a high degree of aggregation of the content of the sub-documents after division and the summary of the same knowledge points. Through this document division algorithm, the system can effectively process document-level unstructured data, convert complex document content into structured and analyzable information, not only improving the processing ability and efficiency of large language models, but also ensuring the accuracy and comprehensiveness of information extraction, providing reliable technical support for information processing in the field of robotics.

[0151] Step 3: Use the producer-consumer mode to integrate the asynchronous architecture of the large model to build a software system, and use the software system to sequentially perform knowledge extraction tasks on the divided sub-documents to extract key information from the unstructured data of the sub-documents. The specific process includes the following:

[0152] 3.1. Deliver the divided sub-documents to the producer thread through the network, and the producer thread stores the feature information of the divided sub-documents into the Redis message queue in the generation order of the sub-documents. In this way, tasks can be queued orderly for processing.

[0153] 3.2. The consumer thread obtains tasks from the Redis message queue and delivers the tasks to the model for knowledge extraction. The overall architecture of the present invention adopts the producer-consumer mode, and each consumer thread runs independently. Therefore, multiple consumer threads can be started as needed to process tasks in parallel, so as to dynamically adjust the hardware load and keep the system at an ideal performance level.

[0154] 3.3. Select a large language model based on specific application requirements and hardware conditions to ensure the balance of performance and accuracy. An open-source or commercial large language model can be selected to search for and extract key information from the divided sub-documents.

[0155] 3.4. Design the prompt words for the model

[0156] Customize the relationship list of the model according to specific application requirements. The relationship list is used to indicate specific relationship types that the model should focus on, such as degrees of freedom, payload, working speed, etc., to ensure that the model can accurately capture the required information.

[0157] Based on the relationship list, extract the relationships between entities in the divided sub-documents and output them in a structured format. Design the output format of the model, requiring the model to output the extracted relationships in a specific format, such as relationship triples represented in json format. This structured output facilitates subsequent data processing and analysis. The relationship triples represented in json format are as follows:

[0158] ["Head entity 1", "Relationship 1", "Tail entity 1"],

[0159] ["Head entity 2", "Relationship 2", "Tail entity 2"],

[0160] ["Head entity 3", "Relationship 3", "Tail entity 3"].

[0161] Large language models are limited by computing power factors when processing tasks and usually require a long waiting time, especially when performing document-level knowledge extraction tasks, which is more obvious. In the case of poor hardware conditions, problems such as out-of-memory errors are even likely to occur. To solve the problems of long waiting time for processing tasks and easy occurrence of out-of-memory errors, the present invention can effectively avoid time-consuming extraction tasks from blocking the entire system, achieve decoupling of producers and consumers, and then release the performance of the producer side by using a producer-consumer and integrating the asynchronous architecture of the large model. It not only hides the waiting time but also can dynamically adjust the hardware load to maintain the balance of server load, thereby improving the stability of the software system. Specifically, when the system load is low, the number of consumer threads can be increased to speed up the processing speed; when the system load is high, the number of consumer threads can be reduced to avoid excessive resource consumption. This dynamic adjustment mechanism ensures that the system can run stably and efficiently under different load conditions, effectively solves the computing power limitation problem of large language models when processing document-level tasks, and achieves efficient knowledge extraction. Through the design of prompt words, the system can provide accurate and fast information extraction services in different application scenarios.

[0162] The process of using the proprietary text in the field of robotics to retrain the model is also included in step three, specifically including:

[0163] 3.5.1. Data preparation: Collect the proprietary text in the field of robotics, which covers relevant terms, technical parameters, usage scenarios, etc. in the field of robotics.

[0164] 3.5.2. Data Cleaning and Annotation: Clean and annotate the robot-specific text collected to ensure that the model can correctly understand and identify domain-specific terms and expressions.

[0165] 3.5.3. Training the Model: Use the annotated data to retrain the large model so that it can identify and process entities and relationships within the domain.

[0166] In the present invention, by using the robot-specific text in the field of robotics to retrain the large model, the large model can better adapt to the terms and expressions in the field of robotics, enhance the entity recognition ability. By introducing specific data in the field of robotics for retraining, the large model can more accurately understand and extract key information related to robots, improving the quality and accuracy of the data extraction results of the large model.

[0167] The process of optimizing the training of the relationship extraction ability of the model after the second training is also included in step three, specifically including:

[0168] 3.6.1. Relationship Definition: In the field of robotics, clarify the types of relationships between entities that need to be recognized, such as the association between components and robot models.

[0169] 3.6.2. Data Annotation: Annotate the entities and their relationships in the data, defining the head entity, relationship, and tail entity.

[0170] 3.6.3. Model Optimization Training: Input the annotated data into the model after the second training for optimization training to make it more accurate in the relationship extraction task. During the model training process, by continuously adjusting the weights, learn how to identify different relationship triples and output these relationships in a structured manner.

[0171] Based on the well-trained model after the second training, further train the model from the perspective of association relationships, especially in the scenario of knowledge graph construction, optimize for tasks such as relationship triple extraction. Through further optimization training, the model can not only identify entities and their attributes but also flexibly handle complex associations between entities. In the field of robotics, the model can identify key entities such as robot models, functions, and performance indicators, and extract the associations between them and manufacturing processes, components, etc. This in-depth training and optimization make the model more intelligent and efficient in processing complex texts, not only improving the understanding and extraction ability of domain-specific information but also enhancing the accuracy of information generation and analysis.

[0172] Through these training methods, the output of the model can not only be used to construct a high-quality knowledge graph but also play a key role in data analysis and decision support. Through continuous optimization and iteration, we can better meet the actual needs in the field of robotics and promote the development of information processing technology.

[0173] Step 4: Integrate and classify all the key information extracted from the same sub-document to obtain regularized data, and then perform knowledge fusion processing on the regularized data. The specific process includes the following:

[0174] 4.1 Analyze each attribute triple in detail, including the extraction of the attribute name and attribute value, to ensure a clear understanding of the information in each triple.

[0175] 4.2 Compare and integrate the extracted attribute triples with the existing knowledge graph. Specifically, a similarity algorithm can be used to remove duplicate information between the attribute triples and the existing knowledge graph to ensure the accuracy and integrity of the final knowledge system.

[0176] 4.3 On the basis of integration, construct new attribute triples to further improve the knowledge system and ensure the consistency and accuracy of the information.

[0177] 4.4 Perform the operation of storing the attribute triple information after knowledge fusion processing into the database.

[0178] Embodiment 2

[0179] This embodiment provides a document-level knowledge extraction and fusion system based on a large language model, including:

[0180] A keyword construction module, which is used to determine the range of required key information and establish a keyword dictionary; specifically including:

[0181] 1.1 Collect key parameters and terms from product manuals, technical specifications, and industry standards to obtain keywords;

[0182] 1.2 Define synonyms and related terms for each keyword;

[0183] 1.3 Dynamically update and maintain the keyword dictionary;

[0184] 1.4 Perform configuration management on the keyword dictionary.

[0185] A document division module, which is used to divide the unstructured data at the document level into paragraphs according to the keyword dictionary to obtain the divided sub-documents; specifically including:

[0186] 2.1 Load the unstructured data at the document level into the system;

[0187] 2.2 Use a document parsing tool to split the unstructured data at the document level so that each paragraph or sentence of the document becomes an independent processing unit;

[0188] 2.3 The system scans the text in each paragraph or sentence to find terms or phrases that match the entries in the keyword dictionary;

[0189] 2.4. Calculate the correlation degree between paragraphs based on the keyword matching results and the structural relationship between paragraphs; the specific calculation of the correlation degree includes the following steps:

[0190] 2.4.1. Keyword matching calculation: The system first calculates the keyword matching situation in each paragraph and assigns weights to the keywords based on the frequency of keyword appearance and their importance in the context.

[0191] Let the keyword sets in paragraph P i and paragraph P j be K i and K j respectively. The weight of keyword k in paragraph P i is W(k, P i ). The calculation formula for the weight W(k, P i ) is:

[0192]

[0193] where f(k, P i ) represents the frequency of keyword k appearing in paragraph P i , represents the total frequency of all keywords appearing in paragraph P i .

[0194] Similarly, the calculation formula for the weight W(k, P j ) in paragraph P j is:

[0195]

[0196] where f(k, P j ) represents the frequency of keyword k appearing in paragraph P j , represents the total frequency of all keywords appearing in paragraph Pj.

[0197] 2.4.2. Keyword weight correction: In order to better reflect the importance of keywords in a specific field, the weight of each keyword is corrected, and some core field vocabulary will be assigned higher weights.

[0198] Let the domain importance weight of keyword k be I(k). Then the corrected weight W′(k, P i ) is:

[0199] W(k, P i ) = W(k, P i ) · I(k)

[0200] The corrected weight W′(k, P j) is as follows:

[0201] W(k, P j ) = W(k, P j ) · I(k)

[0202] Among them, I(k) reflects the importance of keyword k in the field, and the system determines I(k) through pre-training or a domain thesaurus.

[0203] 2.4.3. Paragraph similarity calculation: Based on the weights of keywords after correction, the system calculates the similarity between two paragraphs. The present invention calculates the similarity between paragraphs in the way of weighted cosine similarity, and the calculation formula is:

[0204]

[0205] Among them, Sim(P i , P j ) represents the similarity between paragraph P i and paragraph P j . K i represents the set of keywords in paragraph P i . K j represents the set of keywords in paragraph P j . K i ∩ K j represents the intersection of the set of keywords K i in paragraph P i and the set of keywords K j in paragraph P j . W′(k, P i ), W′(k, P j ) respectively represent the weights of the corrected keyword k in paragraph P i , paragraph P j .

[0206] 2.4.4. Paragraph structure relationship correction: When calculating the similarity between two paragraphs, the system also needs to consider the structural relationship of the paragraphs, such as the relative position and logical order of the paragraphs in the document.

[0207] Suppose the distance between paragraph P i , paragraph P j in the document is d(P i , P j ), and the distance can be calculated by means of paragraph numbers, positions, etc. The calculation formula for the structural relationship weight S(P i , P j ) is:

[0208]

[0209] Among them, α is an adjustment parameter, reflecting the influence of the structural relationship on the similarity. The greater the distance, the smaller the weight of the structural relationship, and vice versa.

[0210] 2.4.5. Calculation of the inter-paragraph correlation: Considering both the keyword similarity and the structural relationship comprehensively, the correlation R(P i and paragraph P j between) is as follows: i ,P j ) is:

[0211] R(P i ,P j ) = β·Sim(P i ,P j )+(1-β)·S(P i ,P j )

[0212] Among them, β is a weighting parameter, used to adjust the influence of the similarity and the structural relationship on the correlation. Usually, the optimal value is determined through experiments.

[0213] 2.5. Divide the document into several sub-documents according to the correlation between paragraphs;

[0214] 2.6. Store the divided sub-documents and generate a structured output.

[0215] The knowledge extraction module is used to build a software system using the asynchronous architecture of the large model in the producer-consumer mode, and use the software system to perform knowledge extraction tasks on the divided sub-documents in turn, and extract key information from the unstructured data of the sub-documents; specifically including:

[0216] 3.1. Deliver the divided sub-documents to the producer thread through the network, and the producer thread stores the feature information of the divided sub-documents into the Redis message queue in the generation order of the sub-documents;

[0217] 3.2. The consumer thread obtains tasks from the Redis message queue and delivers the tasks to the large model for knowledge extraction;

[0218] 3.3. Select a large model based on specific application requirements and hardware conditions;

[0219] 3.4. Customize the relationship list of the model according to specific application requirements, and based on the relationship list, extract the relationships between entities in the divided sub-documents and output them in a structured format.

[0220] The knowledge extraction module also includes:

[0221] The secondary training unit is used to perform secondary training on the model using the proprietary text in the robot field, specifically including the following process:

[0222] 3.5.1. Collect proprietary texts in the field of robotics, which cover relevant terms, technical parameters, and usage scenarios in the field of robotics;

[0223] 3.5.2. Clean and annotate the collected proprietary texts of robots;

[0224] 3.5.3. Use the annotated data to perform secondary training on the large model;

[0225] The relation extraction ability training unit is used to optimize the relation extraction ability of the model after secondary training, specifically including the following processes:

[0226] 3.6.1. In the field of robotics, clarify the types of relationships between entities that need to be recognized;

[0227] 3.6.2. Annotate the entities and their relationships in the data, and define the head entity, relation, and tail entity;

[0228] 3.6.3. Input the annotated data into the model after secondary training for optimization training. During the optimization training process, continuously adjust the weights, learn how to recognize different relation triples, and output these relationships in a structured manner.

[0229] The knowledge fusion module is used to integrate and classify all the key information extracted from the same sub-document to obtain regular data, and then perform knowledge fusion processing on the regular data, specifically including:

[0230] 4.1. Extract the attribute names and attribute values of each attribute triple;

[0231] 4.2. Use a similarity algorithm to remove duplicate information in the attribute triples and the existing knowledge graph, and construct new attribute triples.

[0232] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A document-level knowledge extraction and fusion method based on a large language model, characterized by: Methods include: S1. Determine the scope of key information required and establish a keyword dictionary; S2, dividing the document-level unstructured data into paragraphs according to the keyword dictionary to obtain the divided sub-documents; In S2, the correlation between paragraphs is calculated based on the keyword matching results and the structural relationship between paragraphs, including: 2.4.1 Paragraph and paragraphs The keyword sets in are , , keywords k In paragraph ,paragraph The weight in , for: , in, , Respectively represent keywords k In paragraph ,paragraph The frequency of occurrence in , Represents paragraphs separately ,paragraph The sum of the occurrence frequencies of all keywords in ; 2.4.

2. Based on the importance of keywords in the field, the weight of keywords is modified. The modified keywords k In paragraph ,paragraph The weight in , They are: , 2.4.

3. Calculate the paragraph weight based on the corrected keyword weight and paragraphs Similarity : 2.4.

4. Calculation paragraph and paragraphs The weight of the structural relationship between : 2.4.

5. Based on paragraphs ,paragraph Similarity and structural relationship weight , calculate paragraph and paragraphs The correlation between : in, For keywords k The domain importance weight of is the weighting parameter; S3. Use the asynchronous architecture of the producer-consumer model to build a software system, and use the software system to perform knowledge extraction tasks on the divided sub-documents in turn to extract key information from the unstructured data of the sub-documents; S4. Integrate and classify all key information extracted from the same sub-document to obtain regular data, and then perform knowledge fusion processing on the regular data.

2. The document-level knowledge extraction and fusion method based on a large language model according to claim 1, characterized in that: S1 includes: 1.

1. Collect key parameters and terms from product manuals, technical specifications and industry standards to obtain keywords; 1.

2. Define synonyms and related terms for each keyword; 1.

3. Dynamically update and maintain keyword dictionaries; 1.

4. Perform configuration management on keyword dictionary.

3. The document-level knowledge extraction and fusion method based on a large language model according to claim 1 is characterized in that: S2 includes: 2.

1. Load document-level unstructured data into the system; 2.

2. Use document parsing tools to segment document-level unstructured data so that each paragraph or sentence in the document becomes an independent processing unit; 2.

3. The system scans the text in each paragraph or sentence, looking for terms or phrases that match entries in the keyword dictionary; 2.

4. Divide the document into several sub-documents according to the relevance between paragraphs; 2.

5. Store the divided sub-documents and generate structured output.

4. The document-level knowledge extraction and fusion method based on a large language model according to claim 1, characterized in that: S3 includes: 3.

1. Deliver the divided sub-documents to the producer thread through the network. The producer thread stores the feature information of the divided sub-documents into the Redis message queue according to the generation order of the sub-documents. 3.

2. The consumer thread obtains tasks from the Redis message queue and delivers the tasks to the large model for knowledge extraction; 3.

3. Select a large model based on specific application requirements and hardware conditions; 3.

4. Customize the relationship list of the model according to specific application requirements. Based on the relationship list, extract the relationship between entities in the divided sub-documents and output them in a structured format.

5. The document-level knowledge extraction and fusion method based on a large language model according to claim 4 is characterized in that: S3 also includes a process of performing secondary training on the model using proprietary texts in the field of robotics and a process of optimizing the relationship extraction capability of the secondary trained model. The process of performing secondary training on the model using proprietary texts in the field of robotics includes: 3.5.

1. Collect proprietary texts in the field of robotics, which cover relevant terms, technical parameters, and usage scenarios in the field of robotics; 3.5.

2. Clean and label the collected robot-specific texts; 3.5.

3. Use the labeled data to train the large model again; The process of optimizing the relationship extraction capability of the model after secondary training includes: 3.6.

1. In the field of robotics, clarify the types of relationships between entities that need to be identified; 3.6.

2. Label the entities and their relationships in the data, and define the head entity, relationship, and tail entity; 3.6.

3. Input the labeled data into the model after secondary training for optimization training. During the optimization training process, the weights are continuously adjusted to learn how to identify different relationship triplets and output these relationships in a structured manner.

6. The document-level knowledge extraction and fusion method based on a large language model according to claim 1, characterized in that: S4 includes: 4.

1. Extract the attribute name and attribute value of each attribute triple; 4.

2. Use the similarity algorithm to remove duplicate information between attribute triples and existing knowledge graphs and construct new attribute triples.

7. Document-level knowledge extraction and fusion system based on large language model, characterized by: include: Keyword building module, used to determine the scope of required key information and build a keyword dictionary; The document segmentation module is used to segment the document-level unstructured data into paragraphs according to the keyword dictionary to obtain the segmented sub-documents; The document segmentation module calculates the relevance between paragraphs based on keyword matching results and the structural relationship between paragraphs, including: 2.4.

1. Set paragraphs and paragraphs The keyword sets in are , , keywords k In paragraph ,paragraph The weights in are , , the calculation formula is: , in, , Respectively represent keywords k In paragraph ,paragraph The frequency of occurrence in , Represents paragraphs separately ,paragraph The sum of the occurrence frequencies of all keywords in ; 2.4.

2. Based on the importance of keywords in the field, the weight of keywords is modified. The modified keywords k In paragraph ,paragraph The weight in , They are: , 2.4.

3. Calculate the paragraph weight based on the corrected keyword weight and paragraphs Similarity : 2.4.

4. Calculation paragraph and paragraphs The weight of the structural relationship between : 2.4.

5. Based on paragraphs ,paragraph Similarity and structural relationship weight , calculate paragraph and paragraphs The correlation between : in, For keywords k The domain importance weight of To adjust the parameters, is the weighting parameter; The knowledge extraction module is used to build a software system using the asynchronous architecture of the producer-consumer model to integrate the large model. The software system is used to perform knowledge extraction tasks on the divided sub-documents in turn to extract key information from the unstructured data of the sub-documents. The knowledge fusion module is used to integrate and classify all key information extracted from the same sub-document to obtain regular data, and then perform knowledge fusion processing on the regular data.

8. The document-level knowledge extraction and fusion system based on a large language model according to claim 7 is characterized by: The document partitioning module is also used to: 2.

1. Load document-level unstructured data into the system; 2.

2. Use document parsing tools to segment document-level unstructured data so that each paragraph or sentence in the document becomes an independent processing unit; 2.

3. The system scans the text in each paragraph or sentence, looking for terms or phrases that match entries in the keyword dictionary; 2.

4. Divide the document into several sub-documents according to the relevance between paragraphs; 2.

5. Store the divided sub-documents and generate structured output.

9. The document-level knowledge extraction and fusion system based on a large language model according to claim 7, characterized in that: The knowledge extraction module is also used to: 3.

1. Deliver the divided sub-documents to the producer thread through the network. The producer thread stores the feature information of the divided sub-documents into the Redis message queue according to the generation order of the sub-documents. 3.

2. The consumer thread obtains tasks from the Redis message queue and delivers the tasks to the large model for knowledge extraction; 3.

3. Select a large model based on specific application requirements and hardware conditions; 3.

4. Customize the relationship list of the model according to specific application requirements. Based on the relationship list, extract the relationship between entities in the divided sub-documents and output them in a structured format.

10. The document-level knowledge extraction and fusion system based on a large language model according to claim 9, characterized in that: The knowledge extraction module also includes a secondary training unit and a relationship extraction capability training unit. The secondary training unit is used to perform secondary training on the model using proprietary text in the field of robotics. The process includes: 3.5.

1. Collect proprietary texts in the field of robotics, which cover relevant terms, technical parameters, and usage scenarios in the field of robotics; 3.5.

2. Clean and label the collected robot-specific texts; 3.5.

3. Use the labeled data to train the large model again; The relationship extraction capability training unit is used to optimize the relationship extraction capability of the model after secondary training. The process includes: 3.6.

1. In the field of robotics, clarify the types of relationships between entities that need to be identified; 3.6.

2. Label the entities and their relationships in the data, and define the head entity, relationship, and tail entity; 3.6.

3. Input the labeled data into the model after secondary training for optimization training. During the optimization training process, the weights are continuously adjusted to learn how to identify different relationship triplets and output these relationships in a structured manner.

Citation Information

Patent Citations

  • Unstructured text data knowledge extraction method based on large language model

    CN118036734A

  • Formula searching method and apparatus in text recognition

    CN108133168A

  • Carbon standard knowledge graph construction method based on fusion text

    CN115344712A