Large model-based knowledge base automated knowledge extraction and classification agent method
Patent Information
- Application Number
- CN202610668281.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2046-05-15
AI Technical Summary
[0003]现有技术存在以下缺点:通用信息抽取模型通常仅将文本视为线性词序列,忽略表头缩进、列表层级和跨段落引用等版式结构信息,导致表格属性对和列表约束关系极易被错误拆分或漏识别;现有预处理流程缺少对格式异常程度的量化评估能力,OCR错位、括号不闭合或列表编号断裂的区域与正常文本混同处理,极易在边界定位时产生连锁误判;单一编码器架构难以同时兼顾长程顺序语义依赖和局部非连续结构关联,面对排版复杂的知识库文本时往往顾此失彼,在表格密集或层级较深的文档中实体召回率与关系分类准确度明显下降;全量依赖大语言模型进行知识抽取会带来高昂的计算成本与不可控的响应延迟,而单纯依赖小模型又缺乏对极端格式变异和新术语的有效纠错机制,二者之间缺乏一种按实际结构难度动态调用的协同工作模式
[0071]The beneficial effects of the present invention are as follows: (1) The present invention constructs a structural connection matrix based on the physical layout and hierarchical path of the document, and explicitly encodes the non-linear positional relationship between the header and cell, and between the title and the body into inter-word connection weights, so that the model can directly use the typesetting logic to perform semantic aggregation.
Smart Images

Figure CN122287825B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method for automated knowledge extraction and classification of intelligent agents based on large model knowledge bases. Background Technology
[0002] With the deepening of enterprise digital transformation, the number of various knowledge base documents has exploded. Unstructured or semi-structured texts such as product manuals, operation manuals, standards and specifications, maintenance records, and scanned copies obtained through optical character recognition (OCR) transcription constitute the core data foundation for enterprise internal knowledge management and intelligent retrieval. How to automatically extract structured knowledge units from these massive and diverse documents—such as terminology definitions, equipment parameters, operational constraints, and fault causal relationships—and transform them into standardized data that can be directly consumed by knowledge graphs, intelligent question-answering systems, or retrieval enhancement applications has become a key technical challenge that urgently needs to be solved in the fields of knowledge engineering and natural language processing.
[0003] Existing technologies suffer from the following drawbacks: General information extraction models typically treat text as a linear sequence of words, ignoring formatting information such as header indentation, list hierarchy, and cross-paragraph references, making it easy for table attribute pairs and list constraints to be incorrectly split or missed; existing preprocessing workflows lack the ability to quantitatively assess the degree of formatting anomalies, and areas with OCR misalignment, unclosed brackets, or broken list numbers are processed together with normal text, easily leading to chain-reaction misjudgments during boundary localization; a single encoder architecture struggles to simultaneously handle long-range sequential semantic dependencies and local discontinuous structural relationships, often resulting in a trade-off when dealing with complex knowledge base texts, with entity recall and relation classification accuracy significantly decreasing in documents with dense tables or deep hierarchies; relying entirely on large language models for knowledge extraction incurs high computational costs and uncontrollable response latency, while relying solely on small models lacks effective error correction mechanisms for extreme formatting variations and new terms, and there is a lack of a collaborative working mode that dynamically calls upon models based on actual structural difficulty. Summary of the Invention
[0004] The technical problem to be solved by this invention is to overcome the shortcomings of the prior art and provide an intelligent agent method for automated knowledge extraction and classification based on a large model knowledge base.
[0005] The technical solution adopted to solve the above technical problems is: an intelligent agent method for automated knowledge extraction and classification based on a large model knowledge base, including the following steps:
[0006] S1, Knowledge Base Text Training Sample Construction and Annotation: Construct a knowledge base text sample set for training and validation, and manually perform entity and relation annotation on the training samples;
[0007] S2, perform structured preprocessing on the knowledge base text: perform structure restoration, word-level alignment, structure-enhanced sequence construction, and structurally difficult segment labeling on the knowledge base text to obtain structure-enhanced word sequences. Difficult-to-structure section mask ;
[0008] S3, Constructing and training a structure-first dual-channel knowledge extraction model: The structure-first dual-channel knowledge extraction model includes a sequential context encoding branch, a structural neighborhood aggregation branch, an entity boundary localization module, and a relation classification module, which are used for entity recognition, relation classification, and label prediction;
[0009] S4, Large Language Model-Assisted Data Expansion and Low-Confidence Difficult Example Correction: Automatically generates high-quality labeled data for unlabeled knowledge base text, processes low-confidence difficult example sections in the structure-first dual-channel knowledge extraction model, and improves the model's generalization ability across various text styles and entity types.
[0010] S5, Online Knowledge Base Text Extraction: Deploy the structure-first dual-channel knowledge extraction model to a knowledge base management system, document governance system, or enterprise knowledge platform to automatically extract newly arriving knowledge base text.
[0011] Furthermore, the sources of the knowledge base text sample set in S1 include product instruction documents, operation manuals, Q&A documents, standard specification documents, interface instruction documents, operation and maintenance record documents, and scanned documents obtained through optical character recognition transcription.
[0012] The knowledge base text in S1 includes titles, body paragraphs, list items, table cells, parenthesis descriptions, quotation mark descriptions, and cross-paragraph references;
[0013] The entity annotation categories in S1 include at least one of the following: term entity, equipment entity, parameter entity, operation entity, fault entity, and numerical entity.
[0014] The relationship labeling categories in S1 include at least one of the following: definition relationship, alias relationship, containment relationship, application relationship, reference relationship, and causal relationship.
[0015] Furthermore, the method for performing structure restoration and word-level alignment on the knowledge base text in S2 includes the following steps:
[0016] S20101, Identify structural blocks from knowledge base text. Structural blocks include heading blocks, body blocks, list item blocks, table cell blocks, comment blocks, and reference blocks.
[0017] S20102, Generate a structure path for each structure block. The structure path includes page number, heading level, paragraph number, list item number, and table cell number. The structure path is used to represent the current position in the document's hierarchy.
[0018] S20103, Perform unified word segmentation on the knowledge base text to obtain a unified word sequence. Unified word sequence This represents the set of all word positions arranged in the document reading order;
[0019] S20104 is a unified word sequence. A word-level structure index table is built for each word position in the table. Word-level structure index table Record the following attributes for each word position: character start position, character end position, structural path, page number, horizontal coordinate, vertical coordinate, font type, and whether it is a bold heading.
[0020] Furthermore, the method for constructing the structure-enhancing sequence in S2 includes the following steps:
[0021] S20201 is a unified word sequence. Each word position in the vector generates a basic word vector;
[0022] S20202, based on the word-level structure index table Constructing the structural connection matrix ;
[0023] S20203 generates a layout enhancement vector for each word position in the unified word sequence T. The layout enhancement vector indicates whether the word position has visual emphasis features such as title, bold, italic, monospace font, table header, quotation, or special alignment.
[0024] S20204, using a structural connection matrix Perform two rounds of structural aggregation on the basic word vectors to obtain a structurally enhanced word sequence. Structurally enhanced word sequence Each word position in the text simultaneously contains literal semantics, structural path information, and layout emphasis information.
[0025] Furthermore, the method for marking structurally difficult sections in S2 includes the following steps:
[0026] S20301, Structural Enhancement Word Sequence For each word position, a structural neighborhood is extracted. The structural neighborhood represents the set of word positions that have a strong structural connection with the current position.
[0027] S20302, calculate the deviation between each word position and its structural neighborhood center: first, take the average of the structural enhancement representations within the structural neighborhood to obtain the neighborhood center vector of the current position, and then calculate the L1 distance between the structural enhancement representation of the current position and the neighborhood center vector;
[0028] S20303, Count the boundary anomalies at each word position. , Indicates the first The number of structural anomalous events occurring near the position of each word;
[0029] S20304, Calculate the structural anomaly score based on the degree of deviation and boundary anomaly count. ;
[0030] S20305, Structural anomaly scores for all word positions along a unified word sequence Directional smoothing using moving average;
[0031] S20306, Convert candidate structurally difficult example segments into structurally difficult example segment masks. Output the structure-enhanced word sequence. and structurally difficult example segment mask .
[0032] Furthermore, the method for constructing and training the structure-priority dual-channel knowledge extraction model in S3 includes the following steps:
[0033] S301 captures the linear dependencies of word sequences through the sequential context encoding branch and explicitly models the non-linear connections in the document structure through the structural neighborhood aggregation branch. The outputs of both are used as parallel inputs to subsequent modules, thereby achieving comprehensive encoding of the semantic and structural information of the knowledge base text.
[0034] S30101, structurally enhanced word sequence Input the sequential context encoding branch, which is used to model the local context and cross-sentence context that unfolds in the reading order;
[0035] S30102, structurally enhanced word sequence and structural connection matrix Input the structural neighborhood aggregation branch, which aggregates the related information of the same header block, the same list item, and the same table region. Output the structural aggregation feature sequence. ;
[0036] S302, using structurally difficult example segment masks The outputs of the sequential context encoding branch and the structural neighborhood aggregation branch are differentially fused, and entity boundary localization is completed. The contribution ratio of sequential features and structural features is dynamically adjusted according to the structural stability of the text region.
[0037] S30201, for each word position According to the structurally difficult segment mask Determine the fusion coefficient ;
[0038] S30202, based on the fusion coefficient For sequential context feature sequences and structural aggregation feature sequences Perform weighted fusion to obtain the fused feature sequence. ;
[0039] S30203, from the word-level structure index table Extract boundary cue features and fuse them with the feature sequence. Common input entity label prediction layer;
[0040] S303, filter candidate entity pairs based on structural proximity and perform relation classification;
[0041] S304 improves the training contribution of structurally complex samples by weighting the samples, thus completing the model training.
[0042] Furthermore, the method in S303 that first filters candidate entity pairs based on structural proximity and then performs relation classification includes the following steps:
[0043] S30301, From entity recognition results Extract all entity spans, and for each entity span, read its starting word position features, ending word position features, and average features within the span, and concatenate them to obtain the entity representation vector;
[0044] S30302, Construct candidate entity pairs;
[0045] S30303, construct relation classification features for each candidate entity pair. The relation classification features include at least one of the following: subject entity representation, object entity representation, relative distance between the two entities, structural path similarity, whether they are the same title block, whether they are the same list item, and whether they are a table header to cell mapping.
[0046] S30304: Input the relationship classification features into the relationship classification module, and output the relationship category prediction result. .
[0047] Furthermore, the method in S304 that improves the training contribution of structurally complex samples through sample weighting to complete model training includes the following steps:
[0048] S30401 uses the negative log-likelihood loss of the conditional random field as the entity recognition loss to supervise entity label sequence learning, and uses the cross-entropy loss as the relation classification loss to supervise relation category learning.
[0049] S30402 introduces class weights into the relationship classification loss. The fewer the class samples, the greater the loss weight, thus avoiding the model from only favoring high-frequency relationship classes.
[0050] S30403, the entity recognition loss and relationship classification loss are weighted and summed to form the total loss, and a gradient-based optimization algorithm is used to update all trainable parameters of the structure-first dual-channel knowledge extraction model;
[0051] S30404: After each training cycle, the entity-level F1 score and relation-level F1 score are calculated on the validation set, and the model parameters with the highest comprehensive index are saved as the final deployment model.
[0052] Furthermore, the method for data augmentation assisted by the large language model in S4 is as follows:
[0053] S40101, Select seed samples from manually labeled samples to form a seed label set;
[0054] S40102: Select text fragments to be expanded from unannotated knowledge base text. Prioritize text fragments containing new terms, structurally difficult examples, densely tabulated areas, and multi-level list areas to improve the effectiveness of new training samples.
[0055] S40103 takes the text fragment to be expanded, the corresponding structural path, the entity category to be extracted, and the relation category as input and sends them to the large language model;
[0056] S40104 performs rule-based and manual verification on the annotation results returned by the large language model;
[0057] S40105, the expanded training set and the original training set are used together for training the structure-first dual-channel knowledge extraction model;
[0058] The method for correcting low-confidence difficult cases in S4 is as follows:
[0059] S40201, First, the structure-first dual-channel knowledge extraction model performs entity recognition and relation classification on the current knowledge base text to obtain preliminary extraction results;
[0060] S40202, calculate the extraction confidence score for each candidate segment, and the extraction confidence score shall include at least one of the following: average entity confidence score, average relation confidence score, and proportion of structurally difficult segment;
[0061] S40203 triggers the call to the large language model only for low-confidence difficult example segments. The content sent to the large language model includes the original text fragment, context window, structural path summary, layout prompt information, and preliminary entity and relation results given by the structure-first dual-channel knowledge extraction model.
[0062] S40204 Performs a consistency check on the correction results returned by the large language model;
[0063] S40205 adds the correction results of low-confidence difficult cases that have passed manual sampling to the incremental sample pool, and performs incremental fine-tuning on the structure-first dual-channel knowledge extraction model according to a predetermined period.
[0064] Furthermore, the method for online extraction of knowledge base text in S5 includes the following steps:
[0065] S501, Receive new original knowledge base text ;
[0066] S502, regarding knowledge base text Perform the same structured preprocessing as step S2 to obtain a unified word sequence. Word-level structure index table Structural enhancement word sequence and structurally difficult example segment mask ;
[0067] S503, structurally enhanced word sequence and structurally difficult example segment mask The input structure-first dual-channel knowledge extraction model generates sequential context feature sequences in parallel through a sequential context encoding branch and a structural neighborhood aggregation branch. and structural aggregation feature sequences Differential fusion is performed based on the structurally difficult segment mask to complete entity recognition and relationship classification, and preliminary extraction results are obtained;
[0068] S504 calculates the segment-level confidence level for the preliminary extraction results. If the segment does not belong to the low-confidence difficult example segment, the result is output directly. If the segment belongs to the low-confidence difficult example segment, the large language model is called only to correct the segment, and the consistency check is performed on the correction result.
[0069] S505 organizes the final extraction results into a structured knowledge result set. ;
[0070] S506, structured knowledge result sets Write it into a knowledge graph, knowledge index, or retrieval enhancement database for subsequent retrieval, question answering, review, and data governance.
[0071] The beneficial effects of the present invention are as follows: (1) The present invention constructs a structural connection matrix based on the physical layout and hierarchical path of the document, and explicitly encodes the non-linear positional relationship between the header and cell, and between the title and the body into inter-word connection weights, so that the model can directly use the typesetting logic to perform semantic aggregation.
[0072] (2) In the preprocessing stage, the present invention introduces a structural anomaly score calculation and difficult case segment mask generation mechanism. By quantifying the degree of local format disturbance such as bracket pairing, list jump, and table breakage, unreliable areas caused by OCR or transcription are marked in advance.
[0073] (3) The present invention adopts a structure-first dual-channel parallel coding architecture, which includes sequential context branch and structural neighborhood aggregation branch, and dynamically adjusts the fusion coefficient according to the mask of difficult example section, and automatically increases the structural feature weight in structurally chaotic areas to resist noise interference.
[0074] (4) The present invention adopts a trigger-based correction strategy for large language models only for low-confidence difficult example segments, strictly limiting the scope of large model calls to structural fragments where model prediction is unreliable, and establishing a closed loop of correction result feedback and incremental fine-tuning, which controls computational overhead and response latency while ensuring overall extraction accuracy. Attached Figure Description
[0075] Figure 1 This is a complete flowchart of an embodiment of the intelligent method for automated knowledge extraction and classification based on a large model knowledge base according to the present invention.
[0076] Figure 2 This is a flowchart illustrating the process of constructing and labeling text training samples for a knowledge base.
[0077] Figure 3 This is a flowchart of the process for structured preprocessing of knowledge base text.
[0078] Figure 4 This is a flowchart of the process for building and training a structure-priority dual-channel knowledge extraction model. Detailed Implementation
[0079] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0080] like Figure 1 As shown, the intelligent agent method for automated knowledge extraction and classification based on a large model knowledge base includes the following steps:
[0081] S1, Knowledge Base Text Training Sample Construction and Annotation: Construct a knowledge base text sample set for training and validation. Entity and relation annotations are manually performed on the training samples, such as... Figure 2 As shown.
[0082] Raw knowledge base text was collected from various sources. This text included titles, body paragraphs, list items, table cells, parentheses, quotation marks, and cross-paragraph citations. The original formatting and structural information were preserved during the sample construction phase. The knowledge base text sample set originated from product documentation, user manuals, Q&A documents, standard specification documents, interface documentation, maintenance records, and scanned documents transcribed using optical character recognition (OCR). All documents were uniformly converted to a parsable format. The raw knowledge base text is denoted as... , This represents the complete text of a single document sample and its corresponding structural information. For DOCX, HTML, and Markdown formats, the document structure is read directly; for PDF format, text block coordinates, font size, indentation, and table border information are read; for scanned images or scanned PDFs, the text content is first obtained through optical character recognition, and then the coordinate information of the characters on the page is preserved.
[0083] For each original knowledge base text, character encoding standardization, full-width / half-width character standardization, abnormal whitespace cleanup, and line break merging are performed. Line break merging is not simply deleting line breaks; it considers indentation, punctuation, heading styles, and table boundaries to determine whether the current position should retain the paragraph break. For example, if two lines have the same font size, the same left margin, and the previous line does not end with a paragraph end character, the two lines are merged; if the next line has a deeper indentation or the previous line is a heading, the boundaries are retained.
[0084] Entity annotation categories must include at least one of the following: terminology entities, equipment entities, parameter entities, operation entities, fault entities, and numerical entities. Relationship annotation categories must include at least one of the following: definition relationship, alias relationship, inclusion relationship, application relationship, reference relationship, and causal relationship. Annotation results are represented by both character start and end positions and category labels to ensure strict alignment with the original text.
[0085] In this embodiment, as an example, for the original text "Device parameters: Rated voltage is 220V.", its annotation result may be as follows:
[0086] Entity 1: Text content "rated voltage", entity category "parameter entity", character start position 7, end position 10;
[0087] Entity 2: Text content "220V", entity category "numeric entity", character start position 12, end position 15;
[0088] Relationship 1: Subject entity "Rated Voltage", Object entity "220V", Relationship type "Attribute Relationship".
[0089] All samples are divided into training set, validation set and test set. The total number of samples is denoted as N, which represents the number of data samples used in modeling. Each sample retains the original knowledge base text, manual annotation results and document format information.
[0090] S2, perform structured preprocessing on the knowledge base text: perform structure restoration, word-level alignment, structure-enhanced sequence construction, and structurally difficult segment labeling on the knowledge base text to obtain structure-enhanced word sequences. Difficult-to-structure section mask ,like Figure 3 As shown.
[0091] Knowledge base text often suffers from inconsistent formatting, misaligned table transcriptions, missing list indentation, and incomplete bracket pairing. If the original knowledge base text is directly input into a neural network model, structural noise may be mistaken for semantic features, leading to instability in entity recognition and relationship classification.
[0092] Entities and relationships in knowledge base text often depend on structural position rather than just word order. For example, terms in titles are usually defined objects in the following text, table headers and cell content typically form attribute relationships, and the first half and second half of a list item typically form operational constraint relationships. This embodiment first restores the document structure and then establishes a one-to-one correspondence between word positions and structural positions. The method for performing structural restoration and word-level alignment on knowledge base text in this embodiment includes the following steps:
[0093] S20101 identifies structural blocks from knowledge base text. Structural blocks include heading blocks, body blocks, list item blocks, table cell blocks, comment blocks, and reference blocks.
[0094] In the actual implementation, for DOCX, HTML and Markdown formats, the hierarchical structure is read directly; for PDF and optical character recognition results, the structural blocks are restored based on the text box coordinates, font size, font weight, line spacing, indentation, separator lines and table grid.
[0095] In this embodiment, as an example, in a PDF page, the system detects that text box A has a font of "Hei" and a font size of 16, is located at the top of the page and is horizontally centered, and below it are multiple left-aligned text boxes B, C, etc., with a font size of 10.5. Based on these characteristics, text box A can be deduced to be a "heading block", and text boxes B, C, etc., can be deduced to be "body blocks" belonging to that heading.
[0096] S20102 generates a structure path for each structure block. The structure path includes page number, heading level, paragraph number, list item number, and table cell number. The structure path is used to represent the current position in the document's hierarchy.
[0097] In practice, if a word is located on page 2, in the third paragraph under the first-level heading, its structural path is recorded as "2-1-3"; if a word is located in the cell of the second row and fourth column of the table, it is recorded as "2-1-3-r2c4" on the basis of the above.
[0098] In this embodiment, as an example, the first title on page 3 of the document is "2. Technical Specifications". The word "operating frequency" in the second paragraph under this title can be recorded as: 3-2-1-2, representing "page 3 - second first-level heading - first unnumbered paragraph - second sentence". If "operating frequency" is located in the first row and third column of the first table on this page, its structural path can be expanded to: 3-2-1-2-r1c3.
[0099] S20103, Perform unified word segmentation on the knowledge base text to obtain a unified word sequence. Unified word sequence This represents the set of all word positions arranged in the document reading order, with a length of [length value missing]. .
[0100] In the specific implementation, the code is first segmented by structural blocks, and then further segmented into words within the structural blocks based on punctuation, spaces, and linguistic features. For English abbreviations, hyphens in terminology, abbreviations within parentheses, and numerical expressions with units, a second normalization is performed to avoid erroneous segmentation. For example, "convolutional neural network (CNN)" is preserved as four consecutive word positions: "convolutional neural network", "(", "CNN", ")", instead of being broken down into a meaningless sequence of characters.
[0101] S20104 is a unified word sequence. A word-level structure index table is built for each word position in the table. Word-level structure index table Record the starting and ending positions of characters, structural path, page number, horizontal coordinate, vertical coordinate, font type, and whether it is a bold heading for each word position. Any subsequent entity or relation result can be traced back to its specific position in the original text and original layout.
[0102] In this embodiment, as an example, if the original knowledge base text contains the following content: "1. Installation requirements: The equipment should be grounded. 2. Operational requirements: The voltage should be stable.", it may be transcribed into two consecutive lines of text after optical character recognition. This step will recover two list item blocks based on "1.", "2.", indentation position, and colon pattern, and assign different list item numbers to "Installation requirements" and "Operational requirements" respectively, thereby avoiding the subsequent erroneous merging of the two requirements into a single continuous entity.
[0103] A uniform word sequence alone is insufficient to fully express the structural relationships between "explanatory content under the same title", "cell values corresponding to the same table header", and "short sentences before and after the same list item". This embodiment utilizes a word-level structure index table to construct a structural connection matrix and layout enhancement information, thereby generating a structure-enhanced word sequence. The method for constructing structure-enhancing sequences in this embodiment includes the following steps:
[0104] S20201 is a unified word sequence. Each word position in the algorithm generates a base word vector, which is obtained through a publicly available pre-trained language model encoder, word embedding table, or word-word hybrid encoder.
[0105] In this embodiment, as an example, the base word vectors can be obtained by loading a publicly available Chinese BERT pre-trained model. Specifically, this involves using a unified word sequence... Words (such as "rated" and "voltage") are input into the model, and the last hidden state is taken as the 768-dimensional vector representation of the word. If computational overhead needs to be reduced, a linear layer can be added after this vector to map its dimension to 256 or 384. In a convenient implementation, the basic word vector dimension can be 256.
[0106] S20202, based on the word-level structure index table Constructing the structural connection matrix , structural connection matrix Indicates the strength of the structural association between word positions, with a size of The first in the matrix Line 1 The elements of the column represent the first... The position of the word and the first Structural connection weights between word positions.
[0107] In practice, the structural connection weights are determined as follows: when two words are in the same sentence, a basic connection is assigned; when two words are in the same list item, the connection weight is increased; when two words are in the same table cell or between the table header and the current cell, the connection weight is increased; when two words are in adjacent paragraphs within the same heading block, the connection weight is moderately increased; when two words cross unrelated heading blocks or unrelated table areas, the connection weight is decreased.
[0108] In one implementation, the weight of a connection within the same sentence can be 1.0; the weight of a list item within the same list can be 0.3; the weight of a cell within the same table can be 0.4; the weight from the table header to the cell can be 0.5; the weight of a block within the same heading can be 0.2; and the penalty for crossing irrelevant blocks can be 0.5. After weighting, all connection weights emitted from each word position are normalized to keep the total propagation strength at different word positions within a similar range.
[0109] In this embodiment, as an example, for the phrase "2. Operational requirements: Voltage should be stable," the words "operation" and "voltage" are located in the same list item. The structural connection weight between them is calculated as: base connection 1.0 (same sentence) + list item append 0.3 = 1.3. The normalization process divides this weight by the sum of the connection weights emitted by the word "operation" to all its neighbors, so that the final weight value used for aggregation is within a reasonable range, for example, between 0.1 and 0.8.
[0110] S20203 generates a layout enhancement vector for each word position in the unified word sequence T. The layout enhancement vector indicates whether the word position has visual emphasis features such as title, bold, italic, monospace font, table header, quotation, or special alignment.
[0111] In practice, page numbers, vertical position, horizontal position, heading level, bold text, and table header text are encoded into fixed-length vectors, which are then concatenated or added to the basic word vectors.
[0112] In this embodiment, as an example, the layout enhancement vector generation process for the word "attention" is as follows:
[0113] a> Encode the bold mark as a Boolean value (e.g., 1), the heading level as 0, the table header mark as 0, the page number as 1, and normalize the horizontal / vertical coordinates to the range of 0-1;
[0114] b> These discrete and continuous values are processed through a small embedding network or directly concatenated to form a vector, for example, 16 dimensions;
[0115] Finally, this 16-dimensional vector is concatenated with the basic word vector (e.g., 256-dimensional) to obtain the final input vector of 272 dimensions; or the 16-dimensional vector is mapped to 256 dimensions through a linear layer and then added to the basic word vector.
[0116] S20204, using a structural connection matrix Perform two rounds of structural aggregation on the basic word vectors to obtain a structurally enhanced word sequence. Structurally enhanced word sequence Each word position in the text simultaneously contains literal semantics, structural path information, and layout emphasis information, with a size of [size missing]. , Indicates the number of word positions after unification. Indicates word-level dimension.
[0117] In practice, the first round aggregates information from the same sentence, the same list item, and the same table cell; the second round aggregates information from adjacent structural blocks within the same heading block. After each round of aggregation, the aggregation results are combined with the layout enhancement vector to form a new word-level representation. Based on this, local structural relationships and paragraph-level structural relationships can be preserved without excessively expanding the scope of information dissemination.
[0118] In this embodiment, as an example, during the first round of aggregation, for the term "rated power" in the header cell, it aggregates information from neighboring terms such as "(kW)" within the same cell to obtain an intermediate representation that incorporates unit information. During the second round of aggregation, this intermediate representation of "rated power" further aggregates information from the body paragraphs under the same heading block for "rated power," thereby making the term representation in the header more similar to the term representation in the body explanation in the vector space.
[0119] In this embodiment, as an example, when the term "rated voltage" appears in both the bold table header and the body explanation paragraph, this step will connect them through two types of structures: "table header-cell" and "title-body text," so that the representation of "rated voltage" in these two positions is close to each other, and the bold table header is superimposed to enhance the layout, making it easier for the subsequent model to recognize it as a parameter entity rather than an ordinary noun.
[0120] In knowledge base text, areas with disordered formatting, broken tables, unclosed brackets, and missing list hierarchy are more prone to extraction errors. If subsequent models treat these areas the same as ordinary areas, unstable results can easily occur in boundary localization and relationship classification. This embodiment pre-labels structurally difficult sections by calculating local structural anomalies. The method for labeling structurally difficult sections in this embodiment includes the following steps:
[0121] S20301, Structural Enhancement Word Sequence For each word position, a structural neighborhood is extracted. The structural neighborhood represents the set of word positions that have a strong structural connection with the current position.
[0122] In practical implementation, the structure connection matrix can be used. Word positions with a connection weight greater than 0.2 are considered as the structural neighborhood of the current word position, thus preserving the main structural neighborhood and filtering out weak noise connections.
[0123] In this embodiment, as an example, in a row of the structural connectivity matrix at a certain word position, the connection weights with the other 5 words are 0.8, 0.6, 0.5, 0.1, and 0.05, respectively. Only the three word positions with weights of 0.8, 0.6, and 0.5 are considered as the structural neighborhood of that word.
[0124] S20302, calculate the degree of deviation between each word position and its structural neighborhood center: first, take the average of the structural augmentation representations within the structural neighborhood to obtain the neighborhood center vector of the current position, and then calculate the L1 distance between the structural augmentation representation of the current position and the neighborhood center vector. The larger the L1 distance, the more inconsistent the current position is with the surrounding structural semantics.
[0125] In this embodiment, as an example, the feature vector of a normal text word is very similar to its neighborhood center vector, and its L1 distance may be 0.5. However, the feature vector of a word inserted due to OCR misalignment (such as "2") differs greatly from the neighborhood center vector of the surrounding normal text, and its L1 distance may be as high as 5.0.
[0126] S20303, Count the boundary anomalies at each word position. , Indicates the first The number of structural anomalies occurring near the position of each word.
[0127] In practical implementation, structural anomalies include at least one of the following: mismatched parentheses, mismatched quotation marks, sudden jumps in list hierarchy, interruption of table row and column boundaries, broken inline numbering, and abnormal line breaks caused by optical character recognition. For example, if consecutive parentheses appear before and after the current position without corresponding commas, an additional boundary anomaly count is added near that position; if the same list jumps directly from "1." to "3.", an additional boundary anomaly count is also added.
[0128] S20304, Calculate the structural anomaly score based on the degree of deviation and boundary anomaly count. The calculation method is expressed as follows:
[0129]
[0130] In the above formula, Indicates the first Structural anomaly score for each word position Indicates the first The word position in the structural enhancement word sequence The structure enhancement representation in the middle, Indicates the first The structural neighborhood center vector of each word position Indicates the L1 distance. This represents the boundary anomaly amplification factor, which can range from 0.5 to 1.5. Indicates the first Boundary anomaly count for each word position.
[0131] Based on this, by combining the degree of deviation of feature vectors (L1 distance) and boundary anomaly count, it is possible to identify both the abnormal locations of "sudden semantic deviation" and the abnormal locations of "sudden format break".
[0132] S20305, Structural anomaly scores for all word positions along a unified word sequence The direction is smoothed by sliding average, and the window length can be 3 to 5 to avoid misjudgment caused by a single punctuation mark or a single recognition error.
[0133] The distribution of structural anomaly scores in the current document is statistically analyzed to determine a dynamic threshold, which can be taken as the 85th percentile of the current document's structural anomaly scores. Segments that are continuously above the dynamic threshold and are at least 3 word positions in length are marked as candidate structurally difficult segments, and then extended to the left and right by 1 word position to 2 word positions to preserve necessary context.
[0134] In this embodiment, as an example, the 85th percentile of the structural anomaly score distribution calculated for a document is 2.8. The system marks the segment with anomaly scores [3.1, 4.5, 3.0] for three consecutive word positions, all of which are higher than 2.8, as a candidate structural difficulty segment, and expands it to the left and right by one word position each, finally obtaining a structural difficulty segment of length 5.
[0135] S20306, Convert candidate structurally difficult example segments into structurally difficult example segment masks. Output the structure-enhanced word sequence. and structurally difficult example segment mask , Indicates and One-to-one structural hard example segment mask, length is The value of a structurally difficult example is 1, and the value of a non-structurally difficult example is 0.
[0136] In this embodiment, as an example, a document has 100 word positions (L=100) after word segmentation. Word positions 30 to 34 are marked as structurally difficult segments. The structurally difficult segment mask is then used. It is a vector of length 100, where the values at indices 29 to 33 (assuming counting starts from 0) are 1, and the values at the remaining 95 positions are 0.
[0137] In another embodiment, as an example, the original list of a scanned document, “(1) Check before startup; (2) Check after startup”, is transcribed into “1 Check before startup 2) Check after startup” after optical character recognition. At this time, the list hierarchy information and bracket pairing information before and after “2)” are broken, the boundary anomaly count at this position will increase, and the L1 deviation will also increase due to inconsistency with the surrounding structural neighborhood. Finally, this area is marked as a structurally difficult section.
[0138] It should be noted that ordinary regions are more suitable for sequential semantic modeling, while structurally difficult regions are more suitable for increasing the weight of structural information. This step does not rely on complex trainable parameters; its main purpose is to distinguish between "structurally stable regions" and "structurally unstable regions" in advance. After pre-labeling structurally difficult regions, subsequent models can adopt differentiated fusion strategies, thereby improving stability.
[0139] S3, Constructing and Training a Structure-First Dual-Channel Knowledge Extraction Model: Knowledge base texts have both a linear reading order and obvious structural dependencies. Using only sequential encoding easily overlooks the structural relationships between table headers and cells, titles and body text, and sentences before and after list items; using only structural aggregation easily overlooks local semantic changes in continuous context.
[0140] The structure-first dual-channel knowledge extraction model in this embodiment is used for entity recognition, relation classification, and label prediction. The model includes a sequential context encoding branch, a structural neighborhood aggregation branch, an entity boundary localization module, and a relation classification module. This model independently performs fast extraction on the vast majority of documents, only calling a large language model for auxiliary processing in step S4 for a small number of low-confidence hard examples. Figure 4 As shown.
[0141] S301, In this embodiment, the linear dependencies of word sequences are captured by the sequential context encoding branch, and the non-linear connections in the document structure are explicitly modeled by the structural neighborhood aggregation branch. The outputs of both are used as parallel inputs to subsequent modules to achieve comprehensive encoding of the semantic and structural information of the knowledge base text.
[0142] S30101, structurally enhanced word sequence The input is a sequential context encoding branch, which is used to model local contexts and cross-sentence contexts that unfold in the order of reading.
[0143] In the implementation, the sequential context encoding branch uses a two-layer bidirectional gated recurrent unit network, with each layer having a hidden dimension of 256. After encoding, the sequential context feature sequence is obtained. , This represents the context feature sequence obtained by encoding according to the unified word sequence order, with a size of [size missing]. , This indicates the hidden feature dimension.
[0144] S30102, structurally enhanced word sequence and structural connection matrix Input the structural neighborhood aggregation branch, which aggregates the related information of the same header block, the same list item, and the same table region. Output the structural aggregation feature sequence. , This represents a sequence of structural features obtained by aggregation based on the structural connectivity matrix, with a size of [missing information]. .
[0145] In practice, the structural neighborhood aggregation branch is implemented by processing each word position... Read its structure connection matrix The main neighbor word position in The current location features, neighbor location features, structural connection weights between the two, and layout enhancement information of the current location are concatenated and input into the gating calculation unit to obtain the neighbor transfer coefficient between 0 and 1. Then, the neighbor transfer coefficient is used to perform a weighted summation of the neighbor features. Finally, residual fusion is performed with the current location's own features. This process is repeated twice to absorb both the nearest neighbor structural information and the title block-level structural information.
[0146] In one implementation, if the current position is a term in a bold heading, its layout enhancement information is high, and information from that position is more likely to propagate to the body text explanation section; if the current position is literally similar to a neighboring position but crosses an unrelated heading block, its propagation coefficient will also be low due to the lower structural connection weight.
[0147] It should be noted that the sequential context encoding branch is responsible for capturing the natural language sequence information, while the structural neighborhood aggregation branch is responsible for recovering the structural connections unique to the knowledge base text. Through parallel encoding of the two branches, the model will neither lose continuous semantics nor ignore key hints carried in the typography and structure.
[0148] S302, using structurally difficult example segment masks The outputs of the sequential context encoding branch and the structural neighborhood aggregation branch are differentially fused, and entity boundary localization is completed. The contribution ratio of sequential features and structural features is dynamically adjusted according to the structural stability of the text region.
[0149] Different regions rely on sequential and structural information to varying degrees. For structurally stable ordinary text regions, sequential context is usually sufficient; for structurally complex regions such as misaligned tables, broken lists, and heading references, structural information is more important.
[0150] S30201, for each word position According to the structurally difficult segment mask Determine the fusion coefficient .
[0151] In this embodiment, the ordinary position A value of 0.35 can be taken, which is suitable for structurally difficult locations. A value of 0.7 is acceptable. For ordinary positions, the focus is mainly on sequential context features, while for structurally difficult positions, the proportion of structural aggregation features should be increased.
[0152] S30202, based on the fusion coefficient For sequential context feature sequences and structural aggregation feature sequences Perform weighted fusion to obtain the fused feature sequence. The fusion process is represented as:
[0153]
[0154] In the above formula, Represents the fusion feature sequence The Middle Fusion features of word positions, Represents sequential context feature sequences The Middle Features of word position Represents structural aggregation feature sequences The Middle Features of word position Indicates the first The fusion coefficients for each word position. When the current position belongs to a structurally difficult segment, the model automatically increases the weight of structural features and reduces format noise interference.
[0155] S30203, from the word-level structure index table Extract boundary cue features and fuse them with the feature sequence. Common input entity label prediction layer.
[0156] In practical implementation, boundary hint features include at least one of the following: whether it is located at the boundary of parentheses, whether it is located at the boundary of quotation marks, whether there is a switch from bold to non-bold text, whether it is switching from a title to body text, whether it is switching from a table header to a cell, and whether it is located at the beginning of a list item. Boundary hint features are used to help the model identify the start and end positions of entities.
[0157] In its implementation, the entity labeling prediction layer uses the BIOES annotation system. In BIOES, B represents the entity start position, I represents the internal position of the entity, O represents a non-entity position, E represents the entity end position, and S represents the word entity position. The entity labeling prediction layer first outputs the emission score of each word position for each label. Then, the conditional random field decoder, combined with label transition constraints, calculates the globally optimal label sequence to obtain the entity recognition result. , This represents the set of entity recognition results for the current document. Each entity result includes the start word position, end word position, entity text, and entity category.
[0158] In this embodiment, as an example, "convolutional neural network" in the phrase "convolutional neural network (CNN)" is usually a term entity, while "CNN" in parentheses is usually an alias entity or annotation content. This step combines the boundary cue feature between "network" and "()" to make the conditional random field more inclined to end the main entity boundary at "network", avoiding incorrectly incorporating the left parenthesis and subsequent abbreviations into the main entity.
[0159] It should be noted that in order to maintain high efficiency in stable regions and improve noise resistance in complex regions, the final technical effect is more stable entity boundaries and less sensitive to format anomalies. This step does not use a fixed fusion method for all positions, but dynamically adjusts the proportion of sequential features and structural features according to the mask of structurally difficult example segments.
[0160] S303: First, candidate entity pairs are selected based on structural proximity, and then relation classification is performed.
[0161] After entity recognition is completed, it is also necessary to identify the semantic relationships between entities. If all entities are classified in pairs, the number of candidates is too large and there is a lot of noise.
[0162] S30301, From entity recognition results Extract all entity spans. For each entity span, read its starting word position features, ending word position features, and average features within the span, and concatenate them to obtain the entity representation vector.
[0163] In this embodiment, as an example, for the identified entity "rated voltage", its span is from the 10th to the 11th word position, and its starting word feature is taken from... The end word feature is taken from The average characteristic within the span is By concatenating these three vectors end to end, we obtain the final "rated voltage" entity representation vector.
[0164] S30302, construct candidate entity pairs: the two entities are in the same sentence, or in the same list item, or in the same table row, or satisfy the correspondence of "header entity - cell entity", or are in the same title block and the word position distance does not exceed 80. For entity pairs that cross irrelevant title blocks and whose structural connection weight is lower than the preset threshold, they are directly filtered and do not enter the relationship classification step.
[0165] In this embodiment, as an example, entity A "Installation Steps" (located under heading 1.1) and entity B "Click to Start" (located under heading 3.2) are directly filtered out because they are located under different irrelevant heading blocks and the word position distance exceeds 80, and their structural connection weight is very low. This avoids the waste of computing resources and potential noisy relationships.
[0166] S30303, construct relation classification features for each candidate entity pair. Relationship classification features include at least one of the following: subject entity representation, object entity representation, relative distance between the two entities, structural path similarity, whether they are the same title block, whether they are the same list item, and whether they are a header-to-cell mapping. Structural path similarity is used to reflect the proximity of the two entities in the document structure.
[0167] In this embodiment, as an example, the entities "input voltage" (located at 2-1-1-r1c1) and "220V" (located at 2-1-1-r1c3) have structural paths of 2-1-1-r1c1 and 2-1-1-r1c3, respectively. The calculated structural path similarity will be very high (because they share most of the path prefixes). This high similarity will serve as an important feature input to the relationship classifier, strongly suggesting that there is an "attribute value" relationship between the two entities.
[0168] S30304: Input the relationship classification features into the relationship classification module, and output the relationship category prediction result. .in, This represents the set of relation classification results for the current document. Each relation result includes at least a subject entity, an object entity, and a relation category. In a convenient implementation approach, the relation classification module can employ a 2-layer fully connected network with Softmax output.
[0169] In this embodiment, as an example, if there is a parameter entity "rated voltage" in the table header and a numerical entity "220V" in the corresponding cell, then since the two satisfy the "header entity - cell entity" relationship and are located in the same column mapping, the entity pair will be retained as a candidate entity pair; the relationship classification module will further identify it as a relationship of "attribute value" or "applicable parameter", and will not mistakenly connect it with the content of adjacent unrelated cells.
[0170] S304 improves the training contribution of structurally complex samples by weighting the samples, thus completing the model training.
[0171] When trained using only ordinary supervised loss, the contributions of structurally complex samples and structurally simple samples are similar, which can easily lead the model to be biased towards learning ordinary text and not adapt well to structurally difficult segments.
[0172] S30401 uses the negative log-likelihood loss of conditional random fields as the entity recognition loss to supervise entity label sequence learning; and uses cross-entropy loss as the relation classification loss to supervise relation category learning.
[0173] Based on the structural difficulty section mask Calculate sample weights for each sample The calculation method is expressed as follows:
[0174]
[0175] In the above formula, This represents the training weights of the current sample. This represents the structural difficulty amplification factor, which can be taken from 1.0 to 3.0. This represents the total number of word positions in the current sample. Indicates the first Does the position of each word belong to a structurally difficult segment?
[0176] Based on this, samples with a higher proportion of structurally difficult segments receive a greater loss weight during training.
[0177] S30402 addresses the issue of imbalanced relation categories by introducing category weights into the relation classification loss. The fewer the number of samples in a category, the greater its loss weight, thus preventing the model from favoring only high-frequency relation categories.
[0178] Before training begins, the frequency of all relation categories in the training set is counted. For a rare relation category with a low frequency, its weight is set to the logarithm of the total number of samples divided by the number of samples in that category; that is, the lower the frequency of a category, the greater its weight. When calculating the loss function, this category weight is multiplied by the loss value of the corresponding relation sample. This results in a greater penalty for the model when it mispredicts rare relation categories, thus forcing the model to pay more attention to the feature learning of these minority categories.
[0179] S30403 calculates the total loss by weighted summing of entity recognition loss and relation classification loss, and then uses a gradient-based optimization algorithm to update all trainable parameters of the structure-first dual-channel knowledge extraction model. In one implementation, the optimizer can be AdamW, and the initial learning rate can be set to... The size can be 8 to 16, and the training period can be 20 to 80.
[0180] S30404: After each training cycle, the entity-level F1 score and relation-level F1 score are calculated on the validation set, and the model parameters with the highest comprehensive index are saved as the final deployment model.
[0181] It should be noted that this step adopts a more practical sample weighting training strategy. By directly increasing the training weights of structurally complex samples, the model's adaptability to structurally difficult segments can be improved without significantly increasing training complexity.
[0182] S4, Data augmentation and low-confidence hard example correction assisted by the large language model: The structure-first dual-channel knowledge extraction model undertakes the main work of entity recognition, relation classification and label prediction, while the large language model is only positioned as an auxiliary tool. The functions of the large language model are: automatically generating high-quality labeled data for unlabeled knowledge base text, processing hard example segments with low confidence in the structure-first dual-channel knowledge extraction model, improving the model's generalization ability on various text styles and entity types, and controlling computational cost and response latency while ensuring extraction quality.
[0183] The data augmentation method assisted by the large language model in this embodiment is as follows: by utilizing the understanding and generation capabilities of the large language model, starting from a small number of manually labeled seed samples, high-quality labeled data is automatically generated for a large number of unlabeled knowledge base texts, thereby expanding the training set at a lower cost and improving the model's generalization ability on various text styles and entity types.
[0184] S40101. Select a small number of seed samples from the manually annotated samples to form a seed annotation set. The seed annotation set covers common title styles, common table styles, common list styles and common entity categories to ensure that the prompt examples are representative.
[0185] S40102: Select text fragments to be expanded from unannotated knowledge base text. Prioritize text fragments containing new terms, structurally difficult examples, densely tabulated areas, and multi-level list areas to improve the effectiveness of new training samples.
[0186] S40103 takes the text fragment to be expanded, the corresponding structural path, the entity category to be extracted, and the relation category as input and sends them to the large language model. The large language model is required to output the annotation results according to fixed fields. The output fields include the entity start character position, the entity end character position, the entity category, the relation subject, the relation object, and the relation category.
[0187] S40104 performs rule-based and manual verification on the annotation results returned by the large language model. Rule verification includes: the start and end positions of entities must fall within the range of the original text segment, the subject and object of relations must point to the output entities, and the entity category and relation category must belong to the predefined set. After passing the rule verification, manual sampling or key verification is performed, and the correct samples are added to the expanded training set.
[0188] S40105 allows the expanded training set to be used together with the original training set for training or fine-tuning the structure-first dual-channel knowledge extraction model.
[0189] In one implementation, the output of the large language model is in JSON format. For example, the output fields include start, end, entity_type, head, tail, and relation_type. Using a fixed structure for output facilitates automatic parsing, batch verification, and subsequent manual review.
[0190] In this embodiment, as an example, there is an unlabeled text: "[Environmental Requirements]: Operating temperature: 0℃-40℃; Storage temperature: -20℃-65℃." The system sends this text, its structural path (e.g., identified as a list item block), and the entity types (device entities, parameter entities, numerical entities) and relation types (attribute relations) to be extracted to the large language model.
[0191] The large language model returns a JSON-formatted annotation result: [
[0193] {"start":0","end":3","entity_type":"Parameter Entity"}, / / "Operating Temperature"
[0194] {"start":5,"end":8,"entity_type":"Numerical Entity"}, / / "0℃"
[0195] {"start":9,"end":12,"entity_type":"Numerical Entity"}, / / "40℃"
[0196] / / ...other entities
[0197] {"head":"operating temperature", "tail":"0℃", "relation_type":"attribute relationship"},
[0198] / / ...other relationships ]
[0200] After program verification (all start and end positions are within the original text, and the subjects and objects of the relations are the listed entities) and manual spot checks, this batch of data was added to the expanded training set.
[0201] The monitoring model's output confidence level is used to perform secondary corrections only on a few low-confidence segments where the model's predictions are unreliable. The corrected data is then stored in a sample database for subsequent incremental fine-tuning of the model. This allows for continuous optimization of the model's weaknesses without significantly increasing inference costs. The method for correcting low-confidence difficult examples in this embodiment is as follows:
[0202] S40201, the structure-first dual-channel knowledge extraction model first performs entity recognition and relation classification on the current knowledge base text to obtain preliminary extraction results.
[0203] S40202, calculate the extraction confidence score for each candidate segment. The extraction confidence score includes at least one of the following: average entity confidence score, average relation confidence score, and proportion of structurally difficult segments. In one implementation, when the average entity confidence score of a segment is lower than 0.65, or the average relation confidence score is lower than 0.6, or the proportion of structurally difficult segments is higher than 0.25, the segment is determined to be a low-confidence difficult segment.
[0204] S40203 triggers the call to the large language model only for low-confidence difficult example segments. The content sent to the large language model includes the original text fragment, context window, structural path summary, layout prompt information, and preliminary entity and relation results given by the structure-first dual-channel knowledge extraction model. The large language model does not extract from scratch, but performs corrections based on existing candidate results.
[0205] S40204 Performs consistency check on the correction results returned by the large language model. If the corrected entity span does not fall within the original text range, or the subject and object of the relation do not correspond to the actual entity, the correction result is discarded. If the correction result passes the check, it is used as the final output result of the current segment.
[0206] S40205 adds the corrected results of low-confidence difficult cases that have passed manual sampling to the incremental sample pool, and performs incremental fine-tuning on the structure-first dual-channel knowledge extraction model according to a predetermined period. The predetermined period can be set weekly or monthly, or triggered by the cumulative number of corrected samples.
[0207] It should be noted that the large language model is good at handling new terms, cross-sentence reasoning and extreme structural anomalies, but the calling cost is high; the structure-first dual-channel knowledge extraction model is good at fast and stable extraction and is suitable for undertaking the main process. Therefore, the large language model in this step does not undertake large-scale real-time extraction tasks, but only handles a small number of truly difficult segments. The combination of the two as needed can significantly reduce the overall computing power consumption and reasoning latency while ensuring accuracy.
[0208] In this embodiment, as an example, during inference, the structure-first dual-channel knowledge extraction model extracts text from a section of text where table lines are distorted due to scanning tilt, obtaining the entities "rated voltage" and "220V". However, the confidence level for predicting that the two are "attribute relationships" is only 0.55, lower than the threshold of 0.6. The system determines this section as a low-confidence difficult example section. Subsequently, the system only sends this small section of text, its structural path "located in the second row of the table", and the model's preliminary extraction result ("possibly an attribute relationship of rated voltage - 220V") to the large language model, requesting correction. The large language model, combining context and table structure information, confirms the "attribute relationship" of "rated voltage - 220V" and corrects another entity boundary that was misidentified by the model. After verifying the validity of the correction result, the system finally outputs the corrected result and stores the corrected example in the incremental sample pool.
[0209] S5, Online Knowledge Base Text Extraction: Deploy the structure-first dual-channel knowledge extraction model to a knowledge base management system, document governance system, or enterprise knowledge platform to automatically extract newly arriving knowledge base text.
[0210] The method for online text extraction from the knowledge base in S5 includes the following steps:
[0211] S501, Receive new original knowledge base text , This represents the new document text to be extracted and its original structural information, and the original knowledge base text. The sources are newly uploaded documents, scanned document transcription results, system synchronized documents, or documents scraped from web pages.
[0212] S502, regarding knowledge base text Perform the same structured preprocessing as step S2 to obtain a unified word sequence. Word-level structure index table Structural enhancement word sequence and structurally difficult example segment mask .
[0213] S503, structurally enhanced word sequence and structurally difficult example segment mask The input structure-first dual-channel knowledge extraction model generates sequential context feature sequences in parallel through a sequential context encoding branch and a structural neighborhood aggregation branch. and structural aggregation feature sequences Differential fusion is performed based on the structurally difficult segment mask to complete entity recognition and relationship classification, and preliminary extraction results are obtained.
[0214] S504 calculates the segment-level confidence level for the preliminary extraction results. If the segment does not belong to the low-confidence difficult example segment, the result is output directly. If the segment belongs to the low-confidence difficult example segment, the large language model is called only to correct the segment, and the consistency check is performed on the correction result.
[0215] S505 organizes the final extraction results into a structured knowledge result set. , This represents the final set of structured knowledge results for the current document. Each result includes entity text, entity category, relation category, subject entity, object entity, character start and end positions, and structural path.
[0216] S506, structured knowledge result sets Write it into a knowledge graph, knowledge index, or retrieval enhancement database for subsequent retrieval, question answering, review, and data governance.
[0217] In this embodiment, as an example, for a device instruction document containing a "Parameter Description" title, a parameter table, and a "Precautions" list, the system first restores the title block, table cell block, and list item block. Then, the structure-first dual-channel knowledge extraction model identifies the parameter entity "Rated Power," the numerical entity "2.5kW," and the operational entity "Maintenance after Power Off." Specifically, "Rated Power - 2.5kW" in the table is identified as an attribute relationship, and "Maintenance after Power Off" in the list is identified as operational constraint information. If a table boundary break in a scanned page leads to a decrease in extraction confidence, only the corresponding section of that page is corrected using the large language model, without affecting the rapid processing of other pages.
[0218] Based on this, the structure-first dual-channel knowledge extraction model handles the main tasks of entity recognition, relation classification, and label prediction during system operation, while the large language model intervenes only as needed in two stages: training data augmentation and low-confidence difficult example correction. This significantly reduces overall computational cost and inference latency while ensuring extraction accuracy, and allows the model to continuously absorb new terms, formats, and difficult examples during ongoing use, gradually improving its ability to extract from complex knowledge base texts.
[0219] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention.
Claims
1. A large model-based knowledge base automated knowledge extraction and classification agent method, characterized in that, Includes the following steps: S1, Knowledge Base Text Training Sample Construction and Annotation: Construct a knowledge base text sample set for training and validation, and manually perform entity and relation annotation on the training samples; S2, perform structured preprocessing on the knowledge base text: perform structure restoration, word-level alignment, structure-enhanced sequence construction, and structurally difficult segment labeling on the knowledge base text to obtain structure-enhanced word sequences. Difficult-to-structure section mask ; S3, Constructing and training a structure-first dual-channel knowledge extraction model: The structure-first dual-channel knowledge extraction model includes a sequential context encoding branch, a structural neighborhood aggregation branch, an entity boundary localization module, and a relation classification module, which are used for entity recognition, relation classification, and label prediction; S4, Large Language Model-Assisted Data Expansion and Low-Confidence Difficult Example Correction: Automatically generates high-quality labeled data for unlabeled knowledge base text, processes low-confidence difficult example sections in the structure-first dual-channel knowledge extraction model, and improves the model's generalization ability across various text styles and entity types. S5, Online Knowledge Base Text Extraction: Deploy the structure-first dual-channel knowledge extraction model to a knowledge base management system, document governance system, or enterprise knowledge platform to automatically extract newly arriving knowledge base text.
2. The method for automated knowledge extraction and classification intelligent agents based on large-scale model knowledge bases according to claim 1, characterized in that: The knowledge base text sample set in S1 comes from product specification documents, operation manuals, Q&A documents, standard specification documents, interface specification documents, operation and maintenance record documents, and scanned documents obtained by optical character recognition transcription. The knowledge base text in S1 includes titles, body paragraphs, list items, table cells, parenthesis descriptions, quotation mark descriptions, and cross-paragraph references; The entity annotation categories in S1 include at least one of the following: term entity, equipment entity, parameter entity, operation entity, fault entity, and numerical entity. The relationship labeling categories in S1 include at least one of the following: definition relationship, alias relationship, containment relationship, application relationship, reference relationship, and causal relationship.
3. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for performing structure restoration and word-level alignment on knowledge base text in S2 includes the following steps: S20101, Identify structural blocks from knowledge base text. Structural blocks include heading blocks, body blocks, list item blocks, table cell blocks, comment blocks, and reference blocks. S20102, Generate a structure path for each structure block. The structure path includes page number, heading level, paragraph number, list item number, and table cell number. The structure path is used to represent the current position in the document's hierarchy. S20103, Perform unified word segmentation on the knowledge base text to obtain a unified word sequence. Unified word sequence This represents the set of all word positions arranged in the document reading order; S20104 is a unified word sequence. A word-level structure index table is built for each word position in the table. Word-level structure index table Record the following attributes for each word position: character start position, character end position, structural path, page number, horizontal coordinate, vertical coordinate, font type, and whether it is a bold heading.
4. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for constructing the structure-enhancing sequence in S2 includes the following steps: S20201 is a unified word sequence. Each word position in the vector generates a basic word vector; S20202, based on the word-level structure index table Constructing the structural connection matrix ; S20203 generates a layout enhancement vector for each word position in the unified word sequence T. The layout enhancement vector indicates whether the word position has visual emphasis features such as title, bold, italic, monospace font, table header, quotation, or special alignment. S20204, using a structural connection matrix Perform two rounds of structural aggregation on the basic word vectors to obtain a structurally enhanced word sequence. Structurally enhanced word sequence Each word position in the text simultaneously contains literal semantics, structural path information, and layout emphasis information.
5. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for marking structurally difficult sections in S2 includes the following steps: S20301, Structural Enhancement Word Sequence For each word position, a structural neighborhood is extracted. The structural neighborhood represents the set of word positions that have a strong structural connection with the current position. S20302, calculate the deviation between each word position and its structural neighborhood center: first, take the average of the structural enhancement representations within the structural neighborhood to obtain the neighborhood center vector of the current position, and then calculate the L1 distance between the structural enhancement representation of the current position and the neighborhood center vector; S20303, Count the boundary anomalies at each word position. , Indicates the first The number of structural anomalous events occurring near the position of each word; S20304, Calculate the structural anomaly score based on the degree of deviation and boundary anomaly count. ; S20305, Structural anomaly scores for all word positions along a unified word sequence Directional smoothing using moving average; S20306, Convert candidate structurally difficult example segments into structurally difficult example segment masks. Output the structure-enhanced word sequence. and structurally difficult example segment mask .
6. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for constructing and training a structure-priority dual-channel knowledge extraction model in S3 includes the following steps: S301 captures the linear dependencies of word sequences through the sequential context encoding branch and explicitly models the non-linear connections in the document structure through the structural neighborhood aggregation branch. The outputs of both are used as parallel inputs to subsequent modules, thereby achieving comprehensive encoding of the semantic and structural information of the knowledge base text. S30101, structurally enhanced word sequence Input the sequential context encoding branch, which is used to model the local context and cross-sentence context that unfolds in the reading order; S30102, structurally enhanced word sequence and structural connection matrix Input the structural neighborhood aggregation branch, which aggregates the related information of the same header block, the same list item, and the same table region. Output the structural aggregation feature sequence. ; S302, using structurally difficult example segment masks The outputs of the sequential context encoding branch and the structural neighborhood aggregation branch are differentially fused, and entity boundary localization is completed. The contribution ratio of sequential features and structural features is dynamically adjusted according to the structural stability of the text region. S30201, for each word position According to the structurally difficult segment mask Determine the fusion coefficient ; S30202, based on the fusion coefficient For sequential context feature sequences and structural aggregation feature sequences Perform weighted fusion to obtain the fused feature sequence. ; S30203, from the word-level structure index table Extract boundary cue features and fuse them with the feature sequence. Common input entity label prediction layer; S303, filter candidate entity pairs based on structural proximity and perform relation classification; S304 improves the training contribution of structurally complex samples by weighting the samples, thus completing the model training.
7. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 6, characterized in that, The method for screening candidate entity pairs based on structural proximity and performing relation classification in S303 includes the following steps: S30301, From entity recognition results Extract all entity spans, and for each entity span, read its starting word position features, ending word position features, and average features within the span, and concatenate them to obtain the entity representation vector; S30302, Construct candidate entity pairs; S30303, construct relation classification features for each candidate entity pair. The relation classification features include at least one of the following: subject entity representation, object entity representation, relative distance between the two entities, structural path similarity, whether they are the same title block, whether they are the same list item, and whether they are a table header to cell mapping. S30304: Input the relationship classification features into the relationship classification module, and output the relationship category prediction result. .
8. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 6, characterized in that, The method described in S304, which improves the training contribution of structurally complex samples through sample weighting to complete model training, includes the following steps: S30401 uses the negative log-likelihood loss of the conditional random field as the entity recognition loss to supervise entity label sequence learning, and uses the cross-entropy loss as the relation classification loss to supervise relation category learning. S30402 introduces class weights into the relationship classification loss. The fewer the class samples, the greater the loss weight, thus avoiding the model from only favoring high-frequency relationship classes. S30403, the entity recognition loss and relationship classification loss are weighted and summed to form the total loss, and a gradient-based optimization algorithm is used to update all trainable parameters of the structure-first dual-channel knowledge extraction model; S30404: After each training cycle, the entity-level F1 score and relation-level F1 score are calculated on the validation set, and the model parameters with the highest comprehensive index are saved as the final deployment model.
9. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for data augmentation assisted by the large language model in S4 is as follows: S40101, Select seed samples from manually labeled samples to form a seed label set; S40102: Select text fragments to be expanded from unannotated knowledge base text. Prioritize text fragments containing new terms, structurally difficult examples, densely tabulated areas, and multi-level list areas to improve the effectiveness of new training samples. S40103 takes the text fragment to be expanded, the corresponding structural path, the entity category to be extracted, and the relation category as input and sends them to the large language model; S40104 performs rule-based and manual verification on the annotation results returned by the large language model; S40105, the expanded training set and the original training set are used together for training the structure-first dual-channel knowledge extraction model; The method for correcting low-confidence difficult cases in S4 is as follows: S40201, First, the structure-first dual-channel knowledge extraction model performs entity recognition and relation classification on the current knowledge base text to obtain preliminary extraction results; S40202, calculate the extraction confidence score for each candidate segment, and the extraction confidence score shall include at least one of the following: average entity confidence score, average relation confidence score, and proportion of structurally difficult segment; S40203 triggers the call to the large language model only for low-confidence difficult example segments. The content sent to the large language model includes the original text fragment, context window, structural path summary, layout prompt information, and preliminary entity and relation results given by the structure-first dual-channel knowledge extraction model. S40204 Performs a consistency check on the correction results returned by the large language model; S40205 adds the correction results of low-confidence difficult cases that have passed manual sampling to the incremental sample pool, and performs incremental fine-tuning on the structure-first dual-channel knowledge extraction model according to a predetermined period.
10. The method for automated knowledge extraction and classification intelligent agents based on a large model knowledge base according to claim 1, characterized in that, The method for online text extraction from the knowledge base in S5 includes the following steps: S501, Receive new original knowledge base text ; S502, regarding knowledge base text Perform the same structured preprocessing as step S2 to obtain a unified word sequence. Word-level structure index table Structural enhancement word sequence and structurally difficult example segment mask ; S503, structurally enhanced word sequence and structurally difficult example segment mask The input structure-first dual-channel knowledge extraction model generates sequential context feature sequences in parallel through a sequential context encoding branch and a structural neighborhood aggregation branch. and structural aggregation feature sequences Differential fusion is performed based on the structurally difficult segment mask to complete entity recognition and relationship classification, and preliminary extraction results are obtained; S504 calculates the segment-level confidence level for the preliminary extraction results. If the segment does not belong to the low-confidence difficult example segment, the result is output directly. If the segment belongs to the low-confidence difficult example segment, the large language model is called only to correct the segment, and the consistency check is performed on the correction result. S505 organizes the final extraction results into a structured knowledge result set. ; S506, structured knowledge result sets Write it into a knowledge graph, knowledge index, or retrieval enhancement database for subsequent retrieval, question answering, review, and data governance.
Citation Information
Patent Citations
Government affair service field multi-strategy fusion dialogue method based on knowledge graph
CN116628172A
Electric power knowledge graph construction method based on text enhancement and dual-channel triple extraction
CN118504678A