Method for improving private domain operation and maintenance knowledge retrieval quality based on RAG
Through the RAG-based method, OCR and spatial clustering are used to identify adjacent layout content in the operation and maintenance document, and a cross-document reference relationship diagram is constructed in combination with a large language model, which solves the problems of visual information integration and context loss in the operation and maintenance document, and improves the accuracy and efficiency of operation and maintenance knowledge retrieval.
Patent Information
- Application Number
- CN202510782876.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively integrate visual information and cross-document knowledge in operation and maintenance documents, resulting in semantic breakdown or context loss. The traditional method has low matching rate when dealing with multi-column text, chart mixing and semantic association, which affects operation and maintenance decision efficiency.
Using RAG-based method, the adjacent content is identified through OCR technology and spatial clustering algorithm, and semantic similarity analysis is performed in combination with large language models, cross-document citation relationship diagram is constructed, vector index format is output, and operation and maintenance knowledge retrieval is improved.
It significantly improves the accuracy and practicality of operation and maintenance knowledge retrieval, solves the problems of semantic breakage and context loss in traditional methods, and enhances the semantic understanding and retrieval ability of operation and maintenance scenarios.
Smart Images

Figure CN120277201A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to a method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG. Background Art
[0002] In the field of operation and maintenance document processing, existing systems mainly target complex format materials containing mixed charts and texts, multi-column layouts, and unstructured texts. Such documents usually mix structured data (such as log tables) with unstructured descriptions (such as alarm descriptions), and the content highly depends on the visual layout to convey logical associations. Traditional processing solutions usually adopt step-by-step structured cleaning and keyword retrieval technologies. However, in the face of the unique format diversity (such as multi-column texts and chart interspersions) and semantic association requirements (such as service dependency chains and change tracks) in the operation and maintenance scenarios, it is difficult for existing technologies to effectively integrate scattered visual information and cross-document knowledge, resulting in the risk of semantic breaks or context missing in subsequent analysis.
[0003] Traditional methods mostly rely on regular expressions or simple rules to filter format symbols, but lack the intelligent recognition ability for adjacent text blocks in the layout, resulting in the splitting of key semantic elements (such as table headers and content), or the residue of irrelevant format symbols (such as line breaks and decorative lines), affecting the accuracy of the DOM tree framework.
[0004] Fixed-granularity text cutting (such as segmenting by the number of characters) easily destroys the integrity of core elements such as alarm thresholds and service dependency chains, and the traditional semantic tag system is not customized for the operation and maintenance field, resulting in a low matching rate of professional terms and insufficient context relevance.
[0005] The isolated retrieval mechanism cannot concatenate causal chains or change tracks scattered in different documents. When processing table data, due to the lack of optimization for header alignment and cell merging, data ambiguities often occur due to row and column misalignment.
[0006] Traditional fine-tuning methods either lead to a decline in generalization ability by overly modifying the parameters of the pre-trained model, or have low knowledge injection efficiency and poor clustering effects for similar data. The processing of multi-version logs still relies on manual deduplication, and redundant data directly affects the efficiency of operation and maintenance decisions. Therefore, those skilled in the art provide a method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG to solve the problems raised in the above background art. Summary of the Invention
[0007] The present invention provides the following technical solution: A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG, including the following stages: S1. Document preprocessing and structure cleaning stage: Preprocess the original document, remove non-semantic content in the document, retain elements useful for the semantic structure, and generate a basic DOM tree structure; S2. Visual-Driven Semantic Block Recognition Phase: Using OCR technology and spatial clustering algorithms, based on visual features such as the position coordinates, font attributes, font size, line spacing, etc. of the text, merge adjacent and structurally continuous content on the layout into visual blocks to ensure the integrity and continuity of semantics; S3. Semantic Slicing and Conflict Fusion Phase: Perform semantic subdivision within the visual blocks, conduct semantic similarity analysis through large language models, and combine with the sliding window mechanism to generate knowledge blocks with better semantic granularity, solving the problem of semantic fragmentation during the chunking process; S4. Private Domain Data Structure Recognition Phase: Identify specific structural units in operation and maintenance knowledge, such as large tables, special table structures, etc., and assign semantic labels to enhance the model's understanding ability of the professional field; S5. Contextual Semantic Tracking Modeling Phase: Construct the reference relationships between cross-document and multi-round knowledge blocks, calculate the contextual similarity of adjacent blocks through semantic vectors, and establish a reference relationship graph to improve the discrimination and accuracy of retrieval results; S6. Metadata Binding and Structured Output Phase: Output the content, attributes, semantics, and structure information of all blocks in a format friendly to vector indexing, provide an interface to support the invocation of the vectorization engine, and facilitate subsequent retrieval and generation; Purpose: Output the content, attributes, structure, and semantic relationship information of each semantic block in a structured format friendly to vector retrieval, especially adapted to the semantic tracking requirements of a large number of instruction blocks, exception blocks, and multi-round records in IT operation and maintenance knowledge.
[0008] Specific implementation: Each semantic block outputs the following fields to meet the requirements of structural integrity and operation and maintenance semantic description ability: Text: The original content of the semantic block, such as "systemctl restart app.service"; VisualRegion: The visual position coordinates [x0, y0, x1, y1] obtained by OCR, used to enhance context reasoning; DocumentPosition: The path of the document to which it belongs, such as / doc1 / section3 / step1, facilitating the construction of the document semantic tree; BlockType: The structural type label, such as "exception description", "command block", "repair suggestion"; ParentID / ChildID: Support for multi-granularity nested structure organization of parent and child blocks; ReferenceLinks: Block IDs that reference historical faults or preprocessing steps, in the form of: "ReferenceLinks": [{ "refers_to": "B1.2", "type": "historical processing", "relevance":0.89}].
[0009] Detailed description: Command Block Marker (CommandBlock): Such as Text: "ps -aux | grep java", marked as BlockType: "operation command", which will not participate in the semantic vector later and is only used for auxiliary generation; Metric Block Structure (MetricBlock): Such as Text: "CPU usage has reached 98%", the metric type and value need to be marked, and the additional fields can be: "Metric": { "Name": "CPU usage", "Value": "98%", "Threshold": "85%"} Semantic reference chain tracking: Supports context inheritance in multiple rounds of processing records. For example, a certain "repair suggestion" block is referenced from "historical exception B2.1", forming a fault evolution graph.
[0010] Output format example: { "BlockID": "B3.1", "Text": "systemctl restart app.service", "VisualRegion": [120, 340, 400, 380], "DocumentPosition": " / doc3 / section2 / step1", "BlockType": "operation command", "ParentID": "B3", "ReferenceLinks": { "refers_to": "B2.3", "type": "historical processing", "relevance": 0.91} } S7. Table enhancement processing stage: Perform header alignment and merged cell sinking optimization processing on table data to ensure that the table header and content are always processed as a whole during the retrieval process, improving the quality of table Q&A; S8. Knowledge Transfer Injection Phase: Copy the intermediate Transformer layer with domain knowledge in the large language model and insert it into the same position of the embedding model. Fine-tune the copied layer to achieve knowledge injection and improve the model's performance in domain tasks. S9. Model Training and Optimization Phase: Only train the injected Transformer layer, freeze the original model layer, adopt an offset learning rate, and jointly use TripletLoss and InBatchContrastiveLoss for training to optimize the model performance. The "offset learning rate" in the present invention is essentially a differential learning rate setting strategy based on parameter categories, similar to Layer-wise LR Decay and Adapter Tuning with Custom LR in the prior art. The implementation method is easy to understand, and the name is specific to the present invention.
[0011] The technical solution is as follows: To improve the convergence efficiency of the pluggable Transformer injection layer in the fine-tuning stage and avoid excessive perturbation of the backbone model, the following "offset learning rate" strategy is adopted: Main purpose: Keep the backbone model frozen or update it at a low speed; Set a relatively high independent learning rate for the newly injected Transformer structure (such as Adapter, LoRA module, injection layer); Speed up the learning speed of the new module to ensure that the semantic ability quickly aligns with the original model distribution.
[0012] Technical implementation: 1. Parameter grouping definition When initializing the model, divide the model parameters into two categories according to their functions: base_params: Backbone model parameters (such as the original Transformer layer) injected_params: Newly injected layer parameters (such as the new Transformer sublayer inserted in the Qwen structure).
[0013] 2. Learning rate offset strategy setting When constructing the optimizer, set different learning rate values for different parameter groups: optimizer = AdamW( {"params": base_params, "lr": 1e-5}, {"params": injected_params, "lr": 5e-4} ) Note: In the injection layer, a learning rate significantly higher than that of the backbone layer (such as 10 - 50 times) is used to form an "offset learning rate gradient". 3. Freeze the backbone parameters (optional operation) If the goal is to completely maintain the stability of the backbone, the backbone parameters can be set to requires_grad=False to be completely frozen 4. Dynamic learning rate control (optional operation) Set a Warmup strategy or cosine annealing function for the injection layer to control the dynamic change of its learning rate scheduler = CosineAnnealingLR(optimizer, T_max=100) 5. Pseudocode reference # Define the learning rate ratio for different layers def get_layer_lr(layer_idx, base_lr): if 6 <= layer_idx <= 8: # Injected Transformer layers return base_lr * 2.0 # Increase the learning rate else: return base_lr * 0.1 # Keep other layers at a low learning rate (almost frozen) # Apply to optimizer parameter groups optimizer = AdamW( {'params': model.encoder.layer[i].parameters(), 'lr': get_layer_lr(i, 1e-4)} for i in range(len(model.encoder.layer)) ) S10. Multi-document multi-version context merging and normalization stage: When the input is multi-round processing logs or multiple record versions of the same problem, aggregate the timelines, perform cross-document normalization mapping, and improve the system's data integration and processing capabilities.
[0014] Preferably, in the step S1, document preprocessing and structure cleaning stage, it also includes identifying and converting non-text elements such as pictures and charts in the document to ensure that all useful information is retained and structured.
[0015] Preferably, in the step S2, the visual-driven semantic block recognition stage, the combined use of the OCR engine and the clustering algorithm can more accurately identify the spatial position and semantic relationship of text blocks, laying a foundation for subsequent processing; at the same time, the "small-to-big" chunking method is adopted to provide references from sub-chunks to the parent chunk, retaining the integrity and continuity of semantics.
[0016] Preferably, in the step S3, the semantic slicing and conflict fusion stage, the semantic segmentation technology based on deep learning is adopted to perform more refined semantic partitioning inside visual blocks, while solving the conflict and fusion problems between different semantic blocks; in addition, the private domain knowledge document is used to construct triples (query, pos_doc, neg_doc), and the large language model is used to assist in generating negative samples, improving the generation efficiency and quality of training data.
[0017] Preferably, in the step S4, the private domain data structure recognition stage, a structure unit library unique to operation and maintenance knowledge is also established, facilitating quick identification and invocation during subsequent processing and retrieval; at the same time, for complex table headers containing merged cells, the multi-layer merged cells are sunk to a single-layer structure through algorithm optimization to achieve precise alignment of the table header and data columns.
[0018] Preferably, in the step S5, the context semantic tracking and modeling stage, technologies such as graph neural networks are adopted to model the reference relationships between cross-document and multi-round knowledge blocks, achieving more accurate semantic understanding and retrieval; at the same time, the table header information is dynamically associated with each row of data to ensure that the table header and content are always processed as a whole during the retrieval process.
[0019] Preferably, in the step S6, the metadata binding and structured output stage, rich metadata fields are also defined, such as author, creation time, modification record, etc., facilitating subsequent retrieval and management; at the same time, the middle Transformer layer in the large language model is copied and inserted into the embedding model to achieve knowledge injection through fine-tuning.
[0020] Preferably, in the step S7, the table enhancement processing stage, the table header alignment algorithm adopts technologies such as dynamic programming to ensure the accurate correspondence between the table header and content rows; the optimization of sinking merged cells adopts a hierarchical processing strategy, improving the accuracy of table structure parsing.
[0021] Preferably, in step S8, the knowledge conversion injection stage, techniques such as knowledge distillation are also used to more efficiently migrate the domain knowledge in the large language model to the embedding model, further improving the retrieval quality and efficiency of the model; at the same time, specific freezing strategies, weight sources, learning rate scheduling and loss function selection are adopted to ensure that the learning effect of the injection layer is better than that of the backbone model.
[0022] Preferably, in step S9, the model training and optimization stage, techniques such as early stopping and learning rate decay are also used to prevent overfitting of the model and improve the generalization ability of the model; in addition, the vectorization engine call is supported through the interface, which facilitates subsequent vector retrieval and semantic analysis.
[0023] Preferably, in step S10, the multi-document multi-version context merging and normalization stage, the system can automatically identify and aggregate timeline information from different documents or different versions of the same document, and integrate the context information scattered in multiple documents or versions into a unified and coherent semantic representation through cross-document normalization mapping technology, thereby significantly improving the system's integration capabilities and processing efficiency for multi-document and multi-version data.
[0024] In summary, compared with the prior art, the present invention provides a method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG. This method significantly improves the accuracy and practicality of private domain operation and maintenance knowledge retrieval through multi-stage collaborative optimization, and has the following beneficial effects: 1. At the beginning of document processing, the system first performs structural cleaning on the original data, removes irrelevant format symbols while retaining key semantic elements, and builds a hierarchical DOM tree framework to lay the foundation for subsequent analysis. By combining OCR technology with spatial clustering algorithms, the system can intelligently identify adjacent text blocks and integrate scattered visual information into logically coherent semantic units. It is particularly good at processing complex formats such as mixed chart layouts and multi-column documents that are common in operation and maintenance scenarios; 2. At the semantic parsing level, this method adopts a dynamic slicing mechanism, uses a large language model to analyze the internal associations of the text, and combines the sliding window technology to achieve adaptive adjustment of the semantic granularity, which not only avoids the semantic fracture caused by coarse-grained cutting, but also prevents the context loss caused by excessive segmentation. In view of the unique knowledge structure in the operation and maintenance field, the system uses a dedicated semantic label system to annotate core elements such as alarm thresholds and service dependency chains, which substantially improves the matching accuracy of field-related queries; 3. This method constructs a cross-document context tracking network, breaking through the island effect of traditional retrieval. By calculating the association strength of adjacent knowledge chunks through semantic vectors, it enables the connection of information such as causal chains and change trajectories scattered in different documents. For tabular data, the system implements header alignment and cell merging optimization to ensure that the header and content are always treated as a complete semantic unit during retrieval, effectively solving the common row and column misalignment problems in traditional solutions. 4. In the model optimization stage, this method adopts a lightweight knowledge injection strategy, migrating the intermediate layer parameters of the pre-trained model to the target model, achieving efficient adaptation of domain knowledge through offset learning rate control. By jointly using the TripletLoss and InBatchContrastiveLoss functions, while maintaining the generalization ability of the basic model, it significantly enhances the clustering effect of similar knowledge. When processing multi-version logs or duplicate records, the system can automatically aggregate timeline information and eliminate redundant data through cross-document normalization mapping, providing clear and reliable knowledge support for operation and maintenance decisions. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is the step diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment
[0027] Please refer to Figure 1 , the present invention provides a technical solution, a method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG, including the following steps: Step S1. Document preprocessing and structure cleaning Parse the input operation and maintenance documents in PDF, Word, or HTML format, filter out non-semantic content such as headers, footers, watermarks, and comments through regular expressions, and retain core elements such as the main text, tables, and pictures. Adopt a DOM tree construction algorithm based on BeautifulSoup to generate a basic semantic structure containing tags such as titles, paragraphs, and tables. For non-text elements such as pictures and flowcharts in the document, call the OCR engine for text recognition and embed the recognition results into the corresponding positions of the DOM tree, retaining the original coordinate information for subsequent visual analysis.
[0028] Step S2. Vision-driven semantic chunk recognition Deploy the improved FasterR-CNN object detection model, combine it with the EAST text detection algorithm, and extract the coordinates of the minimum bounding rectangle of the text region and font attributes. Use the density-based spatial clustering algorithm to cluster adjacent text blocks and generate candidate regions for visual blocks. Implement a hierarchical merging strategy to ensure semantic continuity by calculating the similarity between blocks, and finally form a set of semantic blocks containing spatial coordinates and visual features.
[0029] Step S3, Semantic Slicing and Conflict Fusion Apply a semantic segmentation network inside the visual block to generate a pixel-level semantic mask and achieve sub-sentence-level semantic partitioning. Construct a sliding window, calculate the semantic similarity within the window in combination with the BERT model, and trigger the splitting operation. Design a conflict resolution mechanism to calculate the similarity of adjacent splitting results. If it is lower than the threshold, achieve semantic transition by generating a connecting sentence.
[0030] Step S4, Private Domain Data Structure Recognition Establish a structural unit library in the operation and maintenance field, including complex table patterns, fault tree symbol templates, configuration instruction templates, etc. Adopt a hierarchical sinking algorithm for merged cells to achieve precise alignment of the table header and data columns through dynamic programming.
[0031] Step S5, Contextual Semantic Tracking Modeling Construct a cross-document reference graph, where nodes represent knowledge blocks and edge weights are semantic similarities. Use a graph neural network for message passing. Implement dynamic table header association, concatenate the table header vector of the first row of the table with the data vector of each row to generate an enhanced representation vector.
[0032] Step S6, Metadata Binding and Structured Output Define the metadata schema, including fields such as block ID, content, semantic label, spatial information, and structural relationship. The output is a Parquet format vector index, including columns for the original text, semantic vectors, and structural labels.
[0033] Step S7, Table Enhancement Processing Use the dynamic programming algorithm to optimize the table header alignment. After generating the alignment path matrix, backtrack the optimal alignment path.
[0034] Step S8, Knowledge Transformation Injection Copy the intermediate Transformer layer of the large language model to the lightweight model, keeping the dimension alignment. Implement two-stage fine-tuning. In the first stage, freeze all layers and only train the LayerNorm parameters. In the second stage, unfreeze the injection layer and use the knowledge distillation loss.
[0035] Step S9, Model Training and Optimization Construct a triple dataset, including positive sample pairs and negative sample pairs. The training configuration includes optimizer selection, learning rate strategy, loss function, etc. Techniques such as early stopping and learning rate decay are adopted to prevent the model from overfitting.
[0036] Step S10: Multi-document context merging Implement timeline aggregation, extract the document creation timestamps, and construct a temporal graph. Adopt community discovery algorithms to merge version clusters with close time intervals and similar contents. Construct a global entity linking graph to map the descriptions in different documents to a unified entity ID. Embodiment
[0037] System optimization and expansion plan: In step S3, the semantic segmentation network can be replaced with a more advanced architecture to achieve higher performance metrics on the operation and maintenance dataset. In step S8, an adapter can be adopted to replace the full-layer replication, reducing the number of parameters while maintaining stable performance.
[0038] Through the above implementation manners, in the test of a provincial power grid operation and maintenance knowledge base of the present invention, the retrieval accuracy rate is significantly improved, the quality of table question answering is greatly improved, and the multi-document merging efficiency reaches a relatively high level, which is significantly better than the traditional scheme.
[0039] The detailed description of this solution is as follows: I. Block enhancement process: Technical implementation steps: 1. Preprocessing and structural cleaning of the original document Objective: Remove the non-semantic content in the document and retain the elements useful for the semantic structure.
[0040] Specific implementation: Use regular expressions to match fixed templates, and delete the header (such as "company name, time"), footer (such as "page x"), footnotes, and page numbers; For complex documents such as PDF and Word, call third-party libraries (such as ApacheTika, pdfplumber) to extract the layout structure; On the basis of retaining structures such as headings, paragraphs, lists, and code blocks (pre / code), generate a basic DOM tree structure as the input for subsequent parsing.
[0041] 2. Visually-Driven preliminary semantic block recognition Objective: Merge the adjacent and structurally continuous content in the layout into visual blocks Specific implementation: Use an OCR engine (such as PaddleOCR, Tesseract) to identify the position coordinates (x, y), font attributes, font size, and line spacing of each paragraph of text in the document; Use KMeans or DBSCAN clustering algorithm to spatially cluster text blocks, with position, spacing, and alignment as clustering dimensions; Filter the clustering results to find the areas that may belong to the same semantic block and mark their bounding boxes. Output visual area list [BlockID, coordinate range, content paragraph ID].
[0042] Features of operation and maintenance documents: Strong structural fixity: It often adopts the abnormal performance-cause analysis-processing steps format, and can perform semi-supervised recognition based on structural templates.
[0043] Dense domain terminology: Contains a large number of industry abbreviations such as "JVM hang", "CPU soaring", "prometheus alarm", etc. The semantic model needs to be enhanced with industry word vectors or exclusive dictionaries.
[0044] Mixed layout of tables + commands: The text is often interspersed with Linux commands, configuration blocks, tables, and screenshots. OCR needs to enhance its perception of the mixed structure of code and charts.
[0045] Multi-version records / revision chains: The same issue may be recorded, supplemented, and evolved repeatedly, and a version chain or issue thread chain needs to be built to establish a semantic graph.
[0046] Innovative solution points: Structure-driven enhanced semantic classification block Introduce "processing unit template" (TemplateUnit): such as "[phenomenon identification segment]-[log sample segment]-[cause analysis segment]-[processing command segment]".
[0047] After OCR or structural cleaning, semantic segment classification is performed through rule matching or fine-tuning of classifiers (for example, based on the BiLSTM+CRF classification model).
[0048] Code snippet recognition and independent packaging blocks The detected "shell command segment" or "configuration segment" is separately identified as a substructure block, and is set not to participate in semantic vector calculation and is only used for context expansion.
[0049] Merge and normalize multi-document and multi-version contexts When the input is multiple rounds of processing logs or multiple record versions of the same issue, the system should first aggregate the timeline and normalize the graph across documents (establish a unified issue number).
[0050] Example of the mother - child block partitioning structure: The essence of the "mother - child block" mechanism is a multi - granularity information organization method - first coarsely divide large - segment semantics, and then finely divide local semantic points to ensure semantic consistency and vector extraction efficiency.
[0051] First, give an example of an operation and maintenance manual document: Abnormal phenomenon: The application server responds slowly, and the user access latency is higher than 3s.
[0052] Cause analysis: It is initially judged that the database connection pool is exhausted. By checking through the monitoring platform, it is found that the ActiveConnection is continuously at the maximum value.
[0053] Processing steps: 1. Log in to the application server and execute the following command: systemctlrestartapp.service 2. Check the status of the database connection pool: showprocesslist; Fixing suggestions: It is recommended to expand the upper limit of the connection pool from 100 to 300 and configure an automatic alarm policy.
[0054] How to partition: As shown in the following table, Table 1 is used to show the structured form of each "semantic main segment" extracted from the original document. Each mother block represents a clear knowledge semantic category, such as "abnormal phenomenon", "cause analysis", "processing steps", "fixing suggestions", etc., which are used to construct the backbone of the knowledge graph and form large semantic blocks.
[0055] Table 1: Definition table of mother blocks and semantic main structures
[0056] As shown in the following table, Table 2 shows the semantic sub - blocks formed after further disassembling the mother blocks. Each sub - block expresses specific operation commands, configuration suggestions, data queries, etc. Sub - blocks can be used for high - precision vectorization, FAQ construction, and the design of recall units in question - answering systems.
[0057] Table 2: Disassembly table of sub - blocks and fine - grained structures
[0058] 3. Semantic slicing and conflict boundary optimization Goal: Further perform semantic subdivision within the visual block to form knowledge blocks with better semantic granularity.
[0059] Specific implementation: For the text content of each visual block, use a language model to perform semantic similarity analysis to determine whether there is a topic jump (such as multiple semantics within a paragraph); Use a sliding window mechanism (such as a fixed window size + overlap strategy) to perform context sliding pairing on sentences, calculate the similarity of adjacent sentences, and break if it is below the threshold to generate sub-semantic blocks; When there is a conflict between the visual boundary and semantic segmentation, design a fusion strategy: Use the weighted formula "FinalScore = α * visual clustering confidence + β * semantic segmentation intensity" to determine whether further splitting or merging is required; Finally, generate a semantic block structure.
[0060] 4. Private domain structure template recognition Objective: Identify specific structural units in operation and maintenance knowledge and assign semantic labels.
[0061] Specific implementation: Construct regular + keyword recognition rules for identifying common structural segments, such as: Monitoring metric segment: including "CPU usage exceeds..." "Threshold is..."; Fault troubleshooting segment: starting with keywords such as "Possible reasons are as follows" "Please check..."; Operation instruction segment: starting with verb phrases such as "Execute command..." "Click..."; Construct a structure template matching model (supporting a rule engine or a trained classifier) to label the structural type (Type) of each semantic block; The labeling result serves as an important structural index for subsequent query enhancement.
[0062] 5. Context semantic tracking modeling Objective: Construct the reference relationship between cross-document and multi-round knowledge blocks.
[0063] Specific implementation: For a document set originating from multi-round work orders, group them using the "problem number", "processing record ID", and "timestamp" fields; Within the group, calculate the context similarity of adjacent blocks through semantic vectors (Sentence-BERT) to establish a reference relationship graph; Output the semantic tracking graph structure.
[0064] 6. Metadata binding and structured output Purpose: Output the content, attributes, semantics, and structure information of all blocks in a format friendly to vector indexing.
[0065] Specific implementation: Each semantic block outputs the following fields: Text VisualRegion DocumentPosition BlockType ParentID / ChildID ReferenceLinks Stored using JSON, Markdown, or structured tables; Provide an interface to support calls to vectorization engines (such as FAISS, Weaviate).
[0066] II. Table Enhancement: ① Align the table headers with the knowledge content of each row to solve the problem that the knowledge base sharding cannot cover the table headers and the knowledge content, resulting in incorrect retrieval.
[0067] Technical Solution Description: For the optimization of large-scale table data, we adopted a strategy of directly aligning the table header information with the in-row knowledge content, thus significantly improving the semantic understanding ability and query hit rate of the retrieval system. In practical applications, large-scale tables usually have complex header structures, and there may be a long distance between the table headers and the specific data content (for example, the character interval between the content pointed to by the question and the table header text exceeds 2000 characters). This physical distance separation will cause the knowledge base to be unable to cover both the table header information and the corresponding knowledge content during sharding storage, which will in turn lead to retrieval failures or inaccurate retrieval results, ultimately affecting the integrity and correctness of the question-answering system.
[0068] Through our optimization method, the system can dynamically associate the table header information with each row of data during the knowledge base construction phase, ensuring that the table header and the content are always processed as a whole during the retrieval process. This optimization not only solves the retrieval bottleneck caused by the separation of the table header and the content, but also significantly improves the effect of large-table knowledge question answering, enabling the system to more accurately understand user questions and return correct answers. As shown in the figure, this optimization method is particularly prominent in dealing with typical large-table retrieval scenarios, effectively overcoming the retrieval limitations caused by incomplete data sharding in traditional methods.
[0069] ② There are still some problems in the above solution. For the original columns of the merged cell table headers that are not equal or misaligned after conversion, the merged cells are restored to their pre-merged state and sunk to each row of content after the solution is optimized.
[0070] Technical Solution Description: When processing complex table headers with merged cells, before optimization, problems such as inconsistent number of columns or misaligned arrangements often occurred after conversion, resulting in inaccurate parsing of the table structure. To address this challenge, through innovative algorithm optimization, we successfully sank multi-layer merged cells (such as 2-layer merging) to a single-layer structure, thus achieving precise alignment between the table header and data columns.
[0071] Technical implementation principle: Sub-module 1: Alignment algorithm for large table headers Description: In a large table, the header information and the specific question-and-answer data rows are often separated by thousands of characters. Vectorized sharding is difficult to cover the entire header, resulting in the model being unable to correctly "understand the meaning of this row".
[0072] Technical implementation principle: Input: Original table (in formats such as HTML, Excel, JSON, etc.); Header area (usually the first row or the first n rows); Content area (each row of knowledge data); Processing flow: Header extraction: Judge whether the first row of the table is the field name (can be judged by bold font / alignment rules / consistency of the number of columns); If there are multiple rows of headers, adopt the rule: splice the field names from top to bottom (such as ["Index", "CPU"] → Index-CPU); Field expansion: For each row of data, use the header fields spliced in the previous step as keys and the corresponding cells in this row as values to form complete structure pairs; The output structure is as follows: {"Index-CPU": "95%", "Time": "12:00", "Server number": "S123"} Field mapping and merging: Add the fields to the context of each row as supplementary prompt information for vectorization or RAG prompt word generation; For example, convert it to a natural language prompt template: The server number is S123, and the CPU usage rate reaches 95% at 12:00 Output: Each row becomes a verbalizable fragment spliced with "field name + content"; Can be directly sent to the model for vectorization or used as a prompt to participate in the question and answer.
[0073] Sub-module 2: Alignment algorithm for large table headers Description: Merged cells are used in the table header to represent grouping levels. However, if the model does not perform sinking during reading, problems such as "inconsistent number of columns" or "missing fields" will occur.
[0074] Technical implementation principle: Input: There are multiple levels of merging in the table header area (e.g., 2 rows), and the number of columns in the data area is N The initial table structure is as follows:
[0075] Processing flow: Structure positioning: Determine which cells are rowspan / colspan; Build the original two-dimensional table matrix and identify the positions of missing fields (spaces or missing columns); Field sinking algorithm: For all header fields with merging, copy their values to the empty positions in the columns corresponding below; Form a unified single-layer field, such as: [Indicator - CPU Usage, Indicator - Memory, Time,] Column alignment restoration: Recode the original table header into a one-dimensional structure, which is consistent with the number of rows and columns of the data; Update the field index mapping relationship in the data structure to ensure that the field semantics correspond one-to-one with the column data.
[0076] Output: The corrected field structure [Indicator - CPU, Indicator - Memory, Time,]; All row data have a unified set of field names; It can be used for precise retrieval and question-answering matching.
[0077] III. Knowledge conversion injection: Technical solution description: Injection of pluggable Transformer layer structure into the embedding model Principle: Copy and insert the middle Transformer layers (such as the 6th - 9th layers) with domain knowledge in a large language model (such as Qwen2 - 7b) that is structurally compatible into the same position of the embedding model, and fine-tune the copied layers to achieve "knowledge injection".
[0078] Training parameter control and transfer fine-tuning strategy Freezing strategy: Only train the injected Transformer layers, and freeze the rest completely to avoid catastrophic forgetting; Weight source: The parameter initialization of the injection layer comes from the middle layer of the trained Qwen2-7b, not randomly initialized; Learning rate scheduling: Use an offset learning rate to ensure that the learning speed of the injection layer is better than that of the backbone; Loss function selection: Adopt joint training of TripletLoss + InBatchContrastiveLoss to strengthen semantic relevance modeling.
[0079] Training data generation in the private domain task scenario Use private domain knowledge document pairs (questions, work order descriptions, fault knowledge bases) to construct triples (query, pos_doc, neg_doc); Use a large language model to assist in generating negative samples to maintain the property of being semantically relevant but not the answer; Use an automatic labeling strategy to generate the training set to avoid labeling costs: Query: How to handle high CPU usage? PositiveDoc: Check the top command and analyze high CPU processes... NegativeDoc: JVM tuning suggestions: Set the size of the young generation... Task objective adaptation and optimization This solution aims to improve the vector index construction ability of the embedding model, rather than the generative question answering ability; The fine-tuning objective is to improve the Top-K recall accuracy, long document coverage rate, and marginal knowledge hit ability of private domain documents; During the recall process, cooperate with the document structure optimized by the "chunk enhancement module" to form a closed-loop optimization system. It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0080] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG, characterized in that It includes the following stages: S1. Document preprocessing and structure cleaning stage: Preprocess the original document, remove non-semantic content in the document, retain elements useful for the semantic structure, and generate a basic DOM tree structure; S2. Vision-driven semantic block recognition stage: Utilize OCR technology and spatial clustering algorithms, and according to the visual features of the text, merge adjacent and structurally continuous content in the layout into visual blocks; S3. Semantic slicing and conflict fusion stage: Conduct semantic subdivision within the visual blocks, perform semantic similarity analysis through a large language model, and combine with a sliding window mechanism to generate knowledge blocks with better semantic granularity; S4. Private domain data structure recognition stage: Identify specific structural units in the operation and maintenance knowledge and assign semantic labels; S5. Contextual semantic tracking and modeling stage: Construct reference relationships between cross-document and multi-round knowledge blocks, calculate the contextual similarity of adjacent blocks through semantic vectors, and establish a reference relationship graph; S6. Metadata binding and structured output stage: Output the content, attributes, semantics, and structure information of all blocks in a format friendly to vector indexing, and provide an interface to support the invocation of the vectorization engine; S7. Table enhancement processing stage: Perform header alignment and merged cell sinking optimization processing on the table data to ensure that the table header and content are always processed as a whole during the retrieval process; S8. Knowledge transformation injection stage: Copy and insert the intermediate Transformer layer with domain knowledge in the large language model into the same position of the embedding model, and fine-tune the copied layer; S9. Model training and optimization stage: Only train the injected Transformer layer, freeze the original model layer, adopt an offset learning rate, and jointly use TripletLoss and InBatchContrastiveLoss for training; S10. Multi-document and multi-version context merging and normalization stage: When the input is multi-round processing logs or multiple record versions of the same problem, aggregate the timeline and perform cross-document normalization mapping.
2. The method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, wherein: In the step S1, the document preprocessing and structure cleaning stage, it also includes identifying and converting non-text elements in the document to ensure that all useful information is retained and structured.
3. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S2, the vision-driven semantic block recognition stage, it also adopts a "small-to-big" chunking method to provide references from sub-blocks to the parent block, retaining semantic integrity and continuity.
4. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S3, the semantic slicing and conflict fusion stage, it adopts deep learning-based semantic segmentation technology to perform semantic partitioning within the visual blocks, solve the conflict and fusion problems between different semantic blocks; construct triples using private domain knowledge documents and use a large language model to assist in generating negative samples.
5. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S4, the private domain data structure recognition stage, it also establishes a library of structural units unique to operation and maintenance knowledge to facilitate quick identification and invocation during subsequent processing and retrieval; for complex table headers containing merged cells, optimize the algorithm to sink multi-layer merged cells to a single-layer structure.
6. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S5, the context semantic tracking and modeling stage, the graph neural network technology is adopted to model the reference relationships between cross-document and multi-round knowledge blocks; Dynamically associate the header information with each row of data to ensure that the header and the content are always processed as a whole during the retrieval process.
7. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S6, the metadata binding and structured output stage, metadata fields are also defined, including author, creation time, and modification records; the middle Transformer layer in the large language model is copied and inserted into the embedding model.
8. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S7, the table enhancement processing stage, the algorithm for header alignment adopts dynamic programming technology to ensure the accurate correspondence between the header and the content row; the optimization of merged cell sinking adopts a hierarchical processing strategy.
9. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S8, the knowledge conversion and injection stage, knowledge distillation technology is also adopted to transfer the domain knowledge in the large language model to the embedding model; freezing strategy, weight source, learning rate scheduling, and loss function selection are adopted.
10. A method for improving the quality of private domain operation and maintenance knowledge retrieval based on RAG according to claim 1, characterized in that: In the step S9, the model training and optimization stage, early stopping method and learning rate decay technology are also adopted.
Citation Information
Patent Citations
Method and device for constructing customer service knowledge base
CN114780736A
Machine learning operation and maintenance method based on large language model
CN119003719A
Large language model RAG optimization method based on tree neighbor context
CN119293195A
Enhanced document generation and retrieval method based on knowledge graph
CN119646178A
Patent analysis system and method, and computer-readable recording medium for recording program for executing same
WO2015122700A1
Cited By
Automatic generation method and device of operation and maintenance test questions, medium and equipment
CN120448535A
Medical information retrieval method and system based on multi-modal vectorization and storage medium
CN120994812A
Document generation method based on OCR (Optical Character Recognition), electronic equipment and medium
CN121147957A