A power document warehousing method based on large model and visual recognition combination

CN122838488APending Publication Date: 2026-09-29CHINA SOUTHERN POWER GRID DIGITAL GRID GROUP (GUANGDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610947028.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

如果片段只作为向量检索对象存在,而没有与来源文档、章节路径、页码区间和原文位置建立稳定映射,则即便召回结果在语义上相关,也难以实现结构化过滤、精准定位和审计复核,无法作为后续多模块协同调用的统一底层数据对象

Benefits of technology

[0068]与现有技术相比:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838488A_ABST
    Figure CN122838488A_ABST
Patent Text Reader

Abstract

The application discloses to the technical field of information system, specifically to a power document warehousing method based on large model and visual recognition combination, first, the power business documents with dispersed sources, different formats and obvious structural differences are converted into unified warehousing objects; then, different types of documents are converted into unified structural representations; subsequently, the slice granularity is adaptively selected according to the document type, structural level, semantic completeness and information density; then, the generated slices are converted from intermediate results into knowledge base formal management objects; finally, the slice text semantic representations and structured metadata are stored respectively, and a joint index is established; then, based on the constructed joint index, unified slice query, positioning, retrieval and backtracking interfaces are provided externally; based on this, the application effectively solves the core problems in the power knowledge base multi-format document warehousing process, such as non-uniform access, unstable analysis, non-adaptive slicing, non-backtracking objects and difficult coordination of index updating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information system technology, specifically to a method for storing power documents in a database based on a combination of large models and visual recognition. Background Technology

[0002] Power companies have accumulated a wealth of knowledge carriers across various business domains, including regulations, infrastructure, marketing, innovation, and operation and maintenance. These include business documents such as rules and regulations, technical solutions, feasibility studies, work instructions, and meeting minutes, as well as files in different formats such as Word, editable PDFs, formatted PDFs, and scanned copies. These documents vary significantly in chapter hierarchy, paragraph boundaries, table structures, page numbering, and text / image layout. Therefore, the data entry process in building a power knowledge base is not simply a matter of uploading files or extracting text; it requires integrated processing encompassing unified access to multiple formats, content parsing, structure restoration, slice organization, metadata registration, index creation, and subsequent retrieval and backtracking. Therefore, establishing a standardized slice system for the three mainstream document formats—Word, PDF, and scanned copies—and combining vector and relational databases to construct a slice storage and rapid indexing mechanism can provide high-quality data source support for subsequent intelligent retrieval and question answering.

[0003] Among existing technologies, the closest solution to this case is the traditional document management and full-text retrieval solution. This type of solution mainly addresses the issues of centralized storage, classification management, and keyword retrieval of electronic documents. A common approach is to register the format of uploaded documents, extract the full text, and construct an inverted index, then query based on the title, tags, or full-text fields. This type of solution is widely used in enterprise document centers, archive management, and policy retrieval scenarios. Its advantages include mature implementation methods, high efficiency in full-text retrieval, and low engineering deployment costs. However, this type of solution typically treats "documents" as the primary management object, with management granularity limited to the entire document level. This makes it difficult to meet the refined management needs of fragment-level content in the power knowledge base. For subsequent retrieval, question answering, and knowledge reuse, it can only locate "related documents" but cannot reliably locate "related fragments, related page numbers, related clauses, or related table units," easily leading to problems such as an excessively large recall scope, high costs for manual secondary searches, and a lack of traceability of results.

[0004] The second type of solution is the layout document parsing and fixed-rule slicing solution. This type of solution mainly addresses the problem that unstructured documents such as PDFs and scanned copies are difficult to directly enter into the retrieval system. It typically extracts text through OCR recognition, layout analysis, paragraph detection, and table recognition, and then slices the text according to fixed word counts, fixed paragraph lengths, natural paragraph boundaries, or simple heading rules. This type of solution can complete basic text recovery and basic segmentation, and is widely used in scenarios such as scanned document digitization, contract recognition, and report parsing. Its advantage lies in its ability to transform previously difficult-to-process layout documents into searchable text. However, in the context of power knowledge bases, document content often simultaneously contains complex heading levels, body clauses, table indicators, figure descriptions, and scan recognition results. Different business documents vary significantly in structural hierarchy, content themes, and information density. If a uniform window or single slicing method is still used, semantic units are easily truncated, headings are separated from body text, table headers are separated from table data, and low-quality text in scanned copies is mis-sliced. This results in slicing that, while forming slices, fails to balance semantic integrity and retrieval granularity, thus affecting subsequent search hits and the quality of answer generation.

[0005] The third type of solution is vectorized data entry and semantic retrieval enhancement. This type of solution mainly addresses the shortcomings of traditional keyword retrieval in supporting synonymous expressions, colloquial expressions, and fuzzy queries. Common implementations include vectorizing text fragments and writing them into a vector database, or overlaying vector recall capabilities on top of full-text indexing, and combining keyword retrieval, semantic retrieval, and ranking mechanisms during queries to improve recall performance. This type of solution can significantly enhance semantic matching capabilities and has become one of the mainstream technical paths in knowledge base retrieval and retrieval enhancement applications. However, in this case, the power knowledge base not only requires "the ability to recall relevant text" but also "the ability to filter, locate, and trace back according to business domain, document, chapter, and original text position." If the fragments only exist as vector retrieval objects without establishing a stable mapping with the source document, chapter path, page number range, and original text position, even if the recall results are semantically relevant, it will be difficult to achieve structured filtering, accurate location, and audit review, and it cannot serve as a unified underlying data object for subsequent multi-module collaborative calls.

[0006] In summary, existing technologies still have the following shortcomings in the scenario of managing multi-format documents in the power knowledge base: First, there is a lack of a unified access and consistent parsing mechanism for Word, PDF, and scanned documents. Inconsistent management standards after different document formats are entered into the database can easily lead to missing structural information and confused object boundaries. Second, most existing slicing methods use fixed windows or single rules, failing to dynamically adapt to document type, structural hierarchy, and semantic integrity, making it difficult to simultaneously guarantee slice retrievability and semantic fidelity. Third, existing solutions typically do not establish complete metadata for slices as independent management objects, resulting in a lack of stable mapping relationships between slice content, source documents, chapter paths, and original text locations, making it difficult to achieve precise positioning and traceable retrieval. Fourth, a single full-text index or a single vector index cannot simultaneously meet the needs of semantic recall, structural filtering, and rapid positioning, failing to provide unified and efficient data support for subsequent question answering and reuse. Fifth, existing systems lack an incremental processing mechanism bound to slice objects and index records in scenarios of document updates, additions, and replacements, easily leading to problems such as duplicate entries, old index residues, and version mixing. The aforementioned shortcomings correspond to the technical challenges commonly encountered in the scenario of document entry into the power knowledge base, such as heterogeneous multi-format documents, incompatibility with traditional single-slice mode, low efficiency of slice storage and retrieval, and incompatibility of multi-module data interaction.

[0007] Therefore, a method for storing power documents based on a combination of large models and visual recognition is invented. By constructing a unified access mechanism for multi-format documents, a large model-assisted structure parsing mechanism, an adaptive slicing mechanism, a slice-level management object modeling mechanism, a hybrid storage and joint indexing mechanism, and an incremental update and backtracking management mechanism, the method systematically solves the shortcomings of existing technologies in unified parsing, slice adaptation, object management, index collaboration, and continuous updating. This provides manageable, searchable, locatable, and backtrackable underlying data support for subsequent retrieval, question answering, and knowledge reuse. Summary of the Invention

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] A method for storing power documents in a database based on a large model and visual recognition includes the following steps:

[0010] S1, Document Access and Basic Database Modeling: Transform power business documents with scattered sources, different formats, and significant structural differences into unified database objects;

[0011] S2, performs content extraction and large model-assisted structure parsing according to document format: converts different types of documents into a unified structural representation;

[0012] S3, Adaptive slicing decision based on multi-feature scoring: Adaptively selects slicing granularity based on document type, structural hierarchy, semantic integrity and information density;

[0013] S4, Slice Management Object Construction and Backtrackable Encoding Modeling: Transform the slices generated in S3 from intermediate results into formal knowledge base management objects;

[0014] S5, Hybrid Storage, Composite Index and Staged Retrieval Modeling: Separately store the semantic representation of sliced ​​text and structured metadata, and build a composite index;

[0015] S6, based on data interaction standards, provides unified interface for slice query, location, retrieval and backtracking: based on the composite index built by S5, it provides unified interface for slice query, location, retrieval and backtracking.

[0016] S7, Incremental Inbound Processing and Index Update: When documents are added, replaced, or supplemented, the changed parts are rebuilt and the index is updated.

[0017] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S1 are as follows:

[0018] S11, Perform access pre-verification: sequentially verify whether the file can be opened, whether the file size is within the preset limit, whether the extension is consistent with the file header, and whether the uploaded attribute fields are complete;

[0019] S12, Perform format recognition: Combine file extension, file header magic number and parsing capability to make a judgment and achieve classification;

[0020] S13 generates a globally unique document identifier DocID for each document and calculates the content hash value DocHash based on the original file content.

[0021] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of step S2 are as follows:

[0022] S21, Differentiated Content Extraction:

[0023] For Word and editable PDF: Directly extract the main text, headings, paragraphs, tables, font attributes, paragraph spacing, page number information, and original layout information to form a sequence of original text units;

[0024] For formatted PDFs: parse text block coordinates, table borders, reading order, header and footer areas, and image / text description areas;

[0025] For scanned documents: perform OCR recognition, output character or text block sequences, and record the corresponding page coordinates, line numbers, block numbers, and recognition confidence scores;

[0026] S22, Low-quality text correction: For scanned documents and pages with complex layouts, a character neighborhood correction algorithm is introduced for correction.

[0027] S23, Large Model Assisted Structure Analysis: The large model is used to assist in the determination of structure labels. First, preliminary structure identification is performed based on the rule method, and then the rule conflict pages are sent to the large model for auxiliary determination.

[0028] S24, Unified Structure Record and Position Encoding: After completing text extraction and structure determination, a unified structure record is generated, and finally a unified position encoding is generated.

[0029] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S3 are as follows:

[0030] S31, Multi-feature slice scoring model: After candidate units are formed, a multi-feature slice scoring model is constructed to calculate the slice granularity score of each candidate unit. ;

[0031] S32, Granularity Determination and Constraint Rules: Based on Execution granularity determination:

[0032] when At that time, chapter-level slicing is used;

[0033] when When using paragraph-level slicing;

[0034] when At this time, sentence-level slicing is used;

[0035] At the same time, three types of constraint rules are set: hierarchy preservation rules, semantic closure rules, and table semantic binding rules;

[0036] S33, Generation of candidate slice set: After completing the granularity determination, a candidate slice set SegmentCandidateSet is generated.

[0037] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S4 are as follows:

[0038] S41, Slice object definition: During processing, first assign a unique identifier ChunkID to each candidate slice, and write the slice text, source document, chapter path, page range and position code into a unified object;

[0039] S42, Three-layer traceable mapping mechanism: Establish three-layer mapping relationships: document layer to slice layer, structure layer to slice layer, and original text layer to slice layer;

[0040] S43, Slice quality check: Perform slice quality check before writing the object;

[0041] S44, Output Definition: Outputs a standard slice object collection ChunkObjectSet and a slice metadata registry table ChunkMetaTable.

[0042] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S5 are as follows:

[0043] S51, Semantic Vectorization and Hybrid Storage: First, semantic vectorization is performed on ChunkText; after vectorization, the vector records are written to the vector database, and at the same time, the metadata is written to the relational database.

[0044] S52, Composite Index Construction: Construct a composite index, which should include at least: ChunkID primary association index, DocID+SectionPath structure index, BusinessDomain+FileType+ChunkType composite index, and SourcePosCode reverse lookup index;

[0045] S53, Phased retrieval model: A phased retrieval model is pre-built during the database entry stage.

[0046] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S6 are as follows:

[0047] S61, Request and Response Structure Definition:

[0048] The request structure is defined as follows:

[0049] <RequestID,QueryText,QueryDocID,QueryBusinessDomain,QueryFileType,QuerySectionPath,NeedVectorFlag,NeedTraceFlag>

[0050] NeedVectorFlag indicates whether semantic recall is enabled; NeedTraceFlag indicates whether location backtracking information must be returned.

[0051] The response structure is defined as follows:

[0052] <RequestID,ChunkID,DocID,ChunkText,SectionPath,SourcePosCode,PageRange,BusinessDomain,StrategyID,Score_rank,TraceInfo>

[0053] TraceInfo is further expanded as follows:

[0054] <SourceDocName,PageID,SectionTitle,ParagraphIndex,TableIndex> ;

[0055] S62, Exception handling logic:

[0056] When no candidate slices are found after structural filtering, the business domain constraints can be relaxed and the query can be performed again according to the preset strategy.

[0057] When all semantic matching scores are below the threshold of 0.55, the system degenerates to only returning relevant chapter slices that match the structure index.

[0058] If location backtracking fails, at least the ChunkID and DocID should be returned so that the backend can verify the information based on the primary key.

[0059] As a preferred embodiment of the power document database entry method based on a combination of large model and visual recognition described in this invention, the specific steps of S7 are as follows:

[0060] S71, Incremental Processing Flow:

[0061] In terms of processing flow, the current document's DocHash is compared with the historical records:

[0062] If the corresponding DocID does not exist, it is determined to be a newly added document, and the entire process from S1 to S6 is executed;

[0063] If the DocID exists but the DocHash or VersionTag has changed, it is determined to be a modified document and enters the incremental entry process.

[0064] When adding new data to the database, differences between new and old slice objects are identified.

[0065] S72, Version Auditing and Closed Loop: Establishing the VersionAuditRecord table:

[0066] <VersionID,DocID,OldDocHash,NewDocHash,UpdateType,UpdateTime,OperatorID,AffectedChunkCount,RuleVersion>

[0067] UpdateType includes at least addition, replacement, supplementation, and rule reslicing; after incremental processing is completed, the new ParseStatus, VersionTag, and update time are written back to DocumentMasterRecord in S1.

[0068] Compared with existing technologies:

[0069] The method proposed in this invention is designed for the underlying data governance scenario of the power knowledge base. It realizes a complete data entry link for multi-format documents, from unified access, structure restoration, adaptive slicing to joint indexing and incremental updates, providing stable data support for subsequent retrieval, question answering and knowledge reuse.

[0070] First, by adopting a unified access mechanism for multi-format documents and a document master record modeling mechanism, a unified database entry standard has been achieved for heterogeneous documents such as Word, PDF, and scanned documents. This solves the problem of inconsistent access standards for documents from different sources and in different formats in existing technologies, and provides a unified data foundation for subsequent parsing, slicing, and version management.

[0071] Secondly, by adopting a structure recovery mechanism that combines rule-based parsing with large-model-assisted parsing, stable recovery of complex layouts, heading levels, table structures, and low-quality text is achieved. This solves the problems of structural loss and boundary misjudgment that easily occur in scanned documents and layout documents in existing technologies, and improves the completeness and usability of document parsing results.

[0072] Third, by adopting an adaptive slicing mechanism based on structural hierarchy, semantic integrity, and information density, differentiated slicing of different business documents is achieved, solving the problem that the traditional fixed slicing mode is difficult to balance semantic integrity and retrieval granularity, and providing more suitable slice objects for subsequent accurate retrieval and slice hit.

[0073] Fourth, by adopting slice-level management object modeling and a traceable coding mechanism, stable binding between slice content and source documents, chapter paths, page number ranges, and original text positions is achieved, solving the problems of slices not being finely managed, not being able to be located, and not being able to be traced in existing technologies. This enables the verification of search results and rapid location of the original text.

[0074] Fifth, by adopting hybrid storage, combined indexes, and incremental update mechanisms, it achieves collaborative processing of semantic recall, structural filtering, precise positioning, and version synchronization maintenance, solving the problems of insufficient single index capabilities, easy remnants of old indexes after updates, and mixed version usage in existing technologies, thus improving the stability of the power knowledge base for long-term operation and continuous maintenance.

[0075] In summary, this invention effectively solves the core problems of inconsistent access, unstable parsing, incompatible slicing, lack of object backtracking, and difficulty in coordinating index updates during the process of importing multi-format documents into the power knowledge base, providing a new technical solution for the refined management and intelligent application of the power knowledge base. Attached Figure Description

[0076] Figure 1 This is the overall flowchart of the present invention;

[0077] Figure 2 This is a schematic diagram of the document entry and management structure of the present invention;

[0078] Figure 3 This is a block diagram illustrating the adaptive slicing decision-making principle of the present invention. Detailed Implementation

[0079] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0080] This invention provides a method for storing power documents in a database based on a combination of large-scale modeling and visual recognition, the process of which is shown in the attached figure. Figure 1 As shown. This method addresses the unified governance of underlying data objects in the power knowledge base, constructing a complete method chain around multi-format document access, document structure restoration, adaptive slicing, slice object modeling, hybrid storage and joint indexing, unified retrieval and backtracking, and incremental updates. Through the above-mentioned full-process steps, this invention can solve the problems in existing technologies such as inconsistent multi-format document entry standards, poor adaptability of slicing rules, lack of fine-grained management of slice objects, insufficient index collaboration capabilities, and low efficiency of cross-module retrieval and backtracking. To achieve the above-mentioned objectives, this invention adopts the following technical solution:

[0081] S1, Document Access and Basic Database Modeling: This step transforms power business documents, which are scattered in origin, have different formats, and significantly different structures, into a unified database object, solving the problems of inconsistent access standards and subsequent processing entry points for documents from different sources and in different formats in existing technologies. The input for this step is the document file itself to be databased and its accompanying attribute information. The accompanying attribute information includes at least the source file name, file extension, source system, business domain identifier, upload time, document category tag, and storage path. Among them, the file types at least cover Word documents, editable PDFs, formatted PDFs, and scanned copies.

[0082] Perform access pre-check: preferably, check sequentially whether the file can be opened, whether the file size is within the preset upper limit, whether the file extension matches the file header, and whether the upload attribute fields are complete; if the file is damaged, unrecognizable, or has missing key attributes, mark the document as an abnormal document, write it into an abnormal queue, and terminate subsequent parsing;

[0083] Perform format identification: preferably, make determination and classification by combining file extension, file header magic number and parsability: documents from which paragraphs and table structures can be directly extracted are classified as structured documents, documents from which text blocks can be extracted but have strong formatting are classified as semi-structured documents, and documents that can only rely on image recognition to recover text are classified as unstructured documents;

[0084] To ensure consistent data primary keys in subsequent steps, generate a globally unique document identifier DocID for each document, and calculate a content hash value DocHash based on the original file content; DocHash is used for subsequent version change determination and incremental warehousing comparison, and DocID is used as the main association key throughout the entire method chain.

[0085] Preferably, this step constructs a document master record, whose field composition can be defined as:

[0086] <DocID, SourceDocName, FileType, BusinessDomain, SourceSystem,UploadTime, RawFilePath, DocHash, VersionTag (preferably V1.0 / V1.1 / V2.0),ParseStatus, ParseRoute, ParentDocID (record the ID of the source document, blank for the first version)>

[0087] Wherein, ParseStatus includes at least four statuses: to be parsed, parsing, parsing completed, and parsing abnormal; ParseRoute is used to record the specific path for subsequent rule-based parsing, large model-assisted parsing or OCR text recovery; BusinessDomain is used to distinguish business domains such as regulation, infrastructure, marketing, and innovation. For documents with missing source system information, a default source is preferably written; for documents without marked business domain, UnknownDomain can be written first and supplemented in subsequent steps.

[0088] This step outputs a standardized document input object (DocumentInputObject) and a document master record (DocumentMasterRecord). The former includes at least...<DocID, FileType, BusinessDomain,RawFilePath, ParseRoute> This serves as the direct input to S2; the latter serves as the source of the upper-level primary keys for subsequent chapter trees, slice objects, metadata tables, and index mapping tables.

[0089] S2, Content Extraction and Large Model-Assisted Structure Parsing Based on Document Format: This step converts different document types into a unified structural representation, resolving differences in heading levels, paragraph boundaries, table structures, and scanned text quality across various document formats. The inputs for this step are the DocumentInputObject and DocumentMasterRecord output from S1. Unlike traditional methods that rely solely on rules or OCR to recover text, this step introduces a "large model-assisted parsing mechanism." However, the large model is limited to three areas: complex structure recovery, low-confidence fragment correction, and heading level assistance, to avoid generalizing the large model into the core black box of the entire system.

[0090] Differentiated content extraction:

[0091] For Word and editable PDF: Preferably, the main text, headings, paragraphs, tables, font attributes, paragraph spacing, page number information, and original layout information are directly extracted to form a sequence of original text units;

[0092] For formatted PDFs: preferably, the coordinates of text blocks, table borders, reading order, header and footer areas, and image and text description areas should be parsed.

[0093] For scanned documents: first perform OCR recognition, output character or text block sequences, and record the corresponding page coordinates, line numbers, block numbers and recognition confidence scores;

[0094] Low-quality text correction: For scanned documents and pages with complex layouts, OCR output may contain issues such as similar-looking characters, broken characters, and mis-spelling across columns. Therefore, a character neighborhood correction algorithm is preferably introduced for correction. This algorithm is specifically used for OCR correction and position inheritance in this step, and its input is an OCR character sequence. and the position of each character and confidence level The output is the corrected text sequence. and the RelationMap of position inheritance; the reason for using this algorithm in this step is that subsequent chapter recognition and slice boundary judgment strongly depend on text integrity. If OCR noise is not processed first, all subsequent structural judgments will be offset; preferably, the algorithm includes the following steps:

[0095] Based on confidence thresholds, illegal character patterns, dictionary misses, and contextual semantic conflicts, a set of suspected erroneous segments is selected. ;

[0096] Extract the left and right neighboring windows centered on each suspected erroneous segment. A set of candidate error correction items is generated based on a power terminology dictionary, an equipment name thesaurus, and an indicator unit thesaurus. ;

[0097] For each candidate Calculate edit distance similarity Contextual language score and term matching score ;

[0098] Calculate the overall score:

[0099]

[0100] in, This is the original identified segment. The weighting parameters are set to 0.45, 0.35, and 0.20 respectively.

[0101] When the highest score exceeds the preset threshold When the value is greater than or equal to 0.72, the original segment is replaced with the corresponding candidate, and the replacement result is compared with the original position. Establish inheritance mapping; otherwise, retain the original recognition result and mark LowConfFlag=1;

[0102] Large model-assisted structure analysis: For complex pages where rules are difficult to reliably determine, this invention does not directly regenerate the full text using a large model, but instead uses the large model to assist in determining structure tags; preferably, preliminary structure identification is first performed based on the rule-based method, and then pages with rule conflicts are sent to the large model for auxiliary determination.

[0103] The rule-based approach uses a combination of "heading regular expression rules + layout style judgment rules + context continuity verification"; the heading regular expression rules include at least chapter, section, article, clause, numbered heading and table heading rules, and the style judgment rules include at least font size, whether it is bold, whether it is centered, left indentation, paragraph spacing and adjacent block position relationship.

[0104] When a text block simultaneously hits multiple level rules, or when the level jump between adjacent headings exceeds a preset difference, the text block and its context block are combined into a structure determination window. Input the large model auxiliary parsing module; the input of this module is <candidate text block sequence, coordinate information, style features, preceding and following structural context>, and the output is the structural label StructLabel, the level StructLevel, and the boundary confidence StructConf for each text block.

[0105] The reason for using a large model here is that relying solely on numbering and style rules in complex layouts can easily lead to chapter breaks or misjudgments of titles. A large model, on the other hand, can combine contextual semantics to determine whether "this text block is a title, body text, table description, or header / footer," thereby improving stability in complex scenarios.

[0106] Preferably, when StructConf is greater than a preset threshold If the value is not lower than 0.80, the large model output result is used; otherwise, the result is reverted to the rule parsing result and the current page is marked as a "low-structure confidence page".

[0107] Unified Structure Records and Position Encoding: After text extraction and structure determination, unified structure records are generated.

[0108] Construct the SectionTreeRecord table of chapter tree nodes:

[0109] <DocID,SectionID,ParentSectionID,SectionLevel,SectionTitle,StartPage,EndPage,StartParagraph,EndParagraph,StructSource,StructConf>

[0110] StructSource is used to identify whether the node comes from rule-based decision-making or large model-assisted decision-making.

[0111] Constructing a TextUnitRecord:

[0112] <DocID,UnitID,PageID,SectionID,ParagraphID,SentenceID,UnitText,UnitType,CoordBox,OCRConf,LowConfFlag>

[0113] Constructing a TableUnitRecord:

[0114] <DocID,TableID,PageID,RowID,ColID,HeaderText,CellText,CoordBox,MergeFlag>

[0115] Finally, a unified location code is generated:

[0116] The preferred definition for text position encoding is:

[0117]

[0118] The preferred definition for table position coding is:

[0119]

[0120] This step outputs a SectionTree, a TextUnitSet, a TableUnitSet, and a Position Encoding Map (PosMap). If a page is marked as a low-structure confidence page, only the conservative slicing strategy is allowed in S3. All of the above outputs serve as the basis for the input of S3.

[0121] S3, Adaptive Slicing Decision Based on Multi-Feature Scoring: Adaptively selects slice granularity based on document type, structural hierarchy, semantic integrity, and information density to solve the problems of semantic truncation, granularity imbalance, and incompatibility with different business documents caused by fixed window slicing; The input of this step is the SectionTree, TextUnitSet, TableUnitSet, PosMap output by S2, as well as the basic document attribute information.

[0122] This step does not directly segment the entire text with a fixed number of characters, but first forms candidate slice units, and then makes granularity decisions. Preferably, sentences, clauses, paragraphs, heading nodes, table description paragraphs, and table text fragments are used as candidate slice units (CandidateUnit). For table text fragments, it is preferable to merge the table header fields with the corresponding data units into an initial "Header+Cell" unit to avoid semantic anchoring during table retrieval.

[0123] Multi-feature slice scoring model: After candidate units are formed, a multi-feature slice scoring model is constructed to calculate the slice granularity score for each candidate unit. Preferably defined as:

[0124]

[0125] Semantic completeness can be calculated jointly by syntactic closure, conjunction completeness, and semantic continuity of adjacent units;

[0126] Structural hierarchy weights: chapter titles, clause opening paragraphs, and table description paragraphs have higher weights than ordinary narrative paragraphs;

[0127] Information density can be calculated by combining term hit rate, object word density, numerical density, table field density, and constraint statement density.

[0128] Preferably, It can be calculated in the following form:

[0129]

[0130] in, Indicates syntactic closure. Indicates the integrity of the connection indicator. Indicates the semantic similarity between adjacent units.

[0131] Granularity determination and constraint rules: based on Execution granularity determination:

[0132] when At that time, chapter-level slicing is used;

[0133] when When using paragraph-level slicing;

[0134] when At this time, sentence-level slicing is used;

[0135] To avoid slicing results that have scores but lack feasibility, three types of constraint rules are set: hierarchy preservation rules, semantic closure rules, and table semantic binding rules.

[0136] Hierarchical preservation rule: Different SectionLevels should not be merged across levels in principle, unless the current unit length is less than the minimum semantic length threshold of 80 characters and the adjacent units have the same topic;

[0137] Semantic closure rule: If an indicative phrase such as "see table below" appears at the end of the current unit, it must be merged with the immediately following unit;

[0138] Table semantic binding rules: When a column alignment relationship is detected between a table field and a data cell, the HeaderText and the corresponding CellText must be retained together.

[0139] For low-structure confidence pages and scanned pages, a conservative slicing strategy is preferred, which prioritizes natural paragraphs and supplements them with sentence-level slicing, and does not allow cross-page merging, in order to reduce the risk of erroneous slicing.

[0140] Candidate Slice Set Generation: After completing the granularity determination, a candidate slice set SegmentCandidateSet is generated, whose preferred record fields are:

[0141] <DocID,SegmentID,SegmentText,SegmentType,SectionPath,StartPosCode,EndPosCode,SourcePosCode,StrategyID,Score_slice,SemanticBoundaryFlag,TableBindingFlag,ReviewFlag>

[0142] SegmentType: Indicates a combination of chapter, paragraph, sentence, or table level;

[0143] StrategyID: Indicates the strategy version used in the current slice;

[0144] SectionPath: Represents the chapter path;

[0145] SourcePosCode: Represents the source location code or location range code;

[0146] If a candidate slice still cannot be determined to have a reasonable granularity after scoring and rule verification, then the ReviewFlag is set to 1 and the slice is moved to the manual sampling branch or the conservative data entry branch.

[0147] This step outputs a SegmentCandidateSet, which serves as the direct input for constructing the S4 slice management object.

[0148] S4, Segment Management Object Construction and Backtrackable Encoding Modeling: Transform the segments generated in S3 from intermediate results into formal knowledge base management objects, solving the problem of separation between segment content and source, path, page number, and original text location in existing technologies; the inputs for this step are SegmentCandidateSet, SectionTree, and PosMap;

[0149] Slice object definition: During processing, a unique identifier ChunkID is first assigned to each candidate slice, and the slice text, source document, chapter path, page range, and position code are written into a unified object; preferably, the slice object ChunkObject is defined as follows:

[0150] <ChunkID,DocID,SegmentID,ChunkText,ChunkType,FileType,BusinessDomain,SectionPath,PageStart,PageEnd,StartPosCode,EndPosCode,SourcePosCode,StrategyID,TopicTag,SceneTag,QualityFlag,CreateTime>

[0151] ChunkID: Serves as the primary association key for subsequent vector records, metadata tables, and retrieval interfaces;

[0152] SectionPath: Used for structure-level filtering;

[0153] StartPosCode and EndPosCode: used to represent the interval boundaries of the slice in the original text;

[0154] QualityFlag: Used to indicate whether the slice has passed the quality check.

[0155] A three-layer traceable mapping mechanism is established to achieve both "retrievable" and "traceable" capabilities. This involves mapping from the document layer to the slice layer, from the structure layer to the slice layer, and from the original text layer to the slice layer.

[0156] Document layer to slice layer mapping: establish a one-to-many relationship using DocID;

[0157] Mapping from structure layer to slice layer: Establish a one-to-many relationship using SectionPath;

[0158] Mapping from source layer to slice layer: Establish one-to-one or one-to-many relationships using SourcePosCode or PosRange.

[0159] Preferably, a position range encoding is defined for cross-page slices: PosRange=<StartPosCode,EndPosCode>

[0160]

[0161] For slices originating from tables, TablePosCode and its row and column ranges are used for representation. This traceable encoding mechanism is the first core identification mechanism of this invention: slices are no longer just text fragments, but management units bound to their original locations and structural paths, thus enabling not only retrieval but also precise positioning and audit playback.

[0162] Slice quality verification: Before writing the object, slice quality verification must be performed, preferably including the following rules:

[0163] When the length of a ChunkText is less than the minimum length threshold of 50 characters and it does not belong to a title or table key cell, it is marked as a low-quality slice.

[0164] If a slice contains only page numbers, headers and footers, watermarks, or meaningless symbols, it should be removed directly.

[0165] When two slices under the same DocID have the same ChunkTextHash and their position ranges overlap, only the main record is retained;

[0166] For slices that span multiple chapters and have SemanticBoundaryFlag=0, a boundary backoff mechanism is triggered to reassemble the slice back into its original candidate unit.

[0167] Output definition: Outputs a standard slice object collection ChunkObjectSet and a slice metadata registry table ChunkMetaTable, the latter preferably defined as:

[0168] <ChunkID,DocID,SectionPath,PageRange,SourcePosCode,BusinessDomain,FileType,ChunkType,StrategyID,TopicTag,QualityFlag>

[0169] The above output serves as input for S5 hybrid storage and composite index construction.

[0170] S5, Hybrid Storage, Joint Indexing, and Staged Retrieval Modeling: Separately store the semantic representation of the slice text and structured metadata, and establish a joint index to solve the problem that a single full-text index or a single vector index cannot simultaneously satisfy semantic recall, structure filtering, and fast location; the input for this step is ChunkObjectSet and ChunkMetaTable; this solution adopts a hybrid storage architecture of vector database and relational database to meet the needs of refined management and intelligent retrieval of the power knowledge base, and designs a joint indexing mechanism for slice content vectors and metadata;

[0171] Semantic Vectorization and Hybrid Storage: First, the ChunkText is semantically vectorized. The model used is specifically for semantic encoding in this step, not for classification or generation. Its input is ChunkText, and its output is a fixed-dimensional vector, where the vector dimension is 1024. The reason for using this model is that the power knowledge base contains a large number of synonyms, abbreviations, and colloquial questions, which are difficult to cover by keyword matching alone. After vectorization, the vector records are written to the vector database. Preferably, the vector records are defined as follows:

[0172] <VectorID,ChunkID,DocID,ChunkVector,VectorVersion,VectorStatus>

[0173] Simultaneously, metadata is written to a relational database, preferably defined as follows:

[0174] <ChunkID,DocID,ChunkText,SectionPath,SourcePosCode,PageRange,BusinessDomain,FileType,ChunkType,StrategyID,QualityFlag>

[0175] Composite Index Construction: Based on this, construct a composite index, which preferably includes at least: ChunkID main association index, DocID+SectionPath structure index, BusinessDomain+FileType+ChunkType composite index, and SourcePosCode reverse lookup index;

[0176] ChunkID primary associated index: used to connect vector records and metadata records;

[0177] DocID+SectionPath structure index: used for quick filtering by document and section;

[0178] BusinessDomain+FileType+ChunkType combined index: used to filter candidates by business domain and document type;

[0179] SourcePosCode reverse lookup index: used to directly trace back to the original text position based on the slice.

[0180] Phased retrieval model: A phased retrieval model is pre-built during the data entry stage; preferably, subsequent retrieval paths are executed in the following order:

[0181] First, perform structure filtering based on a relational database;

[0182] Then perform vector similarity matching on the candidate set;

[0183] Finally, dynamic rearrangement is performed.

[0184] The formula for calculating the similarity is:

[0185]

[0186] Based on this, define dynamic reordering scores:

[0187]

[0188] Where represents the structure matching weight, represents the source reliability weight, and is the parameter λ1=0.60, λ2=0.25, λ3=0.15.

[0189] The combined index and phased retrieval modeling mechanism is the second core identifying mechanism of this invention: it is not a simple “vectorized input”, but rather it solidifies the subsequent “structural filtering—semantic recall—location backtracking” path into an executable index relationship during the input stage.

[0190] If vectorization of a slice fails, the VectorStatus is marked as Failed, but the metadata record is still retained to ensure that it can be accessed at least through the structure index; the failure record is put into the compensation retry queue.

[0191] This step outputs the VectorStoreSet, MetaStoreSet, and JointIndexMap as inputs for S6 unified retrieval and backtracking.

[0192] S6, based on data interaction standards, provides unified interface for slice query, location, retrieval, and backtracking: based on the joint index built by S5, it solves the problem of inconsistent data protocols between the document slicing module and subsequent retrieval, question answering, or manual review modules; the input of this step includes external requests and JointIndexMap, MetaStoreSet, and VectorStoreSet output by S5;

[0193] Request and response structure definition:

[0194] Preferably, the request structure is defined as follows:

[0195] <RequestID,QueryText,QueryDocID,QueryBusinessDomain,QueryFileType,QuerySectionPath,NeedVectorFlag,NeedTraceFlag>

[0196] NeedVectorFlag indicates whether semantic recall is enabled; NeedTraceFlag indicates whether location backtracking information must be returned.

[0197] Processing flow:

[0198] First, parse the request fields and validate the necessary parameters;

[0199] When the request contains explicit document, business domain, or chapter path, filter candidate slices by structural index first;

[0200] When the request contains natural language query text and NeedVectorFlag=1, a query vector is generated. And perform vector matching and dynamic rearrangement on the candidate slices;

[0201] Finally, based on SourcePosCode or PosRange, we can trace back to the original document's page number, chapter title, and paragraph position.

[0202] Preferably, the response structure is defined as follows:

[0203] <RequestID,ChunkID,DocID,ChunkText,SectionPath,SourcePosCode,PageRange,BusinessDomain,StrategyID,Score_rank,TraceInfo>

[0204] TraceInfo is further expanded as follows:

[0205] <SourceDocName,PageID,SectionTitle,ParagraphIndex,TableIndex>

[0206] Exception handling logic:

[0207] The preferred exception handling logic is:

[0208] When no candidate slices are found after structural filtering, the business domain constraints can be relaxed and the query can be performed again according to the preset strategy.

[0209] When all semantic matching scores are below the threshold of 0.55, the system degenerates to only returning relevant chapter slices that match the structure index.

[0210] If location backtracking fails, at least ChunkID and DocID should be returned so that the backend can verify them based on the primary key;

[0211] This unified retrieval mechanism is the third core identification mechanism of this invention: sliced ​​objects are not only for storage in the database, but also have standard request, standard response, and standard tracking capabilities, enabling subsequent modules to share the same object identifier and the same backtracking caliber. S6 outputs a unified retrieval result ResultSet and provides it to the question-and-answer, retrieval, or manual review modules, while also providing an access interface for index repair and version backtracking in S7.

[0212] S7, Incremental Inbound Processing and Index Update: When adding, replacing, or supplementing documents, only the changed parts are rebuilt and the index is updated, solving the problems of high cost of full reconstruction, old index residue, and mixed version use in the existing technology; the inputs for this step are the new or replaced document, the historical DocumentMasterRecord, ChunkObjectSet, and JointIndexMap;

[0213] Incremental processing flow:

[0214] In terms of processing flow, the current document's DocHash is compared with the historical records:

[0215] If the corresponding DocID does not exist, it is determined to be a newly added document, and the entire process from S1 to S6 is executed;

[0216] If the DocID exists but the DocHash or VersionTag (preferably V1.0 / V1.1 / V2.0) has changed, it is determined to be a changed document and enters the incremental entry process.

[0217] During incremental data import, instead of rebuilding the database for all slices, the differences between new and old slice objects are identified. Preferably, a ChunkTextHash is calculated for each slice and compared in conjunction with its positional encoding. The preferred difference identification rule is:

[0218] When ChunkTextHash_new does not exist in the history collection, it is determined to be a newly added slice;

[0219] When ChunkTextHash_old is missing in the new collection, it is determined that the slice has been deleted;

[0220] If both exist but SourcePosCode or ChunkText changes, it is determined that the slice has been modified.

[0221] Formally, it can be defined as follows:

[0222]

[0223] Among them, Compare returns four states: Add, Delete, Modify, and Keep.

[0224] For newly added and modified slices, re-execute object modeling, vectorization, and index reconstruction from S4 to S6;

[0225] For deleted slices, unregister their ChunkID associations from the JointIndexMap and retain deletion audit traces.

[0226] If a historical slice's StrategyID is detected to be lower than the current slice rule version during the incremental process, the document can be added to the re-slicing candidate queue and re-sliced ​​in offline batch processing.

[0227] Version auditing and closed-loop management:

[0228] Preferably, this step establishes the VersionAuditRecord table:

[0229] <VersionID,DocID,OldDocHash,NewDocHash,UpdateType,UpdateTime,OperatorID,AffectedChunkCount,RuleVersion>

[0230] UpdateType includes at least addition, replacement, supplementation, and rule reslicing; after incremental processing is completed, the new ParseStatus, VersionTag (preferably V1.0 / V1.1 / V2.0) and update time are written back to DocumentMasterRecord in S1.

[0231] This step outputs the updated document master record UpdatedDocumentMasterRecord, the slice object collection UpdatedChunkObjectSet, and the joint index map UpdatedJointIndexMap, thus forming a complete closed-loop method chain.

[0232] In summary, the present invention includes, but is not limited to, the following embodiments:

[0233] To verify the feasibility, effectiveness, and technical advantages of the multi-format document import management method for the power knowledge base based on large-model-assisted parsing, this implementation case is based on anonymized real business data from a power company's knowledge base construction scenario. The selected data comes from materials related to new energy projects, including feasibility study reports, project application forms, and project acceptance materials, covering various formats such as PDF files, editable PDFs, DOCX, and scanned PDFs. Simultaneously, feasibility study reports for newly added energy storage projects are introduced as incremental update samples. Multi-format document import, multi-strategy slicing, hybrid index storage, and unified data interaction are key research contents in this scenario. Therefore, this implementation case focuses on verifying the implementation effects of this invention in unified access of multi-format documents, structure parsing, adaptive slicing, slice object modeling, joint index construction, and incremental updates.

[0234] This embodiment selects 100 anonymized business documents as the initial sample for data entry, including 50 feasibility study reports for new energy projects, 40 project application forms, and 10 project acceptance materials; an additional 10 feasibility study reports for newly added energy storage projects are selected as incremental update verification samples. The implementation environment uses Python 3.9, document parsing tools include python-docx, PyPDF2, and pdfplumber, and scanned document recognition uses PP-OCRv4; vector storage uses an Elasticsearch vector database, and metadata storage uses MySQL. The data scale, tool combinations, and hybrid storage approach are consistent with the design requirements of the aforementioned technical solution.

[0235] Step S1: Document Access and Basic Database Modeling

[0236] The purpose of this step is to achieve unified access and basic modeling of multi-format power documents, establishing a unified entry point for subsequent structure parsing and slicing processing. During implementation, the aforementioned 100 sample documents, along with attributes such as source system identifier, business domain identifier, upload time, and original file path, are input into the access module. The system first identifies the document format based on file extension, file header characteristics, and parsing capability, classifying documents into structured, semi-structured, and unstructured documents. Then, a unique DocID is assigned to each document, and a DocHash is calculated based on the original file content to form a DocumentMasterRecord. Documents lacking the source system field are written to the default source, and unrecognizable or corrupted files are marked as parsing anomalies. The output results are two types of records: DocumentInputObject and DocumentMasterRecord. The latter uniformly registers fields such as DocID, FileType, BusinessDomain, SourceSystem, UploadTime, DocHash, and ParseStatus. After implementation and verification, all 100 sample documents were able to complete unified access and master record registration, forming a consistent data entry point, proving that the present invention is feasible in the multi-source heterogeneous document access process.

[0237] Step S2: Perform content extraction and large model auxiliary structure parsing according to document format.

[0238] The purpose of this step is to parse documents of different formats into a unified structural representation and restore heading levels, paragraph boundaries, table structures, and positional information. During implementation, DOCX and editable PDFs are directly extracted using python-docx and PyPDF2 respectively; layout PDFs are analyzed using pdfplumber to extract text blocks, table blocks, and coordinate relationships; scanned PDFs are processed using PP-OCRv4 for character recognition, and low-confidence text is corrected using a character neighborhood correction algorithm. For complex pages where rule-based methods cannot reliably determine the text, a large model-assisted parsing module is invoked to determine the structural labels, hierarchical relationships, and boundary confidence of candidate text blocks, outputting a SectionTree, TextUnitSet, TableUnitSet, and a PosMap. Existing sample implementation results show that scanned documents achieve high recognition accuracy after correction. The vast majority of sample documents can complete chapter structure restoration and location index construction, with good table structure restoration. Only a small number of complex merged tables require manual verification, thus verifying the effectiveness of this step in complex power document layout scenarios.

[0239] Step S3: Adaptive Slicing Decision Based on Multi-Feature Scoring

[0240] The purpose of this step is to generate slice results adapted to business semantics based on document type, structural hierarchy, and content density. During implementation, the chapter tree, text units, and table units output by S2 are used as input. Candidate slice units such as sentences, clauses, paragraphs, and table textual fragments are first constructed. Then, semantic completeness, structural hierarchy weight, and information density are calculated according to preset rules to form a slice score (Score_slice). For units with high scores and complete chapter structures, chapter-level or paragraph-level slicing is used; for units with low information density but clear semantic boundaries, sentence-level slicing is used; for table content, table slices are generated using a composite method of "header fields + data units." A conservative slicing strategy is used for low-structure confidence pages, without cross-page merging. The output is a SegmentCandidateSet, which records information such as SegmentID, SegmentText, SegmentType, SectionPath, SourcePosCode, and StrategyID. During implementation, a multi-granularity slice set covering main text clauses, chapter descriptions, and table fields can be formed, demonstrating that this invention can avoid the semantic fragmentation caused by traditional fixed-window slicing and provide more reasonable slice objects for subsequent indexing and retrieval. The "content theme, structural hierarchy, and information density" slicing criteria used in this step are consistent with those of Project Task 1.

[0241] Step S4: Construction of Slice Management Object and Backtrackable Coding Modeling

[0242] The purpose of this step is to transform the slice results into formal management objects in the knowledge base. During implementation, a unique ChunkID is assigned to each candidate slice output by S3, and the slice text, source document, chapter path, page number range, SourcePosCode, and slice strategy number are written into a ChunkObject. Simultaneously, a ChunkMetaTable is constructed to register fields such as ChunkID, DocID, SectionPath, PageRange, BusinessDomain, and QualityFlag. Slices containing only page numbers, headers / footers, or meaningless symbols are directly discarded; slices with highly repetitive content and overlapping position ranges are deduplicated and retained. The output results are a ChunkObjectSet and a ChunkMetaTable. After implementation, each slice can establish a stable mapping with the original document, chapter node, and position range, directly supporting original text backtracking and result verification, demonstrating the clear operability of this invention in slice object-oriented management.

[0243] Step S5: Hybrid storage, composite index, and phased retrieval modeling

[0244] The purpose of this step is to write the sliced ​​text and metadata into suitable storage systems and establish a composite index relationship. During implementation, semantic vectorization is performed on the ChunkText in the ChunkObjectSet to generate a ChunkVector, and the vector records are written to the Elasticsearch vector database. Simultaneously, metadata such as ChunkID, DocID, ChunkText, SectionPath, SourcePosCode, BusinessDomain, and StrategyID are written to MySQL. Subsequently, a primary associative index for ChunkID, a DocID+SectionPath structural index, a BusinessDomain+FileType+ChunkType composite index, and a SourcePosCode reverse lookup index are established, and a retrieval path of "structural filtering—semantic recall—location backtracking" is pre-defined. The output results are VectorStoreSet, MetaStoreSet, and JointIndexMap. This implementation result demonstrates that the present invention can uniformly incorporate the semantic representation and structural constraints of sliced ​​content into a composite index system, consistent with the "dynamic rule slicing + hybrid index storage" technical approach proposed in the project documents.

[0245] Step S6: Slice query, location, retrieval, and backtracking based on data interaction standards

[0246] The purpose of this step is to verify that the constructed composite index can be called by subsequent modules using a unified interface. During implementation, standardized request records are constructed, with request fields including RequestID, QueryBusinessDomain, QuerySectionPath, NeedVectorFlag, and NeedTraceFlag. The system first performs structural filtering based on the business domain and chapter path, then performs semantic matching and dynamic rearrangement on candidate slices, finally returning a standard response with ChunkID, DocID, ChunkText, SectionPath, PageRange, and TraceInfo. For requests with NeedTraceFlag of 1, the system further returns the original document page number and chapter title, enabling backtracking from the slice to the original text. Implementation results show that the retrieved results can be directly provided to subsequent search or question-answering modules, and the returned objects have a unified primary key and a unified location field, indicating that this invention can effectively solve the problem of inconsistent calling standards among multiple modules.

[0247] Step S7: Incremental data entry processing and index update

[0248] The purpose of this step is to verify the continuous evolution capability under scenarios of document addition, replacement, and supplementation. During implementation, 10 new energy storage project feasibility study reports were input into the system as incremental samples. First, the DocHash of the new document was compared with the historical DocumentMasterRecord. If the DocID did not exist, the entire process of adding the document was performed. If the DocID existed and the DocHash changed, the difference identification process was initiated. The ChunkTextHash and positional encoding of the new and old slices were compared to identify added, deleted, and modified slices. Only the differing parts were reconstructed, including the ChunkObject, vector records, and composite indexes. Simultaneously, the update type, update time, and number of affected slices were registered in the VersionAuditRecord. The output results are the updated document master record, slice object set, and composite index mapping table. Implementation shows that this invention can complete new document addition and index repair without reconstructing all historical data, verifying the feasibility of the incremental update chain.

[0249] Although the present invention has been described above with reference to embodiments, various modifications can be made and components can be replaced with equivalents without departing from the scope of the invention. In particular, as long as there is no structural conflict, the features in the disclosed embodiments can be combined with each other in any manner. The lack of an exhaustive description of these combinations in this specification is merely for the sake of brevity and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A method for storing power documents in a database based on a combination of large-scale models and visual recognition, characterized in that, Includes the following steps: S1, Document Access and Basic Database Modeling: Transform power business documents with scattered sources, different formats, and significant structural differences into unified database objects; S2, performs content extraction and large model-assisted structure parsing according to document format: converts different types of documents into a unified structural representation; S3, Adaptive slicing decision based on multi-feature scoring: Adaptively selects slicing granularity based on document type, structural hierarchy, semantic integrity and information density; S4, Slice Management Object Construction and Backtrackable Encoding Modeling: Transform the slices generated in S3 from intermediate results into formal knowledge base management objects; S5, Hybrid Storage, Composite Index and Staged Retrieval Modeling: Separately store the semantic representation of sliced ​​text and structured metadata, and build a composite index; S6, based on data interaction standards, provides unified interface for slice query, location, retrieval and backtracking: based on the composite index built by S5, it provides unified interface for slice query, location, retrieval and backtracking. S7, Incremental Inbound Processing and Index Update: When documents are added, replaced, or supplemented, the changed parts are rebuilt and the index is updated.

2. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S1 are as follows: S11, Perform access pre-verification: sequentially verify whether the file can be opened, whether the file size is within the preset limit, whether the extension is consistent with the file header, and whether the uploaded attribute fields are complete; S12, Perform format recognition: Combine file extension, file header magic number and parsing capability to make a judgment and achieve classification; S13 generates a globally unique document identifier DocID for each document and calculates the content hash value DocHash based on the original file content.

3. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S2 are as follows: S21, Differentiated Content Extraction: For Word and editable PDF: Directly extract the main text, headings, paragraphs, tables, font attributes, paragraph spacing, page number information, and original layout information to form a sequence of original text units; For formatted PDFs: parse text block coordinates, table borders, reading order, header and footer areas, and image / text description areas; For scanned documents: perform OCR recognition, output character or text block sequences, and record the corresponding page coordinates, line numbers, block numbers, and recognition confidence scores; S22, Low-quality text correction: For scanned documents and pages with complex layouts, a character neighborhood correction algorithm is introduced for correction. S23, Large Model Assisted Structure Analysis: The large model is used to assist in the determination of structure labels. First, preliminary structure identification is performed based on the rule method, and then the rule conflict pages are sent to the large model for auxiliary determination. S24, Unified Structure Record and Position Encoding: After completing text extraction and structure determination, a unified structure record is generated, and finally a unified position encoding is generated.

4. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S3 are as follows: S31, Multi-feature slice scoring model: After candidate units are formed, a multi-feature slice scoring model is constructed to calculate the slice granularity score of each candidate unit. ; S32, Granularity Determination and Constraint Rules: Based on Execution granularity determination: when At that time, chapter-level slicing is used; when When using paragraph-level slicing; when At this time, sentence-level slicing is used; At the same time, three types of constraint rules are set: hierarchy preservation rules, semantic closure rules, and table semantic binding rules; S33, Generation of candidate slice set: After completing the granularity determination, a candidate slice set SegmentCandidateSet is generated.

5. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S4 are as follows: S41, Slice object definition: During processing, first assign a unique identifier ChunkID to each candidate slice, and write the slice text, source document, chapter path, page range and position code into a unified object; S42, Three-layer traceable mapping mechanism: Establish three-layer mapping relationships: document layer to slice layer, structure layer to slice layer, and original text layer to slice layer; S43, Slice quality check: Perform slice quality check before writing the object; S44, Output Definition: Outputs a standard slice object collection ChunkObjectSet and a slice metadata registry table ChunkMetaTable.

6. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S5 are as follows: S51, Semantic Vectorization and Hybrid Storage: First, semantic vectorization is performed on ChunkText; after vectorization, the vector records are written to the vector database, and at the same time, the metadata is written to the relational database. S52, Composite Index Construction: Construct a composite index, which should include at least: ChunkID primary association index, DocID+SectionPath structure index, BusinessDomain+FileType+ChunkType composite index, and SourcePosCode reverse lookup index; S53, Phased retrieval model: A phased retrieval model is pre-built during the database entry stage.

7. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S6 are as follows: S61, Request and Response Structure Definition: The request structure is defined as follows: <RequestID,QueryText,QueryDocID,QueryBusinessDomain,QueryFileType,QuerySectionPath,NeedVectorFlag,NeedTraceFlag> NeedVectorFlag indicates whether semantic recall is enabled; NeedTraceFlag indicates whether location backtracking information must be returned. The response structure is defined as follows: <RequestID,ChunkID,DocID,ChunkText,SectionPath,SourcePosCode,PageRange,BusinessDomain,StrategyID,Score_rank,TraceInfo> TraceInfo is further expanded as follows: <SourceDocName,PageID,SectionTitle,ParagraphIndex,TableIndex> ; S62, Exception handling logic: When no candidate slices are found after structural filtering, the business domain constraints can be relaxed and the query can be performed again according to the preset strategy. When all semantic matching scores are below the threshold of 0.55, the system degenerates to only returning relevant chapter slices that match the structure index. If location backtracking fails, at least the ChunkID and DocID should be returned so that the backend can verify the information based on the primary key.

8. The method for power document database entry based on a combination of large model and visual recognition as described in claim 1, characterized in that, The specific steps of S7 are as follows: S71, Incremental Processing Flow: In terms of processing flow, the current document's DocHash is compared with the historical records: If the corresponding DocID does not exist, it is determined to be a newly added document, and the entire process from S1 to S6 is executed; If the DocID exists but the DocHash or VersionTag has changed, it is determined to be a modified document and enters the incremental entry process. When adding new data to the database, differences between new and old slice objects are identified. S72, Version Auditing and Closed Loop: Establishing the VersionAuditRecord table: <VersionID,DocID,OldDocHash,NewDocHash,UpdateType,UpdateTime,OperatorID,AffectedChunkCount,RuleVersion> UpdateType includes at least addition, replacement, supplementation, and rule reslicing; after incremental processing is completed, the new ParseStatus, VersionTag, and update time are written back to DocumentMasterRecord in S1.