Professional knowledge corpus processing method and system

By employing a systematic approach to processing professional knowledge corpora, the issues of completeness, consistency, and accuracy of professional knowledge corpora have been resolved, generating high-quality data suitable for training professional models and supporting various downstream application scenarios.

CN122019674APending Publication Date: 2026-05-12QIANGWEI BIAOTANG (SHANGHAI) INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610184810.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-09
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

When existing technologies struggle to effectively process professional knowledge corpora, they suffer from issues such as difficulty in ensuring completeness, maintaining consistency, and the tendency for the processing to become distorted. This results in semantically incomplete knowledge fragments that fail to meet the requirements for training professional models.

Method used

A systematic approach is adopted, through content identification, transformation and reorganization steps, to transform unstructured professional knowledge corpora into structured knowledge units, including content identification, structured processing, verification and assembly, to ensure the integrity, consistency and correctness of the corpus.

Benefits of technology

It achieves accurate parsing and lossless conversion of professional knowledge corpora, generating semantically complete and logically coherent knowledge units that meet the high standards required for professional model training, and improves processing efficiency and controllability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019674A_ABST
    Figure CN122019674A_ABST
Patent Text Reader

Abstract

The invention provides a professional knowledge corpus processing method and system. The method comprises the following steps: acquiring a to-be-processed professional knowledge corpus; the professional knowledge corpus is subjected to structured processing so as to convert non-structured professional knowledge content into structured knowledge units, and the structured processing comprises the steps of content recognition, content recognition and classification are carried out on the professional knowledge corpus according to page units, and the content recognition and classification are carried out on the professional knowledge corpus according to the page units; determining types and positions of various content elements in the page; a content conversion and recombination step: based on a result of the content identification step, converting and recombining the classified content elements according to a preset structured rule, and generating an intermediate document with a uniform structure mark; and checking and assembling the knowledge content subjected to the structured processing to form a standardized knowledge corpus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of predictive processing technology, and more specifically, to a method and system for processing professional knowledge corpora. Background Technology

[0002] In the field of artificial intelligence, high-quality professional knowledge corpora are crucial for training domain-specific models. 1.2 Professional knowledge corpora are systematically organized and labeled domain-specific language materials used to support research, teaching, or technological applications in a professional field. These corpora originate from specific industries such as medicine and law and contain a large amount of professional terminology, complex logic, and multimodal content (such as text, tables, and charts). For example, in the medical field, the corpus "three elements and one soup" refers to mixed drugs such as antibiotics.

[0003] Currently, automated processing technology for general text is relatively mature, but it still faces significant shortcomings when applied to specialized knowledge corpora. First, completeness is difficult to guarantee. Professional documents have complex structures, often including elements such as tables of contents, appendices, and tables spanning multiple pages. Common processing methods can easily disrupt their inherent structure, leading to missing elements or breaks in context.

[0004] Second, maintaining consistency is difficult. Professional fields require strict uniformity in the expression of concepts, but existing methods lack the ability to deeply identify and maintain the internal logical structure of documents (such as chapter hierarchy and chart reference relationships), which can easily lead to inconsistencies in terminology or misplaced references.

[0005] Third, the processing is prone to distortion. Preprocessing based on simple rules or formats makes it difficult to identify the semantic boundaries and functional attributes of the content (such as distinguishing between titles and body text, and determining whether tables span multiple pages), resulting in incomplete semantics in the generated knowledge fragments, which cannot meet the data fidelity requirements of professional model training.

[0006] Therefore, a systematic method for extracting professional knowledge corpora is needed to achieve accurate parsing, lossless conversion, and standardized reorganization of multimodal professional documents, thereby ensuring the integrity, consistency, and correctness of the output corpora and providing a reliable foundation for building high-quality domain knowledge bases and models. Summary of the Invention

[0007] In view of the deficiencies in the prior art, the purpose of this invention is to provide a method and system for processing professional knowledge corpora.

[0008] A method for processing professional knowledge corpora according to the present invention includes the following steps: Acquire the professional knowledge corpus to be processed; The professional knowledge corpus undergoes structured processing to transform unstructured professional knowledge content into structured knowledge units. This structured processing includes: Content recognition steps: The professional knowledge corpus is classified and identified by page unit to determine the type and position of various content elements on the page; Content conversion and reorganization steps: Based on the results of the content recognition steps, the classified content elements are converted and reorganized according to preset structured rules to generate an intermediate document with unified structure tags; The knowledge content that has undergone the structured processing is verified and assembled to form a standardized knowledge corpus.

[0009] Preferably, the content recognition step specifically includes: The professional knowledge corpus is segmented by page; Identify content elements in each page, wherein the content elements include at least one of the following: text, table, image, title, header, and footer; Based on the identified content element type, the content elements are batched and their locations are identified.

[0010] Preferably, the identification of content elements in each page is achieved through optical character recognition technology.

[0011] Preferably, the content recognition step further includes: identifying the document structure of the professional knowledge corpus, wherein the document structure includes at least one of a table of contents, main text, and appendix.

[0012] Preferably, the content conversion and reorganization steps specifically include: The identified unstructured content elements are arranged into predefined formatted content blocks according to page and content type; Based on the document structure, the formatted content blocks are logically organized and marked to generate the intermediate document.

[0013] Preferably, the intermediate document is a document with structured tags.

[0014] Preferably, the structured processing, prior to the content recognition step, further includes: Initialization steps: Check the access permissions and time synchronization status of the storage container; create a separate storage directory for each processing item; start the processing process to monitor the task status; Verification steps: Verify the identity and permissions of the processing request.

[0015] Preferably, the identity and permission verification is implemented by verifying a JSON Web Token.

[0016] Preferably, the structured processing further includes a task execution and monitoring step, used to execute processing tasks concurrently and record the status.

[0017] A system for processing professional knowledge corpora according to the present invention includes: The corpus acquisition module is used to acquire professional knowledge corpora to be processed. A structured processing module is used to perform structured processing on the professional knowledge corpus, so as to transform unstructured professional knowledge content into structured knowledge units. The structured processing module includes: The content recognition unit is used to perform content recognition and classification on the professional knowledge corpus according to page units, and to determine the type and position of various content elements on the page. The content conversion and reorganization unit is used to convert and reorganize the classified content elements according to preset structured rules based on the results of the content recognition unit, and generate an intermediate document with unified structure tags. The quality control module is used to verify and assemble the knowledge content processed by the structured processing module to form a standardized knowledge corpus.

[0018] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention effectively preserves the complex structure of the original professional document (such as table of contents, appendices, cross-page tables, etc.) by performing fine-grained content identification and classification by page unit and establishing logical connections between cross-page content, avoiding missing elements and context breaks, and ensuring the integrity of the output corpus.

[0019] 2. During the structured processing, by identifying the macro structure of the document (such as the table of contents and the main text) and the relationship between micro elements, and by logically reorganizing it according to preset rules, the consistency of professional terminology, concept expression and citation relationship can be strictly maintained, which significantly reduces the risk of terminology confusion or citation misplacement.

[0020] 3. This invention adopts a strategy that combines content prediction, classification and deep structured transformation, which can accurately identify the semantic boundaries and functional attributes of content (such as distinguishing between titles and body text, and determining the belonging of tables), thereby generating semantically complete and logically coherent knowledge units, meeting the high standard requirements of professional model training for data fidelity.

[0021] 4. This invention integrates end-to-end control from initialization verification and concurrent task execution to multi-level quality checks, and achieves standardization and automation of the processing through modular design. This not only improves processing efficiency but also ensures the controllability, monitorability, and traceability of the processing, making it suitable for large-scale engineering deployment.

[0022] 5. The standardized and modular knowledge corpus produced by this invention can directly support various downstream application scenarios such as high-quality question-answer pair extraction, fine-tuning of large language models, knowledge base / graph construction, and intelligent agent invocation, providing a high-quality and ready-to-use data foundation for artificial intelligence applications in professional fields. Attached Figure Description

[0023] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the hierarchical management process of incoming professional knowledge corpus in this embodiment of the invention. Figure 2 This is a logic diagram for hierarchical processing of professional knowledge corpus in an embodiment of the present invention; Figure 3 This is a flowchart illustrating the hierarchical inspection process for incoming professional knowledge corpus in this embodiment of the invention. Figure 4 This is a diagram of the assembly of minimal units of professional knowledge corpus in an embodiment of the present invention; Figure 5 This is a diagram illustrating the scenario-based application of the assembled professional knowledge corpus in an embodiment of the present invention. Detailed Implementation

[0024] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0025] This invention provides an engineering method for extracting professional knowledge corpora in specialized knowledge domains. The core of this method lies in transforming raw, unstructured multimodal professional knowledge documents into high-quality, structured standard corpora that can be directly used for AI model training or knowledge base construction through a systematic engineering process. This method particularly emphasizes maintaining the integrity, consistency, and correctness of the original knowledge during the transformation process. The following describes the process in conjunction with... Figures 1-5 The various steps of the method of the present invention will be described in detail below.

[0026] Step 1: Management of Incoming Professional Knowledge Corpus This step aims to standardize and initially organize raw professional knowledge corpora from diverse sources. Professional knowledge corpora refer to systematic and practical knowledge carriers within specific industries (such as medicine, finance, law, and engineering), including professional terminology, industry jargon (e.g., the medical term "three elements and one soup" refers to mixed antibiotics and other medications), and the complex sentence structures and logical connections characteristic of academic language. Their sources include publicly available external materials (books, manuals) and non-public materials (internal standards, specifications, system documents, etc.), as well as documents stored in internal application systems, archives, or scattered throughout the system.

[0027] First, incoming materials are classified in multiple dimensions, including: document type classification based on file extension (e.g., .pdf, .docx, .xlsx); business type classification based on the business scope to which the corpus belongs; management type classification based on management requirements (e.g., confidentiality level, version status); and knowledge type classification based on knowledge attributes (e.g., concept definition, operating procedure, case analysis).

[0028] Secondly, the categorized corpora are identified and fused. After being identified and categorized, the external corpora are fused with the categorized corpora from internal sources to unify the entry point for subsequent processing.

[0029] Finally, the fused corpus is stored as files to prepare for the next stage of processing. The management workflow for this process can be found in the appendix. Figure 1 The flowchart shown is for hierarchical management of incoming professional knowledge corpus.

[0030] Step Two: Processing of Professional Knowledge Corpus This step is the core of the method, responsible for converting the unstructured raw corpus into an intermediate form with clearly defined structural markers. The processing follows a hierarchical logic, mainly including content prediction and recognition, preprocessing and postprocessing, and tokenization combination.

[0031] 2.1 Initialization and Permission Verification Before processing begins, the system performs initialization operations. First, it checks the system time synchronization status and access permissions of the storage container (such as MinIO object storage). Next, it creates a separate directory structure based on the current processing project. Simultaneously, a background task monitoring process is started to manage and monitor the lifecycle of the processing tasks.

[0032] When a processing component is invoked, the request must first be validated using a token to ensure its legitimacy. The system verifies the validity of the passed-in JSON Web Token (JWT), and only after successful verification is the subsequent processing logic allowed to proceed.

[0033] 2.2 Content Recognition The goal of the identification phase is to deeply analyze the structure and composition of the document content.

[0034] Content Prediction and Classification: The system first segments the document by physical pages. Then, it uses technologies such as Optical Character Recognition (OCR) to identify the content of each page and predict its content type. The system counts the number of various elements on each page, such as: total number of pages, number of text blocks, number of tables, number of images, number of titles, number of page headers, number of page footers, and number of footnotes.

[0035] Batch processing: Based on the above classification statistics, the system divides similar content elements into processing batches. For example, all text blocks are divided into batches such as Text1, Text2, ..., TextN; all tables are divided into batches such as Table1, Table2, ..., TableN, and so on.

[0036] Position and Relationship Identification: The system identifies the start (PageN_Start) and end (PageN_End) positions of each page's content, using a page-by-page approach. Furthermore, based on the page order and the order in which elements appear within a page, the system infers the continuity relationship between content (Page_Content_Next), thereby establishing logical connections between content across pages. For example, by analyzing the position of elements within a page, semantic coherence, and page number order, logical links are established between tables, consecutive paragraphs, charts, and explanatory text across pages. Document structure identifiers: The system identifies the macro structure of the document, for example, identifying the table of contents as "M" and the main body as "Z". The main body is further identified by detailed identifiers, such as the main content (Z_Content), main title (Z_Title), references (Z_fNote), appendices (Z_fContent), and tables (Z_fTable).

[0037] 2.3 Content Conversion The transformation phase is responsible for reorganizing the identified non-normalized content into a normalized structure.

[0038] Preprocessing: This involves initially organizing the unstructured content of each page according to a predefined format template. For example, for Page1, a block from [Page1_Start] to [Page1_End] is generated, and the header, text, headings, tables, footers, and other elements and their content are arranged sequentially within it.

[0039] Post-processing: Based on the results of preprocessing and the structural information identified in the early stages, deep structural transformation is performed. For example, determining whether it is a table of contents or the main text, and then performing logical merging across pages based on relational functions.

[0040] For the table of contents (identified as M), the system calculates and determines its position in the final document based on its page number and the content relationship function f(X, Page_Content_Next).

[0041] For the main text (identified as Z), the system calls the content processing function Content_Process(Z_Content, Z_Title, Z_fNote, Z_fContent, Z_fTable) to process each component of the main text. Then, combined with the page number and content relationship function F(X, Page_Content_Next, Content_Process), the final conversion from unstructured to structured text is completed.

[0042] 2.3.3 MD Combination: The standardized content, after transformation and containing complete batch information and structured tags, is combined to generate a unified, structured intermediate document. For the specific processing logic, please refer to the appendix. Figure 2 .

[0043] 2.4 Task Execution and Monitoring This stage is responsible for scheduling and executing processing tasks. The system is based on concurrent programming models such as Go, utilizing mechanisms like `sync.WaitGroup` to manage multiple concurrent processing tasks, ensuring efficient and orderly execution. The system records the execution status, progress, and results of each task throughout the process, facilitating monitoring and problem tracing.

[0044] Step 3: Verification of professional knowledge corpus To ensure processing quality, this step involves multi-level manual and systematic verification of the processed corpus.

[0045] Task allocation: The processed corpus is assigned to authorized reviewers in the form of review tasks. Reviewers include internal experts, external suppliers (requiring additional authorization and approval), and commissioned third parties.

[0046] Verification process: Verifiers log into the system, receive tasks, and review the corpus, focusing on its completeness, correctness, and consistency.

[0047] Status Transition and Logs: After verification is completed and the results are submitted, the corpus status changes from "Under Inspection" to "Pending Review." Once the review is passed, the entire verification process ends. The system records a complete inspection log from task acquisition and verification operation to result submission. See the appendix for the detailed verification process. Figure 3 .

[0048] Step 4: Assembly of professional knowledge corpus This step aims to build the smallest application-oriented knowledge unit based on the verified structured intermediate documents.

[0049] Minimum Knowledge Unit Segmentation: The system intelligently segments intermediate documents with batch and category identifiers based on their actual semantic content. The composition of the minimum knowledge unit is dynamic and determined by the content itself. For example, a mathematical formula spanning multiple pages and its explanatory text are consolidated into one unit after removing headers and footers; a multi-page table is consolidated into another unit after removing irrelevant page elements.

[0050] Knowledge Reorganization: The assembly process involves hierarchically and logically reorganizing these smallest knowledge units. The aim is to restore and optimize the knowledge structure, hierarchy, and relationships of the original document, ultimately forming a knowledge package that retains the original meaning while being suitable for computational processing. This achieves completeness, accuracy, and consistency in corpus extraction. A diagram illustrating the assembly process can be found in the appendix. Figure 4 .

[0051] Step 5: Application of Professional Knowledge Corpus The resulting standardized and modular knowledge corpus can flexibly support a variety of downstream scenarios.

[0052] In a preferred embodiment, this invention can be directly used to extract high-quality question-answer pairs (QAPairs) from documents. It can serve as high-quality training data for fine-tuning large language models in specialized domains. It can be imported into knowledge base systems or used to construct domain knowledge graphs, providing knowledge support for intelligent question answering, decision support, and other systems. It can be invoked by various intelligent agents as the knowledge basis for performing their specialized tasks. Corpora from different scenarios can be used individually or fused together as needed. See the appendix for application scenario diagrams. Figure 5 .

[0053] In summary, this invention, through a systems engineering method consisting of the above five steps, systematically solves the problems of distortion, inconsistency, and fragmentation in the processing of professional knowledge corpora, providing a reliable data foundation for the in-depth application of artificial intelligence in professional fields.

[0054] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for processing professional knowledge corpora, characterized in that, Includes the following steps: Acquire the professional knowledge corpus to be processed; The professional knowledge corpus undergoes structured processing to transform unstructured professional knowledge content into structured knowledge units. This structured processing includes: Content recognition steps: The professional knowledge corpus is classified and identified by page unit to determine the type and position of various content elements on the page; Content conversion and reorganization steps: Based on the results of the content recognition steps, the classified content elements are converted and reorganized according to preset structured rules to generate an intermediate document with unified structure tags; The knowledge content that has undergone the structured processing is verified and assembled to form a standardized knowledge corpus.

2. The method for processing professional knowledge corpora according to claim 1, characterized in that, The content recognition steps specifically include: The professional knowledge corpus is segmented by page; Identify content elements in each page, wherein the content elements include at least one of the following: text, table, image, title, header, and footer; Based on the identified content element type, the content elements are batched and their locations are identified.

3. The method for processing professional knowledge corpora according to claim 2, characterized in that, The identification of content elements in each page is achieved through optical character recognition technology.

4. The method for processing professional knowledge corpora according to claim 2, characterized in that, The content recognition step further includes: identifying the document structure of the professional knowledge corpus, wherein the document structure includes at least one of a table of contents, main text, and appendix.

5. The method for processing professional knowledge corpora according to claim 1, characterized in that, The content conversion and reorganization steps specifically include: The identified unstructured content elements are arranged into predefined formatted content blocks according to page and content type; Based on the document structure, the formatted content blocks are logically organized and marked to generate the intermediate document.

6. The method for processing professional knowledge corpora according to claim 5, characterized in that, The intermediate document is a document with structured tags.

7. The method for processing professional knowledge corpora according to any one of claims 1 to 6, characterized in that, The structured processing, prior to the content recognition step, also includes: Initialization steps: Check the access permissions and time synchronization status of the storage container; create a separate storage directory for each processing item; start the processing process to monitor the task status; Verification steps: Verify the identity and permissions of the processing request.

8. The method for processing professional knowledge corpora according to claim 7, characterized in that, The identity and permission verification is achieved by verifying the JSON Web Token.

9. The method for processing professional knowledge corpora according to any one of claims 1 to 6, characterized in that, The structured processing also includes task execution and monitoring steps, which are used to execute processing tasks concurrently and record the status.

10. A system for processing professional knowledge corpora, characterized in that, include: The corpus acquisition module is used to acquire professional knowledge corpora to be processed. A structured processing module is used to perform structured processing on the professional knowledge corpus, so as to transform unstructured professional knowledge content into structured knowledge units. The structured processing module includes: The content recognition unit is used to perform content recognition and classification on the professional knowledge corpus according to page units, and to determine the type and position of various content elements on the page. The content conversion and reorganization unit is used to convert and reorganize the classified content elements according to preset structured rules based on the results of the content recognition unit, and generate an intermediate document with unified structure tags. The quality control module is used to verify and assemble the knowledge content processed by the structured processing module to form a standardized knowledge corpus.