A long document-oriented intelligent information extraction and accurate labeling method and system

CN122779019APending Publication Date: 2026-09-18江西省科技基础条件平台中心(江西省计算中心) +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610781143.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

(1)处理效率低下,难以满足多场景需求:传统的信息提取方法多采用同步处理模式,未对任务时效性进行差异化区分,对于需要实时响应的场景(如自动驾驶决策、实时客服问答、紧急公文关键信息提取),该模式存在严重的延迟问题,无法满足毫秒级的响应要求,易导致业务中断或决策失误;而对于非实时处理的场景(如文档归档、批量数据统计、历史文献整理),又缺乏高效的批量处理和结果缓存机制,大量重复任务需重复调用算力资源,导致系统资源浪费,处理效率低下,无法适配海量长文档的批量处理需求

Benefits of technology

1、高效的任务调度能力,兼顾实时与非实时需求:本发明通过动态识别任务类型,对实时任务采用大模型实时接口同步调用模式,对非实时任务采用队列排队+批量处理的策略,既满足了实时任务对毫秒级响应速度的要求,又提高了非实时任务的处理效率,避免了系统资源的浪费,实现了系统资源的最优利用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779019A_ABST
    Figure CN122779019A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence and information processing, and discloses an intelligent information extraction and accurate labeling method and system for long documents, which comprises the following steps: obtaining and preprocessing input data of a long document to be processed, wherein the preprocessing comprises document cutting and position marking; performing information extraction processing and data cleaning processing on the preprocessed input data of the long document to be processed to obtain a structured extraction result; and performing integration processing on the structured extraction result to obtain a final labeling output result. The application avoids the influence of complex information and task types, and improves the accuracy and generalization ability of the long document intelligent information extraction and labeling method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence and information processing, and in particular to an intelligent information extraction and accurate annotation method and system for long documents. Background Technology

[0002] With the rapid development of information technology, the number of long documents (such as academic papers, technical reports, laws and regulations, and corporate annual reports) has exploded. How to efficiently and accurately extract key information from massive long documents and generate structured and annotated documents has become a research hotspot in the fields of artificial intelligence and information processing.

[0003] Existing methods for information extraction from long documents mainly suffer from the following defects and shortcomings: (1) Low processing efficiency and difficulty in meeting the needs of multiple scenarios: Traditional information extraction methods mostly adopt synchronous processing mode and do not differentiate the timeliness of tasks. For scenarios that require real-time response (such as autonomous driving decision-making, real-time customer service Q&A, and extraction of key information from urgent documents), this mode has serious latency problems and cannot meet the millisecond-level response requirements, which can easily lead to business interruption or decision-making errors. For non-real-time processing scenarios (such as document archiving, batch data statistics, and historical document organization), there is a lack of efficient batch processing and result caching mechanisms. A large number of repetitive tasks need to repeatedly call computing resources, resulting in waste of system resources, low processing efficiency, and inability to adapt to the batch processing needs of massive long documents.

[0004] (2) Poor information extraction accuracy and easy to produce "illusion": Long documents have a large amount of content, and direct whole extraction can easily lead to information fragmentation and semantic loss. At the same time, existing technologies lack effective context association and physical location verification mechanisms during the extraction process, which may cause the extraction results to be inconsistent with the original text, i.e., produce "illusion" and reduce the credibility of the information.

[0005] (3) Difficulties in task management and result traceability: When faced with a large number of concurrent tasks, existing technologies lack effective task priority sorting and status tracking mechanisms, resulting in task chaos and low execution efficiency. In addition, the storage and traceability of extracted results rely on complex manual management, making it difficult to quickly locate and reuse historical processing results, which increases the maintenance cost of the system.

[0006] In summary, existing technologies cannot simultaneously meet the requirements of processing efficiency, extraction accuracy, and task management. Therefore, there is an urgent need for a method for intelligent information extraction and annotation of long documents that can support both real-time and non-real-time tasks, and possesses high accuracy and good scalability. Summary of the Invention

[0007] Based on this, the present invention proposes an intelligent information extraction and accurate annotation method and system for long documents. The method involves acquiring input data of the long document to be processed and preprocessing it, including document segmentation and location marking; performing information extraction and data cleaning on the preprocessed input data to obtain structured extraction results; and integrating the structured extraction results to obtain the final annotation output. This invention avoids the influence of complex information and task types, improving the accuracy and generalization ability of the intelligent information extraction and annotation method for long documents.

[0008] This invention proposes an intelligent information extraction and accurate annotation method for long documents, comprising: The input data of the long document to be processed is obtained and preprocessed, including document segmentation and position marking. Information extraction and data cleaning are performed on the preprocessed long document input data to obtain structured extraction results. The information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal. The structured extraction results are integrated to obtain the final labeled output. The integration process is based on the global starting position of the position lookup table and the semantic order of the original long document.

[0009] Furthermore, the step of acquiring and preprocessing the long document input data specifically includes: Obtain input data for a long document to be processed, wherein the input data for the long document to be processed includes unstructured long text of different format types; The long document input data to be processed is divided into multiple text blocks by a segmentation algorithm. The segmentation algorithm is based on semantic integrity and is used to segment complete semantic units according to semantic boundaries. A position lookup table is generated, which includes a chunk ID, a starting character index, and text chunk content. The chunk ID is based on the chunk-X naming rules, and each text chunk has a unique corresponding chunk ID. The starting character index is used to record the starting character position of each text chunk in the original long document. The text chunk content is a snapshot of the original text of each text chunk. Output the location lookup table and text blocks.

[0010] Furthermore, the step of performing information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results specifically includes: Obtain a location lookup table, text blocks, and task-related parameters, wherein the task-related parameters are extraction instructions that indicate the type of the target task. Perform a cloud cache query operation, extract the core features of each text block and calculate the hash value to determine whether there is a cached result with the same hash value in the cloud database; If it is determined that there is a cached result with the same hash value in the cloud database, the extraction result of the text block is directly obtained; if it is determined that there is no cached result with the same hash value in the cloud database, the task type is determined. Based on the characteristics in the task-related parameters, determine the task type of the current extraction task, which includes real-time tasks and non-real-time tasks; If the task type of the current extraction task is determined to be a real-time task, the large model real-time interface synchronous call mode is adopted; if the task type of the current extraction task is determined to be a non-real-time task, the queue queuing + batch processing mode is adopted. The synchronous call mode of the large model real-time interface embeds the segment ID of the text block into the large model metadata field, calls the large model real-time interface, and obtains the extraction results. The queue queuing + batch processing mode embeds the segment ID of the text block into the metadata field of the large model, constructs a complete large model extraction request instruction, and submits it to the large model queue to obtain the extraction result. The extraction results include text content and corresponding tags, and the extraction results are associated with the segment ID; The extracted results are normalized with labels, and then intra-block location retrieval and coordinate calculation are performed. The intra-block location retrieval and coordinate calculation are used to perform character-level retrieval in the original content of the corresponding text block to obtain the relative character coordinates of each extracted entity within the text block. Then, according to the position lookup table, the starting character index corresponding to the text block is extracted, and the global physical coordinates of each extracted entity are calculated using the coordinate conversion formula. An anti-illusion cleaning process is performed, in which the validity of each calculated global physical coordinate is determined. If the relative coordinates within the block of the currently extracted entity are not found and cannot be converted to global coordinates, then the currently extracted entity is determined to be illusory data without original text support, and the illusory data is removed from the extraction results. Original text annotation is performed based on the extraction results after anti-illusion cleaning and the corresponding global physical coordinates. The regions corresponding to small text blocks in the original long document are annotated. The annotation content includes extracted entities, standardized labels and global physical coordinates to obtain the labeled text blocks.

[0011] Furthermore, the step of determining the task type of the current extraction task based on the features in the task-related parameters specifically includes: Receive task-related parameters, including extraction instructions, few-shot examples, text chunks and corresponding chunk IDs. The extraction instructions are used to indicate the specific requirements for information extraction, and the few-shot examples are used to guide the large model to perform extraction. The system performs request instruction concatenation and task type determination. It concatenates extraction instructions, few sample examples, and text chunks to construct a complete large model to extract request instructions, thereby determining extraction requirements and extraction criteria. Based on the real-time nature of the request task requirements, the system automatically identifies the task type. If the request task requires real-time feedback results, the task type is determined to be a real-time task; if the request task requires non-real-time feedback results, the task type is determined to be a non-real-time task. When the task type is determined to be a real-time task, real-time task scheduling is performed. The segment ID in the location lookup table is embedded into the metadata field of the large model call request. Then, a complete large model extraction request instruction is constructed and passed into the large model. The real-time interface of the large model is directly called to initiate a real-time extraction task to obtain the original extraction results corresponding to the text segments. When a task is determined to be a non-real-time task, non-real-time task scheduling is performed. The segment ID from the location lookup table is embedded into the metadata field of the large model call request. A complete large model extraction request instruction is constructed and passed to the large model, which is then submitted to the large model task queue. The large model task queue sorts tasks according to the priority sorting rules specified by the large model service provider. After the current extraction task enters the queue, a polling request is constructed to periodically query the status and ranking of the current extraction task in the service provider's queue. The current extraction task will wait for the large model call tasks ranked higher to complete their execution. When it is the current extraction task's turn, the large model service provider automatically calls the large model to perform the extraction operation. After the extraction operation is completed, the large model service provider sets the status of the current extraction task to complete. When a polling request is initiated to query the task status, the extraction result corresponding to the current extraction task will be obtained. The format of the extraction result is consistent with that of the current extraction task.

[0012] Furthermore, the step of integrating the structured extraction results to obtain the final labeled output specifically includes: Obtain the extraction results and position lookup table of text blocks, wherein the structured extraction results correspond one-to-one with the block IDs in the position lookup table; The extraction results are marked. Based on the segment ID in the position lookup table, the starting character index corresponding to each extraction result is queried. The starting character index and segment ID are used as marking information and added to the header of the corresponding extraction result. The extraction results are associated with the position of the original long document to trace the position of each extraction result. The extraction results are sorted by sorting them in ascending order according to the starting character index of each text block in the position lookup table, so as to restore the order of the extraction results in the original long document. The extracted results are combined by concatenating the sorted results in order and deleting duplicate marker information to optimize the document format and obtain a complete annotated document. The complete annotated document is used as the final annotated output.

[0013] This invention proposes an intelligent information extraction and precise annotation system for long documents, comprising: The document segmentation and position marking module is used to acquire long document input data to be processed and perform preprocessing, the preprocessing including document segmentation and position marking; The text block-oriented intelligent extraction module is used to perform information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results. The information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal. The short text block sorting and combination module is used to integrate the structured extraction results to obtain the final labeled output results. The integration process is based on the global starting position of the position lookup table and the semantic order of the original long document.

[0014] The present invention also provides a storage medium that stores one or more programs, which, when executed by a processor, implement the intelligent information extraction and precise annotation method for long documents as described above.

[0015] The present invention also provides a computer device, the computer device including a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the intelligent information extraction and precise annotation method for long documents as described above.

[0016] Compared with the prior art, the present invention has the following advantages: 1. Highly efficient task scheduling capability, balancing real-time and non-real-time requirements: This invention dynamically identifies task types and adopts a large-model real-time interface synchronous call mode for real-time tasks, while employing a queue queuing + batch processing strategy for non-real-time tasks. This satisfies the millisecond-level response speed requirements of real-time tasks and improves the processing efficiency of non-real-time tasks, avoiding the waste of system resources and achieving optimal utilization of system resources.

[0017] 2. Highly accurate information extraction, effectively avoiding "illusions": By segmenting documents and marking locations, a precise mapping relationship is established between text blocks and the original document. During the extraction process, segment IDs are used for contextual association and physical location verification, ensuring the accuracy of the extraction results, effectively avoiding "illusions," and enhancing the credibility of the information.

[0018] 3. Excellent scalability and maintainability: This invention adopts a modular design, with each functional module being relatively independent, facilitating system development, testing, and maintenance. Simultaneously, a cloud caching mechanism enables the reuse of historical task results, reducing the cost of repetitive calculations.

[0019] 4. Easy to integrate and deploy: The technical solution of this invention has good versatility and compatibility, and can be easily integrated into existing information processing systems. It is suitable for various types of long document information extraction scenarios and has broad application prospects. Attached Figure Description

[0020] Figure 1 The flowchart shows the intelligent information extraction and precise annotation method for long documents proposed in the first embodiment of the present invention. Figure 2 This is a schematic diagram of the intelligent information extraction and precise annotation method for long documents proposed in the second embodiment of the present invention.

[0021] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0022] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0023] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0025] Please see Figure 1 The diagram shows a flowchart of the intelligent information extraction and precise annotation method for long documents proposed in the first embodiment of the present invention. This intelligent information extraction and precise annotation method for long documents includes steps S01 to S03, wherein: Step S01: Obtain the input data of the long document to be processed and perform preprocessing; It should be noted that in this embodiment, the preprocessing includes document segmentation and position marking. The original long document to be processed (unstructured long text, such as long reports, papers, and legal provisions) is used as the core input data. No pre-formatting is required; it is directly passed to the document segmentation and position marking module. This module differs from existing technologies that only perform text segmentation. It focuses on completing semantic segmentation and accurate position mapping of long documents. The final output includes a position lookup table containing unique identifiers, global starting positions, and original text content, as well as multiple semantically complete and independently processable text blocks. This lays the core data foundation for subsequent accurate extraction, illusion verification, and result tracing.

[0026] Addressing the pain points of directly processing long documents: Directly inputting long documents into large models can easily lead to problems such as input limits, semantic loss, and slow response. Semantic block processing can achieve parallel processing of blocks and accurate extraction of blocks, which can significantly improve the overall processing efficiency, unlike the existing technical purpose of only achieving retrieval and recall through block processing. It provides core support for subsequent processes: Recording the global starting character index of each text block provides the core data foundation for subsequent dedicated operations such as global physical coordinate positioning, anti-illusion verification, and original text annotation, ensuring that the extraction results are traceable and verifiable, and making up for the shortcomings of existing technologies that the extraction results cannot be accurately bound to the original text and are difficult to verify.

[0027] In this embodiment, the document segmentation and position marking module is used to realize the semantic segmentation and position mapping of the original long document, construct a complete preprocessing process, and provide basic support for subsequent accurate extraction, sorting and combination, and original text tracing. A specific process example is as follows: The module receives the input raw long document to be processed. This raw long document can be unstructured long text in various formats such as txt, md, and OCR recognition results, without the need for pre-processing of the format. For example, the received original long document to be processed is in txt format, and its content is: "Research in the field of artificial intelligence includes robotics and speech recognition, which are currently the core hot topics in the field of information technology. Facial recognition, as an important branch of artificial intelligence, is widely used in security, payment and other scenarios. Speech recognition technology is mainly used in voice interaction, intelligent translation and other fields to improve the efficiency of human-computer interaction." A semantic integrity-based segmentation algorithm (prioritizing the identification of semantic boundaries of line breaks, periods, and semicolons) is adopted to segment the original long document into several ordered independent text blocks. During the segmentation process, complete semantic units (such as sentences and paragraphs) are avoided, ensuring the semantic coherence of each text block. For example, a semantic integrity-based segmentation algorithm identifies two periods in the original text and divides the document into three semantically complete text chunks: chunk_0: "Research in the field of artificial intelligence includes robotics, speech recognition, etc., and is currently a core hot topic in the field of information technology."; chunk_1: "Facial recognition, as an important branch of artificial intelligence, is widely used in security, payment, and other scenarios."; and chunk_2: "Speech recognition technology is mainly used in voice interaction, intelligent translation, and other fields to improve the efficiency of human-computer interaction." No complete sentences are split into chunks.

[0028] Simultaneously with the segmentation, a position lookup table is generated. This lookup table contains three core fields: Segment ID (assigning a unique identifier to each text chunk using the "chunk-X" naming convention to ensure uniqueness), Starting Character Index (recording the starting character position of each text chunk in the original long document, starting with the first character of the original document as index 0 and incrementing sequentially), and Chunk Content (fully preserving the original text snapshot of each text chunk without any modifications). For example, chunk_0 is assigned a chunk ID of "chunk_0", which starts from the 0th character in the original document (starting character index = 0), and the content of the chunk is "Research in the field of artificial intelligence includes robotics, language recognition, etc., which are the core hot topics in the current information technology field."; chunk_1 starts from the 52nd character (starting character index = 52), and the content of the chunk is the corresponding text, and so on to generate a complete position lookup table.

[0029] The module outputs a location lookup table and multiple text blocks after segmentation. The location lookup table can be stored in tabular form for easy querying and retrieval by subsequent modules. The text blocks are numbered according to the segmentation order and correspond one-to-one with the segment ID in the location lookup table. For example, the output text chunks are chunk_0, chunk_1, and chunk_2, and the position lookup table is: [ { "chunk_id": "chunk_0", "start_offset": 0, "content": "Research in the field of artificial intelligence includes robotics, speech recognition, etc., and is currently a core hot topic in the field of information technology."}, { "chunk_id": "chunk_1", "start_offset": 52, "content": "Face recognition, as an important branch of artificial intelligence, is widely used in security, payment, and other scenarios."}, { "chunk_id": "chunk_2", "start_offset": 98, "content": "Speech recognition technology is mainly used in voice interaction, intelligent translation, and other fields to improve the efficiency of human-computer interaction."} ], with each text chunk number corresponding to a chunk ID, which can be directly called by subsequent modules.

[0030] This module clarifies the core requirements of the segmentation algorithm and the rules for generating the position lookup table, ensuring that the position of each text block is traceable and the content is reusable. It provides core data support for the cache query and coordinate verification of the subsequent text block-oriented intelligent extraction module, as well as the sorting mark of the short text block sorting and combination module, and solves the problem of position loss after long document segmentation.

[0031] Step S02: Perform information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results; It should be noted that in this embodiment, the information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal for accurate extraction and data cleaning. The position comparison table and text blocks output by the document segmentation and position marking module are simultaneously transmitted to the text block-oriented intelligent extraction module. As the core processing module of the system, this module differs from existing extraction technologies that lack caching, differentiated scheduling, and handling of illusions. Combining preset extraction instructions and few sample examples, it first queries and reuses historical extraction results through cloud caching, significantly reducing the number of repeated calls to the large model, lowering computing costs, and improving processing efficiency. If no relevant cached results are found, a differentiated scheduling strategy is used to call the large model to perform information extraction according to the real-time requirements of the task. During the extraction process, label normalization (unifying the output format), intra-block and global coordinate calculation (achieving accurate binding between the extracted results and the original text position), and illusion data removal are completed simultaneously (ensuring that the extracted results are 100% supported by the original text). Finally, a cleaned, uniformly formatted, traceable, and illusion-free structured extraction result is output, providing high-quality data support for subsequent result integration.

[0032] In this embodiment, the text block-oriented intelligent extraction module is used to specifically illustrate the implementation process of the module receiving text blocks and location lookup tables, combining cloud caching and calling large models to complete information extraction. It is the core module for realizing differentiated processing of real-time and non-real-time tasks.

[0033] The specific process example is as follows: It receives input text blocks and a location lookup table, and also receives task-related parameters transmitted by the system (for subsequent task type determination). For example, the received text chunks are chunk_0, chunk_1, and chunk_2, the position lookup table is the lookup table generated in step three above, the task-related parameters are "no timeout requirement, feedback can be delayed" (used to determine if it is a non-real-time task), and the received extraction instructions are "identify relevant research content in the field of artificial intelligence" and a few sample examples "robot".

[0034] First, perform a cloud cache query operation to extract the core features of each text block (text content, extraction command, model version), calculate the hash value, and query the cloud database to see if there is a cached result with the same hash value. For example, the core features of chunk_0 are extracted as "text content: research in the field of artificial intelligence includes robotics, language recognition, etc...; extraction instruction: identify relevant research content in the field of artificial intelligence; model version: V1.0", and a unique hash value "abc123" is calculated. However, when querying the cloud database, no cached result corresponding to this hash value is found (cache miss).

[0035] If the cache query hits (i.e., the same task's extraction result exists in the cloud), the extraction result of that text block is obtained directly without calling the large model, thus improving processing efficiency; if the cache query misses, the task type determination step is entered. For example, if the cache lookup for chunk_0 fails, the task type determination process begins. If a subsequent chunk_0 extraction task with the same characteristics is found and the cache lookup succeeds, the extraction result "tech: [robot, language recognition]" can be obtained directly without calling the large model.

[0036] Automatic task type identification: Based on the characteristics in the task parameters (such as whether timeout requirements are included, whether real-time feedback is required), determine whether the extraction task is a real-time task or a non-real-time task. For example, based on the task parameter "no timeout requirement, feedback can be delayed", the extraction task (for chunk_0, chunk_1, and chunk_2) is determined to be a non-real-time task; if the task parameter is "timeout requirement 100ms, feedback required in real time", it is determined to be a real-time task.

[0037] For different task types, the large model is called to complete the extraction: real-time tasks use the real-time interface synchronous call mode of the large model, and non-real-time tasks use the queue queuing + batch processing mode. In both modes, the segment ID needs to be embedded in the metadata field of the large model call request to ensure that the extraction result corresponds one-to-one with the text segment. For example, since this is a non-real-time task, the chunk ID "chunk_0" of chunk_0 is embedded in the metadata field, and a complete extraction request is constructed and submitted to the large model service provider's queue; if it is a real-time task, the large model real-time interface is called directly to obtain the extraction results synchronously, and "chunk_0" is also embedded in the metadata field.

[0038] After extraction, the extraction results for each text block are output. The extraction results include the text content, corresponding tags, and other information, and are associated with the corresponding block ID for easy sorting and combination in the future. For example, after chunk_0 is extracted, the original extraction result is output as { result: [robot, speech recognition, face recognition], chunk_id: chunk_0}, associated with the chunk ID "chunk_0" to ensure that it corresponds one-to-one with the text chunk.

[0039] Label normalization and adaptation: The output extraction results are standardized and regulated by labeling. The non-standardized labels and messy result formats returned by the large model are adapted to the system's preset standardized label format. For example, the output of the large model, {result: [robot, speech recognition, face recognition]}, is normalized to {tech: [robot, speech recognition, face recognition]}. This ensures that the labels of all extraction results are consistent and the format is standardized, which is convenient for subsequent coordinate calculation and result combination. For example, the original extraction result of chunk_0, { result: [robot, speech recognition, face recognition], chunk_id: chunk_0}, is normalized to { "chunk_id": "chunk_0", "tech": ["robot", "speech recognition", "face recognition"]}, and "tech" is uniformly used as the standardized label.

[0040] Intra-block location retrieval and coordinate calculation: Based on the normalized extraction results, character-level retrieval is performed in the original content of the corresponding text block (i.e., the "block content" in the location lookup table) to obtain the relative character coordinates of each extracted entity within the text block. For example, {Robot: (12-14), Language Recognition: (16-19), Face Recognition: null} is obtained, where null indicates that the entity did not find matching content in the current text block. Subsequently, combined with the location lookup table generated by the document segmentation and location marking module, the starting character index (i.e., Global_Offset) corresponding to the text block is extracted. The global physical coordinates of each entity are calculated using the coordinate conversion formula. The conversion logic is "Global starting coordinate = starting character index + block starting coordinate, global ending coordinate = starting character index + block ending coordinate", for example, {Robot: (60-62), Language Recognition: (64-67), Face Recognition: null} is obtained, realizing the accurate association between the extracted entities and the physical location of the original long document. For example, for chunk_0, whose starting character index is 0, searching for "robot" in its original content yields the relative coordinates (12-14) within the chunk, which are converted to global coordinates (0+12, 0+14) = (12-14); searching for "face recognition" does not find it, the relative coordinates are null, and the global coordinates are also null; finally, the coordinate results of chunk_0 are {robot: (12-14), speech recognition: (16-19), face recognition: null}.

[0041] Anti-illusion cleaning: The validity of each calculated global physical coordinate is checked. If the global physical coordinate of a certain extracted entity is null (i.e., no relative coordinates within the block are found and global coordinates cannot be converted), the entity is determined to be illusory data without original text support and is automatically removed from the extraction results to ensure the accuracy of the extraction results. For example, the global coordinates of "face recognition" in chunk_0 are null, which is determined to be hallucination data and is removed from the extraction results. After removal, the extraction results of chunk_0 are { "chunk_id": "chunk_0", "tech": ["robot","language recognition"], "positions": { "robot": [12, 14], "language recognition": [16, 19]}}.

[0042] Original text annotation: Based on the effective extraction results and corresponding global physical coordinates after anti-illusion cleaning, the region corresponding to the text block in the original long document is accurately annotated. The annotation content includes extracted entities, standardized labels and global physical coordinates, and finally the labeled text block is obtained, which provides the annotated basic data for the subsequent sorting and combination of short text block sorting and combination modules. For example, in the original long document, the "robot" (global coordinates 12-14) corresponding to chunk_0 is labeled "tech-robot, coordinates (12-14)", and the "language recognition" (global coordinates 16-19) is labeled "tech-language recognition, coordinates (16-19)", thus obtaining the labeled chunk_0 text block, which is used for subsequent sorting and combination.

[0043] The text block-oriented intelligent extraction module clearly distinguishes the logical order of cache query and model call, and clarifies the differentiated extraction path for real-time / non-real-time tasks. It reduces redundant calculations through caching mechanisms, and ensures the traceability and accuracy of extraction results through block ID association, label normalization, coordinate calculation and illusion cleaning, effectively improving extraction efficiency and quality.

[0044] In this embodiment, the text block-oriented intelligent extraction module also includes a real-time and non-real-time task large model scheduling module, which is used to implement the large model call scheduling process according to the task type. It is the core sub-module of the text block-oriented intelligent extraction module and solves the task scheduling contradiction with different timeliness requirements.

[0045] The specific process example is as follows: The system receives task information from the text block-guided intelligent extraction module, including extraction instructions, few-shot examples, text blocks and their corresponding chunk IDs (e.g., chunk_1). The extraction instructions specify the requirements for information extraction (e.g., "identify relevant research content in the field of artificial intelligence"), the few-shot examples are used to guide the large model to extract accurately (e.g., few-shot example: robot), and the text blocks are the specific content to be extracted (e.g., chunk_1: research in the field of artificial intelligence includes robots, language recognition, etc.). For example, the received task information is: extraction instruction "identify relevant research content in the field of artificial intelligence", few-sample example "robot", text chunk_1 (content: "Face recognition, as an important branch of artificial intelligence, is widely used in security, payment and other scenarios."), and chunk ID "chunk_1".

[0046] Request instruction assembly and task type determination: First, the extraction instruction, few sample examples, and corresponding text chunks are assembled to construct a complete large model extraction request instruction, ensuring that the large model clearly defines the extraction requirements and extraction criteria; then, based on the real-time requirements of the request task, the task type is automatically identified—if the task requires real-time feedback, it is determined to be a real-time task; if the task does not require real-time feedback and the return of results can be delayed, it is determined to be a non-real-time task. For example, the extraction request instruction for the concatenated large model is "Extraction instruction: Identify relevant research content in the field of artificial intelligence; Few sample examples: Robots; Small text chunk content: Face recognition, as an important branch of artificial intelligence, is widely used in security, payment and other scenarios."; Based on the task requirement "No real-time feedback required, return can be delayed", the task (chunk_1) is determined to be a non-real-time task; if the task requires "real-time feedback within 100ms", it is determined to be a real-time task.

[0047] Real-time task scheduling: For requests identified as real-time tasks, the chunk ID from the location lookup table is embedded into the metadata field (such as the chunk_id field) of the large model call request, ensuring that the chunk ID is bound one-to-one with the call request. Then, the concatenated complete extraction request instruction is passed to the large model, directly calling the large model's real-time interface to initiate the real-time extraction task. The large model's real-time interface responds synchronously, directly returning the original extraction result corresponding to the text chunk. An example of the original result format is { result: [robot, speech recognition, face recognition], chunk_id: chunk_1}. There is no need for task queuing and polling operations, ensuring that the timeliness requirements of real-time tasks are met. For example, if chunk_1 is a real-time task, the chunk ID "chunk_1" is embedded in the metadata field, the concatenated request instruction is passed in, the real-time interface of the large model is called, and the original extraction result { result: [face recognition], chunk_id: chunk_1} is returned synchronously. The whole process takes 50ms, which meets the real-time efficiency requirements.

[0048] Non-real-time task scheduling: For requests determined to be non-real-time tasks, the chunk ID is embedded into the metadata field of the large model call request. A call request containing complete extraction request instructions is constructed, and then the task request is submitted to the task queue of the large model service provider. The tasks are sorted according to the priority sorting rules specified by the large model service provider. After the task enters the queue, the system constructs a polling request to periodically query the status and ranking of the task in the service provider's queue. The task will wait one by one for the large model call tasks ranked higher to complete their execution. When it is the turn of the task, the large model service provider automatically calls the large model to perform the extraction operation. After the extraction operation is completed, the large model service provider sets the status of the task to "completed". When the system initiates a polling request to query the task status, it will obtain the final raw extraction result corresponding to the task. The format of the raw result is consistent with that of the real-time task (e.g., { result: [robot, speech recognition, face recognition], chunk_id: chunk_1}). For example, if chunk_1 is a non-real-time task, embed "chunk_1" into the metadata field, submit the task to the service provider's queue, and sort it as the 3rd task. The system initiates a poll every 10 seconds. After the first two tasks have been completed, this task will start executing. Once the extraction is complete, the status will change to "completed". When the system polls, it will obtain the original extraction result { result: [face recognition], chunk_id: chunk_1}, with the same format as the real-time task.

[0049] The real-time and non-real-time task large model scheduling module clarifies the scheduling logic and processing path of real-time and non-real-time tasks. It ensures the accuracy of extraction by extracting instructions and splicing a small number of sample examples. Through differentiated scheduling strategies, it not only meets the rapid response requirements of real-time tasks, but also achieves orderly and efficient processing of non-real-time tasks through queue sorting and polling. It optimizes the allocation of system resources and solves the problems of low task scheduling efficiency and insufficient extraction accuracy in existing technologies.

[0050] Step S03: Integrate the structured extraction results to obtain the final labeled output results; It should be noted that in this embodiment, the integration process is based on the global starting position of the position lookup table and the semantic order of the original long document. The cleaned and structured extraction results are fed into the short text block sorting and combination module in conjunction with the position lookup table output by the document segmentation and position marking module. This module differs from existing technologies that sort by relevance score and focus on cross-document fusion. It uses the global starting position in the position lookup table as the core basis, and completes the sorting, splicing, and original text annotation of all text block extraction results according to the semantic order of the original long document. It deletes duplicate marking information, optimizes document format, and finally outputs a well-organized, logically coherent, directly usable, and original text-locating complete annotated document, completing the closed-loop process of long document information extraction and annotation.

[0051] In this embodiment, the short text block sorting and combination module is used to sort and combine the extracted results of each text block with the position lookup table to form a complete labeled document. It is a key module for achieving the final labeled output.

[0052] The specific process example is as follows: The system receives input text blocks and extracts the answer results and a position lookup table. The extracted answer results are associated with corresponding block IDs, which correspond one-to-one with the block IDs in the position lookup table. For example, the received extraction results are chunk_0: { "chunk_id": "chunk_0", "tech": ["robot", "language recognition"], "positions": { "robot": [12, 14], "language recognition": [16, 19]}}, chunk_1: { "chunk_id": "chunk_1", "tech": ["face recognition"], "positions": { "face recognition": [52, 55]}}, chunk_2: { "chunk_id": "chunk_2", "tech": ["language recognition"],"positions": { "language recognition": [98, 101]}}. The received position lookup table is the lookup table generated above that contains information on the three text chunks. The chunk IDs of the extraction results correspond one-to-one with the lookup table.

[0053] Extraction result labeling: Using the segment ID in the position lookup table, query the starting character index corresponding to each extraction result, and add the starting character index and segment ID as labeling information to the header of the corresponding extraction result to achieve accurate association between the extraction result and the position of the original long document, ensuring that the position of each extraction result is traceable; For example, by querying the position lookup table using the chunk ID "chunk_0", we find that its starting character index is 0. We then add "chunk_0 (starting index 0)" as a tag to the header of the chunk_0 extraction result. Similarly, we add the tag "chunk_1 (starting index 52)" to chunk_1 and "chunk_2 (starting index 98)" to chunk_2.

[0054] Sorting of extraction results: Based on the starting character index of each segment in the position lookup table, all marked extraction results are sorted in ascending order to restore the order of each extraction result in the original long document and avoid the problem of disordered extraction results; For example, the starting character indices of the three extracted results are 0 (chunk_0), 52 (chunk_1), and 98 (chunk_2), respectively. After sorting in ascending order, the order is chunk_0→chunk_1→chunk_2, which is completely consistent with the text order of the original long document.

[0055] Extraction result combination: The sorted extraction results are concatenated in order, duplicate marker information is removed, the document format is optimized, and the concatenated document is semantically coherent and logically clear to form a complete annotation document. For example, the sorted extraction results of chunk_0, chunk_1, and chunk_2 are concatenated in order, duplicate and redundant label information is deleted, and the format is optimized to form a coherent annotation content containing all extracted entities and coordinate information.

[0056] The module outputs a complete annotation document, which contains the extraction results of all text blocks and their corresponding location markers, and can be directly used for subsequent review and use. For example, the output of the complete annotated document is { "original_document": "Research in the field of artificial intelligence includes robotics, speech recognition, etc., and is currently a core hot topic in the field of information technology. Face recognition, as an important branch of artificial intelligence, is widely used in security, payment, and other scenarios. Speech recognition technology is mainly used in voice interaction, intelligent translation, and other fields to improve the efficiency of human-computer interaction.", "annotations": [ { "chunk_id": "chunk_0", "start_offset": 0, "tech": ["robotics", "speech recognition"], "positions": { "robotics": [12,14], "speech recognition": [16, 19]}}, { "chunk_id": "chunk_1", "start_offset": 52, "tech": ["face recognition"], "positions": { "face recognition": [52, 55]}}, { "chunk_id": "chunk_2", "start_offset": 98, "tech": ["speech recognition"], "positions": { "speech recognition": [98, 101]}} ]} can be used directly for browsing and archiving.

[0057] The short text block sorting and combination module clarifies the core steps of marking, sorting, and combining. Relying on the position lookup table, it achieves accurate sorting and position association of the extracted results, solving the pain point of disordered results and inability to restore the original logic after long document segmentation and extraction, and ensuring that the final output labeled document is complete, accurate, and usable.

[0058] Please see Figure 2 The diagram shows a schematic of the intelligent information extraction and accurate annotation system for long documents proposed in the second embodiment of the present invention. The system includes: The document segmentation and position marking module 10 is used to acquire input data of a long document to be processed and perform preprocessing, the preprocessing including document segmentation and position marking. The text block-oriented intelligent extraction module 20 is used to perform information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results. The information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal. The short text block sorting and combination module 30 is used to integrate the structured extraction results to obtain the final annotation output results. The integration process is based on the global starting position of the position lookup table and the semantic order of the original long document.

[0059] The present invention also proposes a computer storage medium storing one or more programs, which, when executed by a processor, implement the above-described intelligent information extraction and precise annotation method for long documents.

[0060] The present invention also proposes a computer device, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory to realize the above-mentioned intelligent information extraction and accurate annotation method for long documents.

[0061] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can mean any means that can contain stored, communicated, propagated, or transmitted programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0062] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0063] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0064] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0065] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of this patent should be determined by the appended claims.

Claims

1. A method for intelligent information extraction and precise annotation of long documents, characterized in that, include: The input data of the long document to be processed is obtained and preprocessed, including document segmentation and position marking. Information extraction and data cleaning are performed on the preprocessed long document input data to obtain structured extraction results. The information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal. The structured extraction results are integrated to obtain the final labeled output. The integration process is based on the global starting position of the position lookup table and the semantic order of the original long document.

2. The intelligent information extraction and precise annotation method for long documents according to claim 1, characterized in that, The steps of acquiring and preprocessing the input data of the long document to be processed specifically include: Obtain input data for a long document to be processed, wherein the input data for the long document to be processed includes unstructured long text of different format types; The long document input data to be processed is divided into multiple text blocks by a segmentation algorithm. The segmentation algorithm is based on semantic integrity and is used to segment complete semantic units according to semantic boundaries. A position lookup table is generated, which includes a chunk ID, a starting character index, and text chunk content. The chunk ID is based on the chunk-X naming rules, and each text chunk has a unique corresponding chunk ID. The starting character index is used to record the starting character position of each text chunk in the original long document. The text chunk content is a snapshot of the original text of each text chunk. Output the location lookup table and text blocks.

3. The intelligent information extraction and precise annotation method for long documents according to claim 1, characterized in that, The steps of performing information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results specifically include: Obtain a location lookup table, text blocks, and task-related parameters, wherein the task-related parameters are extraction instructions that indicate the type of the target task. Perform cloud cache query operations, extract the core features of each text block and calculate the hash value to determine whether there is a cached result with the same hash value in the cloud database; If it is determined that there is a cached result with the same hash value in the cloud database, the extraction result of the text block is directly obtained; if it is determined that there is no cached result with the same hash value in the cloud database, the task type is determined. Based on the characteristics in the task-related parameters, determine the task type of the current extraction task. The task type includes real-time tasks and non-real-time tasks. If the task type of the current extraction task is determined to be a real-time task, the large model real-time interface synchronous call mode is adopted; if the task type of the current extraction task is determined to be a non-real-time task, the queue queuing + batch processing mode is adopted. The synchronous call mode of the large model real-time interface embeds the segment ID of the text block into the large model metadata field, calls the large model real-time interface, and obtains the extraction results. The queue queuing + batch processing mode embeds the segment ID of the text block into the metadata field of the large model, constructs a complete large model extraction request instruction, and submits it to the large model queue to obtain the extraction result. The extraction results include text content and corresponding tags, and the extraction results are associated with the segment ID; The extracted results are normalized with labels, and then intra-block location retrieval and coordinate calculation are performed. The intra-block location retrieval and coordinate calculation are used to perform character-level retrieval in the original content of the corresponding text block to obtain the relative character coordinates of each extracted entity within the text block. Then, according to the position lookup table, the starting character index corresponding to the text block is extracted, and the global physical coordinates of each extracted entity are calculated using the coordinate conversion formula. An anti-illusion cleaning process is performed, in which the validity of each calculated global physical coordinate is determined. If the relative coordinates within the block of the currently extracted entity are not found and cannot be converted to global coordinates, then the currently extracted entity is determined to be illusory data without original text support, and the illusory data is removed from the extraction results. Original text annotation is performed based on the extraction results after anti-illusion cleaning and the corresponding global physical coordinates. The regions corresponding to small text blocks in the original long document are annotated. The annotation content includes extracted entities, standardized labels and global physical coordinates to obtain the labeled text blocks.

4. The intelligent information extraction and precise annotation method for long documents according to claim 1, characterized in that, The step of determining the task type of the current extraction task based on the features in the task-related parameters specifically includes: Receive task-related parameters, including extraction instructions, few-shot examples, text chunks and corresponding chunk IDs. The extraction instructions are used to indicate the specific requirements for information extraction, and the few-shot examples are used to guide the large model to perform extraction. The system performs request instruction concatenation and task type determination. It concatenates extraction instructions, few sample examples, and text chunks to construct a complete large model to extract request instructions, thereby determining extraction requirements and extraction criteria. Based on the real-time nature of the request task requirements, the system automatically identifies the task type. If the request task requires real-time feedback results, the task type is determined to be a real-time task; if the request task requires non-real-time feedback results, the task type is determined to be a non-real-time task. When the task type is determined to be a real-time task, real-time task scheduling is performed. The segment ID in the location lookup table is embedded into the metadata field of the large model call request. Then, a complete large model extraction request instruction is constructed and passed into the large model. The real-time interface of the large model is directly called to initiate a real-time extraction task to obtain the original extraction results corresponding to the text segments. When a task is determined to be a non-real-time task, non-real-time task scheduling is performed. The segment ID from the location lookup table is embedded into the metadata field of the large model call request. A complete large model extraction request instruction is constructed and passed to the large model, which is then submitted to the large model task queue. The large model task queue sorts tasks according to the priority sorting rules specified by the large model service provider. After the current extraction task enters the queue, a polling request is constructed to periodically query the status and ranking of the current extraction task in the service provider's queue. The current extraction task will wait for the large model call tasks ranked higher to complete their execution. When it is the current extraction task's turn, the large model service provider automatically calls the large model to perform the extraction operation. After the extraction operation is completed, the large model service provider sets the status of the current extraction task to complete. When a polling request is initiated to query the task status, the extraction result corresponding to the current extraction task will be obtained. The format of the extraction result is consistent with that of the current extraction task.

5. The intelligent information extraction and precise annotation method for long documents according to claim 1, characterized in that, The step of integrating the structured extraction results to obtain the final labeled output results specifically includes: Obtain the extraction results and position lookup table of text blocks, wherein the structured extraction results correspond one-to-one with the block IDs in the position lookup table; The extraction results are marked. Based on the segment ID in the position lookup table, the starting character index corresponding to each extraction result is queried. The starting character index and segment ID are used as marking information and added to the header of the corresponding extraction result. The extraction results are associated with the position of the original long document to trace the position of each extraction result. The extraction results are sorted by sorting them in ascending order according to the starting character index of each text block in the position lookup table, so as to restore the order of the extraction results in the original long document. The extracted results are combined by concatenating the sorted results in order and deleting duplicate marker information to optimize the document format and obtain a complete annotated document. The complete annotated document is used as the final annotated output.

6. A smart information extraction and precise annotation system for long documents, characterized in that, include: The document segmentation and position marking module is used to acquire long document input data to be processed and perform preprocessing, the preprocessing including document segmentation and position marking; The text block-oriented intelligent extraction module is used to perform information extraction and data cleaning on the preprocessed long document input data to obtain structured extraction results. The information extraction process includes label normalization, intra-block and global coordinate calculation, and illusion data removal. The short text block sorting and combination module is used to integrate the structured extraction results to obtain the final labeled output results. The integration process is based on the global starting position of the position lookup table and the semantic order of the original long document.

7. A storage medium, characterized in that, The storage medium stores one or more programs, which, when executed by a processor, implement the intelligent information extraction and precise annotation method for long documents as described in any one of claims 1-5.

8. A computer device, characterized in that, The computer device includes a memory and a processor, wherein: The memory is used to store computer programs; When the processor executes the computer program stored in the memory, it implements the intelligent information extraction and precise annotation method for long documents as described in any one of claims 1-5.