Intelligent PDF document retrieval method and system combining OCR recognition

By combining OCR recognition technology with an information platform to perform term conversion and OCR enhancement, a dynamic search chain is generated, which solves the problem that traditional PDF document search methods have difficulty in handling image-based text, and achieves more accurate and efficient search results.

CN121030070BActive Publication Date: 2026-04-03BEIJING GUANGLIANDA YUNTU DREAM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Traditional PDF document retrieval methods struggle to handle image-based text content in multi-source, heterogeneous PDFs, resulting in incomplete and weakly correlated search results that fail to meet the need for accurate and efficient retrieval.

Method used

By combining OCR recognition technology, search terms are input into an information platform, term conversion and OCR enhancement are performed, a dynamic search chain based on concept change paths is generated, and dynamic OCR retrieval is performed in the document database to determine the PDF search form, thus realizing intelligent retrieval.

Benefits of technology

It enables effective processing of image-based text content in multi-source heterogeneous PDFs, improving the comprehensiveness and relevance of search results and meeting the needs for accurate and efficient retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121030070B_ABST
    Figure CN121030070B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent PDF document retrieval method and system combining OCR recognition, belonging to the field of optical character recognition technology. The method includes: inputting search terms into an information platform, performing term conversion and OCR enhancement to determine the enhanced term system; setting a look-skip retrieval mechanism to generate a dynamic retrieval chain based on concept change paths; and finally writing the results to a register via an interactive thread, performing dynamic OCR retrieval in the document database, and displaying the PDF retrieval form in a pop-up window. This invention solves the technical problem that traditional PDF document retrieval methods struggle to handle image-based text and other content in multi-source heterogeneous PDFs, resulting in one-sided and weakly correlated retrieval results that fail to meet the requirements of accurate and efficient retrieval. It achieves effective processing of image-based text and other content in multi-source heterogeneous PDFs, making the retrieval results more comprehensive and correlated, thus achieving the technical effect of accurate and efficient retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of optical character recognition technology, and in particular to a method and system for intelligent retrieval of PDF documents that combines OCR recognition. Background Technology

[0002] In the information age, PDF documents are widely used as important information carriers, and their efficient retrieval is crucial for data processing and information acquisition. Current technologies primarily employ traditional keyword-based retrieval methods for PDF documents. While these methods are effective for structured data or simple documents, with the surge in document volume and increasing content complexity, especially when processing multi-source PDF documents containing scanned images, screenshots, formulas, and other visual content, traditional methods lack mature character recognition technology. This makes them unable to effectively extract textual information from images and convert it into searchable text data, resulting in the omission of a large amount of crucial content. Furthermore, traditional methods do not incorporate OCR recognition for term enhancement and dynamic retrieval design, limiting retrieval to the raw text within the document and failing to access the information carried by the visual text. This leads to incomplete and weakly correlated search results, failing to meet the needs for accurate and efficient PDF document data processing and retrieval. Summary of the Invention

[0003] This application provides a method and system for intelligent PDF document retrieval that combines OCR recognition, which addresses the technical problem that traditional PDF document retrieval methods struggle to handle image-based text and other content in multi-source heterogeneous PDFs, resulting in one-sided and weakly correlated retrieval results after data processing, failing to meet the requirements for accurate and efficient retrieval.

[0004] The first aspect of this application provides an intelligent PDF document retrieval method combining OCR recognition. The method includes: inputting search terms into an information platform; performing term conversion and OCR enhancement; determining an enhanced term system, wherein the enhancement dimensions include extended enhancement and paradoxical enhancement, with paradoxical enhancement guided by the semantic opposition or complementarity of the terms; setting a look-skip retrieval mechanism for the enhanced term system to generate a dynamic retrieval chain based on concept change paths, wherein dynamic optical character recognition based on focus movement is the setting principle; writing the dynamic retrieval chain into a register via an interaction thread between the document database and the retrieval terminal; performing dynamic OCR retrieval of PDF documents in the document database; determining a PDF retrieval form; and displaying it as a pop-up in the retrieval window of the information platform.

[0005] The second aspect of this application provides an intelligent PDF document retrieval system combined with OCR recognition. The system includes: an enhanced term system acquisition module, used to input search terms into an information platform, perform term conversion and OCR enhancement, and determine the enhanced term system, wherein the enhancement dimensions include extended enhancement and paradoxical enhancement, with paradoxical enhancement guided by the semantic opposition or complementarity of the terms; a dynamic retrieval chain construction module, used to set a look-skip retrieval mechanism for the enhanced term system, generating a dynamic retrieval chain based on concept change paths, wherein the setting principle is based on dynamic optical character recognition based on focus movement; and a PDF retrieval form acquisition module, used to write the dynamic retrieval chain into a register through an interaction thread between the document database and the retrieval terminal, perform dynamic OCR retrieval of PDF documents in the document database, determine the PDF retrieval form, and display it in a pop-up window in the retrieval window of the information platform.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application achieves intelligent retrieval of PDF documents by inputting search terms into an information platform, performing term conversion and OCR enhancement to determine an enhanced term system, setting a look-skip retrieval mechanism to generate a dynamic retrieval chain based on concept change paths, and then writing the dynamic retrieval chain into a register through an interaction thread between the document database and the retrieval terminal. Dynamic OCR retrieval of PDF documents is then performed in the database to determine the PDF retrieval form and display it in a pop-up window, thereby achieving intelligent retrieval of PDF documents. This makes the PDF document retrieval results more accurate and efficient, effectively processing image text and other content in multi-source heterogeneous PDFs, making the processed retrieval results more comprehensive and relevant, and meeting the technical requirements of accurate and efficient retrieval. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart illustrating the intelligent PDF document retrieval method combining OCR recognition provided in the embodiments of this application.

[0010] Figure 2 This is a schematic diagram of the structure of the intelligent PDF document retrieval system that combines OCR recognition, provided in an embodiment of this application.

[0011] Figure labeling: Enhanced term system acquisition module 1, dynamic search chain construction module 2, PDF search form acquisition module 3. Detailed Implementation

[0012] This application provides a method and system for intelligent PDF document retrieval that combines OCR recognition, which addresses the technical problem that traditional PDF document retrieval methods struggle to handle image-based text and other content in multi-source heterogeneous PDFs, resulting in one-sided and weakly correlated retrieval results after data processing, failing to meet the requirements for accurate and efficient retrieval.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] Example 1, as Figure 1 As shown, the intelligent PDF document retrieval method combining OCR recognition includes:

[0016] Step A100: Input the search terms into the information platform, perform term conversion and OCR enhancement, and determine the enhanced term system. Among them, the enhancement dimensions include expansion enhancement and contradiction theory enhancement. Contradiction theory enhancement is based on the opposition or complementarity of the semantics of the terms as the enhancement guide.

[0017] In this embodiment, the information platform is a comprehensive platform that integrates hardware devices, software systems, and network resources, enabling information collection, storage, processing, transmission, and sharing. Through data interaction and business collaboration, it provides users with unified information management and application support, such as an enterprise office automation platform or a government information management platform, meeting the needs for efficient information utilization in multiple scenarios. OCR (Optical Character Recognition) is a technology that uses optical technology to scan text information in images, then uses computer algorithms to analyze and recognize the shape and structure of the text, converting image-formatted text into editable and searchable text format.

[0018] Specifically, the process involves performing term conversion and OCR enhancement to determine the enhanced term system. This includes building a unified data interface that performs multi-dimensional heterogeneous data structure conversion to the platform, training an adversarial network architecture OCR enhancement engine using parallel threads of the first extended generation and the second contradictory generation under the smallest element unit segmentation, and then preprocessing and retrieving terms based on both to determine the enhanced term system. The specific steps are explained in detail in A110-A130.

[0019] Step A200: Set a look-skip retrieval mechanism for the enhanced term system to generate a dynamic retrieval chain based on the concept change path, wherein the setting principle is dynamic optical character recognition based on focus movement.

[0020] In this embodiment, the look-skip retrieval mechanism refers to a retrieval strategy that, during the information retrieval process, skips redundant or low-relevance intermediate nodes and directly locates highly relevant information based on the semantic relevance and coreness between elements. The concept evolution path refers to the trajectory of a concept or element unit as it transforms from its original form to its derived or related forms during the evolution of the terminology system.

[0021] Optionally, the enhanced term system can be divided into multiple local systems based on the homology of the same element units. Based on these local systems, cascaded foci containing the first and second foci can be determined, and then a dynamic retrieval chain based on the concept change path can be generated. The dynamic optical character recognition based on the focus movement is used as the setting principle. The specific steps are explained in detail in A210-A230.

[0022] Step A300: Using the interaction thread between the document database and the retrieval terminal, the dynamic retrieval chain is written into the register, dynamic OCR retrieval of PDF documents is performed in the document database, a PDF retrieval form is determined, and a pop-up window is displayed in the retrieval window of the information platform.

[0023] In this embodiment, the document database refers to a database system that uses unstructured or semi-structured documents as the basic storage unit, and typically stores data in formats such as JSON, BSON, and XML. The retrieval terminal refers to the terminal or interface that initiates an information retrieval request; it is the entry point for interaction between the user or system and the data storage layer.

[0024] In one embodiment of this application, firstly, dynamic OCR retrieval of PDF documents is performed in the document database. The dynamic retrieval chain is written into the register through the interaction thread between the document database and the retrieval terminal. A first retrieval signal based on the first concept node is generated according to the look-skip retrieval mechanism. The signal is then sent down to perform a one-step retrieval to determine the first retrieval column. The specific steps are described in detail in A310-A320.

[0025] Next, the polling retrieval quantity is set. When the retrieval quantity in one step is satisfied, the first retrieval column is determined. The register generates and releases retrieval signals according to the dynamic retrieval chain until the Nth retrieval column is determined. Then, these columns are integrated and sorted to determine the PDF retrieval list. The specific steps are explained in detail in A321-A323.

[0026] Furthermore, step A100 in the method provided in this application embodiment includes:

[0027] A110: Based on the platform data structure, construct a unified data interface, wherein the unified data interface performs the conversion from multi-dimensional heterogeneous data to the platform data structure.

[0028] A120: An OCR enhancement engine is constructed by parallel thread training of the first extended generation under the smallest feature unit segmentation and the second contradictory generation, wherein the generation architecture is an adversarial network architecture.

[0029] A130: Based on the unified data interface and the OCR enhancement engine, the search terms are preprocessed to determine the enhanced term system.

[0030] In this embodiment, the OCR enhancement engine is a tool built on an adversarial network architecture, which is formed by parallel thread training of the first extended generation training and the second contradiction generation under the minimum feature unit segmentation.

[0031] Specifically, after entering search terms into the information platform, the search terms may come from different data sources, exhibiting diverse and heterogeneous formats and structures. For example, some terms exist as text strings, some are embedded in image metadata, and others are semi-structured table content. These differences can affect the consistency of subsequent processing. Therefore, it is necessary to build a unified data interface based on the platform's data structure. This interface can convert these diverse and heterogeneous data into a unified structure compatible with the platform. For example, it can convert XML-formatted terms into structured data containing semantic fields and character encoding as defined by the platform, ensuring that all search terms are processed under the same data standard.

[0032] Furthermore, when constructing a unified data interface, the specific specifications of the platform's data structure are first clarified, including pre-defined data fields such as core terms, semantic categories, and character encoding formats; logical relationships between fields, such as the mapping relationship between core terms and associated descriptions; and data validation standards, such as field length limits and mandatory field definitions. Subsequently, a comprehensive review is conducted of the diverse and heterogeneous data types that may be accessed, such as multi-encoded text data from different systems, incomplete text from OCR recognition embedded in images, semi-structured tabular data, and specially formatted tagged text.

[0033] For each of the aforementioned heterogeneous data types, conversion rules were developed individually: for multi-encoded text, a unified encoding conversion algorithm was used to convert it to the platform's UTF-8 encoding; for incomplete text recognized by OCR, semantic completion logic was added to correct punctuation breaks or garbled characters; for tabular data, row and column keywords were extracted and mapped to corresponding fields; for tagged text, tags were removed while retaining the core content. Based on these rules, an interface functional module was developed, including a data access port, an automatic type recognition component, a rule matching engine, and a post-conversion verification unit. This ensures that the incoming heterogeneous data can be automatically identified in type, the corresponding rules are invoked to complete the conversion, and the data is output after verification to conform to the platform's data structure. Through multiple rounds of testing and iteration, such as testing with 1000 heterogeneous data samples of different types, the conversion success rate was gradually increased to over 98%, ultimately forming a stable unified data interface that achieves accurate conversion of diverse heterogeneous data to the platform's data structure.

[0034] Next, to improve the richness and adaptability of search terms, an OCR enhancement engine needs to be built. This engine adopts an adversarial network architecture and performs two parallel thread training processes: first, expansion generation training; and second, contradiction generation training, both based on minimum feature unit segmentation. Specifically, sample terms are first segmented into minimum feature units according to semantic and character features. For example, PDF retrieval is segmented into two units: PDF and retrieval. For these units, the first expansion generation training generates related expansion elements, such as portable document format and information search, while the second contradiction generation training generates opposing or contradictory elements, such as non-PDF and non-retrieval. Through iterative training of the generator and discriminator in the adversarial network, the relevance and accuracy of the generated expansion and contradictory elements are continuously improved. After multiple rounds of training, the matching degree between the generated elements and the original units reaches over 90%.

[0035] Once the unified data interface and OCR enhancement engine are ready, the search terms are preprocessed, including identifying the search terms and triggering the unified data interface to make judgments and transformations based on the platform's data structure to determine the transformed search terms. Then, for the transformed search terms, the OCR enhancement engine is triggered to perform parallel enhancement processing of the first extension generation and the second contradiction generation threads to build an enhanced term system. The specific steps are explained in detail in A131-A132.

[0036] By constructing a unified data interface to achieve standardized conversion of diverse and heterogeneous data, and by using an adversarial network architecture OCR enhancement engine for parallel training to generate extended and contradictory elements, the final preprocessing yields an enhanced term system, achieving the effect of enriching the dimensions of search terms and improving term adaptability.

[0037] Furthermore, step A130 in the method provided in this application embodiment includes:

[0038] A131: Identify the search terms, trigger the unified data interface, perform judgment and transformation based on the platform data structure, and determine the transformed search terms.

[0039] A132: For the transformed search terms, the OCR enhancement engine is triggered to perform parallel enhancement processing of the first extended generation thread and the second contradictory generation thread to construct the enhanced term system.

[0040] Optionally, when preprocessing search terms, the recognition of search terms is completed first. When a user enters a search term on the information platform, the multimodal recognition module is activated to parse different forms of input content: for plain text terms, the character sequence is directly extracted; for text terms embedded in images, basic OCR recognition tools are used to convert the image information into text format, such as recognizing characters like "artificial intelligence" from scanned image areas in PDFs; for terms in semi-structured tables or lists, core keywords are located and extracted using field extraction algorithms to ensure that the raw data for subsequent processing is accurate and usable.

[0041] After recognition is complete, the system automatically triggers the unified data interface. This interface first determines the recognized search terms based on the platform's data structure. Specifically, it checks the data format of the terms, such as whether it is JSON, XML, or a plain string; the character encoding, such as UTF-8, GBK, etc.; and the completeness of the fields, such as whether it contains core words, semantic tags, and other necessary fields specified by the platform. For example, if a term is determined to be a plain string encoded in GBK and lacks a semantic tag field, the interface will call a preset conversion rule: first, it will be converted to the platform's standard UTF-8 encoding using an encoding conversion algorithm, and then semantic tags will be automatically added based on the built-in semantic mapping library, such as adding the tag "AI" to "artificial intelligence technology".

[0042] After judgment and transformation, the output is a transformed search term that conforms to the platform's data structure. Testing was conducted using 1000 search terms of different types as a sample, including 300 multi-encoded texts, 200 OCR-recognized texts, 300 semi-structured table datas, and 200 tagged texts. After processing through a unified data interface, the format compliance rate of the transformed search terms increased to 98%, and the field completeness rate reached 100%, providing a standardized data foundation for subsequent enhancement processing.

[0043] Finally, parallel enhancement processing of the first extension generation and the second contradiction generation threads is performed. This includes triggering the OCR enhancement engine for the transformed retrieval terms, interpreting their semantic elements and segmenting them into the smallest units to determine the segmented element set, performing the first extension generation based on extension enhancement to determine the extended retrieval elements, and the second contradiction generation based on contradiction theory enhancement to determine the contradiction theory enhanced elements. The three are then integrated to determine the enhanced term system. The specific steps are explained in detail in A132-1-A132-3.

[0044] By first identifying different forms of search terms, and then triggering a unified data interface for structure determination and transformation, the transformed search terms that conform to the platform specifications are finally obtained. This achieves the effect of unifying the search term data standard and providing reliable input for subsequent enhancement processing.

[0045] Furthermore, step A132 in the method provided in this application embodiment includes:

[0046] A132-1: Perform semantic element interpretation and minimum unit segmentation on the transformed retrieval terms to determine the segmentation element set.

[0047] A132-2: Perform a first extension generation based on extension enhancement on the segmented feature set to determine the extended retrieval features, and perform a second contradiction generation based on contradiction theory enhancement to determine the contradiction theory enhanced features.

[0048] A132-3: Integrate the segmented element set, expanded retrieval elements, and contradiction-theoretic enhancement elements to determine the enhanced term system.

[0049] Specifically, once the converted search terms enter the processing flow, a trigger command is automatically sent to the OCR enhancement engine. At this time, the engine, which is in standby mode, receives the command and activates its internal dual-thread processing module, preparing for parallel enhancement processing. This triggering mechanism is implemented through a preset signal interaction protocol, with a response latency controlled within 100 milliseconds, ensuring that the enhancement processing stage starts quickly.

[0050] Next, semantic element interpretation and minimum unit segmentation are performed on the transformed search terms. The semantic element interpretation process relies on multi-level natural language processing technology, and the specific process is as follows: First, the transformed search terms are preprocessed by removing interference information such as special symbols and redundant spaces through regularization algorithms, and the terms are standardized into plain text sequences. For example, "machine learning? paper PDF" is processed into "machine learning paper PDF".

[0051] The word segmentation module was then activated, using a bidirectional LSTM word segmentation model to cut the standardized text. The segmentation accuracy was optimized by combining a domain dictionary covering computer science, document processing and other fields. The word "machine learning paper PDF" was split into three core word blocks: machine learning, paper, and PDF. The bidirectional LSTM word segmentation model was constructed by building a bidirectional long short-term memory network structure, using a large-scale text corpus covering multiple fields, including vocabulary from computer science, document processing and other fields, for training, and integrating domain dictionaries for parameter optimization.

[0052] Next, part-of-speech tagging is performed. A BERT pre-trained model is used to tag each word segment with its part of speech, such as machine learning as a noun phrase, papers as nouns, and PDFs as proper nouns. Implicit action pointers are also identified, such as the action of searching for related articles. Based on this, a Named Entity Recognition (NER) model is used to locate domain entities. Machine learning is accurately identified as belonging to the field of artificial intelligence, and PDF as a document format. The aforementioned BERT pre-trained model is based on the general BERT model, fine-tuned using massive amounts of text data related to the retrieval domain (including vocabulary from PDF document retrieval scenarios), and optimized for output layer parameters related to part-of-speech tagging. The NER model uses a BiLSTM-CRF architecture, trained with 100,000 retrieval term samples labeled with domain entities (such as technical fields, document types, etc.), and its entity recognition accuracy is improved through iterative adjustments to the model weights.

[0053] Finally, the semantic role labeling module is activated to analyze the logical relationships between word blocks: identifying machine learning as the domain qualifier, academic papers as the core search object, and PDF as the carrier qualifier, while also extracting the implicit core action of searching. Integrating this information yields a structured interpretation result, such as: technology field: artificial intelligence; core action: searching; carrier type: PDF document. This provides accurate semantic basis for subsequent minimum unit segmentation.

[0054] Subsequently, based on the semantic interpretation results, the smallest unit segmentation is performed. According to the principle of semantic indivisibility, the term is split into independent elements. For example, the above term is divided into three smallest units: machine learning, paper, and PDF, forming a set of segmented elements.

[0055] After the segmented feature set is generated, the OCR enhancement engine simultaneously launches the first expansion generation thread and the second contradiction generation thread for parallel processing. The first expansion generation thread, based on the expansion enhancement logic, generates associated expanded retrieval elements for each segmented feature. These include synonyms (e.g., PDF expanded to portable document format); hypernyms (e.g., machine learning expanded to artificial intelligence algorithms); hyponyms (e.g., paper expanded to journal articles and dissertations); and related domain terms (e.g., PDF expanded to document format conversion). On average, 3-5 expanded elements are generated for each feature.

[0056] The second contradiction generation thread, based on the logic of contradiction theory enhancement, generates contradiction-theoretic enhancement elements that have an opposing or complementary relationship with the segmented elements. For example, a segmented element represents the opposing or complementary arguments of a point, generating corresponding extended search elements to ensure that the elements cover both positive and negative dimensions. Due to the parallel processing architecture, the processing time of dual-threaded processing is basically the same as that of single-threaded processing, which is more efficient than serial processing.

[0057] Finally, the segmented element set, expanded retrieval elements, and contradiction-theoretic enhancement elements are integrated. Duplicate items are removed using a deduplication algorithm to keep the duplication rate below 5%, and the elements are sorted according to semantic relevance to form an enhanced term system containing original elements, expanded elements, and contradiction elements.

[0058] Furthermore, the semantic relevance is obtained by using the BERT pre-trained model to convert the segmented element set, expanded search elements, contradiction-based enhancement elements, and original transformed search terms into high-dimensional semantic vectors. By calculating the cosine similarity between each element vector and the original term vector, the basic relevance score is obtained, with a value range of 0-1. The higher the score, the stronger the relevance.

[0059] Meanwhile, a domain knowledge graph is introduced to supplement the association dimension: the hierarchical relationship between elements and core terms is extracted from the graph, such as hypernyms, hyponyms, and synonyms, and a hierarchical correction coefficient is assigned to the extended elements, for example, synonyms +0.2, hypernyms +0.1, and hyponyms -0.1; through domain tag matching, such as when an element in the field of artificial intelligence matches a core term in the same field, the score is adjusted by +0.15.

[0060] Finally, the semantic relevance score of each element is obtained by combining the basic similarity and the correction coefficient. The elements are then sorted from high to low scores to ensure that the enhanced term system is presented in an orderly manner according to the degree of relevance to the retrieval needs.

[0061] By triggering the OCR enhancement engine to perform semantic interpretation and unit segmentation on the converted search terms, and generating and integrating extended and contradictory elements in parallel, a multi-dimensional enhanced term system is formed, which achieves the effect of making up for the deficiencies of user input terms and improving the completeness and relevance of the search.

[0062] Furthermore, step A200 in the method provided in this application embodiment includes:

[0063] A210: For the enhanced term system, multiple local systems are divided, wherein the homology of the same element units is the basis for segmentation.

[0064] A220: Based on the plurality of local systems, cascaded focus is determined, wherein the cascaded focus includes a first focus and a second focus, the first focus is at least one local system among the plurality of local systems, and the second focus is at least one element unit among the first focus.

[0065] A230: Generate a dynamic retrieval chain using the multiple local systems and the cascaded focus.

[0066] In this embodiment of the application, the cascaded focus is a set of focuses determined step by step based on multiple local systems when setting the skip-look retrieval mechanism for the enhanced term system, including the first focus and the second focus.

[0067] Specifically, when dividing the enhanced term system into multiple local systems, the core segmentation criterion is the homology of elements. This homology is determined through a dual approach: tracing the element generation path and assessing semantic relevance. Specifically, the generation source information for each element is first extracted, such as whether it originated from the original segmentation, extended generation, or contradictory generation. Then, a BERT pre-trained model is used to calculate the semantic vector similarity between elements. Elements with consistent generation sources and semantic similarity greater than 0.7 are grouped into the same local system. For example, elements such as portable document format and document format conversion, generated by extending PDF segmentation elements, are grouped into one local system due to their homology and semantic relevance exceeding 0.85. Elements generated by machine learning extensions form another local system, thus laying the foundation for subsequent focus determination.

[0068] Next, the fundamental implementation step of the dynamic optical character recognition principle based on focus movement is carried out. When determining the cascaded focus based on multiple local systems, the dynamic optical character recognition principle based on focus movement is followed, i.e., the dynamic OCR retrieval mechanism. First, the first focus is determined: the comprehensive relevance between each local system and the original search term is calculated, i.e., the proportion of fused elements and the average semantic similarity. The local systems with the highest relevance are selected as the first focus. For example, in a scenario containing 10 local systems, the 3 local systems with the highest relevance are determined as the first focus.

[0069] Subsequently, a second focus is determined within each primary focus: by analyzing the retrieval frequency and coreity score of elements in the OCR recognition scenario, at least one key element unit is selected as the second focus. Specifically, the retrieval frequency is obtained based on the fusion statistics of the platform's historical retrieval logs and real-time OCR recognition data: the occurrence frequency of each element in the OCR recognition scenario within the past 30 days is extracted and classified and statistically analyzed according to the retrieval scenario, such as academic literature, office documents, etc. The frequency weight of the PDF document retrieval scenario is set to 1.2, higher than the 1.0 of other scenarios, to highlight the relevance of the target scenario. For example, PDF appeared 1200 times in PDF document retrieval within the past 30 days, and the weighted frequency score is 1440; portable document format appeared 800 times, and the weighted score is 960. At the same time, the count of elements that appear repeatedly in a single retrieval is removed to avoid redundant data interference and ensure that the frequency data reflects the real retrieval needs.

[0070] The highest frequency is the maximum value obtained by sorting the weighted frequencies of all elements in the target search scenario for PDF document retrieval within a statistical period of nearly 30 days. Specifically, the system collects the original occurrence counts of all elements in the PDF document retrieval scenario within this period, calculates the weighted frequency with a scenario weight of 1.2 (original count × 1.2), and then extracts the maximum value from the weighted frequencies of all elements. For example, if the statistics show that the element "document" appears 1500 times in the original logs, after weighting by 1500 × 1.2, it becomes 1800. If this value is higher than the weighted frequencies of other elements in the same period, such as 1440 for PDF and 960 for portable document formats, then 1800 is determined as the highest frequency within this period and used for the standardized calculation of subsequent frequency scores.

[0071] The core score is calculated using a multi-dimensional weighted average: the primary dimension is the semantic similarity between the element and the original search term, with a weight of 0.4. Vectors are generated using a BERT pre-trained model, and cosine similarity is calculated. For example, the similarity between PDF and the original term is 0.92. Secondary dimensions include the element's proportion within the local system, with a weight of 0.3 (the proportion of the element to the total number of elements in the local system; for example, PDF's proportion is 35%, resulting in a weight of 0.35), and the priority of the generation path, with a weight of 0.3. Original segmentation elements receive a weight of 1.0, expanded elements receive 0.7, and contradictory elements receive 0.5, with PDF, as an original segmentation element, receiving 1.0. The overall calculated score is 0.92×0.4+0.35×0.3+1.0×0.3=0.368+0.105+0.3=0.773. After fine-tuning and adding a domain fit coefficient of 0.05, the final core score is 0.823, meeting the standard of ≥0.8.

[0072] When determining the second focus, first select the top 20% of search frequencies and elements with a core score ≥ 0.8. Then calculate the combined score of the two, with the frequency score (standardized) accounting for 0.6 and the core score accounting for 0.4. Select at least one element with the highest combined score as the second focus. For example, the combined score for PDF is (1440 / highest frequency 1800) × 0.6 + 0.823 × 0.4 = 0.48 + 0.329 = 0.809, and the combined score for portable document format is (960 / 1800) × 0.6 + 0.81 × 0.4 = 0.32 + 0.324 = 0.644. Therefore, PDF is determined as the second focus, ensuring that it is both frequently occurring and corely relevant to the search needs.

[0073] Finally, a dynamic retrieval chain is generated based on multiple local systems and cascaded focuses. This includes reconstructing and weighting the local systems according to the first cascaded focus to determine the first concept node, traversing the cascaded focuses to determine the Nth concept node, setting the look-skip retrieval mechanism under node transition for these nodes, and generating a dynamic retrieval chain. The specific steps are explained in detail in A231-A232.

[0074] By dividing local systems according to homology, determining cascaded focal points based on dynamic OCR principles, and generating dynamic retrieval chains that adapt to the concept change path, the retrieval mechanism can follow the movement of the recognition focal point and improve the matching degree between the retrieval path and the OCR recognition process.

[0075] Furthermore, step A230 in the method provided in this application embodiment includes:

[0076] A231: Based on the first cascade focus, the multiple local systems are reconstructed and weighted to determine the first concept node.

[0077] A232: Traverse the cascaded focal points, complete the determination of the Nth concept node based on the Nth cascaded focal point, and set the skip-look retrieval mechanism under node transition for the first concept node up to the Nth concept node to generate the dynamic retrieval chain.

[0078] Specifically, the process of implementing the focus movement rule is as follows: when reconstructing and weighting multiple local systems based on the first-level focus, the core features of the first focus (local system) and the semantic vectors of the second focus (element unit) in the first-level focus are extracted first. Based on this, the association structure of the local system is adjusted. Local systems with different semantic similarity to the first-level focus are divided into corresponding association groups and assigned weights, as shown in Table 1.

[0079] Table 1: Classification and Weight Range of Local Systems in the First-Cascade Focal Association

[0080]

[0081] For example, when the first cascade focus is a local system related to PDF format, local systems such as document conversion and electronic storage are classified into the core association group because their similarity reaches 0.82, and their weight is set to 0.7. The three local systems with the highest weight in the core association group and their second focus elements are integrated to form the first concept node. The element coverage of this node must reach more than 85% of the total elements of the core association group.

[0082] Then, a dynamic execution phase based on the dynamic optical character recognition principle of focus movement is implemented. This involves traversing cascaded focuses to determine the Nth concept node, processing them sequentially according to their importance, decreasing from the first to the Nth, with weight percentages of 30%, 25%, 20%, 15%, and 10%, respectively, suitable for scenarios with N=5. For the Nth cascaded focus, the reconstruction and weighting process is repeated: using the local system and elements of the Nth cascaded focus as a benchmark, the association groups are re-divided and their weights adjusted. The top two local systems and key elements in the core association groups are selected to form the Nth concept node. For example, when the second cascaded focus is a machine learning-related system, the generated second concept node must include core elements such as algorithm models and data training, and its semantic relevance to the first concept node must be maintained above 0.6 to ensure logical coherence between nodes.

[0083] Next, a look-ahead retrieval mechanism based on node transitions is implemented for the first to Nth concept nodes. First, the node transition trigger condition is defined: when the frequency of an element in a node's OCR recognition result fluctuates by more than 20%, or its semantic similarity with the next-level node is less than 0.5, a path adjustment is triggered. The mechanism includes dynamic weight updates (adjusting node weights based on frequency changes, e.g., a 15% increase in frequency increases the weight by 0.1) and path jump rules (skipping the current node and directly associating with a more relevant node when the similarity is below a threshold). For example, if the recognition frequency of deep learning elements in the third concept node decreases by 25%, its weight drops from 0.4 to 0.3, and the retrieval path jumps directly from the third node to the fifth node. Through these settings, nodes are linked together according to the transition rules, forming a dynamic retrieval chain that includes node sequences, dynamic weight adjustment rules, and jump logic.

[0084] By constructing concept nodes based on cascading focus and setting a skip search mechanism in conjunction with node transition rules, a dynamic search chain is generated, which achieves the effect of enabling the search path to be dynamically adjusted with concept changes, thereby improving the flexibility and accuracy of the search.

[0085] Furthermore, step A300 in the method provided in this application embodiment includes:

[0086] A310: Write the dynamic retrieval chain into the register, and generate a first retrieval signal based on the first concept node according to the skip-look retrieval mechanism.

[0087] A320: Based on the first retrieval signal issued by the interaction thread, perform a one-step retrieval in the document database to determine the first retrieval column.

[0088] In one embodiment, the final retrieval stage of the dynamic optical character recognition principle based on focus movement is performed. When writing the dynamic retrieval chain to the register, the concept node sequence, weight rules, and jump logic contained in the retrieval chain are first structurally transformed and encapsulated into a binary instruction stream that can be directly parsed by the register, containing a 16-bit opcode and a 32-bit data segment. A checksum is also embedded to ensure data integrity. The writing process is completed through the bus interface, using DMA (Direct Memory Access) to reduce CPU intervention. The average write time for a single dynamic retrieval chain containing 5-8 concept nodes is controlled within 30 milliseconds. After 1000 write tests, if the data verification pass rate reaches 99% or higher, it ensures that the retrieval chain information is accurately stored in the register's temporary storage area.

[0089] Next, a first retrieval signal based on the first concept node is generated according to the skip-look retrieval mechanism. Core elements from the first concept node, such as PDF format and machine learning papers, are extracted and converted into 128-dimensional semantic vectors using a BERT pre-trained model, serving as the core matching basis for the retrieval signal. Simultaneously, a threshold parameter from the skip-look mechanism is introduced: a lower limit of element matching similarity is set to 0.8; documents with similarities below this value are skipped, and the weight of the first concept node (0.6-0.8) is assigned as a signal priority identifier. The signal generation process utilizes the built-in vector operation unit in the register. The average generation time for a single retrieval signal is 50 milliseconds, with semantic vector conversion accounting for 60% and threshold parameter encapsulation accounting for 40%. The generated signal contains three parts: target vector, priority, and matching threshold, and can be directly recognized by the document database.

[0090] Then, the first retrieval signal is released through the interaction thread between the document database and the retrieval terminal. The interaction thread establishes a long-lived connection based on the TCP / IP protocol, encapsulates the retrieval signal into a data packet containing a signal identifier, target database sharding information, and a verification field, and distributes it to the corresponding shard node in the document database using a load balancing algorithm. During transmission, a sliding window mechanism is used to control traffic, with a single packet transmission latency of ≤20 milliseconds and a packet loss rate of less than 0.5%. Automatic retransmission is triggered in case of packet loss. After receiving the signal, the database node extracts the core element vector and matching threshold through the signal parsing module, completing signal parsing and preparing for the subsequent retrieval step.

[0091] Finally, after a search in the document database, the initially matched documents will be sorted in descending order of matching degree to form the first search column. The data in this column is temporarily stored in the database cache in a linked list structure, waiting for subsequent polling and verification of the search volume. The specific steps are explained in detail in A321-A323.

[0092] By writing the dynamic retrieval chain into the register, generating and releasing the first retrieval signal, the first-step retrieval of the document database is initiated and the first retrieval column is determined, thus achieving the effect of quickly activating the dynamic OCR retrieval process and providing initial matching results for subsequent polling retrieval.

[0093] Furthermore, step A320 in the method provided in this application embodiment includes:

[0094] A321: Set the polling retrieval quantity. When the number of PDF documents retrieved in one step meets the polling retrieval quantity, integrate and determine the first retrieval column.

[0095] A322: The register performs polling and retrieval of retrieval signals according to the dynamic retrieval chain until the determination of the Nth retrieval column is completed.

[0096] A323: Integrate and sort the first search column up to the Nth search column to determine the PDF search list.

[0097] Optionally, the polling retrieval count can be dynamically adjusted based on the total number of PDF documents in the document database and the required retrieval precision: if the database storage is less than 10,000 documents, the polling retrieval count is typically set to 500 documents, accounting for 5% of the total; if it exceeds 10,000 documents, it is set at 3%, such as 900 documents for 30,000 documents, while a minimum threshold of 300 documents is set to avoid insufficient retrieval results. This value is optimized based on historical interaction data between the retrieval client and the database. In 80% of retrieval scenarios, this range ensures result coverage while controlling the computational resource consumption of a single retrieval.

[0098] When the number of PDF documents retrieved in one step reaches the polling retrieval limit, the integration process of the first retrieval column is initiated: first, the similarity of the initially matched documents is checked twice, and documents with a semantic similarity of ≥0.7 with the first concept node are retained. Then, duplicates are removed using a hash algorithm to keep the duplication rate below 3%. Finally, the documents are arranged in descending order of matching degree to form a single column structure.

[0099] Next, the register generates retrieval signals in a polling manner according to the dynamic retrieval chain, triggered sequentially according to the transition order of concept nodes: when transitioning from the first concept node to the second concept node, a retrieval signal based on the core elements of the second node is generated, such as switching the signal associated with PDF format to the signal associated with machine learning algorithm. The signal generation process reuses the vector transformation function of the BERT pre-trained model to ensure semantic consistency. Decentralized retrieval is achieved by reusing the long-connection interaction thread between the document database and the retrieval end. The interval between each round of signal transmission is controlled at 200 milliseconds to avoid excessive database load, and a polling identifier is carried, such as "polling for the second time," to track progress.

[0100] The polling process continues until the Nth search column is determined, where the value of N is determined by the number of conceptual nodes in the dynamic search chain. Typically, N is the same as the number of nodes, such as 5 nodes corresponding to 5 columns. For each search signal generated, the register records the coverage of the searched nodes. The coverage must be above 90%. If a node is not effectively searched (i.e., coverage < 90%), a re-examination mechanism is triggered to regenerate the signal.

[0101] Finally, the first search column up to the Nth search column are integrated and sorted, the segmented feature set is retrieved and the feature weights are set, and each column is traversed to perform cross-correlation coefficient matching based on the segmented feature set and weight calculation based on the feature weights to determine M search correlation coefficients. Then, based on the display quantity constraints of the search window and these correlation coefficients, strongly related content is selected to generate a PDF search sheet. The specific steps are explained in detail in A323-1-A323-3.

[0102] By dynamically setting the polling retrieval volume, integrating retrieval results that meet the conditions, and generating and distributing signals by polling concept nodes, the orderly determination of single columns for multi-round retrieval is achieved, thereby improving document matching coverage and result coherence while controlling retrieval resource consumption.

[0103] Furthermore, step A322 in the method provided in this application embodiment includes:

[0104] A323-1: Retrieve the segmentation feature set and set the feature weights.

[0105] A323-2: Traverse the first search column up to the Nth search column, and perform cross-correlation coefficient matching based on the segmented feature set and weighting calculation based on feature weights for each column to determine M search correlation coefficients, where M is a positive integer greater than or equal to N.

[0106] A323-3: Based on the display quantity constraint of the search window, according to the M search correlation coefficients, strong correlation filtering is performed in the first search column up to the Nth search column to generate the PDF search form.

[0107] In one embodiment, when retrieving the segmented feature set, core feature units generated earlier are extracted, such as PDF format, machine learning, and academic papers. Feature weights are then set based on the analytic hierarchy process (AHP): through domain expert annotation and back-calculation using historical search data, the weights of core features (e.g., PDF format) are determined to be 0.6, secondary features (e.g., academic papers) to be 0.3, and auxiliary features (e.g., publication date) to be 0.1. The total weights are normalized to 1. This weight setting is validated after 1000 sets of search samples, ensuring a match of over 90% with actual user needs, thus providing a benchmark for subsequent association calculations.

[0108] Next, the first to Nth search columns are traversed, and two operations are performed on each document in each column: First, based on the cross-correlation coefficient matching of the segmented feature set, the Pearson correlation coefficient formula is used to calculate the linear correlation between the features contained in the document and the segmented feature set, with values ​​ranging from -1 to 1, and positive values ​​indicating positive correlation; Second, a weighted calculation is performed by multiplying the cross-correlation coefficient by the corresponding feature weight and then summing the results to obtain the retrieval association coefficient for a single document. For example, if a document has a correlation coefficient of 0.9 on PDF format features and 0.8 on machine learning features, then its association coefficient is 0.9 × 0.6 + 0.8 × 0.3 = 0.54 + 0.24 = 0.78. If each column contains a total of 800 documents, N = 5, and each column contains 160 documents, then M = 800 retrieval association coefficients are ultimately generated, where M = 800 ≥ 5, and M is a positive integer greater than or equal to N.

[0109] Finally, based on the constraint of the number of items displayed in the search window, typically set to 50, which can be dynamically adjusted according to the device screen size (30 on mobile devices), the M search relevance coefficients are sorted in descending order. A strong relevance threshold of 0.6 is set, and documents with a coefficient ≥ 0.6 are selected. If the number of documents meeting the criteria exceeds the display limit, the first 50 are selected; if the number is insufficient, the actual number is retained and supplemented with documents with the second highest coefficient, with a minimum of 0.4. For example, if 75 out of 800 coefficients are ≥ 0.6, the first 50 are selected to form a PDF search list. The average time for filtering and sorting a single document is 5 milliseconds, and the overall generation efficiency meets the requirements of real-time retrieval, i.e., the total time is ≤ 1 second.

[0110] By retrieving the segmented feature set and setting weights, the retrieval correlation coefficient between documents and features is calculated. Combined with display constraints, strongly related documents are filtered, achieving the effect of integrating multiple columns of search results, improving output relevance and display adaptability.

[0111] In summary, the intelligent PDF document retrieval method combining OCR recognition provided in this application has the following technical effects:

[0112] This application achieves intelligent retrieval of PDF documents by inputting search terms into an information platform, performing term conversion and OCR enhancement to determine an enhanced term system including extended enhancement and contradiction theory enhancement, setting a look-skipping retrieval mechanism based on focus movement dynamic optical character recognition, generating a dynamic retrieval chain based on concept change paths, and then writing the dynamic retrieval chain into a register through an interaction thread between the document database and the retrieval terminal, performing dynamic OCR retrieval of PDF documents in the database, confirming the PDF retrieval form and displaying it in a pop-up window, thereby realizing intelligent retrieval of PDF documents, making the PDF document retrieval results more accurate and efficient, and achieving effective processing of image text and other content in multi-source heterogeneous PDFs, making the data-processed retrieval results more comprehensive and more relevant, and meeting the technical effect of accurate and efficient retrieval.

[0113] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a PDF document intelligent retrieval system combining OCR recognition, the system comprising:

[0114] The enhanced term system acquisition module 1 is used to input search terms into the information platform, perform term conversion and OCR enhancement, and determine the enhanced term system. The enhancement dimensions include extended enhancement and contradiction theory enhancement. The contradiction theory enhancement is based on the opposition or complementarity of term semantics as the enhancement orientation.

[0115] The dynamic retrieval chain construction module 2 is used to set the skip-look retrieval mechanism for the enhanced term system and generate a dynamic retrieval chain based on the concept change path, wherein the setting principle is based on dynamic optical character recognition based on focus movement.

[0116] PDF search form acquisition module 3 is used to write the dynamic search chain into a register through the interaction thread between the document database and the search terminal, perform dynamic OCR search of PDF documents in the document database, determine the PDF search form, and display it in a pop-up window in the search window of the information platform.

[0117] Furthermore, the enhanced term acquisition module 1 is used to perform the following steps:

[0118] Based on the platform's data structure, a unified data interface is constructed, wherein the unified data interface performs the conversion from multi-dimensional heterogeneous data to the platform's data structure; an OCR enhancement engine is constructed by parallel thread training of the first extended generation training and the second contradiction generation under the smallest element unit segmentation, wherein the generation architecture is an adversarial network architecture; based on the unified data interface and the OCR enhancement engine, the search terms are preprocessed to determine the enhanced term system.

[0119] Furthermore, the enhanced term acquisition module 1 is used to perform the following steps:

[0120] The search terms are identified, triggering the unified data interface to perform judgment and transformation based on the platform data structure, and determining the transformed search terms. For the transformed search terms, the OCR enhancement engine is triggered to perform parallel enhancement processing of the first extended generation thread and the second contradiction generation thread to construct the enhanced term system.

[0121] Furthermore, the enhanced term acquisition module 1 is used to perform the following steps:

[0122] The semantic elements of the transformed search terms are interpreted and segmented into the smallest units to determine the segmented element set; the segmented element set is subjected to a first extension generation based on extension enhancement to determine the extended search elements; the second contradiction generation based on contradiction theory enhancement is performed to determine the contradiction theory enhancement elements; the segmented element set, the extended search elements, and the contradiction theory enhancement elements are integrated to determine the enhanced term system.

[0123] Furthermore, the dynamic retrieval chain construction module 2 is used to perform the following steps:

[0124] For the enhanced term system, multiple local systems are divided, with homology of the same element unit as the segmentation criterion; based on the multiple local systems, cascaded focus is determined, wherein the cascaded focus includes a first focus and a second focus, the first focus being at least one local system among the multiple local systems, and the second focus being at least one element unit among the first focus; a dynamic search chain is generated using the multiple local systems and the cascaded focus.

[0125] Furthermore, the dynamic retrieval chain construction module 2 is used to perform the following steps:

[0126] Based on the first cascade focus, the multiple local systems are reconstructed and weighted to determine the first concept node; the cascade focus is traversed to complete the determination of the Nth concept node based on the Nth cascade focus; the skip-look retrieval mechanism under node transition is set for the first concept node up to the Nth concept node to generate the dynamic retrieval chain.

[0127] Furthermore, the PDF retrieval form acquisition module 3 is used to perform the following steps:

[0128] Write the dynamic retrieval chain into the register, generate a first retrieval signal based on the first concept node according to the skip retrieval mechanism, and perform a one-step retrieval in the document database according to the first retrieval signal released by the interaction thread to determine the first retrieval column.

[0129] Furthermore, the PDF retrieval form acquisition module 3 is used to perform the following steps:

[0130] A polling retrieval count is set. When the number of PDF documents retrieved in one step meets the polling retrieval count, the first retrieval column is determined. The register performs polling generation and retrieval of retrieval signals according to the dynamic retrieval chain until the determination of the Nth retrieval column is completed. The first retrieval column up to the Nth retrieval column are integrated and sorted to determine the PDF retrieval column.

[0131] Furthermore, the PDF retrieval form acquisition module 3 is used to perform the following steps:

[0132] Retrieve the segmented feature set and set feature weights; traverse the first search column up to the Nth search column, perform cross-correlation matching based on the segmented feature set and weight calculation based on feature weights, and determine M search correlation coefficients, where M is a positive integer greater than or equal to N; based on the display quantity constraint of the search window, perform strong correlation filtering in the first search column up to the Nth search column according to the M search correlation coefficients, and generate the PDF search form.

[0133] The intelligent PDF document retrieval system combined with OCR recognition provided in this embodiment of the invention can execute the intelligent PDF document retrieval method combined with OCR recognition provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0134] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0135] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A PDF document intelligent retrieval method combining OCR recognition, characterized in that, The method includes: Input search terms into the information platform, perform term conversion and OCR enhancement, and determine the enhanced term system. Among them, the enhancement dimensions include expansion enhancement and contradiction theory enhancement. Contradiction theory enhancement is based on the opposition or complementarity of term semantics as the enhancement guide. A look-skipping retrieval mechanism is set for the enhanced term system to generate a dynamic retrieval chain based on the concept change path, wherein the setting principle is based on dynamic optical character recognition based on focus movement; Using the interaction thread between the document database and the retrieval terminal, the dynamic retrieval chain is written into the register, and dynamic OCR retrieval of PDF documents is performed in the document database to determine the PDF retrieval form, which is then displayed in a pop-up window in the retrieval window of the information platform. A look-skip retrieval mechanism is set up for the enhanced term system to generate a dynamic retrieval chain based on concept change paths, including: For the enhanced term system, multiple local systems are divided, with the homology of the same element units serving as the segmentation criterion; Based on the multiple local systems, cascaded focal points are determined, wherein the cascaded focal points include a first focal point and a second focal point, the first focal point is at least one local system among the multiple local systems, and the second focal point is at least one element unit among the first focal points; A dynamic retrieval chain is generated using the multiple local systems and the cascaded focus; A dynamic retrieval chain is generated using the multiple local systems and the cascaded focus, including: Based on the first cascade focus, the multiple local systems are reconstructed and weighted to determine the first concept node; Traverse the cascaded focal points to determine the Nth concept node based on the Nth cascaded focal point, and set the skip-look retrieval mechanism under node transition for the first concept node up to the Nth concept node to generate the dynamic retrieval chain; Dynamic OCR retrieval of PDF documents within the document database includes: Write the dynamic retrieval chain into the register, and generate a first retrieval signal based on the first concept node according to the look-skip retrieval mechanism; Based on the first retrieval signal issued by the interaction thread, a one-step retrieval is performed in the document database to determine the first retrieval column; Set a polling retrieval count. When the number of PDF documents retrieved in one step meets the polling retrieval count, integrate and determine the first retrieval column. The register performs polling and retrieval of retrieval signals according to the dynamic retrieval chain until the determination of the Nth retrieval column is completed; The first search column up to the Nth search column are integrated and sorted to determine the PDF search list.

2. The intelligent PDF document retrieval method combining OCR recognition as described in claim 1, characterized in that, Perform term conversion and OCR enhancement, and determine the enhanced term system, including: Based on the platform's data structure, a unified data interface is constructed, wherein the unified data interface performs the conversion from multi-source heterogeneous data to the platform's data structure; An OCR enhancement engine is constructed by parallel thread training of the first extended generation under the smallest element unit segmentation and the second contradictory generation, wherein the generation architecture is an adversarial network architecture. Based on the unified data interface and the OCR enhancement engine, the search terms are preprocessed to determine the enhanced term system.

3. The intelligent PDF document retrieval method combining OCR recognition as described in claim 2, characterized in that, Preprocessing the search terms includes: The search terms are identified, the unified data interface is triggered, and judgment and transformation based on the platform data structure are performed to determine the transformed search terms. For the transformed search terms, the OCR enhancement engine is triggered to perform parallel enhancement processing of the first extended generation thread and the second contradictory generation thread to construct the enhanced term system.

4. The intelligent PDF document retrieval method combining OCR recognition as described in claim 3, characterized in that, Perform parallel enhancement processing of the first extended generation thread and the second contradiction generation thread, including: The semantic elements of the transformed retrieval terms are interpreted and the smallest unit is segmented to determine the segmentation element set; Perform a first extension generation based on extension enhancement on the segmented feature set to determine the extended retrieval features, and perform a second contradiction generation based on contradiction theory enhancement to determine the contradiction theory enhanced features; By integrating the segmentation element set, the expanded retrieval elements, and the contradiction theory enhancement elements, the enhanced term system is determined.

5. The intelligent PDF document retrieval method combining OCR recognition as described in claim 1, characterized in that, The first search column up to the Nth search column are integrated and sorted, including: Retrieve the segmented feature set and set the feature weights; Traverse the first search column up to the Nth search column, and perform cross-correlation coefficient matching based on the segmented feature set and weighting calculation based on the feature weight to determine M search correlation coefficients, where M is a positive integer greater than or equal to N; Based on the display quantity constraint of the search window, and according to the M search correlation coefficients, strong correlation filtering is performed in the first search column up to the Nth search column to generate the PDF search form.

6. A PDF document intelligent retrieval system combining OCR recognition, characterized in that, The system is used to implement the intelligent PDF document retrieval method combining OCR recognition as described in any one of claims 1-5, the system comprising: The enhanced term acquisition module is used to input search terms into the information platform, perform term conversion and OCR enhancement, and determine the enhanced term system. Among them, the enhancement dimensions include expansion enhancement and contradiction theory enhancement. Contradiction theory enhancement is based on the opposition or complementarity of term semantics as the enhancement orientation. The dynamic retrieval chain construction module is used to set the look-skip retrieval mechanism for the enhanced term system and generate a dynamic retrieval chain based on the concept change path, wherein the setting principle is based on dynamic optical character recognition based on focus movement; the PDF retrieval form acquisition module is used to write the dynamic retrieval chain into the register through the interaction thread between the document database and the retrieval terminal, perform dynamic OCR retrieval of PDF documents in the document database, determine the PDF retrieval form, and display it in a pop-up window in the retrieval window of the information platform.

Citation Information

Patent Citations

  • Intelligent retrieval method and system for association of images and contents in PDF (Portable Document Format) document

    CN120407819A

  • Privacy enhanced intelligent search method and system based on multi-round iteration

    CN120596652A