RAG-based pdf intelligent retrieval and generation method and system
By extracting and integrating multimodal features from documents using a RAG-based method, a retrieval index library is constructed, which solves the problems of low efficiency and insufficient accuracy in the processing of complex documents in existing technologies, and achieves efficient and accurate document retrieval and generation.
Patent Information
- Application Number
- CN202511517633.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-23
AI Technical Summary
Existing technologies struggle to efficiently extract multimodal information and generate user-friendly results when processing complex documents, resulting in insufficient completeness and accuracy of information extraction and failing to meet diverse user needs.
The method employs RAG-based approach, which parses document data using a pre-established classification model, extracts text and image features by combining deep learning models and natural language processing techniques, performs multimodal feature fusion and integration, constructs a retrieval index library containing a classification index structure, and performs similarity matching and semantic expansion based on user query conditions.
It enables intelligent parsing and efficient retrieval of multimodal documents, improving the accuracy and comprehensiveness of document retrieval and ensuring that the generated results meet user needs.
Smart Images

Figure CN120994845B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of document processing and intelligent retrieval technology, and in particular discloses a PDF intelligent retrieval and generation method and system based on RAG. Background Technology
[0002] In the context of rapid information digitization, document processing and intelligent retrieval have become particularly crucial. This field directly relates to how to efficiently extract valuable information from massive amounts of unstructured data to meet users' ever-growing knowledge acquisition needs, especially when dealing with widely used document formats, where its importance is self-evident. However, many current solutions often suffer from low processing efficiency, insufficient extraction accuracy, and limited output formats when faced with complex document content, failing to fully meet users' diverse information acquisition needs. Specifically, the main challenges in this field lie in how to accurately extract information from complex documents and generate user-friendly results.
[0003] The primary problem lies in the diversity and unstructured nature of document content. Documents of different formats and qualities often contain multiple elements such as text and images. Traditional processing methods are difficult to uniformly process these heterogeneous data, which limits the completeness of information extraction.
[0004] Due to incomplete information extraction, subsequent retrieval processes are also affected, making it difficult to accurately match relevant content to user needs, resulting in a significant deviation between the final generated results and user expectations.
[0005] This chain of problems, from extraction to retrieval to output, puts the efficiency and accuracy of the entire information processing flow to a severe test.
[0006] Therefore, how to efficiently extract and integrate multimodal information from complex documents, accurately retrieve relevant content based on user needs, and ultimately generate intuitive and easy-to-understand output has become a critical issue that urgently needs to be addressed. Solving this problem will directly determine whether the goal of intelligently extracting information from massive amounts of documents and meeting diverse user needs can be truly achieved. Summary of the Invention
[0007] This invention provides a method and system for intelligent PDF retrieval and generation based on RAG, aiming to solve at least one of the defects existing in the prior art.
[0008] One aspect of this invention relates to a PDF intelligent retrieval and generation method based on RAG (Retrieval-Augmented Generation), comprising the following steps:
[0009] The input document data is obtained, and a pre-established classification model is used to parse the document data, extracting text content and image content to form the first dataset;
[0010] A deep learning model is used to extract features from the image content in the first dataset, and natural language processing techniques are applied to the text content in the first dataset for semantic analysis to obtain a multimodal feature set.
[0011] Based on the multimodal feature set, a unified encoding process is performed using an information integration algorithm to generate a second dataset. If the completeness of the fused feature vector in the second dataset is found to be lower than a preset threshold, contextual semantic analysis is used to fill in the missing information.
[0012] A pre-defined index building mechanism is used to cluster the fused feature vectors in the second dataset to generate a retrieval index library containing a classification index structure.
[0013] Based on the user's input query conditions, a similarity algorithm is used to match the search index. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded and a new match is made.
[0014] Further, the steps of acquiring input document data, parsing the document data using a pre-established classification model, and extracting text and image content to form the first dataset include:
[0015] The original content is obtained from the input document data, and the document is initially analyzed using a pre-established classification tool to extract the text content and image content to form an initial set.
[0016] For the initial set, the completeness is determined by a content recognition tool. If the text content is missing, the complete data is obtained again by a document scanning tool to obtain the first text set.
[0017] If the resolution of the image content is lower than a preset threshold, it is processed by an image enhancement tool to obtain a first set of images;
[0018] For the first text set, a text segmentation tool is used to segment the text to obtain the segmented text units. Based on the text units, a semantic matching tool is used to label the categories to determine their business categories, thus obtaining the second text set.
[0019] For the first image set, an image segmentation tool is used to split it into multiple image units. The visual features of the image units are obtained through a feature extraction tool. If the feature value is lower than a preset threshold, it is adjusted through an image optimization tool to obtain the second image set.
[0020] Based on the second text set and the second image set, a data fusion tool is used to perform association mapping to form a unified business dataset. The consistency of the dataset is checked by a content verification tool to determine whether it meets the classification target, thus obtaining the final first dataset.
[0021] Furthermore, the steps of using a deep learning model to extract features from the image content in the first dataset, and simultaneously applying natural language processing techniques to perform semantic analysis on the text content in the first dataset to obtain a multimodal feature set include:
[0022] For the image content in the first dataset, an image segmentation tool is used to split it into multiple visual element units. The visual element units are then processed by a feature extraction tool to obtain the corresponding image features. If the value of the image feature is lower than a preset threshold, it is adjusted by an image enhancement tool to obtain an adjusted set of image features.
[0023] Based on the adjusted image feature set, a feature mapping tool is used to compare the image features with the preset classification criteria. If they do not meet the criteria, a feature optimization tool is used for secondary processing to determine the subset of image features that meet the criteria.
[0024] For the text content in the first dataset, a text segmentation tool is used to split it into multiple language units. The language units are then processed by a semantic analysis tool to obtain the corresponding text semantic features, resulting in a text semantic set.
[0025] By using data integration tools, a subset of image features is fused with a set of text semantics in a multimodal manner. A content classification tool is then used to perform business matching on the fused data to determine whether it meets the preset goals, thus obtaining the final multimodal feature set.
[0026] Furthermore, based on the multimodal feature set, a unified encoding process is performed using an information integration algorithm to generate a second dataset. If the completeness of the fused feature vector in the second dataset is detected to be lower than a preset threshold, the steps for supplementing missing information through contextual semantic analysis include:
[0027] Based on the multimodal feature set, the feature encoding is uniformly processed using an information integration tool to generate a preliminary fusion dataset. The encoding consistency of the preliminary fusion dataset is detected. If the detected consistency is lower than a preset threshold, it is adjusted using a data calibration tool to obtain a fusion dataset with consistent encoding.
[0028] For the fused dataset with consistent encoding, a feature detection tool is used to evaluate vector integrity. If the integrity of the fused dataset is found to be lower than a preset threshold, supplementary information is extracted using a context analysis tool to determine the fused dataset with improved integrity.
[0029] By using information filling tools to fill in the missing parts of the fused dataset after the integrity improvement, and combining them with semantic completion tools, the fused dataset after filling is obtained.
[0030] A business matching tool is used to classify and compare the populated fused dataset. Based on pre-established rules, the fused dataset is judged to obtain the final dataset that meets the business objectives.
[0031] Furthermore, the steps of clustering the fused feature vectors in the second dataset using a pre-defined indexing mechanism to generate a retrieval index library containing a classification index structure include:
[0032] For the fusion features in the second dataset, a preset mechanism is used to perform preliminary grouping of the feature vectors, and preliminary classification results are obtained through clustering operations;
[0033] Based on the preliminary classification results obtained from the clustering operation, an index building tool is used to structure the grouped feature vectors and determine the hierarchical distribution of the classification index during the construction process of the classification index.
[0034] By optimizing and adjusting the index structure through the hierarchical distribution of the categorized index and using a retrieval index generation tool, a retrieval index framework with efficient query capabilities is obtained.
[0035] For the retrieval index framework, the feature vectors in the second dataset are batch mapped in conjunction with the data processing module. If data deviation is detected during the mapping process, it is corrected by a calibration tool to determine the mapping result that meets the standard.
[0036] Based on the standard-compliant mapping results, for the construction process of the retrieval library, a storage management tool is used to associate and store the optimized index structure with the feature vector to obtain the complete retrieval library content;
[0037] By using the complete search database content, the matching degree between the category index and the search index is detected. If the matching degree is found to be lower than the preset threshold, the index structure is locally optimized by adjustment tools to determine the final search database architecture.
[0038] Based on the final retrieval database architecture, a verification tool is used to test the query response capability of the retrieval database, and the verified retrieval database data is obtained.
[0039] Furthermore, based on the user's input query conditions, a similarity algorithm is used to match the retrieved index. If the relevance of the matching result is lower than a preset threshold, the step of expanding the semantic query scope and rematching includes:
[0040] Based on the user's input query conditions, the input content is structured using a parsing tool to obtain the decomposed query elements;
[0041] Based on the decomposed query elements, the cosine similarity algorithm is used to perform a preliminary matching operation on the index structure in the retrieval index library to obtain the initial matching results;
[0042] For the initial matching results, the relevance evaluation module calculates the degree of fit between the matching results and the query conditions to determine whether the degree of fit reaches the preset threshold.
[0043] If the fit is lower than the preset threshold, the range of the query conditions will be adjusted by the semantic expansion tool to obtain the expanded semantic content.
[0044] Based on the expanded semantic content, the search index is re-matched to determine the updated matching results.
[0045] For the updated matching results, the verification module checks the completeness of the matching results. If missing data is detected, the supplementation tool extracts relevant information from the index structure to obtain complete matching data.
[0046] Based on the complete matching data, the storage management module associates and saves the matching results with the user-input query conditions to determine the final query record.
[0047] This invention relates to a RAG-based intelligent PDF retrieval and generation system, used to implement the aforementioned RAG-based intelligent PDF retrieval and generation method. The RAG-based intelligent PDF retrieval and generation system includes:
[0048] The forming module is used to acquire the input document data, parse the document data using a pre-established classification model, and extract text content and image content to form the first dataset;
[0049] The acquisition module is used to extract features from the image content in the first dataset using a deep learning model, and to perform semantic analysis on the text content in the first dataset using natural language processing techniques to obtain a multimodal feature set.
[0050] The first generation module is used to generate a second dataset by applying an information integration algorithm to perform unified encoding processing based on the multimodal feature set. If the completeness of the fused feature vector in the second dataset is detected to be lower than a preset threshold, contextual semantic analysis is used to fill in the missing information.
[0051] The second generation module is used to perform clustering processing on the fused feature vectors in the second dataset using a preset index building mechanism, and generate a retrieval index library containing a classification index structure.
[0052] The matching module is used to match the search index database based on the query conditions input by the user, using a similarity algorithm. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded and a new match is made.
[0053] Furthermore, the forming modules include:
[0054] The forming unit is used to obtain the original content from the input document data, and to perform preliminary analysis of the document using a pre-established classification tool, extracting the text content and image content to form an initial set respectively;
[0055] The first acquisition unit is used to determine the completeness of the initial set using a content recognition tool. If the text content is missing, the complete data is re-acquired using a document scanning tool to obtain the first text set.
[0056] The second acquisition unit is used to process the image content resolution by an image enhancement tool to obtain a first image set if the image content resolution is lower than a preset threshold.
[0057] The third acquisition unit is used to segment the first text set using a text segmentation tool, acquire the segmented text units, and then use a semantic matching tool to label the text units to determine their business categories, thus obtaining the second text set.
[0058] The fourth acquisition unit is used to split the first image set into multiple image units using an image segmentation tool, obtain the visual features of the image units using a feature extraction tool, and adjust the feature values using an image optimization tool if the feature values are lower than a preset threshold to obtain the second image set.
[0059] The fifth acquisition unit is used to perform association mapping based on the second text set and the second image set using a data fusion tool to form a unified business dataset. The consistency of the dataset is checked by a content verification tool to determine whether it meets the classification target, thus obtaining the final first dataset.
[0060] Furthermore, the acquisition module includes:
[0061] The sixth acquisition unit is used to split the image content in the first dataset into multiple visual element units using an image segmentation tool, process the visual element units using a feature extraction tool to obtain the corresponding image features, and if the value of the image feature is lower than a preset threshold, it is adjusted using an image enhancement tool to obtain an adjusted set of image features.
[0062] The first determining unit is used to compare the image features with the preset classification criteria using a feature mapping tool based on the adjusted image feature set. If the features do not meet the criteria, a secondary processing is performed using a feature optimization tool to determine the subset of image features that meet the criteria.
[0063] The seventh acquisition unit is used to split the text content in the first dataset into multiple language units using a text segmentation tool, process the language units using a semantic analysis tool, obtain the corresponding text semantic features, and obtain a text semantic set.
[0064] The eighth acquisition unit is used to perform multimodal fusion of image feature subsets and text semantic set through data integration tools, and to perform business matching on the fused data using content classification tools to determine whether it meets the preset target, so as to obtain the final multimodal feature set.
[0065] Furthermore, the first generation module includes:
[0066] The ninth acquisition unit is used to uniformly process the feature encoding based on the multimodal feature set using an information integration tool to generate a preliminary fusion dataset. The encoding consistency of the preliminary fusion dataset is detected. If the detected consistency is lower than a preset threshold, it is adjusted using a data calibration tool to obtain a fusion dataset with consistent encoding.
[0067] The second determining unit is used to evaluate the vector integrity of the fused dataset with consistent encoding using a feature detection tool. If the integrity of the fused dataset is found to be lower than a preset threshold, supplementary information is extracted through a context analysis tool to determine the fused dataset with improved integrity.
[0068] The tenth acquisition unit is used to fill in the missing parts of the fused dataset after the integrity improvement by using information filling tools and semantic supplementation tools to obtain the filled fused dataset.
[0069] The eleventh acquisition unit is used to classify and compare the populated fused dataset using a business matching tool, and to judge the fused dataset according to pre-established rules to obtain the final dataset that meets the business objectives.
[0070] The beneficial effects achieved by this invention are as follows:
[0071] This invention provides a method and system for intelligent PDF retrieval and generation based on RAG (Reference Object Group). It parses input documents using a pre-established classification model, extracting text and image content to form an initial dataset. Then, deep learning and natural language processing techniques are used to extract features and perform semantic analysis on the images and text respectively, resulting in a multimodal feature set. Subsequently, an information integration algorithm is applied for unified encoding to generate a fused feature vector, supplementing contextual semantics to fill in missing information when necessary. Finally, clustering is used to construct a retrieval library containing a classification index, and similarity matching is performed based on user query conditions, expanding the semantic query scope as needed. This invention achieves intelligent parsing, feature extraction, information fusion, and efficient retrieval of multimodal documents, improving the accuracy and comprehensiveness of document retrieval. Attached Figure Description
[0072] Figure 1 This is a flowchart illustrating an embodiment of the RAG-based intelligent PDF retrieval and generation method of the present invention. Detailed Implementation
[0073] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0074] like Figure 1 As shown, the first embodiment of the present invention proposes a PDF intelligent retrieval and generation method based on RAG, including the following steps:
[0075] Step S100: Obtain the input document data, parse the document data using a pre-established classification model, and extract the text content and image content to form the first dataset.
[0076] Document data refers to the digitization process of transmitting text, charts, symbols, and other information contained in a document to a computer system for processing via input devices (such as keyboards and scanners) or a file system. Essentially, it transforms structured or unstructured content from external storage media (such as electronic documents or paper documents) into a computer-readable data stream. Document data can be in the form of PDF documents.
[0077] Text content is an information carrier based on character sequences, containing a set of human-readable symbols such as letters, numbers, and symbols, which are converted into binary data that can be processed by computers through encoding (such as ASCII, UTF-8, etc.).
[0078] Image content is a visual information carrier composed of a pixel matrix as the basic unit. It expresses features such as shape, texture, and color through color channels (such as RGB and CMYK) and spatial distribution, and relies on specific encoding formats (such as JPEG and PNG) to achieve digital storage and transmission.
[0079] The first dataset is the initial set of data used for data processing or analysis in a specific scenario. In essence, it is a collection of structured or semi-structured data elements organized according to established rules.
[0080] Step S200: Use a deep learning model to extract features from the image content in the first dataset, and apply natural language processing techniques to perform semantic analysis on the text content in the first dataset to obtain a multimodal feature set.
[0081] Deep learning models are multi-layered nonlinear computational models built on artificial neural networks (ANNs). They achieve the recognition and prediction of complex patterns by extracting hierarchical features from large amounts of data and performing nonlinear transformations.
[0082] Natural Language Processing (NLP) is an interdisciplinary field of computer science and linguistics that aims to enable computers to understand, generate, and manipulate human natural language through algorithms and models. Its core is to analyze the syntactic structure, semantic relationships, and pragmatic context of text to transform language into a machine-computable form (such as vector representations or probability distributions) and perform tasks such as classification, translation, and generation. Typical technologies include lexical analysis, syntactic parsing, semantic understanding, and dialogue systems.
[0083] Feature extraction is the process of identifying, transforming, and extracting key information relevant to the target task from raw data. Its purpose is to transform high-dimensional, redundant, or unstructured data into low-dimensional, highly discriminative feature representations to improve the efficiency and performance of machine learning models. Essentially, it reveals the fundamental structure of data through mathematical or statistical methods, removing noise and irrelevant interference to form numerical features that can be directly processed by algorithms. In this embodiment, feature extraction refers to extracting features from image content.
[0084] Semantic analysis is the process of understanding the logic, intent, or sentiment of linguistic expressions (text or code) by parsing their deeper meaning. In this embodiment, semantic analysis refers to parsing the text content.
[0085] A multimodal feature set refers to a combination of feature vectors extracted from different modalities (such as images, text, and speech). It aims to integrate complementary information from multiple data sources to form a more comprehensive and robust data representation. Its core objective is to provide highly discriminative input features for subsequent tasks (such as classification, detection, and generation) by preserving the unique information of each modality and eliminating redundancy.
[0086] Step S300: Based on the multimodal feature set, apply the information integration algorithm to perform unified encoding processing to generate a second dataset. If the completeness of the fused feature vector in the second dataset is detected to be lower than a preset threshold, supplementary contextual semantic analysis is used to fill in the missing information.
[0087] Information integration algorithms are a set of algorithms that extract, associate, and fuse features from multi-source heterogeneous data (such as text, images, and speech) to generate a unified and semantically consistent high-order representation.
[0088] Encoding processing is the process of converting raw data (such as text, images, and speech) into a structured representation that a computer can recognize, store, or transmit. Its purpose is to map complex, heterogeneous input data into standardized sequences of numerical or symbolic representations suitable for specific tasks (such as machine learning modeling or data transmission) using specific rules or algorithms.
[0089] The second dataset refers to a supplementary dataset used in conjunction with the main dataset in multimodal research. Its core function is to enhance the diversity and joint modeling capabilities of the main dataset by introducing additional modal or scenario information.
[0090] Feature vector fusion is a unified and highly semantically relevant vectorized representation generated by integrating feature representations from multi-source, multi-modal, or heterogeneous data. Its core objective is to combine complementary information from different data sources (such as text, images, and sensor data) to eliminate redundancy and noise, forming a more comprehensive and robust representation of the target object, thereby improving the performance of machine learning models in downstream tasks (such as classification, retrieval, and generation).
[0091] Contextual semantic analysis is a technique that deconstructs and infers the deeper meaning of text or data by combining linguistic environment, background knowledge, and multimodal information. Its core lies in overcoming the limitations of isolated semantics and achieving dynamic and contextualized understanding.
[0092] Filling in missing information refers to using technical means to repair incomplete modal information in multimodal datasets caused by collection errors, transmission loss, or human factors. The aim is to restore data integrity and improve the reliability of downstream tasks (such as classification and retrieval).
[0093] Step S400: Cluster the fused feature vectors in the second dataset using a preset index building mechanism to generate a retrieval index library containing a classification index structure.
[0094] Index building mechanisms are a systematic approach used to transform raw data (such as text, images, and multimodal feature vectors) into efficient retrieval structures. By establishing a mapping between data identifiers and storage locations, they accelerate query response speed and reduce computational complexity. Their core objective is to balance storage costs and query efficiency, supporting real-time or near real-time access to large-scale datasets.
[0095] Clustering is an unsupervised learning technique used to divide unlabeled data into subsets (clusters) with similar intrinsic structures. Its goal is to maximize intra-cluster similarity and minimize inter-cluster similarity to reveal data distribution patterns or latent category characteristics. This process does not rely on prior knowledge and is based entirely on the data's own attributes for grouping.
[0096] Categorical indexing is a technique that organizes and indexes data hierarchically through a predefined category system. It aims to achieve efficient targeted retrieval and range queries based on data's categorical attributes (such as topics, tags, and hierarchical relationships). Its core principle is to optimize data distribution under category constraints, thereby reducing retrieval space complexity and improving query response speed.
[0097] A retrieval index library is a structured collection and storage system designed for efficient data querying. It establishes a mapping between data identifiers and physical storage locations through pre-built multidimensional indexes (such as inverted indexes, vector indexes, and categorical indexes), supporting rapid retrieval, sorting, and recall of multimodal data such as text, images, and videos. Its core objective is to trade space for time, reducing the computational complexity of online queries through offline index building while ensuring throughput and low latency response in high-concurrency scenarios.
[0098] Step S500: Based on the query conditions input by the user, a similarity algorithm is used to match the search index. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded and rematching is performed.
[0099] A similarity algorithm is a mathematical model or computational method used to quantify the degree of consistency or correlation between two data objects. It evaluates their proximity or semantic relevance in a feature space by defining the distance, angle, or probability distribution relationship between objects. Its core objective is to serve tasks such as clustering, retrieval, and recommendation, filtering the most relevant candidate results through similarity ranking.
[0100] Expanding the semantic query scope is a strategy that enhances the semantic expressiveness of queries, breaks through the limitations of literal matching, and improves the retrieval system's breadth of understanding user intent and its ability to recall related content. Its core objective is to expand the candidate result set from the perspective of semantic relevance rather than keyword matching, covering content that is implicitly related but differs in its expression.
[0101] Furthermore, the RAG-based intelligent PDF retrieval and generation method provided in this embodiment includes step S100 as follows:
[0102] Step S110: Obtain the original content from the input document data, use a pre-established classification tool to perform preliminary analysis of the document, and extract the text content and image content to form an initial set.
[0103] When processing an internal financial report document, a pre-established classification tool can be used to perform preliminary analysis of the document, extracting textual descriptions such as financial data and explanatory text into an initial text set, while extracting image content such as charts and signatures into an initial image set.
[0104] Step S120: For the initial set, use a content recognition tool to determine its completeness. If the text content is missing, use a document scanning tool to re-obtain the complete data to obtain the first text set.
[0105] If some data is found to be missing from the text content, such as the table data on a certain page not being recognized, the page will be rescanned using a document scanning tool to ensure that complete data is obtained and the first text set is formed.
[0106] Step S130: If the resolution of the image content is lower than a preset threshold, then the image is processed by an image enhancement tool to obtain a first image set.
[0107] For image content, if it is found that the resolution of a certain chart is lower than a preset threshold such as 300 dpi, the clarity is adjusted by the image enhancement tool to improve it to a level that meets the requirements, forming the first image set.
[0108] Step S140: For the first text set, use a text segmentation tool to perform word segmentation processing, obtain the segmented text units, and use a semantic matching tool to perform category labeling based on the text units to determine their business category, thereby obtaining the second text set.
[0109] For the first text set, a text segmentation tool was used to break the text down into independent word units. For example, "annual financial report" was broken down into units such as "annual," "financial," and "report." Then, a semantic matching tool was used to analyze the meaning of these units and label them as "financial" business categories, ultimately forming the second text set. This approach effectively improves the accuracy of text classification, laying the foundation for subsequent data analysis.
[0110] Step S150: For the first image set, use an image segmentation tool to split it into multiple image units, and use a feature extraction tool to obtain the visual features of the image units. If the feature value is lower than a preset threshold, adjust it using an image optimization tool to obtain the second image set.
[0111] For the first image set, an image segmentation tool is used to break down a complex financial chart into multiple smaller image units, such as pie charts and bar charts. Then, a feature extraction tool is used to analyze the visual features of each unit. If a unit's feature value is found to be below a preset threshold—for example, if the color contrast is insufficient to clearly display data differences—the brightness and contrast are adjusted using image optimization tools to ensure the image information is intuitive and readable, forming the second image set. This step significantly improves the usability of the image data.
[0112] Step S160: Based on the second text set and the second image set, use a data fusion tool to perform association mapping to form a unified business dataset. Use a content verification tool to check the consistency of the dataset and determine whether it meets the classification target to obtain the final first dataset.
[0113] When processing the second text and image sets, data fusion tools were used for correlation mapping. For example, the phrase "first quarter revenue increased by 10%" in the text was matched with the corresponding bar chart in the image to form a unified business dataset. Subsequently, content validation tools were used to check for consistency, ensuring that the text descriptions and image data were consistent, such as verifying whether the growth rates were the same. If inconsistencies were found, the original data could be traced back for correction, ultimately forming the first dataset. This correlation mapping and consistency verification effectively ensured the integrity and reliability of the data, providing solid support for subsequent business decisions.
[0114] In practical applications, assuming a financial report contains multiple pages of data and complex charts, the above process ensures that both text and image content are accurately extracted and categorized, missing data is supplemented, low-quality images are optimized, and the final dataset comprehensively reflects the financial situation. This method not only improves data processing efficiency but also reduces the cost of manual proofreading, demonstrating significant practical value.
[0115] Furthermore, the RAG-based intelligent PDF retrieval and generation method provided in this embodiment includes step S200 as follows:
[0116] Step S210: For the image content in the first dataset, use an image segmentation tool to split it into multiple visual element units, process the visual element units with a feature extraction tool to obtain the corresponding image features, and if the value of the image feature is lower than a preset threshold, adjust it with an image enhancement tool to obtain an adjusted image feature set.
[0117] When processing internal financial report data, the processing of the first dataset can be analyzed in detail from multiple perspectives, and specific implementation methods can be discussed around the processing of image content and text content.
[0118] For image content processing, image segmentation tools are used to break down complex financial charts into multiple visual element units, such as dividing a chart containing multiple data points into a line chart section and a table area.
[0119] In principle, image segmentation tools divide charts into different regions by identifying boundaries and color differences. Assuming a chart contains multiple data display areas, after segmentation, the feature extraction tool for each unit analyzes features such as line thickness and color distribution to obtain specific image feature values. If the feature value of a unit is below a preset threshold, such as a resolution of less than 100 pixels, an image enhancement tool adjusts its resolution and contrast, forming an adjusted set of image features.
[0120] Step S220: Based on the adjusted image feature set, use a feature mapping tool to compare the image features with the preset classification criteria. If they do not meet the criteria, use a feature optimization tool for secondary processing to determine the subset of image features that meet the criteria.
[0121] Subsequently, feature mapping tools are used to compare these features with preset classification standards. If they do not meet the requirements, such as insufficient color contrast, feature optimization tools are used for secondary processing to ensure that the final subset of image features meets the business display requirements.
[0122] Step S230: For the text content in the first dataset, use a text segmentation tool to split it into multiple language units, process the language units through a semantic analysis tool, obtain the corresponding text semantic features, and obtain a text semantic set.
[0123] For text processing, text segmentation tools can break down long sentences in financial reports into multiple linguistic units. For example, "revenue growth data for this quarter" can be broken down into units such as "this quarter," "revenue," "growth," and "data." Semantic analysis tools then analyze the meaning of these units based on contextual relationships, such as identifying a positive correlation between "growth" and financial data, thereby forming a semantic set of the text. This approach helps to identify key information points in the text, providing a foundation for subsequent classification.
[0124] Step S240: Using a data integration tool, the image feature subset and the text semantic set are fused in a multimodal manner. A content classification tool is used to perform business matching on the fused data to determine whether it meets the preset target, and the final multimodal feature set is obtained.
[0125] In the multimodal fusion stage, data integration tools are used to associate subsets of image features with sets of text semantics. For example, the phrase "5% income growth" mentioned in the text is matched with the corresponding upward trend of the line graph in the image.
[0126] Content classification tools then determine whether the fused data meets the needs of financial analysis based on preset objectives, such as whether it can clearly reflect quarterly trends, thus obtaining the final multimodal feature set. This fusion method ensures the consistency of data across different formats. For example, from a business perspective, financial reports often contain complex tables and descriptive text; relying solely on single-modal data can easily lead to overlooking certain information.
[0127] By employing the multimodal fusion described above, data features are comprehensively captured from both image and text dimensions. For example, images visually display trends, while text provides specific numerical explanations. The combination of the two offers a more complete picture of the financial situation. This approach has significant advantages in data presentation and business understanding.
[0128] Furthermore, in the RAG-based intelligent PDF retrieval and generation method provided in this embodiment, step S300 includes:
[0129] Step S310: Based on the multimodal feature set, use an information integration tool to uniformly process the feature encoding to generate a preliminary fusion dataset. Detect the encoding consistency of the preliminary fusion dataset. If the detected consistency is lower than a preset threshold, adjust it using a data calibration tool to obtain a fusion dataset with consistent encoding.
[0130] When processing internal financial reporting data, we conduct specific analyses focusing on the integration and optimization of multimodal feature sets, providing detailed implementation methods and scenario examples for each technical topic to ensure that the content is closely aligned with the field of financial data analysis.
[0131] The unified processing of feature encoding by information integration tools can be understood as the process of converting image and text features into a unified format.
[0132] In principle, image features may be stored as vectors, while text features may be semantic labels. The integration tool needs to map both to the same dimensional space. Suppose we are processing a quarterly financial report, where the image feature vector has a dimension of 128 and the text feature vector has a dimension of 64. The integration tool will unify the two into a 100-dimensional vector through dimensionality reduction or expansion operations, forming a preliminary fused dataset, which will help with subsequent consistency detection.
[0133] Step S320: For the fused dataset with consistent encoding, use a feature detection tool to evaluate vector integrity. If the integrity of the fused dataset is found to be lower than a preset threshold, extract supplementary information through a context analysis tool to determine the fused dataset with improved integrity.
[0134] The principle behind the application of code consistency detection and data calibration tools is to ensure the consistency of features across different parts of a dataset. Assuming a preset consistency threshold of 85%, if detection reveals that a dataset has only 70% consistency, this might be due to significant discrepancies between image and text features in certain dimensions. Data calibration tools adjust these outlier dimensions, for example, by using a weighted average method to balance the discrepancies, ultimately improving the consistency to 88%, resulting in a fused dataset with consistent code.
[0135] Evaluating vector integrity for feature detection tools can be understood as checking whether the fused dataset covers the necessary information. If the integrity threshold is 90% but the detection result is 80%, it may indicate that some image features are missing. Contextual analysis tools extract supplementary information from relevant text descriptions, such as inferring image trend features from the description "revenue growth of 10%", and after the integrity is improved to 92%, a new fused dataset is formed.
[0136] Step S330: Using information filling tools, fill in the missing parts of the fused dataset after the integrity improvement with semantic supplementation tools to obtain the filled fused dataset.
[0137] The principle behind the application of information imputation and semantic completion tools is to fill in the missing content in the dataset. Assuming that some textual descriptions in the improved dataset are still incomplete, such as lacking information related to "cost changes," semantic completion tools will infer possible content based on the context, such as inferring a cost-decreasing trend from "profit increase," thus creating a more comprehensive dataset.
[0138] Step S340: Using a business matching tool, classify and compare the populated fused dataset, and judge the fused dataset according to the pre-established rules to obtain the final dataset that meets the business objectives.
[0139] The principle behind the classification and comparison of business matching tools is to filter datasets based on financial analysis objectives. Assuming the rules require the dataset to reflect quarterly trends, the tool compares the merged data to see if it includes key indicators such as revenue and costs. If any part of the data does not meet the requirements, such as missing expense details, it is removed or flagged, ensuring that the final dataset meets business objectives. This approach effectively improves the business relevance of the data.
[0140] Furthermore, in the RAG-based intelligent PDF retrieval and generation method provided in this embodiment, step S400 includes:
[0141] Step S410: For the fusion features in the second dataset, a preset mechanism is used to perform preliminary grouping of the feature vectors, and preliminary classification results are obtained through clustering operations.
[0142] In scenarios involving the analysis of internal financial data of an enterprise, preliminary grouping of the fusion features in the second dataset can be understood as performing preliminary clustering of complex feature vectors based on similarity.
[0143] In principle, clustering operations group similar financial indicators into one category based on the distribution characteristics of feature vectors. For example, revenue-related features and cost-related features are clustered into different groups. Suppose we are processing an annual financial report dataset. Revenue-related feature vectors may be concentrated in one numerical range, while cost-related features are distributed in another range. Clustering operations can yield two preliminary classification results, laying the foundation for subsequent index construction.
[0144] Step S420: Based on the clustering operation, a preliminary classification result is obtained. For the construction process of the classification index, an index building tool is used to structurally organize the grouped feature vectors and determine the hierarchical distribution of the classification index.
[0145] For the construction process of the categorized index, an index building tool is used to structure the grouped feature vectors, aiming to form a clear hierarchical distribution. In principle, the hierarchical distribution is divided according to the priority or importance of the financial data, such as placing core indicators at the top level and secondary indicators at the bottom. Assuming the core indicators are total revenue and net profit, and the secondary indicators are detailed cost items, the index building tool will place total revenue at the first level and cost details at the second level, forming a tree structure that facilitates rapid data location later.
[0146] Step S430: By optimizing and adjusting the index structure through the hierarchical distribution of the classification index, a retrieval index generation tool is used to obtain a retrieval index framework with efficient query capabilities.
[0147] The retrieval index frame is derived using the following formula:
[0148] (1)
[0149] In formula (1), This indicates the optimized index structure configuration. Represents the set of all possible index structure schemes. Indicates the number of samples queried. Indicates the first One query request, This indicates the corresponding index structure response. The function represents a query for distance metrics. Function representation is a measure of structural complexity. and For balancing parameters. Formula (1) obtains the optimal retrieval index frame by minimizing the objective function.
[0150] The optimization and adjustment of the index structure by the retrieval index generation tool aims to build a framework for efficient query capabilities. In principle, the optimization reduces redundant paths between index levels, improving query speed. For example, if the original index structure required three hops to query a cost detail, the optimized structure only requires two, significantly shortening response time.
[0151] Step S440: For the retrieval index framework, the feature vectors in the second dataset are batch mapped in conjunction with the data processing module. If data deviation is detected during the mapping process, it is corrected by a calibration tool to determine the mapping result that meets the standard.
[0152] The linear transformation process of batch mapping of feature vectors by the processing module is described by the following formula:
[0153] (2)
[0154] In formula (2), Represents a mapping function. Indicates the second dataset, the first... 1 eigenvector Represents the mapping weight matrix. This represents the bias vector.
[0155] The degree of data deviation detected during the mapping process is quantified by the following formula:
[0156] (3)
[0157] In formula (3), This indicates the data deviation detected during the mapping process. This indicates the total number of mapping results. Indicates the first Each mapping output result This indicates the reference standard value.
[0158] The following formula describes the calculation process by which the calibration tool corrects for deviation data:
[0159] (4)
[0160] In formula (4), This indicates the calibrated mapping result. This represents the original mapping result. Indicates the calibration intensity coefficient. Represents the calibration matrix. This represents the target standard value.
[0161] For the batch mapping of feature vectors by the data processing module, calibration is performed if data deviation is detected. The principle is to ensure the accuracy of the mapping results. For example, if a revenue feature vector deviates from the normal range during the mapping process, the calibration tool will correct it based on historical data trends to ensure the mapping results conform to financial analysis logic. Regarding the application of storage management tools during the retrieval library construction process, the principle is to associate and store the optimized index structure with the feature vectors to form a complete retrieval library.
[0162] Step S450: Based on the mapping results that conform to the standard, for the construction process of the retrieval library, use a storage management tool to associate and store the optimized index structure and feature vectors to obtain the complete retrieval library content.
[0163] The complete search database content is derived using the following formula:
[0164] (5)
[0165] In formula (5), This represents the complete content of the search database. This represents the associated storage functions of storage management tools. This represents the optimized index structure. Represents the set of feature vectors. Indicates the first 1 index element, Indicates the first 1 eigenvector This indicates the total amount of data stored in the retrieval database.
[0166] Suppose that the storage management tool binds revenue metrics to corresponding vectors for storage, ensuring that relevant data can be quickly matched during queries.
[0167] Step S460: Using the complete search database content, the matching degree between the category index and the search index is detected. If the matching degree is found to be lower than the preset threshold, the index structure is locally optimized by the adjustment tool to determine the final search database architecture.
[0168] The matching degree between the category index and the retrieval index is:
[0169] (6)
[0170] In formula (6), Represents a category index With search index The degree of matching between them Indicates the total number of index feature dimensions. Indicates the first The weight coefficients of each feature, Represents a category index In the Values on each feature Indicates the search index In the Values on each feature The function represents the similarity calculation function.
[0171] The preset threshold for matching degree is:
[0172] (7)
[0173] In formula (7), The preset threshold representing the matching degree. This represents the mean of historical matching scores. This represents the adjustment coefficient. The standard deviation of historical matching degree This indicates the total number of indexes in the search database. Indicates the number of valid matches in the index.
[0174] The optimal structure for local optimization of the index structure is:
[0175] (8)
[0176] In formula (8), Indicates the index structure The optimal structure after local optimization. This represents the candidate optimized structure. This indicates the number of index nodes that need optimization. Indicates the first Optimization weights for each node, Functions represent measures of structural differences. Indicates the original number A node structure Indicates the optimized first... A node structure Represents the regularization parameter. This represents the structural complexity penalty term.
[0177] The tool detects the matching degree between the category index and the retrieval index. If it falls below a preset threshold, local optimization is performed to improve the coordination between the indexes. For example, if the preset threshold is 80% and the actual detection result is only 75%, the adjustment tool will optimize the abnormal hierarchical structure, increasing the matching degree to over 82%.
[0178] Step S470: Based on the final search database architecture, use a verification tool to test the query response capability of the search database to obtain the verified search database data.
[0179] The validation tool used to test the query response capability of the retrieval database aims to ensure its usability. For example, if the response time for a query on a certain financial indicator is 0.5 seconds, meeting the expected standard, then the validation passes. This approach helps ensure the efficiency of financial data analysis.
[0180] Furthermore, in the RAG-based intelligent PDF retrieval and generation method provided in this embodiment, step S500 includes:
[0181] Step S510: Based on the query conditions input by the user, the input content is processed in a structured manner using a parsing tool to obtain the decomposed query elements.
[0182] In the context of internal financial data analysis within enterprises, parsing tools play a crucial role in processing user-input queries. In principle, parsing tools break down complex user queries into multiple identifiable elements for subsequent matching operations. For example, if a user's query is about changes in net profit for a specific year, the parsing tool will decompose it into specific elements such as time range and financial indicator categories, laying the foundation for subsequent retrieval.
[0183] Step S520: Based on the decomposed query elements, the cosine similarity algorithm is used to perform a preliminary matching operation on the index structure in the retrieval index library to obtain the initial matching results.
[0184] The core of applying the cosine similarity algorithm lies in measuring the similarity between query elements and the index structure. Conceptually, this algorithm judges relevance by the difference in the angle of vectors; the smaller the angle, the higher the similarity. Assuming a user queries total revenue data, the algorithm compares this element with the relevant vectors stored in the index, initially filtering out the closest index nodes to form the initial matching results.
[0185] Specifically, cosine similarity is calculated using the vector dot product and the modulus to obtain the query vector. With each index vector The similarity, cosine similarity is used to measure the similarity of query vectors. and index vector The similarity between two vectors is determined by the consistency of their directions (the smaller the angle between them, the higher the similarity). The formula is as follows:
[0186] (9)
[0187] In formula (9), Represents the query vector and index vector The cosine of the angle between two vectors, i.e., cosine similarity. Represents the query vector and index vector The dot product (inner product) of . and These represent the query vectors respectively. and index vector The modulus (also known as the modulus) Norm, which is the "length" of a vector.
[0188] The calculation method is the query vector. and index vector The sum of the products of corresponding components of two vectors is:
[0189] (10)
[0190] In formula (10), To query the n-dimensional components of the vector, For index vectors dimensional components, For vector dimensions.
[0191] vector The modulus length is:
[0192] (11)
[0193] In formula (11), Represents the query vector The length of the mold, To query the n-dimensional components of the vector, For vector dimensions.
[0194] vector The modulus length is:
[0195] (12)
[0196] In formula (12), Represents an index vector The length of the mold, For index vectors dimensional components, For vector dimensions.
[0197] Step S530: For the initial matching results, the relevance evaluation module calculates the degree of fit between the matching results and the query conditions, and determines whether the degree of fit reaches the preset threshold.
[0198] The relevance score between the query criteria and the matching results is calculated using the following formula:
[0199] (13)
[0200] In formula (13), Indicates query conditions With the Match results The correlation score between them This indicates the number of feature dimensions for the query conditions. Indicates the first The weight coefficients of each feature, Indicates the first in the query conditions 1 eigenvalue, Indicates the first In the matching results, the first... 1 eigenvalue, The function represents the function for calculating the similarity between two feature values.
[0201] Whether the compatibility reaches the preset threshold is determined by the following formula:
[0202] (14)
[0203] In formula (14), Indicates the first The final decision on each matching result. Indicates the first The score representing the fit of each matching result. This indicates the preset matching threshold. When the matching degree is greater than or equal to the threshold, it is judged as passing and returns 1; otherwise, it is judged as failing and returns 0.
[0204] The relevance assessment module is introduced to further verify the accuracy of the matching results. In principle, the relevance assessment module scores the initial results based on multiple dimensions, such as content coverage and semantic consistency. For example, if the preset fit threshold is 85%, but the initial matching result scores only 78%, further adjustments are needed to improve the matching quality.
[0205] Step S540: If the fit is lower than the preset threshold, the range of the query conditions is adjusted by the semantic expansion tool to obtain the expanded semantic content.
[0206] The expanded semantic content is derived using the following formula:
[0207] (15)
[0208] In formula (15), This represents the semantically expanded query vector. Represents the original query vector. Indicates the extended strength control parameters. Indicates the number of extended semantic words. Indicates the first The weight coefficients of each extended word, Indicates the first Vector representation of extended semantic words.
[0209] The purpose of using semantic expansion tools is to broaden the scope of queries when the match is insufficient. Specifically, if the initial match is not satisfactory, the semantic expansion tool will expand the query conditions to include synonyms or related indicators based on the semantic relevance of the financial domain.
[0210] Step S550: Based on the expanded semantic content, perform a new matching operation on the retrieval index to determine the updated matching results.
[0211] The updated matching results are derived using the following formula:
[0212] (16)
[0213] In formula (16), This indicates the final matching result after the update. Represents the original set of matching results. This represents a new set of matching results based on extended semantics. This represents the retention weight factor of the original result, with a value between 0 and 1.
[0214] If a user queries net profit, we can expand the query by adding related concepts such as total revenue and costs, and then perform a matching process again to obtain a more comprehensive result.
[0215] Step S560: For the updated matching results, the integrity of the matching results is checked by the verification module. If data is missing, the relevant information is extracted from the index structure by the supplementation tool to obtain the complete matching data.
[0216] The integrity check of the updated matching results is obtained using the following formula:
[0217] (17)
[0218] In formula (17), This represents the integrity check value of the matching result. Represents the current set of matching results. Indicates a reference to the complete set of matches. This indicates the size of the intersection between the current match and the reference match. This indicates the total size of the reference matching set. Indicates the number of non-empty matches. Indicates the expected total number of matches. This represents the integrity weighting coefficient.
[0219] To ensure the completeness of the updated matching results, the verification module checks for missing data. In principle, if some financial indicator data is found to be missing, the supplementation tool will extract the relevant information from the index structure. For example, if the matching result lacks a specific cost detail, the supplementation tool will retrieve it from the secondary index level to ensure the final result is complete.
[0220] Step S570: Based on the complete matching data, the storage management module is used to associate and save the matching results with the user-input query conditions to determine the final query record.
[0221] The correlation strength between query conditions and matching results is calculated using the following formula:
[0222] (18)
[0223] In formula (18), Indicates query conditions Matching results The correlation score between them Indicates the first The weight coefficient of each matching attribute, Indicates the query condition. The first attribute and the matching result A similarity function for each attribute. This indicates the total number of attributes involved in the matching.
[0224] The final query records are obtained using the following formula:
[0225] (19)
[0226] In formula (19), This represents the final set of query records. Indicates the first The query criteria entered by each user. Indicates the first One matching result, Indicates the query timestamp. Indicates the correlation threshold. Representation and query conditions The relevant matching result index set, where the difference represents the total number of query conditions. Formula (19) defines how the query records that meet the conditions form the final result set.
[0227] The storage management module is used to associate and save the final matching results with the user's query conditions for later reuse or traceability. For example, if a user queries quarterly financial data, the system will bind and store the matched indicator data with the query record, forming a traceable query log. This approach helps improve the standardization of data management.
[0228] Through the above multi-faceted processing, from query parsing to result saving, each step is closely centered around the needs of financial data analysis, ensuring the rigor of the query process and the reliability of the results, while also providing users with an efficient data acquisition experience.
[0229] This invention relates to a RAG-based intelligent PDF retrieval and generation system, used to implement the aforementioned RAG-based intelligent PDF retrieval and generation method. The RAG-based intelligent PDF retrieval and generation system includes a forming module, an acquisition module, a first generation module, a second generation module, and a matching module. The forming module acquires input document data, parses the document data using a pre-established classification model, and extracts text and image content to form a first dataset. The acquisition module uses a deep learning model to extract features from the image content in the first dataset, and simultaneously applies natural language processing techniques to perform semantic analysis on the text content in the first dataset. The system analyzes and obtains a multimodal feature set. A first generation module generates a second dataset by applying an information integration algorithm to perform unified encoding processing based on the multimodal feature set. If the completeness of the fused feature vectors in the second dataset is lower than a preset threshold, contextual semantic analysis is used to fill in the missing information. A second generation module clusters the fused feature vectors in the second dataset using a preset indexing mechanism to generate a retrieval index library containing a classification index structure. A matching module matches the retrieval index library using a similarity algorithm based on the user-input query conditions. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded for re-matching.
[0230] Furthermore, the RAG-based PDF intelligent retrieval and generation system provided in this embodiment includes a forming module comprising a forming unit, a first acquisition unit, a second acquisition unit, a third acquisition unit, a fourth acquisition unit, and a fifth acquisition unit. The forming unit is used to acquire the original content from the input document data, perform preliminary analysis of the document using a pre-established classification tool, and extract the text content and image content to form an initial set. The first acquisition unit is used to determine the completeness of the initial set using a content recognition tool. If the text content is missing, the complete data is re-acquired using a document scanning tool to obtain a first text set. The second acquisition unit is used to process the image content if its resolution is lower than a preset threshold using an image enhancement tool to obtain a first image set. The system consists of five parts: a first text set and a second image set. The first text set is segmented into multiple image units using an image segmentation tool. Visual features of these image units are extracted using a feature extraction tool. If the feature values are below a preset threshold, they are adjusted using an image optimization tool to obtain the second image set. The second text set and the second image set are then correlated using a data fusion tool to form a unified business dataset. The consistency of the dataset is checked using a content verification tool to determine if it meets the classification objective, thus obtaining the final first dataset.
[0231] Preferably, the RAG-based PDF intelligent retrieval and generation system provided in this embodiment includes a sixth acquisition unit, a first determination unit, a seventh acquisition unit, and an eighth acquisition unit. The sixth acquisition unit is used to segment the image content in the first dataset into multiple visual element units using an image segmentation tool, process the visual element units using a feature extraction tool to obtain corresponding image features, and adjust the image feature set using an image enhancement tool if the image feature value is lower than a preset threshold. The first determination unit is used to compare the image features with preset classification standards using a feature mapping tool based on the adjusted image feature set. If the features do not meet the standards, a secondary processing is performed using a feature optimization tool to determine a subset of image features that meet the standards. The seventh acquisition unit is used to segment the text content in the first dataset into multiple language units using a text segmentation tool, process the language units using a semantic analysis tool to obtain corresponding text semantic features, and obtain a text semantic set. The eighth acquisition unit is used to perform multimodal fusion of the image feature subset and the text semantic set using a data integration tool, perform business matching on the fused data using a content classification tool, determine whether it meets preset targets, and obtain the final multimodal feature set.
[0232] Furthermore, the RAG-based intelligent PDF retrieval and generation system provided in this embodiment includes a first generation module comprising a ninth acquisition unit, a second determination unit, a tenth acquisition unit, and an eleventh acquisition unit. The ninth acquisition unit is used to uniformly process feature encodings using an information integration tool based on a multimodal feature set to generate a preliminary fused dataset. It then detects the encoding consistency of the preliminary fused dataset; if the detected consistency is lower than a preset threshold, it adjusts the dataset using a data calibration tool to obtain a fused dataset with consistent encoding. The second determination unit is used to evaluate the vector integrity of the fused dataset with consistent encoding using a feature detection tool. If the integrity of the fused dataset is found to be lower than a preset threshold, it extracts supplementary information using a context analysis tool to determine the fused dataset with improved integrity. The tenth acquisition unit is used to fill in the missing parts of the fused dataset with an information filling tool, combined with a semantic supplementation tool, to obtain a filled fused dataset. The eleventh acquisition unit is used to classify and compare the filled fused dataset using a business matching tool, judging the fused dataset according to pre-established rules to obtain a final dataset that meets business objectives.
[0233] The RAG-based intelligent PDF retrieval and generation method and system provided in this embodiment, compared with existing technologies, parses the input document using a pre-established classification model to extract text and image content to form an initial dataset. Then, deep learning and natural language processing techniques are used to extract features and perform semantic analysis on the images and text respectively, resulting in a multimodal feature set. Subsequently, an information integration algorithm is applied for unified encoding to generate a fused feature vector, and contextual semantics are supplemented to fill in missing information when necessary. Finally, clustering processing is used to construct a retrieval library containing a classification index, and similarity matching is performed based on user query conditions, expanding the semantic query scope when necessary. This embodiment achieves intelligent parsing, feature extraction, information fusion, and efficient retrieval of multimodal documents, improving the accuracy and comprehensiveness of document retrieval. The specific beneficial effects achieved by the RAG-based intelligent PDF retrieval and generation method and system provided in this embodiment are as follows:
[0234] I. Efficiently process unstructured data and improve data availability
[0235] Traditional PDF files contain unstructured data such as text and images, making them difficult and inefficient to process. This embodiment uses a pre-established classification model to parse document data, accurately extracting text and image content to form the first dataset. This process quickly breaks down complex PDF files into processable basic data units, significantly reducing data processing time and improving efficiency compared to manual processing or simple data extraction methods. Simultaneously, complete data extraction avoids information omissions, ensuring subsequent processing is based on a comprehensive data foundation, significantly improving the usability of PDF data and providing strong support for further in-depth analysis and applications.
[0236] II. Multimodal feature fusion enhances information representation capabilities
[0237] Deep learning models are used to extract features from image content, and natural language processing techniques are employed to perform semantic analysis on text content, resulting in a multimodal feature set. This multimodal processing approach overcomes the limitations of single-modal information, enabling the capture of information from multiple perspectives within PDF files. The semantic information of the text and the visual features of the image complement each other, making the information representation richer and more comprehensive.
[0238] III. Information Integration and Gap Filling to Ensure Data Quality
[0239] An information integration algorithm is applied to uniformly encode the multimodal feature set to generate a second dataset. For cases where the completeness of the fused feature vector is below a preset threshold, contextual semantic analysis is used to fill in the missing information. This mechanism ensures the accuracy and completeness of the data during the integration process. In actual PDF files, some images or text may be blurry or incomplete. Information filling through contextual semantic analysis effectively corrects data defects and avoids retrieval errors or inaccurate answers caused by missing information. This guarantees the quality of the final data used for retrieval and answer generation, improving the reliability of the system's output.
[0240] IV. Clustering to build indexes and speed up retrieval.
[0241] A pre-defined indexing mechanism is used to cluster the fused feature vectors in the second dataset, generating a retrieval index library with a categorized index structure. This clustering-based indexing method categorizes and stores massive amounts of PDF data according to features, allowing for quick location of relevant data areas based on the index, rather than traversing all the data during retrieval. When a user enters query conditions, the system can rapidly narrow down the search scope based on the index, significantly improving retrieval speed.
[0242] V. Dynamically adjust search strategies to improve search accuracy.
[0243] Based on the user's input query criteria, a similarity algorithm is used to match the search index. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded for rematching. This dynamic adjustment of the retrieval strategy fully considers the diversity and ambiguity of user query intent. When the initial matching results are unsatisfactory, expanding the semantic query scope can uncover more potential relevant information, avoiding the omission of important content due to limitations in the expression of query criteria.
[0244] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.
Claims
1. A method for intelligent PDF retrieval and generation based on RAG, characterized in that, Includes the following steps: The input document data is obtained, and a pre-established classification model is used to parse the document data, extracting text content and image content to form the first dataset; A deep learning model is used to extract features from the image content in the first dataset, and natural language processing techniques are applied to the text content in the first dataset to perform semantic analysis, resulting in a multimodal feature set. Specifically, it includes: For the image content in the first dataset, an image segmentation tool is used to split it into multiple visual element units. The visual element units are then processed by a feature extraction tool to obtain the corresponding image features. If the value of the image feature is lower than a preset threshold, it is adjusted by an image enhancement tool to obtain an adjusted set of image features. Based on the adjusted image feature set, a feature mapping tool is used to compare the image features with a preset classification standard. If the features do not meet the standard, a feature optimization tool is used for secondary processing to determine a subset of image features that meet the standard. For the text content in the first dataset, a text segmentation tool is used to split it into multiple language units, and a semantic analysis tool is used to process the language units to obtain the corresponding text semantic features, thus obtaining a text semantic set. The image feature subset and the text semantic set are fused in a multimodal manner using a data integration tool. The fused data is then matched with the business data using a content classification tool to determine whether it meets the preset target, thus obtaining the final multimodal feature set. Based on the multimodal feature set, a unified encoding process is performed using an information integration algorithm to generate a second dataset. If the completeness of the fused feature vector in the second dataset is detected to be lower than a preset threshold, contextual semantic analysis is used to fill in the missing information. Specifically, it includes: Based on the multimodal feature set, the feature encoding is uniformly processed using an information integration tool to generate a preliminary fusion dataset. The encoding consistency of the preliminary fusion dataset is detected. If the detected consistency is lower than a preset threshold, it is adjusted using a data calibration tool to obtain a fusion dataset with consistent encoding. For a fused dataset with consistent encoding, a feature detection tool is used to evaluate vector integrity. If the integrity of the fused dataset is found to be lower than a preset threshold, supplementary information is extracted using a context analysis tool to determine the fused dataset with improved integrity. By using information filling tools to fill in the missing parts of the fused dataset after the integrity improvement, and combining them with semantic completion tools, the fused dataset after filling is obtained. A business matching tool is used to classify and compare the populated fused dataset. The fused dataset is judged according to pre-established rules to obtain the final dataset that meets the business objectives. A pre-defined index building mechanism is used to cluster the fused feature vectors in the second dataset to generate a retrieval index library containing a classification index structure. Specifically, it includes: For the fusion features in the second dataset, a preset mechanism is used to perform preliminary grouping of the feature vectors, and preliminary classification results are obtained through clustering operations; Based on the preliminary classification results obtained from the clustering operation, for the construction process of the classification index, the index building tool is used to organize the grouped feature vectors in a structured manner to determine the hierarchical distribution of the classification index; By optimizing and adjusting the index structure through the hierarchical distribution of the categorized index and using a retrieval index generation tool, a retrieval index framework with efficient query capabilities is obtained. For the retrieval index framework, the feature vectors in the second dataset are batch mapped using the data processing module. If data deviation is detected during the mapping process, it is corrected using a calibration tool to determine the mapping result that meets the standard. Based on the standard-compliant mapping results, for the construction process of the retrieval library, a storage management tool is used to associate and store the optimized index structure with the feature vector to obtain the complete retrieval library content; By using the complete search database content, the matching degree between the category index and the search index is detected. If the matching degree is found to be lower than the preset threshold, the index structure is locally optimized by adjustment tools to determine the final search database architecture. The matching degree between the category index and the retrieval index is: in, Represents a category index With search index The degree of matching between them Indicates the total number of index feature dimensions. Indicates the first The weight coefficients of each feature Represents a category index In the Values on each feature Indicates the search index In the Values on each feature The function represents the similarity calculation function; The optimal structure for local optimization of the index structure is: in, Indicates the index structure The optimal structure after local optimization. This represents the candidate optimized structure. This indicates the number of index nodes that need optimization. Indicates the first Optimization weights for each node, Functions represent measures of structural differences. Indicates the original number A node structure Indicates the optimized first... A node structure Represents the regularization parameter. This represents the structural complexity penalty term; Based on the final retrieval database architecture, a verification tool is used to test the query response capability of the retrieval database, and the verified retrieval database data is obtained. Based on the user's input query conditions, a similarity algorithm is used to match the retrieval index. If the relevance of the matching result is lower than a preset threshold, the semantic query scope is expanded and a new match is made.
2. The RAG-based intelligent PDF retrieval and generation method as described in claim 1, characterized in that, The steps of acquiring input document data, parsing the document data using a pre-established classification model, and extracting text and image content to form the first dataset include: The original content is obtained from the input document data, and the document is initially parsed using a pre-established classification tool to extract the text content and image content to form an initial set. For the initial set, its completeness is determined by a content recognition tool. If the text content is missing, the complete data is obtained again by a document scanning tool to obtain the first text set. If the resolution of the image content is lower than a preset threshold, then the image is processed by an image enhancement tool to obtain a first image set; For the first text set, a text segmentation tool is used to segment the text to obtain segmented text units. Based on the text units, a semantic matching tool is used to perform category labeling to determine their business category, thus obtaining the second text set. For the first image set, an image segmentation tool is used to split it into multiple image units, and a feature extraction tool is used to obtain the visual features of the image units. If the feature value is lower than a preset threshold, it is adjusted by an image optimization tool to obtain a second image set. Based on the second text set and the second image set, a data fusion tool is used to perform association mapping to form a unified business dataset. The consistency of the dataset is detected by a content verification tool to determine whether it meets the classification target, thus obtaining the final first dataset.
3. The RAG-based intelligent PDF retrieval and generation method as described in claim 1, characterized in that, The step of matching the retrieval index database using a similarity algorithm based on the user-input query conditions, and expanding the semantic query scope for rematching if the relevance of the matching result is lower than a preset threshold, includes: Based on the user's input query conditions, the input content is structured using a parsing tool to obtain the decomposed query elements; Based on the decomposed query elements, the cosine similarity algorithm is used to perform a preliminary matching operation on the index structure in the retrieval index library to obtain the initial matching results; For the initial matching results, the relevance evaluation module calculates the degree of fit between the matching results and the query conditions to determine whether the degree of fit reaches the preset threshold. If the fit is lower than the preset threshold, the range of the query conditions will be adjusted by the semantic expansion tool to obtain the expanded semantic content. Based on the expanded semantic content, the search index is re-matched to determine the updated matching results. For the updated matching results, the verification module checks the completeness of the matching results. If missing data is detected, the supplementation tool extracts relevant information from the index structure to obtain complete matching data. Based on the complete matching data, the storage management module associates and saves the matching results with the user-input query conditions to determine the final query record.
4. A RAG-based intelligent PDF retrieval and generation system, used to implement the RAG-based intelligent PDF retrieval and generation method as described in any one of claims 1 to 3, characterized in that, The RAG-based intelligent PDF retrieval and generation system includes: The forming module is used to acquire the input document data, parse the document data using a pre-established classification model, and extract text content and image content to form the first dataset; The acquisition module is used to extract features from the image content in the first dataset using a deep learning model, and simultaneously perform semantic analysis on the text content in the first dataset using natural language processing techniques to obtain a multimodal feature set. The first generation module is used to generate a second dataset by applying an information integration algorithm to perform unified encoding processing based on the multimodal feature set. If the completeness of the fused feature vector in the second dataset is detected to be lower than a preset threshold, then supplementary contextual semantic analysis is used to fill in the missing information. The second generation module is used to perform clustering processing on the fused feature vectors in the second dataset using a preset index building mechanism to generate a retrieval index library containing a classification index structure. The matching module is used to match the search index database with the query conditions input by the user using a similarity algorithm. If the relevance of the matching result is lower than a preset threshold, the semantic query range is expanded and rematching is performed.
5. The RAG-based intelligent PDF retrieval and generation system as described in claim 4, characterized in that, The forming module includes: The forming unit is used to obtain the original content from the input document data, perform preliminary analysis of the document using a pre-established classification tool, and extract the text content and image content to form an initial set respectively; The first acquisition unit is used to determine the completeness of the initial set using a content recognition tool. If the text content is missing, the complete data is re-acquired using a document scanning tool to obtain the first text set. The second acquisition unit is used to obtain a first image set by processing the image content resolution with an image enhancement tool if the image content resolution is lower than a preset threshold. The third acquisition unit is used to perform word segmentation processing on the first text set using a text segmentation tool, acquire the segmented text units, and perform category labeling on the text units using a semantic matching tool to determine their business categories, thereby obtaining the second text set. The fourth acquisition unit is used to split the first image set into multiple image units using an image segmentation tool, acquire the visual features of the image units using a feature extraction tool, and adjust the feature values using an image optimization tool if the feature values are lower than a preset threshold to obtain a second image set. The fifth acquisition unit is used to perform association mapping based on the second text set and the second image set using a data fusion tool to form a unified business dataset, and to detect the consistency of the dataset using a content verification tool to determine whether it meets the classification target, thereby obtaining the final first dataset.
6. The RAG-based intelligent PDF retrieval and generation system as described in claim 4, characterized in that, The acquisition module includes: The sixth acquisition unit is used to split the image content in the first dataset into multiple visual element units using an image segmentation tool, process the visual element units using a feature extraction tool to obtain corresponding image features, and if the value of the image feature is lower than a preset threshold, adjust it using an image enhancement tool to obtain an adjusted image feature set. The first determining unit is used to compare the image features with a preset classification standard using a feature mapping tool based on the adjusted image feature set. If the features do not meet the standard, a secondary processing is performed using a feature optimization tool to determine a subset of image features that meet the standard. The seventh acquisition unit is used to split the text content in the first dataset into multiple language units using a text segmentation tool, process the language units using a semantic analysis tool, obtain the corresponding text semantic features, and obtain a text semantic set. The eighth acquisition unit is used to perform multimodal fusion of the image feature subset and the text semantic set through a data integration tool, and to perform business matching on the fused data using a content classification tool to determine whether it meets the preset target, so as to obtain the final multimodal feature set.
7. The RAG-based intelligent PDF retrieval and generation system as described in claim 4, characterized in that, The first generation module includes: The ninth acquisition unit is used to perform unified processing on the feature encoding according to the multimodal feature set using an information integration tool to generate a preliminary fusion dataset, and to detect the encoding consistency of the preliminary fusion dataset. If the detected consistency is lower than a preset threshold, it is adjusted by a data calibration tool to obtain a fusion dataset with consistent encoding. The second determining unit is used to evaluate the vector integrity of the fused dataset with consistent encoding using a feature detection tool. If the integrity of the fused dataset is found to be lower than a preset threshold, supplementary information is extracted using a context analysis tool to determine the fused dataset with improved integrity. The tenth acquisition unit is used to fill in the missing parts of the fused dataset after the integrity improvement by using information filling tools and semantic supplementation tools to obtain the filled fused dataset. The eleventh acquisition unit is used to classify and compare the filled fusion dataset using a business matching tool, and to judge the fusion dataset according to pre-established rules to obtain the final dataset that meets the business objectives.
Citation Information
Patent Citations
Multi-mode-based data retrieval enhancement method
CN119961461A
Archive management system based on artificial intelligence
CN120216747A
Multi-modal data processing method and device and storage medium
CN120541274A
Cited By
An inference enhancement system based on document parsing
CN122616702A