RAG multi-mode document analysis method and device and medium
By adopting technologies such as Kubernetes and large language models in the RAG system, the problem of insufficient complexity and dynamics of multimodal document analysis is solved, and a multimodal document processing system with high efficiency, strong scalability and real-time performance is realized.
Patent Information
- Application Number
- CN202510571284.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-06
AI Technical Summary
The existing RAG systems have high resolution complexity when processing multimodal documents, lack unified processing capabilities, insufficient scalability and resource utilization, and are difficult to deal with high concurrent requests and large-scale document processing, and the update process is not dynamic, affecting system availability and real-time.
The multimodal document analysis system based on Kubernetes is adopted to identify document types through classification models and assign parsing tasks, adjust the parsing process of Kubernetes task nodes, use large language models to perform semantic slicing, and convert semantic chunking into embedded vectors and store them in vector database to support fast retrieval.
It realizes efficient analysis and processing of multimodal documents, improves the scalability and resource utilization of the system, supports high concurrent requests and large-scale document processing, and enhances the dynamic and real-time nature of the system.
Smart Images

Figure CN120087357A_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the field of natural language processing technology, and in particular to a RAG multimodal document parsing method, device and medium. Background Art
[0002] In the field of artificial intelligence and natural language processing, Retrieval-Augmented Generation (RAG) technology significantly improves the processing capabilities of knowledge-intensive tasks by combining retrieval systems and generation models. With the development of natural language technology, the existing RAG system needs to process multimodal documents such as PDF, tables, images, audio, and video. At this time, how to parse multimodal documents has become an important part of the processing process.
[0003] The current parsing complexity of multimodal documents is relatively high. It is difficult to efficiently extract structured information such as text, tables, codes, and formulas based on traditional parsing methods. Secondly, existing parsing methods lack the ability to uniformly process multimodal data and cannot simultaneously support the parsing and retrieval of text, images, audio, and video, resulting in low data utilization. In addition, the system scalability and resource utilization of existing parsing methods are insufficient, making it difficult to cope with high concurrent requests and large-scale document processing requirements. Especially in the process of document parsing and segmentation, computing resources are unevenly distributed, which can easily become a performance bottleneck. Moreover, traditional RAG systems lack dynamism in document updates and knowledge base maintenance. The update process usually requires downtime or manual intervention, which affects the availability and real-time performance of the system. Summary of the invention
[0004] In order to solve the above technical problems, one or more embodiments of the present specification provide a RAG multimodal document parsing method, device and medium.
[0005] One or more embodiments of this specification adopt the following technical solutions: One or more embodiments of this specification provide a RAG multimodal document parsing method, the method comprising: Acquire the multimodal document uploaded by the user terminal based on the multimodal document parsing system, and identify the document type of the multimodal document based on the classification model; Allocating the parsing task of the multimodal document to a corresponding Kubernetes task node based on the document type; wherein each of the Kubernetes task nodes has a corresponding parser; Based on the parsing requirements of the multimodal document, the parsing process of each of the Kubernetes task nodes is adjusted to execute the parsing task of the multimodal document based on the updated parsing process to obtain parsed data; Preprocess the parsed data, and perform semantic slicing on the processed parsed data based on the dynamic text windowing method of a pre-set large language model to obtain semantic chunks corresponding to the processed parsed data; Convert the semantic chunks into embedding vectors and store them in a pre-set vector database for fast retrieval based on the pre-set vector database.
[0006] Optionally, in one or more embodiments of this specification, before obtaining the multi-modal document uploaded by the client, the method further includes: Install and configure the container orchestration environment of the Kubernetes cluster on the multi-modal document parsing system, and install and initialize the pre-set distributed storage system to implement the environment initialization of the multi-modal document parsing system; Obtain the metadata attributes corresponding to the document upload task created by the client to automatically generate a storage path based on the metadata attributes; Shard the multi-modal document based on a pre-set shard size, and parallelly upload each shard to the pre-set distributed storage system based on the storage path, and generate a unique document identifier corresponding to the multi-modal document; wherein, the pre-set distributed storage system is a MinIO distributed storage system.
[0007] Optionally, in one or more embodiments of this specification, identifying the document type of the multi-modal document based on a deep learning model specifically includes: Determine the document field corresponding to the multi-modal document, collect the publicly available multi-modal documents corresponding to the document field, and fine-tune the existing deep learning model based on the publicly available multi-modal documents and the ResNet-50 architecture to obtain a corresponding classification model; If it is determined that the multi-modal document has been stored in the pre-set distributed storage system, trigger the classification task of the multi-modal document based on the pre-set distributed storage system, extract the corresponding preamble content in the multi-modal document, and use the preamble content as the classification basis content for the classification task; Perform format standardization processing on the classification basis content, and send the processed classification basis content to the classification model based on a pre-set protocol to output the type prediction result of the multi-modal document; If the confidence level corresponding to the type prediction result is greater than the pre-set confidence level threshold, use the type prediction result as the document type. If the confidence level is less than the pre-set confidence level threshold, trigger an artificial review process to determine the document type of the multi-modal document.
[0008] Optionally, in one or more embodiments of this specification, extracting the corresponding preamble content in the multi-modal document and using the preamble content as the classification basis content for the classification task specifically includes: Determine the constraint conditions for the multi-modal document extraction content based on the historical classification task records in the document field; wherein, the constraint conditions include: latency budget, computing resource limitation; Determine the key information distribution information of different types of documents according to the historical parsing data corresponding to the historical classification task records; Based on the key information distribution information of different types of documents, determine the key information concentration area corresponding to the multi-modal document; Based on the constraint conditions and the key information concentration area, define the objective function corresponding to the current extraction content; Perform simulated annealing on the objective function within the analyzable capacity range of the classification model to iteratively determine the size of the pre-extracted content that should be extracted from the multi-modal document.
[0009] Optionally, in one or more embodiments of this specification, allocate the parsing tasks of the multi-modal documents to the corresponding Kubernetes task nodes based on the document type, specifically including: Obtain the scalable scheduler framework of the Kubernetes cluster, and perform secondary development on the scalable scheduler framework based on the corresponding parsing scenario of the user side to construct the scheduling decision engine for the parsing tasks; Receive the batch parsing tasks created by the user side, and generate the parsing tasks for each multi-modal document based on the document type and storage path of each multi-modal document corresponding to the batch parsing tasks, and submit each parsing task to the task queue of Kubernetes; Obtain the metric data of each task node in the Kubernetes cluster, and evaluate the resource status of each task node in the Kubernetes cluster based on the metric data; Based on the scheduling decision engine, obtain each parsing task in the task queue, and perform scheduling on the parsing tasks based on the average parsing resources corresponding to each document type and the resource status of each node to determine the Kubernetes task node corresponding to each parsing task.
[0010] Optionally, in one or more embodiments of this specification, adjust the parsing processes of each Kubernetes task node based on the parsing requirements of the multi-modal document, and execute the parsing tasks of the multi-modal document based on the updated parsing processes to obtain parsing data, specifically including: Based on the parsing task of the multi-modal document uploaded by the user side, determine whether the parsing task has custom parsing content; Otherwise, obtain the fixed parsing processes corresponding to each document type in each of the Kubernetes task nodes, and perform the parsing task of the multimodal document based on the fixed parsing processes to obtain parsing data; If so, obtain the parsing operations required for the custom parsing content, and determine whether the parsing operations conflict with the fixed parsing processes; If so, determine the conflict type of the conflict, and update the fixed parsing processes based on the conflict type to add the parsing operations, so as to obtain the updated parsing processes; If not, determine the dependent operations corresponding to the parsing operations, and add the parsing operations to the fixed parsing processes based on the dependent operations to obtain the updated parsing processes; Perform the parsing task of the multimodal document based on the updated parsing processes to obtain parsing data.
[0011] Optionally, in one or more embodiments of this specification, preprocess the parsing data, and perform semantic slicing on the processed parsing data based on the dynamic text windowing method of a pre-set large language model to obtain semantic chunks corresponding to the processed parsing data, specifically including: Perform encoding detection on the parsing data based on a pre-set encoding library, and execute a text cleaning pipeline to clean the detected parsing data to obtain the cleaned parsing data, and perform dependency syntactic analysis on the cleaned parsing data to implement sentence segmentation to obtain the processed parsing data; Initialize the parameters of the dynamic text windowing method, and window each of the processed parsing data based on the initialized parameters; Through semantic boundary detection, determine the perplexity mutation points of each of the processed parsing data within the window, determine the semantic boundary according to a preset number of consecutive perplexity mutation points, and determine the initial semantic chunks corresponding to the processed parsing data based on the semantic boundary; Merge the initial semantic chunks based on the similarity between adjacent initial semantic chunks to obtain the merged initial semantic chunks; Construct a semantic relationship graph corresponding to the merged initial semantic chunks, perform clustering optimization on the semantic relationship graph, and determine the semantic chunks corresponding to the processed parsing data according to the optimized semantic relationship graph and a pre-set rule engine.
[0012] Optionally, in one or more embodiments of this specification, convert the semantic chunks into embedding vectors and store them in a pre-set vector database for fast retrieval based on the pre-set vector database, specifically including: Preprocess the semantic chunks based on a preset indexing module, convert the preprocessed semantic chunks into embedding vectors based on a preset large language model, and store them in a preset vector database; Receive the user's query content, and convert the query content into a query embedding vector based on a preset large language model; Retrieve the most similar semantic chunks based on the similarity between the query embedding vector and each of the embedding vectors in the preset vector database.
[0013] One or more embodiments of this specification provide a RAG multi-modal document parsing device, which includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute any of the above methods.
[0014] A non-volatile computer storage medium provided by one or more embodiments of this specification stores computer-executable instructions, and the computer-executable instructions are configured to: be able to execute any of the above methods.
[0015] The above at least one technical solution adopted by the embodiments of this specification can achieve the following beneficial effects: Quickly and accurately identify the document type of the multi-modal document through a classification model, and based on this, assign the parsing task to the corresponding Kubernetes task node, realizing the automatic assignment and parallel processing of tasks, improving the efficiency of the entire parsing process, being able to process multiple different types of document parsing tasks simultaneously, and reducing the processing time. According to the parsing requirements of the multi-modal document, the parsing process of the Kubernetes task node can be adjusted, and this flexibility enables the system to adapt to changes in different document types and parsing requirements. Using the dynamic text windowing method of the preset large language model for semantic slicing can better understand the semantic structure of the text, divide the parsing data into meaningful semantic chunks, and help to deeply mine the semantic information in the document. Converting the semantic chunks into embedding vectors and storing them in a preset vector database facilitates fast retrieval based on the vector database and improves the retrieval efficiency. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. In the drawings: Figure 1 It is a schematic flowchart of a method for RAG multi-modal document parsing provided by an embodiment of this specification; Figure 2 It is a schematic structural diagram of a device for RAG multi-modal document parsing provided by an embodiment of this specification; Figure 3 It is a schematic structural diagram of a non-volatile storage medium provided by an embodiment of this specification. Specific embodiments
[0017] The embodiments of this specification provide a method, device and medium for RAG multi-modal document parsing.
[0018] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this specification, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.
[0019] As Figure 1 shown, the embodiments of this specification provide a schematic flowchart of a method for RAG multi-modal document parsing. As Figure 1 can be seen, in one or more embodiments of this specification, a method for RAG multi-modal document parsing includes: S101: Obtain a multi-modal document uploaded by a user terminal based on a multi-modal document parsing system, so as to identify the document type of the multi-modal document based on a classification model.
[0020] The multi-modal document parsing system receives the multi-modal document uploaded by the user through the interaction interface of the user terminal. These documents may contain various data forms such as text, images, audio, and video. To facilitate the parsing of multi-modal documents, it is necessary to identify the document type of the multi-modal document according to the classification model.
[0021] Further, in one or more embodiments of this specification, before obtaining the multi-modal document uploaded by the user terminal, the method further includes the following process: First, install and configure the container orchestration environment of the Kubernetes cluster on the multi-modal document parsing system, and install and initialize the pre-configured distributed storage system to realize the environment initialization of the multi-modal document parsing system. That is, in a certain scenario, the Kubernetes cluster or other container orchestration environments will be installed and configured, and the system and other middleware such as Minio, Elasticsearch, PostgreSQL, Argo workflow, etc. will be installed and initialized. Then, when the user creates a document upload task on the multi-modal document parsing system, the system will obtain the metadata attributes corresponding to the task. That is, when the user creates a document upload task through the system, PostgreSQL will be used to record metadata attributes such as the file MD5, size, and upload time, so as to automatically generate a storage path such as " / tenant_id / year / month / day / uuid_filename.ext" based on the metadata attributes. Then, the multi-modal document is fragmented based on the pre-configured fragmentation size, and each fragment is parallelly uploaded to the pre-configured distributed storage system according to the storage path, and a unique document identifier corresponding to the multi-modal document is generated. Among them, the pre-configured distributed storage system is the MinIO distributed storage system. The use of the MinIO distributed storage system in this process ensures the high availability and fault tolerance of the data. Even if some storage nodes fail, the data will not be lost and can still be normally accessed and processed, providing a reliable guarantee for the long-term storage and subsequent use of multi-modal documents. Moreover, whether it is document fragmentation processing, when faced with a large number of multi-modal document uploads, the system can run stably and efficiently without performance bottlenecks due to too many files or too large files.
[0022] Specifically, in one or more embodiments of the present specification, identifying the document type of a multi-modal document based on a deep learning model specifically includes: Determine the document domain corresponding to the multi-modal document, collect publicly available multi-modal documents corresponding to the document domain, and then fine-tune the existing deep learning model based on the publicly available multi-modal documents and the ResNet-50 architecture to obtain the corresponding classification model. Through fine-tuning, the model learns the features of different types of documents within a specific document domain, thereby obtaining a classification model for multi-modal documents in that domain, making it more adaptable to the specific classification task requirements. If it is determined that the multi-modal document has been stored in the pre-set distributed storage system, trigger the classification task of the multi-modal document based on the pre-set distributed storage system to extract the corresponding preamble content within the multi-modal document, for example, extract the first 1MB of the document content, and use the preamble content as the classification basis content for the classification task. Then, perform format standardization processing on the classification basis content, and send the processed classification basis content to the classification model based on the pre-set protocol to output the type prediction result of the multi-modal document. If the confidence level corresponding to the type prediction result is greater than the pre-set confidence threshold, use the type prediction result as the document type; if the confidence level is less than the pre-set confidence threshold, trigger the manual review process for manual review to determine the document type of the multi-modal document.
[0023] In this process, the model is fine-tuned using publicly available multi-modal documents in a specific domain, making the classification model more suitable for the actual application scenario, better able to capture the subtle features of different types of documents, and improving the accuracy of classification. Compared with the general model, the model optimized for a specific domain can more accurately identify the document type. In addition, the system automatically triggers the classification task and makes predictions based on the model, improving the classification efficiency and realizing the automation of document type recognition. At the same time, the introduction of manual review for low model confidence levels makes up for the possible deficiencies of the model, ensures the reliability of the classification results, and achieves a balance between efficiency and accuracy.
[0024] Furthermore, in one or more embodiments of this specification, extracting the corresponding preamble content within the multi-modal document to use the preamble content as the classification basis content for the classification task specifically includes: First, refer to the historical classification task records in the document field and analyze the constraints that need to be followed when extracting content from multimodal documents. Among them, it should be noted that the constraints include: latency budget and computing resource limitations. The latency budget refers to the maximum time that can be tolerated from the start of content extraction to the completion of the classification task, which is to ensure the system response speed and user experience; the computing resource limitations involve the available computing resources such as CPU and memory when performing content extraction and classification tasks, ensuring that the task can still be completed normally under limited resources. Based on the historical analysis data corresponding to the historical classification task records, conduct a detailed analysis of different types of documents to find out the distribution rules of key information in various types of documents. For example, in some technical documents, the key information may be concentrated in the abstract part at the beginning and the core technology elaboration part in the middle; while in news reports, the key information is often in the lead part at the beginning. Combining the key information distribution information of different types of documents obtained above, determine the key information concentration area for the current multimodal document. According to the constraints and the key information concentration area, define the objective function corresponding to the current extracted content. For example, define the objective function as: processing latency ≤ latency budget, resource consumption ≤ resource limitation, key information coverage rate ≥ preset threshold. Within the analyzable capacity range of the classification model, use the simulated annealing algorithm to optimize the objective function to iteratively determine the size of the pre-extracted content that should be extracted from the multimodal document.
[0025] S102: Allocate the parsing task of the multimodal document to the corresponding Kubernetes task nodes based on the document type; where each of the Kubernetes task nodes has a corresponding parser.
[0026] Multimodal documents contain various data forms, and the parsing methods for different types of documents vary greatly. Without a targeted parsing mechanism, it is difficult for a general parser to process all types of documents efficiently and accurately. Therefore, in the multimodal document parsing system in the embodiments of this specification, first, a deep learning model or other classification methods will be used to identify the type of the multimodal document. Kubernetes task nodes are the working units of the multimodal document parsing system in the Kubernetes cluster, and each task node is equipped with a corresponding parser. These parsers are designed specifically for processing specific types of documents. When the system identifies the type of the multimodal document, it will send the parsing task to the corresponding Kubernetes task node according to the preset allocation rules. If it is a PDF document, the parsing task will be allocated to the task node with a PDF parser; if it is an audio document, it will be allocated to the node equipped with an audio parser. In this way, each task node can focus on processing the document type it is good at, improving the efficiency and accuracy of parsing.
[0027] Specifically, in one or more embodiments of this specification, the parsing task of the multimodal document is assigned to the corresponding Kubernetes task node based on the document type, which specifically includes the following process: The resource requirements for parsing different types of documents vary greatly, and the general scheduling strategy of the native Kubernetes scheduler is difficult to meet the specific needs of multimodal document parsing. Therefore, first, obtain the extensible scheduler framework of the Kubernetes cluster, which provides the infrastructure for subsequent development. Based on the specific parsing scenario of the user side, the framework is developed secondarily to build a scheduling decision engine suitable for multimodal document parsing tasks. For example, if the user mainly processes multimodal documents in the financial field, including a large number of contracts in PDF format and data reports in Excel format, the scheduling decision engine should be developed according to the parsing characteristics of such documents. After receiving the batch parsing tasks created by the user side, the system generates corresponding parsing tasks according to the document type and storage path of each multimodal document, and submits each parsing task to the task queue of Kubernetes, waiting for scheduling and execution. Then, obtain the metric data of each task node in the Kubernetes cluster, including CPU usage, memory occupancy, disk I / O rate, etc. By analyzing these metric data, the system can comprehensively evaluate the resource status of each task node, understand the current workload of each node and the remaining resources available for new tasks.
[0028] Based on the scheduling decision engine, obtain each parsing task in the task queue, and perform scheduling on the parsing tasks based on the average parsing resources corresponding to each document type and the resource status of each node to determine the Kubernetes task node corresponding to each parsing task. In this process, the scheduling decision engine constructed by secondary development of the extensible scheduler framework based on the user parsing scenario can accurately match the user needs. In different industries and different business scenarios, the types and parsing requirements of multimodal documents vary greatly. The customized scheduler can flexibly adjust the scheduling strategy to improve the adaptability and efficiency of scheduling. By obtaining node metric data in real time to evaluate the resource status and combining the average parsing resource requirements of the document type for task scheduling, tasks can be assigned to the most suitable nodes, avoiding resource waste and node overload, and improving the resource utilization rate of the entire Kubernetes cluster.
[0029] In a certain application scenario, the process of allocating the parsing task to the corresponding Kubernetes task node can also be as follows: Use the Kubernetes Scheduler Framework for secondary development of the scheduling decision engine, implement queue management based on Kueue, and use Prometheus to collect node metrics in real time; when the user clicks to create a batch document parsing task, obtain document information from Postgresql and enter the task queue; the scheduling decision engine obtains the task from the task queue, then performs real-time resource evaluation through Prometheus, and at the same time performs multi-dimensional scheduling strategies on the task. For example, for PDF file parsing, if the resources are sufficient, 5 documents are allocated to each node job. If the resources are insufficient, resource expansion is performed through HPA.
[0030] S103: Based on the parsing requirements of the multi-modal document, adjust the parsing processes of the respective Kubernetes task nodes, so as to execute the parsing task of the multi-modal document based on the updated parsing processes and obtain parsing data.
[0031] In a multi-modal document parsing system, different multi-modal documents and different application scenarios will generate diverse parsing requirements. However, the traditional fixed parsing process is difficult to adapt to the diverse parsing requirements of multi-modal documents in different scenarios. Therefore, in the embodiments of this specification, the parsing processes of each Kubernetes task node will be adjusted according to the parsing requirements of the multi-modal document, so as to execute the parsing task of the multi-modal document based on the updated parsing processes and obtain parsing data.
[0032] Specifically, in one or more embodiments of this specification, based on the parsing requirements of the multi-modal document, adjusting the parsing processes of the respective Kubernetes task nodes, so as to execute the parsing task of the multi-modal document based on the updated parsing processes and obtain parsing data, specifically includes: First, based on the parsing task of the multi-modal document uploaded by the user terminal, determine whether the parsing task has custom parsing content. If not, if the parsing task has no custom parsing content, the system will obtain the fixed parsing processes preset for different document types in each Kubernetes task node. Then, parse the multi-modal document according to these fixed parsing processes, and finally obtain parsing data.
[0033] For example, when the Kubernetes task node is a PDF parser, first extract the original text, use ApachePDFBox to parse the PDF native text layer, and retain metadata such as fonts and positions; then use the DeepDoc text purification module: fix garbled characters (such as "##" → "company") based on rules + BiLSTM model; finally output text blocks with coordinates (each block contains: x / y coordinates, font size, semantic tags). After that, scans are processed and the image page is processed using the Tesseract 5.0 OCR engine (LSTM+Attention model); then deepdoc is used based on the DRF version for analysis to distinguish between the main text, header, footer, etc.; table area is re-detected to avoid OCR misrecognition; finally, image content is understood, the CLIP-ViT-L / 14 model is called to generate image descriptions, and Faster R-CNN is used to detect key elements in the image (such as logos / charts); finally, the description text is matched based on spatial coordinates to associate text with images.
[0034] When the Kubernetes task node is a table parser: perform table detection, first use the CascadeTabNet model to locate the table area (output a rotated rectangular box); then use the Transformer-based sequence classifier (to determine whether it is a real table); finally filter the decorative border (accuracy 99.2%). Structural reconstruction: First use OpenCV's Hough transform to detect straight lines, and use morphological processing to repair broken borders; then use the deepdoc cell merging algorithm based on Graph NeuralNetwork to identify cross-row / column cells; dynamically adjust the separator weight (process dotted / light borders). Finally, perform content extraction: first use PDFBox to extract native table text, and then use DeepDoc table semantic annotation: identify table headers / data units (BERT+CRF) and derive cell relationships (such as "same as above" → auto-fill).
[0035] When the Kubernetes task node is an audio parser: first, speech recognition is performed, and automatic language detection (supporting 50+ languages) and speaker separation (based on voiceprint clustering) are performed through Whisper-large-v3 model processing; then, domain terminology correction (medical / legal and other professional dictionaries) and timestamp calibration (MFCC feature alignment) are performed based on deepdoc.
[0036] When the Kubernetes task node is a video parser: First, use the FFmpeg frame extraction strategy to extract frames, and use the key frame priority strategy for dynamic sampling. Use DINOv2 feature extraction to generate a 1024-dimensional visual embedding for each frame, and then filter based on DeepDoc key frames (based on semantic importance scores).
[0037] If it is determined that the parsing task has custom parsing content, obtain the parsing operations required for the custom parsing content, and determine whether there is a conflict between the parsing operations and the fixed parsing process. The conflict may manifest as contradictions in aspects such as the order of parsing operations and resource occupancy. For example, the custom operation needs to be completed before a certain step, but in the inherent process, this step has already occurred after other operations. If there is a conflict, the system will determine the type of conflict, such as an order conflict or a resource conflict. Then, based on the type of conflict, update the fixed parsing process and reasonably add the custom parsing operation to the process to resolve the conflict and obtain the updated parsing process. For example, if it is an order conflict, it may be necessary to adjust the execution order of certain steps. If the parsing operation does not conflict with the fixed parsing process, the system will determine the dependent operations corresponding to the parsing operation, that is, other operations that need to be completed before executing this parsing operation. Based on these dependent operations, add the parsing operation to the fixed parsing process to form the updated parsing process. Then, based on the updated parsing process, execute the parsing task of the multimodal document to obtain the parsing data.
[0038] S104: Preprocess the parsing data to perform semantic slicing on the processed parsing data based on the dynamic text windowing method of the preset large language model, and obtain the semantic chunks corresponding to the processed parsing data.
[0039] Traditional text chunking methods often divide based on fixed lengths or simple rules, making it difficult to accurately capture the semantic boundaries of the text, resulting in incomplete or incoherent semantics after chunking. Therefore, in the embodiments of this specification, the parsing data will be preprocessed, and then semantic slicing will be performed on the processed parsing data according to the dynamic text windowing method of the preset large language model to obtain the semantic chunks corresponding to the processed parsing data.
[0040] Specifically, in one or more embodiments of this specification, preprocess the parsing data to perform semantic slicing on the processed parsing data based on the dynamic text windowing method of the preset large language model, and obtain the semantic chunks corresponding to the processed parsing data, which specifically includes the following processes: Use a preset encoding library (such as chardet) to detect the encoding of the parsing data, and execute a text cleaning pipeline to clean the detected parsing data to obtain the cleaned parsing data. Then, perform dependency syntactic analysis on the cleaned parsing data to achieve sentence segmentation and obtain the processed parsing data. It should be noted that the text cleaning pipeline includes operations such as removing invisible control characters (characters with ASCII < 32), standardizing quotation marks and hyphens, and merging consecutive whitespace characters.
[0041] Initialize parameters for the dynamic text windowing method, setting parameters such as the initial window size (e.g., 256 tokens), the minimum effective chunk size (e.g., 64 tokens), and the overlap region size between adjacent windows (e.g., 64 tokens). Based on these initialization parameters, perform windowing on each processed parsed data, dividing the text into multiple windows, each window containing a certain amount of text content, and there is partial overlap between adjacent windows to ensure the coherence of context information. Then, through the semantic boundary detection method, calculate the perplexity mutation points of the processed parsed data within each window. The perplexity mutation points reflect the changes in text semantics. When a preset number (e.g., 3) of perplexity mutation points appear continuously and meet certain significance conditions (e.g., P < 0.01), determine that this is the semantic boundary. According to these semantic boundaries, divide the processed parsed data into initial semantic chunks, each chunk having relatively independent semantics. Then, calculate the similarity between adjacent initial semantic chunks. If the cosine similarity > 0.85 and the time continuity < 5 seconds, the initial semantic chunks are merged to obtain the merged initial semantic chunks.
[0042] Construct the semantic relationship graph corresponding to the merged initial semantic chunks to optimize the clustering of the semantic relationship graph. According to the optimized semantic relationship graph and the preset rule engine, determine the semantic chunks corresponding to the processed parsed data. In a certain scenario, it is possible to optimize the chunk quality based on the graph structure representation and the Louvain community discovery algorithm, combined with rule-based post-processing. First, construct the semantic relationship graph, where the nodes represent candidate chunks, and the weight of the edge is determined by the co-reference resolution score calculated by the Coreferee tool. Subsequently, perform community partitioning, use the Louvain method for clustering optimization, and by setting the resolution parameter γ = 1.25, continuously iterate and optimize until the modularity Q > 0.6. Finally, use the rule engine for post-processing, including forcibly splitting long paragraphs such as those greater than 512 tokens, merging short paragraphs such as those less than 32 tokens and with a similarity greater than 0.9, and processing dialogue scenarios, thereby ensuring the coherence and rationality of the chunks.
[0043] In this process, the combination of dynamic text windowing and semantic boundary detection can accurately capture the semantic changes in the text and divide the text into chunks with clear semantic boundaries. This method can better reflect the semantic structure of the text than traditional fixed-length chunking, improving the accuracy of semantic understanding. In addition, the initial semantic chunk merging step merges chunks with similar and coherent semantics through similarity calculation and time continuity judgment, avoiding the problem of semantic incoherence caused by overly fragmented chunks, making the final semantic chunks more in line with the logical structure of the text. In addition, the construction and clustering optimization of the semantic relationship graph, as well as the use of the preset rule engine, enable this process to adapt to different types of texts and diverse analysis requirements.
[0044] S105: Convert the semantic chunk into an embedding vector and store it in a pre-set vector database for fast retrieval based on the pre-set vector database.
[0045] Traditional text retrieval methods are mainly based on keyword matching, which is difficult to understand the semantics of text. For synonyms, near-synonyms, and texts with similar semantics but different expressions, the retrieval effect is poor. Therefore, in the embodiments of this specification, the semantic chunk is converted into an embedding vector and a vector database is used for retrieval, so that matching can be performed from the semantic level, greatly improving the accuracy and recall rate of retrieval. Then, after converting the semantic chunk into an embedding vector and storing it in the pre-set vector database, fast retrieval will be performed according to the pre-set vector database.
[0046] Specifically, in one or more embodiments of this specification, converting the semantic chunk into an embedding vector and storing it in the pre-set vector database for fast retrieval based on the pre-set vector database specifically includes: Preprocess the semantic chunk based on a pre-set indexing module, such as text standardization, adding domain prefixes, etc., and convert the preprocessed semantic chunk into an embedding vector based on a pre-set large language model and store it in the pre-set vector database. In this process, batch requests can be constructed and then written to the repository in parallel using multi-threading. Then, receive the user's query content, convert the query content into a query embedding vector based on the pre-set large language model, and retrieve the most similar semantic chunk based on the similarity between the query embedding vector and each of the embedding vectors in the pre-set vector database. This vector-based retrieval method significantly improves the speed and accuracy of semantic retrieval.
[0047] As Figure 2 shown, in the embodiments of this specification, a schematic structural diagram of a RAG multi-modal document parsing device is provided. From Figure 2 it can be seen that in one or more embodiments of this specification, a RAG multi-modal document parsing device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute any of the above methods.
[0048] As Figure 3 shown, the embodiments of this specification provide a schematic structural diagram of a non-volatile storage medium. From Figure 3It can be seen that in one or more embodiments of this specification, a non-volatile storage medium stores computer-executable instructions 301, and the computer-executable instructions 301 can execute any of the above-described methods.
[0049] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and non-volatile computer storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0050] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended specification. In some cases, the actions or steps recorded in the specification can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or consecutive order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0051] The above is only one or more embodiments of this specification and is not used to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of this specification.
Claims
1. A RAG multimodal document parsing method, characterized in that: The method comprises: Acquire the multimodal document uploaded by the user terminal based on the multimodal document parsing system, and identify the document type of the multimodal document based on the classification model; Allocating the parsing task of the multimodal document to a corresponding Kubernetes task node based on the document type; wherein each of the Kubernetes task nodes has a corresponding parser; Based on the parsing requirements of the multimodal document, the parsing process of each of the Kubernetes task nodes is adjusted to execute the parsing task of the multimodal document based on the updated parsing process to obtain parsed data; Preprocessing the parsed data, semantically slicing the processed parsed data in a dynamic text windowing manner based on a preset large language model, and obtaining semantic blocks corresponding to the processed parsed data; The semantic blocks are converted into embedded vectors and stored in a preset vector database so as to be quickly retrieved based on the preset vector database.
2. A RAG multimodal document parsing method according to claim 1, characterized in that: Before obtaining the multimodal document uploaded by the user, the method further includes: Install and configure the container orchestration environment of the Kubernetes cluster on the multimodal document parsing system, and install and initialize the pre-set distributed storage system to achieve environment initialization of the multimodal document parsing system; Acquire the metadata attribute corresponding to the document upload task created by the user end, so as to automatically generate a storage path based on the metadata attribute; The multimodal document is segmented based on a preset segment size, so that each segment is uploaded to the preset distributed storage system in parallel based on the storage path, and a unique document identifier corresponding to the multimodal document is generated; wherein the preset distributed storage system is a MinIO distributed storage system.
3. A RAG multimodal document parsing method according to claim 1, characterized in that: Identifying the document type of the multimodal document based on a deep learning model specifically includes: Determine the document domain corresponding to the multimodal document to collect public multimodal documents corresponding to the document domain, and fine-tune the existing deep learning model based on the public multimodal documents and the ResNet-50 architecture to obtain a corresponding classification model; If it is determined that the multimodal document has been stored in the preset distributed storage system, a classification task of the multimodal document is triggered based on the preset distributed storage system to extract the corresponding pre-content in the multimodal document, so as to use the pre-content as the classification basis content of the classification task; Performing format unification processing on the classification basis content, so as to send the processed classification basis content to the classification model based on a preset protocol, and outputting the type prediction result of the multimodal document; If the confidence corresponding to the type prediction result is greater than a preset confidence threshold, the type prediction result is used as the document type; if the confidence is less than the preset confidence threshold, a manual review process is triggered to determine the document type of the multimodal document.
4. A RAG multimodal document parsing method according to claim 3, characterized in that: Extracting the corresponding preceding content in the multimodal document so as to use the preceding content as the classification basis content of the classification task specifically includes: Based on the historical classification task records of the document field, determining the constraints of the multimodal document content extraction; wherein the constraints include: delay budget and computing resource limitations; Determine key information distribution information of different types of documents according to the historical parsing data corresponding to the historical classification task records; Determining a key information concentration area corresponding to the multimodal document based on key information distribution information of the different types of documents; Based on the constraint conditions and the key information concentration area, define an objective function corresponding to the current extracted content; The objective function is subjected to simulated annealing within the analyzable capacity of the classification model to iteratively determine the size of the preceding content that should be extracted from the multimodal document.
5. A RAG multimodal document parsing method according to claim 2, characterized in that: Allocating the parsing task of the multimodal document to the corresponding Kubernetes task node based on the document type specifically includes: Obtain an extensible scheduler framework of the Kubernetes cluster, perform secondary development on the extensible scheduler framework based on the corresponding parsing scenario of the user end, and build a scheduling decision engine for the parsing task; Receive the batch parsing tasks created by the user end, generate parsing tasks for each multimodal document based on the document type and storage path of each multimodal document corresponding to the batch parsing task, and submit each parsing task to the task queue of Kubernetes; Obtaining indicator data of each task node in the Kubernetes cluster to evaluate the resource status of each task node in the Kubernetes cluster based on the indicator data; Based on the scheduling decision engine, each parsing task in the task queue is obtained, and based on the average parsing resources corresponding to each document type and the resource status of each node, the parsing tasks are scheduled to determine the Kubernetes task node corresponding to each parsing task.
6. A RAG multimodal document parsing method according to claim 1, characterized in that: Based on the parsing requirements of the multimodal document, the parsing process of each of the Kubernetes task nodes is adjusted to execute the parsing task of the multimodal document based on the updated parsing process to obtain parsed data, specifically including: Based on the parsing task of the multimodal document uploaded by the user, determining whether the parsing task has custom parsing content; If not, obtain the fixed parsing process corresponding to each document type in each Kubernetes task node, so as to perform the parsing task of the multimodal document based on the fixed parsing process to obtain parsed data; If yes, obtaining the parsing operation required for the custom parsing content, and determining whether the parsing operation conflicts with the fixed parsing process; If yes, determining the conflict type of the conflict, and updating the fixed parsing process based on the conflict type to add the parsing operation to obtain an updated parsing process; If not, determining the dependent operation corresponding to the parsing operation, so as to add the parsing operation to the fixed parsing process based on the dependent operation to obtain an updated parsing process; The parsing task of the multimodal document is performed based on the updated parsing process to obtain parsed data.
7. A RAG multimodal document parsing method according to claim 1, characterized in that: Preprocessing the parsed data, semantically slicing the processed parsed data in a dynamic text windowing manner based on a preset large language model, and obtaining semantic blocks corresponding to the processed parsed data, specifically includes: Performing encoding detection on the parsed data based on a preset encoding library, and executing a text cleaning pipeline to clean the parsed data after the detection, to obtain cleaned parsed data, and performing dependency syntactic analysis on the cleaned parsed data to achieve sentence segmentation to obtain processed parsed data; Initializing parameters of the dynamic text windowing method to window each processed parsed data based on the initialization parameters; Determine the perplexity mutation point of each processed parsed data in the window through semantic boundary detection, determine the semantic boundary according to a preset number of continuous perplexity mutation points, and determine the initial semantic block corresponding to the processed parsed data based on the semantic boundary; Based on the similarity between adjacent initial semantic blocks, the initial semantic blocks are merged to obtain merged initial semantic blocks; A semantic relationship graph corresponding to the merged initial semantic blocks is constructed to perform clustering optimization on the semantic relationship graph, and the semantic blocks corresponding to the processed parsed data are determined based on the optimized semantic relationship graph and a preset rule engine.
8. A RAG multimodal document parsing method according to claim 1, characterized in that: The semantic blocks are converted into embedded vectors and stored in a preset vector database so as to be quickly retrieved based on the preset vector database, specifically including: Preprocessing the semantic blocks based on a preset index module, converting the preprocessed semantic blocks into embedded vectors based on a preset large language model, and storing the embedded vectors in a preset vector database; Receiving user query content, and converting the query content into a query embedding vector based on a preset large language model; Based on the similarity between the query embedding vector and each embedding vector in the preset vector database, the most similar semantic block is retrieved.
9. A RAG multimodal document parsing device, characterized in that: The device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: execute any of the methods described in claims 1-8.
10. A non-volatile storage medium storing computer executable instructions, characterized in that: The computer executable instructions can: execute the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Method for rapidly analyzing multi-source marine business observation data
CN113111140A
Distributed storage method and system for image-text data based on judicial industry
CN116595226A
Scheduling method for analysis and calculation resources of unstructured documents of power grid
CN118260082A
Multi-modal document analysis method and device, electronic equipment and storage medium
CN118917300A
Document classification splitting method and device and storage medium
CN119226381A
Cited By
Multi-modal data processing method and device and electronic equipment
CN120688021A
RAG-oriented embedded service flexible deployment method
CN120872615A
Large language model dynamic adaptation method and system based on Java
CN121012756A
Coal industry multi-modal data intelligent blocking and label fusion decision-making method
CN121030636A
Document analysis method and device, storage medium and electronic device
CN121052240A