RAG document splitting optimization method and system in operator field
By building a document loader and splitter, combined with the BAAI/bge-m3 model and the ParadeDB database, we solved the complex issues of image extraction and conversion and multi-type document processing in RAG, achieved RAG document splitting optimization in the operator field, reduced operation and maintenance costs, and improved processing efficiency.
Patent Information
- Application Number
- CN202510689267.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-16
AI Technical Summary
Existing RAG technology has problems such as missing document images, complex processing of multiple document types, and high maintenance costs. In particular, it is difficult to extract vector graphics and convert them into markdown format, and users need to manually select text splitting logic.
This paper adopts a RAG document splitting optimization method for the operator field, uses minio and ParadeDB to store files, builds a document loader to automatically process different types of files, parses images and converts them into base64 format, uses MarkdownSplitter and CharacterSplitter to split text, combines the BAAI/bge-m3 model for vector embedding and ParadeDB database retrieval, and realizes vector, full-text and hybrid retrieval.
It realizes the unified processing of multiple types of files, reduces manual participation, lowers operation and maintenance costs, is suitable for the processing of complex unstructured files, and improves the accuracy and efficiency of RAG.
Smart Images

Figure CN120654656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model optimization, and in particular to a RAG document splitting optimization method and system in the operator field. Background Art
[0002] RAG (Retrieval-Augmented Generation) is a key direction for the implementation of large models. However, making RAG document parsing and splitting more reasonable and accurate has become a pain point for RAG.
[0003] The following defects exist in the existing technology:
[0004] 1. The problem of missing images in documents, especially the difficulty in extracting vector graphics and converting them into an image format that can be displayed in markdown, reduces the accuracy of RAG.
[0005] 2. The processing methods for various types of documents are relatively complicated, and users need to choose the text splitting logic themselves.
[0006] 3. Vector retrieval and full-text retrieval use different databases, which increases maintenance costs. Summary of the Invention
[0007] The technical task of the present invention is to address the above shortcomings and provide a RAG document splitting optimization method and system in the operator field to solve the current problems of complex image extraction, conversion and processing in RAG. One database can realize vector retrieval, full-text retrieval and hybrid retrieval functions, reducing operation and maintenance costs.
[0008] The technical solution adopted by the present invention to solve its technical problem is:
[0009] A RAG document splitting optimization method in the operator field, the implementation of the method includes the following steps:
[0010] S1: Upload documents and use minio and paradeDB to store source files and file information respectively;
[0011] S2: Build a document loader, automatically select the corresponding loader according to the file type to process the file, convert it into a unified markdown text, parse the images in the document and convert them into base64 format;
[0012] S3: Build an image processor to extract the Base64 string of the image in the text, process the bitmap and vector images, and convert them into the Markdown image reference format;
[0013] S4: Build MarkdownSplitter and CharacterSplitter document splitters to split markdown text and custom text and table text respectively;
[0014] S5: Vector conversion: Use the BAAI / bge-m3 model for vector embedding to convert the text blocks output in step S4 into vectors. Store the text blocks and vectors in the ParadeDB database.
[0015] S6: Text recall, using ParadeDB's `<=>` (vector search) and `@@@` (BM25 search) syntax to search for text blocks. After the search, the BAAI / bge-reranker-v2-m3 model is used to rerank the text and output text blocks that are highly similar to the user's question.
[0016] This method is based on converting various types of files into markdown format, extracting bitmaps and vector graphics and converting them into markdown image elements, and splitting text according to paragraphs. It is suitable for processing complex unstructured files of operators.
[0017] Furthermore, in step S1, after uploading the document, a unique ID is generated for the file, and the file is uploaded to minio for file preview and re-parsing; the file name, file type, and file size information are generated and stored in the ParadeDB database.
[0018] Furthermore, the document loader includes:
[0019] PptLoader: Use the soffice command line tool to convert ppt files into pptx files, and then use PptxLoader to process the converted files;
[0020] PptxLoader: Uses the pptx2md tool to convert pptx to markdown files. After conversion, it generates a markdown file with the same name as the source file and an images folder for storing images. It reads the markdown file, uses regular expressions to extract local image links in the text, reads the corresponding image files in base64 format, replaces the image links in the text, and outputs the converted text string.
[0021] XlsLoader: Use the soffice command line tool to convert xls files to xlsx files, and then use XlsxLoader to process them.
[0022] XlsxLoader: decompresses the xlsx file in zip format. After decompression, the images in the table can be extracted by parsing the rels file. There are two ways to extract images: embedded image extraction and non-embedded image extraction.
[0023] Embedded image extraction: If the source file contains xl / _rels / cellimages.xml.rel and xl / cellimages.xml files, it means that the source file contains embedded images. Parse the xl / cellings.xml file and extract the image path and image ID;
[0024] Extract non-embedded images: If the xl / drawings / drawing1.xml file is included, it means that the source file contains non-embedded images. Parse all rels files under xl / worksheets / sheet and extract the image path and the sheet, row, and column number of the image in the table.
[0025] First, use pandas to read the source file, read all sheet pages, traverse all sheet pages, and replace the corresponding cell contents in the table based on the parsed embedded image and non-embedded image information; convert each row in the table into a markdown table, use [SEQ] as the delimiter for each markdown table, and output text strings;
[0026] CsvLoader: Converts each row of a CSV table into a Markdown table, uses [SEQ] as the delimiter for each Markdown table, and outputs a text string;
[0027] DocLoader: Use the soffice command line tool to convert doc files to docx files, and then use DocxLoader to process the converted files;
[0028] DocxLoader: Uses Mammoth to read docx files into HTML text, and then uses html2text to convert it into Markdown text. Images are automatically converted into base64 format and output as Markdown text.
[0029] PdfLoader: Use the pdf2docx tool to convert the pdf file into a docx file, and then use DocxLoader to process it;
[0030] ImageLoader: It is necessary to extract information from the image. Here, OCR is used to extract the text in the image, and the Qwen2-VL model is used to generate an image description. The text, image description, and image base64 string are integrated into a text and the text is output.
[0031] HtmlLoader: Use the html2text tool to read HTML content into Markdown format, identify image URLs in the text, download the images and convert the image URLs into image base64 strings, and output the modified text;
[0032] XmlLoader: Use lxml to parse XML files, use regular expressions to match image URLs in the text, download the images and convert the image URLs into image base64 strings, and then convert them into markdown text strings;
[0033] TxtLoader: There are two types of txt files: traditional txt files and custom files separated by [SEQ]. It uses regular expressions to match image URLs in text, downloads images, converts image URLs into base64 strings, and outputs text strings.
[0034] JsonLoder: json processing is divided into two types, list format and dictionary format;
[0035] List format processing: convert each object in the list into a key:value string, and separate the objects with [SEQ];
[0036] Dictionary format: converted to key:value format text;
[0037] Use regular expressions to match image URLs in the text, download the image, convert the image URL into a base64 string, and then convert it into a markdown text string;
[0038] JsonlLoader: reads jsonl files, converts the file contents into json files, and then processes them using JsonLoader;
[0039] MarkdownLoader: Reads a markdown file as text, uses regular expressions to match image URLs in the text, downloads the image and converts the image URL into a base64 string, and outputs the text string.
[0040] Furthermore, in step S2, the format after conversion is: [](data:image / {image type};base64,{image base64 string}.
[0041] Furthermore, in step S3,
[0042] For the text string output in step S2, use the regular expression (data:image\ / [\w+-]+;base64,[A-Za-z0-9+ / =]+) to extract the base64 string, generate a unique ID for the image, upload the image to minio for storage, and replace the base64 string with {image bucket} / {knowledge base id} / {file id} / {image file name}, and output the processed text;
[0043] The images to be processed include bitmaps (png, jpg, webp, gif) and vector images (emf, wmf, svg):
[0044] For bitmap images, use base64 to save;
[0045] For vector images, convert the vector images in the document, including flowcharts, icons, and inserted attachments, into bitmaps so that markdown can reference them normally; first use the inkscape tool to convert the vector images into png. If the conversion is unsuccessful, use soffice to convert it to pdf first, and then use inkscape to convert the image area to png.
[0046] Furthermore, the document segmenter is used to segment the text into text blocks and output the segmented text blocks;
[0047] MarkdownSplitter: Applicable to splitting Markdown text blocks. It uses the title splitting method, treating the content within the same title as a text block. It also sets the maximum text block length and the number of text block overlaps to prevent the text block content from being out of focus due to excessive text splitting. Due to the limitation of the large model context length, the text block size is generally 1024 characters or 768 characters.
[0048] CharacterSplitter: Suitable for text splitting after custom txt files and table conversion, using [SEQ] as the paragraph separator.
[0049] Furthermore, the ParadeDB database is a vector database developed based on postgres, which integrates functions such as vector search and BM25 search; one database can realize vector search, full-text search, and hybrid search functions.
[0050] The present invention also claims protection for a RAG document splitting optimization system in the operator field, comprising:
[0051] The document upload module is used to store source files and file information through minio and paradeDB respectively;
[0052] The document conversion module is used to build a document loader, automatically select the corresponding loader according to the file type to process the file, convert it into a unified markdown text, parse the images in the document and convert them into base64 format;
[0053] The document processing module is used to build an image processor, extract the base64 string of the image in the text, process the bitmap and vector images, and convert them into the markdown image reference format;
[0054] Document splitting module, used to build MarkdownSplitter and CharacterSplitter document splitters, which split markdown text and custom text and table text respectively;
[0055] The vector conversion module uses the BAAI / bge-m3 model for vector embedding, converts the text blocks output by the document splitting module into vectors, and stores the text blocks and vectors in the ParadeDB database;
[0056] The text recall module uses ParadeDB's `<=>` (vector search) and `@@@` (BM25 search) syntax to search for text blocks. After the search, it uses the BAAI / bge-reranker-v2-m3 model to rerank the text and output text blocks that are highly similar to the user's question.
[0057] The system specifically implements RAG document splitting in the operator field through the above method.
[0058] The present invention also claims protection for a RAG document splitting optimization implementation device in the operator field, comprising: at least one memory and at least one processor;
[0059] The at least one memory is configured to store a machine-readable program;
[0060] The at least one processor is configured to call the machine-readable program to implement the above method.
[0061] The present invention also claims protection for a computer-readable medium, characterized in that the computer-readable medium stores computer instructions, and when the computer instructions are executed by a processor, the above method is implemented.
[0062] Compared with the prior art, the RAG document splitting optimization method and system of the present invention in the operator field has the following beneficial effects:
[0063] The present invention provides a unified file processing interface that converts various file types into Markdown text format, extracts bitmaps and vector graphics from text and converts them into Markdown image elements, and splits text according to Markdown paragraphs. This method is suitable for processing operators' complex unstructured files and reduces manual intervention. By using the ParadeDB database, a single database can implement vector search, full-text search, and hybrid search functions, reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 This is a schematic diagram illustrating the principles of a RAG document splitting optimization method in an operator field provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0066] A RAG document splitting optimization method in the operator field, the implementation of the method includes the following steps:
[0067] S1: Upload documents and use minio and paradeDB to store source files and file information respectively.
[0068] After uploading the document, a unique ID is generated for the file and the file is uploaded to Minio for file preview and re-parsing; the file name, file type, and file size information are generated and stored in the ParadeDB database.
[0069] S2: Build a document loader, automatically select the corresponding loader according to the file type to process the file, convert it into a unified markdown text, parse the images in the document and convert them into base64 format for S3 processing. The converted format is:
[0070] ` to extract the base64 string, generate a unique ID for the image, upload the image to minio for storage, and replace the base64 string with `{image bucket} / {knowledge base id} / {file id} / {image file name}`, and output the processed text.
[0095] The images to be processed are mainly bitmaps (png, jpg, webp, gif) and vector images (emf, wmf, svg):
[0096] For bitmap images, use base64 to save them.
[0097] For vector images, the vector images in the document, including flowcharts, icons, and inserted attachments, need to be converted into bitmaps so that Markdown can reference them normally; first use the inkscape tool to convert the vector image to png. If the conversion is unsuccessful, use soffice to convert it to pdf first, and then use inkscape to convert the image area to png.
[0098] S4: Build MarkdownSplitter and CharacterSplitter document splitters to split markdown text and custom text and table text respectively.
[0099] The document splitter is used to split the text into text blocks and output the split text blocks.
[0100] MarkdownSplitter: Applicable to splitting markdown text blocks. It adopts the title splitting method, treating the content within the same title as a text block. It also sets the maximum text block length and the number of text block overlaps to prevent the text block content from being out of focus due to excessive text splitting. Due to the limitation of the large model context length, the text block is generally 1024 characters or 768 characters.
[0101] CharacterSplitter: Suitable for text splitting after custom txt files and table conversion, using `[SEQ]` as the paragraph separator.
[0102] S5: Vector conversion: Use the BAAI / bge-m3 model for vector embedding, convert the text block output in step S4 into a vector, and store the text block and vector in the ParadeDB database.
[0103] The ParadeDB database is a vector database developed based on postgres, which integrates functions such as vector search and BM25 search; one database can realize vector search, full-text search, and hybrid search functions.
[0104] S6: Text recall, using ParadeDB's `<=>` (vector search) and `@@@` (BM25 search) syntax to search for text blocks. After the search, the BAAI / bge-reranker-v2-m3 model is used to rerank the text and output text blocks that are highly similar to the user's question.
[0105] This method converts various file types into markdown format, extracts bitmaps and vector graphics and converts them into markdown image elements, and splits text according to paragraphs. It solves the problems of image extraction, conversion, and complex processing in existing RAGs proposed in the background technology, and is suitable for processing complex unstructured files of operators.
[0106] An embodiment of the present invention further provides a RAG document splitting optimization system in the operator field, which specifically implements RAG document splitting in the operator field through the RAG document splitting optimization method in the operator field described in the above embodiment.
[0107] The system includes:
[0108] 1. The document upload module is used to store source files and file information using minio and paradeDB, respectively. After uploading a document, a unique ID is generated for the file and the file is uploaded to minio for preview and re-parsing. The file name, file type, and file size are generated and stored in the ParadeDB database.
[0109] 2. Document conversion module, used to build a document loader, automatically select the corresponding loader for file processing according to the file type, convert it into a unified markdown text, parse the images in the document and convert them into base64 format for S3 processing. The converted format is:
[0110] ` to extract the base64 string, generate a unique ID for the image, upload the image to minio for storage, and replace the base64 string with `{image storage bucket} / {knowledge base id} / {file id} / {image file name}`, and output the processed text.
[0135] The images to be processed are mainly bitmaps (png, jpg, webp, gif) and vector images (emf, wmf, svg):
[0136] For bitmap images, use base64 to save them.
[0137] For vector images, the vector images in the document, including flowcharts, icons, and inserted attachments, need to be converted into bitmaps so that Markdown can reference them normally; first use the inkscape tool to convert the vector image to png. If the conversion is unsuccessful, use soffice to convert it to pdf first, and then use inkscape to convert the image area to png.
[0138] 4. Document splitting module, used to build MarkdownSplitter and CharacterSplitter document splitters, which split markdown text and custom text and table text respectively.
[0139] The document splitter is used to split the text into text blocks and output the split text blocks.
[0140] MarkdownSplitter: Applicable to splitting markdown text blocks. It adopts the title splitting method, treating the content within the same title as a text block. It also sets the maximum text block length and the number of text block overlaps to prevent the text block content from being out of focus due to excessive text splitting. Due to the limitation of the large model context length, the text block is generally 1024 characters or 768 characters.
[0141] CharacterSplitter: Suitable for text splitting after custom txt files and table conversion, using `[SEQ]` as the paragraph separator.
[0142] 5. The vector conversion module uses the BAAI / bge-m3 model for vector embedding, converts the text blocks output by the document splitting module into vectors, and stores the text blocks and vectors in the ParadeDB database.
[0143] The ParadeDB database is a vector database developed based on postgres, which integrates functions such as vector search and BM25 search; one database can realize vector search, full-text search, and hybrid search functions.
[0144] 6. The text recall module uses ParadeDB's `<=>` (vector search) and `@@@` (BM25 search) syntax to search for text blocks. After the search, it uses the BAAI / bge-reranker-v2-m3 model to rerank the text and output text blocks that are highly similar to the user's question.
[0145] An embodiment of the present invention further provides a device for implementing RAG document splitting optimization in an operator field, comprising: at least one memory and at least one processor;
[0146] The at least one memory is configured to store a machine-readable program;
[0147] The at least one processor is configured to call the machine-readable program to implement the RAG document splitting optimization method in the operator field described in the above embodiment.
[0148] An embodiment of the present invention further provides a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions implement the RAG document splitting optimization method for the operator domain described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided. The storage medium stores software program code that implements the functions of any of the above embodiments, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.
[0149] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0150] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0151] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0152] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0153] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.
Claims
1. A RAG document splitting optimization method in the operator field, characterized in that: The implementation of this method includes the following steps: S1: Upload documents and use minio and paradeDB to store source files and file information respectively; S2: Build a document loader, automatically select the corresponding loader according to the file type to process the file, convert it into a unified markdown text, parse the images in the document and convert them into base64 format; S3: Build an image processor to extract the Base64 string of the image in the text, process the bitmap and vector images, and convert them into the Markdown image reference format; S4: Build MarkdownSplitter and CharacterSplitter document splitters to split markdown text and custom text and table text respectively; S5: Vector conversion: Use the BAAI / bge-m3 model for vector embedding to convert the text blocks output in step S4 into vectors. Store the text blocks and vectors in the ParadeDB database. S6: Text recall, using ParadeDB's vector search and BM25 search syntax to retrieve text blocks. After retrieval, the BAAI / bge-reranker-v2-m3 model is used to rerank the text and output text blocks that are highly similar to the user's question.
2. The RAG document splitting optimization method in the operator field according to claim 1, characterized in that: In step S1, after uploading the document, a unique ID is generated for the file, and the file is uploaded to minio for file preview and re-parsing; the file name, file type, and file size information are generated and stored in the ParadeDB database.
3. The RAG document splitting optimization method in the operator field according to claim 1, characterized in that: The document loader includes: PptLoader: Use the soffice command line tool to convert ppt files into pptx files, and then use PptxLoader to process them. PptxLoader: Uses the pptx2md tool to convert pptx files into markdown files. After conversion, it generates a markdown file with the same name as the source file and an images folder for storing images. It reads the markdown file, uses regular expressions to extract local image links in the text, reads the corresponding image files in base64 format, replaces the image links in the text, and outputs the converted text string. XlsLoader: Use the soffice command line tool to convert xls files to xlsx files, and then use XlsxLoader to process them. XlsxLoader: decompresses the xlsx file in zip format and extracts images from the table by parsing the rels file. The image extraction methods include embedded image extraction and non-embedded image extraction. Embedded image extraction: If the source file contains xl / _rels / cellimages.xml.rel and xl / cellimages.xml files, it means that the source file contains embedded images. Parse the xl / cellings.xml file and extract the image path and image ID; Extract non-embedded images: If the xl / drawings / drawing1.xml file is included, it means that the source file contains non-embedded images. Parse all rels files under xl / worksheets / sheet and extract the image path and the sheet, row, and column number of the image in the table. First, use pandas to read the source file, read all sheet pages, traverse all sheet pages, and replace the corresponding cell contents in the table based on the parsed embedded image and non-embedded image information; convert each row in the table into a markdown table, use [SEQ] as the delimiter for each markdown table, and output text strings; CsvLoader: Converts each row of a CSV table into a Markdown table, uses [SEQ] as the delimiter for each Markdown table, and outputs a text string; DocLoader: Use the soffice command line tool to convert doc files to docx files, and then use DocxLoader to process the converted files; DocxLoader: Uses Mammoth to read docx files into HTML text, and then uses html2text to convert it into Markdown text. Images are automatically converted into base64 format and output as Markdown text. PdfLoader: Use the pdf2docx tool to convert the pdf file into a docx file, and then use DocxLoader to process it; ImageLoader: Uses OCR to extract text from images, uses the Qwen2-VL model to generate an image description, integrates the text, image description, and image base64 string into a text file, and outputs the text. HtmlLoader: Use the html2text tool to read HTML content into Markdown format, identify image URLs in the text, download the images and convert the image URLs into image base64 strings, and output the modified text; XmlLoader: Use lxml to parse XML files, use regular expressions to match image URLs in the text, download the images and convert the image URLs into image base64 strings, and then convert them into markdown text strings; TxtLoader: TXT file types include traditional TXT files and custom files separated by [SEQ]. It uses regular expressions to match image URLs in text, downloads images, converts image URLs into image base64 strings, and outputs text strings. JsonLoder: json processing types include list format and dictionary format; List format processing: convert each object in the list into a key:value string, and separate the objects with [SEQ]; Dictionary format: converted to key:value format text; Use regular expressions to match image URLs in the text, download the image, convert the image URL into a base64 string, and then convert it into a markdown text string; JsonlLoader: reads jsonl files, converts the file contents into json files, and then processes them using JsonLoader; MarkdownLoader: Reads a markdown file as text, uses regular expressions to match image URLs in the text, downloads the image and converts the image URL into a base64 string, and outputs the text string.
4. The RAG document splitting optimization method in the operator field according to claim 1, characterized in that: In step S2, the converted format is: [](data:image / {image type};base64,{image base64 string}).
5. The RAG document splitting optimization method in the operator field according to claim 1 or 4, characterized in that: Step S3, For the text string output in step S2, use the regular expression (data:image\ / [\w+-]+;base64,[A-Za-z0-9+ / =]+) to extract the base64 string, generate a unique ID for the image, upload the image to minio for storage, and replace the base64 string with {image bucket} / {knowledge base id} / {file id} / {image file name}, and output the processed text; The images to be processed include bitmaps and vector images: For bitmap images, use base64 to save; For vector images, convert the vector images in the document, including flowcharts, icons, and inserted attachments, into bitmaps: first use the inkscape tool to convert the vector image to png. If the conversion is unsuccessful, use soffice to convert it to pdf first, and then use inkscape to convert the image area to png.
6. The RAG document splitting optimization method in the operator field according to claim 1, characterized in that: The document splitter is used to split the text into text blocks and output the split text blocks; MarkdownSplitter: Applicable to splitting markdown text blocks. It adopts the title splitting method. The content of the same title is regarded as a text block. At the same time, the maximum text block length and the number of text block overlaps are set to prevent the text block content from being out of focus due to excessive text splitting. CharacterSplitter: Suitable for text splitting after custom txt files and table conversion, using [SEQ] as the paragraph separator.
7. The RAG document splitting optimization method in the operator field according to claim 1, characterized in that: The ParadeDB database is a vector database developed based on postgres, which integrates vector search and BM25 search functions; one database can realize vector search, full-text search, and hybrid search functions.
8. A RAG document splitting and optimization system in the operator field, characterized by: include: The document upload module is used to store source files and file information through minio and paradeDB respectively; The document conversion module is used to build a document loader, automatically select the corresponding loader according to the file type to process the file, convert it into a unified markdown text, parse the images in the document and convert them into base64 format; The document processing module is used to build an image processor, extract the base64 string of the image in the text, process the bitmap and vector images, and convert them into the markdown image reference format; Document splitting module, used to build MarkdownSplitter and CharacterSplitter document splitters, which split markdown text and custom text and table text respectively; The vector conversion module uses the BAAI / bge-m3 model for vector embedding, converts the text blocks output by the document splitting module into vectors, and stores the text blocks and vectors in the ParadeDB database; The text recall module uses ParadeDB's vector search and BM25 search syntax to retrieve text blocks. After retrieval, it uses the BAAI / bge-reranker-v2-m3 model to rerank the text and output text blocks that are highly similar to the user's question. The system specifically implements RAG document splitting in the operator field through the method described in any one of claims 1 to 7.
9. A RAG document splitting optimization implementation device in the operator field, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to implement the method according to any one of claims 1 to 7.
10. A computer-readable medium, characterized in that The computer-readable medium stores computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.
Citation Information
Cited By
RAG-oriented document analysis method and system and computer equipment
CN120849350A