Data Processing and Storage Method and Device Based on Graph Database and Vector Database
By combining the graph database Neo4j and vector database Elasticsearch, combined with technologies such as LayoutLMv3 and Transformer models, the accuracy and reliability problems of unstructured data processing and storage in the existing technology are solved, and efficient and accurate data preprocessing and retrieval are achieved.
Patent Information
- Application Number
- CN202411421476.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-10-12
AI Technical Summary
The prior art has problems such as low data structure, inaccurate retrieval or noise when processing and storing unstructured and semi-structured data (such as documents, pictures, tables, etc.), which affects the final output quality.
The data processing and storage method combined with the graph database Neo4j and the vector database Elasticsearch is adopted, and the LayoutLMv3 document layout analysis model, the table conversion model based on Transformer and OCR technology are used to perform document layout analysis and data conversion processing to ensure high-quality data input.
The layout analysis and data conversion processing of document pictures are realized, the structure and accessibility of the data are improved, the high-quality data input required by RAG technology is ensured, and the accuracy and reliability of the final output are improved.
Smart Images

Figure CN118964514B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data processing, and particularly to a data processing and storage method and device based on a graph database and a vector database. Background Art
[0002] In recent years, with the rapid development of LLM (Large Language Model), applications based on large language models have emerged in an endless stream. However, LLM has problems such as "talking nonsense" and hallucinations. Therefore, researchers have proposed a language model that combines retrieval and generation, namely the RAG model. The RAG model first needs to build a knowledge base, then retrieve relevant information from the knowledge base according to the user's question, and combine the generation function of the LLM to output an answer, thereby improving the accuracy of LLM Q&A. However, the effect of RAG technology depends to a large extent on the quality of the structured data in the knowledge base. If the degree of data structuring is low, the retrieved information may be inaccurate or contain noise, which may lead to the generation module generating results based on unreliable contexts, thus affecting the final output quality.
[0003] To solve this problem, the present invention proposes a data processing and storage method that combines the graph database Neo4j and the vector database Elasticsearch, aiming to achieve layout analysis and data conversion processing of document pictures. In this method, Neo4j is good at storing and processing data containing complex relationships, and intuitively represents data and its mutual relationships in the form of nodes and edges. As a vector database, Elasticsearch provides an efficient way to store and retrieve vector-based feature data for processing search tasks of unstructured data, such as similarity search of text and images.
[0004] To better analyze the layout of documents, the present invention adopts the LayoutLMv3 document layout analysis model, which can identify and analyze different elements in documents, such as text, tables, image areas, etc., making document processing more intelligent and efficient. Combining a table conversion model based on the Transformer model and OCR technology, these tools can automatically convert unstructured data (such as images, tables, text, etc.) in documents into text data that can be structured for processing, thereby greatly improving the accessibility and utilization of document information.
[0005] In summary, the present invention proposes a data processing and storage method based on the graph database Neo4j and the vector database Elasticsearch. Combining LayoutLMv3, a Transformer-based table conversion model, and OCR technology, it realizes the layout analysis and data conversion processing of document images. In this method, the graph database, with its flexible data model and powerful semantic expression ability, can effectively store and query complex relational networks; the vector database, through vectorization technology, converts data such as text and images into numerical vectors to support efficient similarity search and data clustering, thus ensuring high-quality data input required by the RAG technology. Summary of the Invention
[0006] Aiming at the deficiencies of the prior art, namely the limitations in the effective storage of unstructured and semi-structured data (such as documents, pictures, tables, etc.), the present invention proposes a data processing and storage method and device based on a graph database and a vector database, providing users with a more efficient, accurate, and reliable data preprocessing method.
[0007] The method proposed by the present invention is realized through the following technical solutions: A data processing and storage method based on a graph database and a vector database, the method comprising the following steps:
[0008] Step 1: Identify the document layout and perform table format conversion, convert the recognized content into Markdown format and store it;
[0009] Step 2: Extract the theme, reference documents, and appendix key information of the document in the Markdown format file based on the large language model;
[0010] Step 3: Convert the Markdown format file into unstructured data, divide the data into blocks and store them in the vector database, and based on the key information of the document, construct a set of relationship types between documents and generate a visual knowledge graph.
[0011] Further, in Step 1, by constructing a layout detection model to analyze the document layout, the construction steps are as follows:
[0012] 1) Obtain the data sets for model training, including the Publaynet data set and the self-built data set manually collected and annotated, and divide them into training sets and test sets;
[0013] 2) Design the layout detection model architecture, and based on the LayoutLMv3 architecture, use the document image to identify the structure of the document, realizing the function of dividing the document image into three types of regions: text, image, and table;
[0014] 3) Train the layout detection model, set the learning rate, measure the difference between the model prediction and the actual label through the cross-entropy loss function, and update the model parameters through the Adam optimizer.
[0015] Further, in step 1, by constructing a table conversion model, convert the table into json format. The construction steps are as follows:
[0016] 1) Obtain the datasets for model training, including the FeTaQA, TAT-QA datasets and the self-built Chinese table dataset collected and annotated manually, and divide them into training sets and test sets;
[0017] 2) Design the architecture of the table detection model, add a trainable prompt learning module based on the Transformer architecture to assist the model's generation task;
[0018] 3) Train the table conversion model, set the learning rate, measure the difference between the model prediction and the actual label through the cross-entropy loss function, and use the AdamW optimizer to improve the training stability.
[0019] Further, after converting the document into a document image, use the layout detection model to analyze the content layout of the document image, divide the document image into three types of regions: text, image, and table, and represent them with different colors; the layout detection model outputs the four boundary values of all region boxes, indicating the position of the region in the image, and each detected region will be assigned a class label; extract the content of the text, image, and table regions, and save the extracted data in Markdown format.
[0020] Further, in step 3, construct a data storage module, structurally divide the converted Markdown-formatted document, and save it in the graph database neo4j and the vector database elasticsearch. The construction steps are as follows:
[0021] 1) Structurally divide the Markdown document; for text data, use the title as the basis for paragraph division, and for the text within the paragraph, use the sliding window technique to divide the data blocks; for table and image data, take each entity as a unit, and regard each table or image as an independent data block;
[0022] 2) Define the entity type set and the relationship type set, where the entity type set includes the document name and document type; the relationship type set generates the connections between documents based on the theme, reference documents, and appendix key information of the document; import the obtained entity-relationship binary tuples into the neo4j database to form a visual document relationship knowledge graph;
[0023] 3) Create an index of an associated type in Elasticsearch. This index stores the specific data of the documents. The structure of the index consists of attribute information, and the attributes include the unique ID of the index, the unique file_id of the document, the theme, content, embedding, type, and level of the document;
[0024] 4) According to the divided data blocks, use the BGE-embedding vectorization model to expand all data blocks and save them in the Elasticsearch vector database.
[0025] Furthermore, the layout detection model uses a text-image multimodal Transformer to learn cross-modal features, and captures text information and image information through three modules: Masked Language Modeling (MLM), Masked Image Modeling (MIM), and Word-Patch Alignment (WPA); among them, Masked Language Modeling randomly masks a part of the text word vectors, but retains the corresponding two-dimensional position information, and the task objective is to restore the masked words in the text according to the unmasked text, image, and layout information; Masked Image Modeling randomly masks a part of the image patches, and the task objective is to restore the discretized ID of the masked image patches according to the unmasked text and image information; Word-Patch Alignment learns the fine-grained alignment relationship between the language and visual modalities by explicitly predicting whether the corresponding image patch of a text word is masked.
[0026] Furthermore, the table conversion model stacks a 12-layer Transformer. Specifically, a special symbol "instruction:" is added before the input table content to prompt text generation, and trainable vectors are added to the multi-head attention module of each layer to prompt the pre-training task; during the training process, first flatten the table into a sequence for direct input into the model; special markers are inserted to represent the boundaries of the table.
[0027] Furthermore, construct a graph database with "theme" as the core to form an interconnected information network. In this network, the "theme" nodes occupy the central position and are connected to multiple "file" nodes. Each "file" node is a rich information collection, containing "reference file", "es index", and "attachment" sub-nodes; all nodes are connected to other nodes through the relationship types of reference files and appendices to form a complex semantic network.
[0028] Further, the data blocks are represented in a vectorized manner. For image data, after being converted into a vector representation, they are stored separately in Elasticsearch. Specifically, the BGE-embedding vectorization model is used to vectorize the divided data blocks. In the Elasticsearch database, files are used as indexes. Each file contains multiple data blocks, and each data block has its own data type. In particular, the original version of the file data block, that is, the version without vectorization, is also saved in the Elasticsearch database to achieve hybrid retrieval.
[0029] On the other hand, the present invention also provides a data processing and storage device based on a graph database and a vector database, including a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, the data processing and storage method based on the graph database and the vector database as described above is implemented.
[0030] The beneficial effects of the present invention are as follows:
[0031] 1. By using a variety of professional parsers to parse different data types in the document, the invention can comprehensively identify and accurately extract multi-modal data such as text, tables, and images.
[0032] 2. Construct a multi-modal knowledge graph, effectively convert unstructured data into a structured form, store it in the graph database, and at the same time retain the hierarchical structure and semantic information of the data.
[0033] 3. Save the original data and vectorized data in Elasticsearch at the same time, support hybrid retrieval of the original data and vectorized data, and enhance the consistency and flexibility of retrieval.
[0034] 4. Combining the advantages of the graph database and the vector database, the invention provides an efficient information retrieval mechanism, which can achieve fast and accurate semantic search. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic flowchart of a data processing and storage method based on a graph database and a vector database provided by an embodiment of the present invention.
[0036] Figure 2 It is an example diagram of custom data and labels for layout detection provided by an embodiment of the present invention.
[0037] Figure 3 It is an example diagram of a knowledge graph provided by an embodiment of the present invention.
[0038] Figure 4 It is a structural diagram of text division provided by an embodiment of the present invention.
[0039] Figure 5 Schematic diagram of the table conversion model provided by the embodiment of the present invention.
[0040] Figure 6 Structural diagram of a data processing and storage device based on a graph database and a vector database provided by the embodiment of the present invention. Detailed implementation manners
[0041] To more clearly elaborate the purpose, technical solutions and their advantages of this specification, we will describe it in detail through specific embodiments and related drawings. It should be clear that the described embodiments only represent some cases and do not cover all possible implementation manners. Based on these embodiments, other embodiments that can be deduced by those skilled in the art without creative work all fall within the protection scope of this specification.
[0042] In summary, this specification aims to comprehensively display the details of the technical solutions through specific embodiments and drawings, while reserving the protection for other possible embodiments. All operations are within the scope permitted by law and respect the rights and interests of data owners.
[0043] The present invention solves the data processing and storage problems in the prior art by constructing a multi-modal unified graph database and vector database. On the one hand, by constructing a knowledge graph, the structured information inside the documents is stored, reflecting the correlation between the documents, and rich semantic information can be provided; on the other hand, by constructing a vector database, efficient retrieval of document information is realized, converting data such as text and images into numerical vectors, greatly improving the speed and accuracy of retrieval. Combining these two database technologies, the present invention can comprehensively consider other documents related to the query document when retrieving knowledge, realizing the expansion of the retrieval range and the enhancement of the retrieval depth.
[0044] As Figure 1 shown, the present invention is a data processing and storage method based on a graph database and a vector database, and the method includes the following steps:
[0045] (1) Train a document layout detection model, and analyze the layout of the document based on the layout detection model LayoutLMv3.
[0046] First, obtain the dataset for model training. Based on the Publaynet dataset and the self-built dataset collected and annotated manually, divide them into a training set and a test set. The Publaynet contains a large number of PDF documents and document page images, but lacks other types of documents, such as word, html, etc., and cannot handle the content missing problem in document conversion. Moreover, the lack of Chinese data results in a low Chinese expression in this dataset. Therefore, other types of Chinese documents are collected, converted into images, and then combined with the Publaynet dataset as the training dataset. As Figure 2 shown, after collecting Chinese samples, add a label with a blue background to represent text, including the title, abstract, and keywords of the document. Add a label with a purple background to represent images, and each image has an image frame. Add a label with a yellow background to represent each table.
[0047] Then, design the layout detection model architecture. The model is based on the LayoutLMv3 architecture and uses document images to identify the structure of the document, realizing the function of dividing the document image into three types of regions: text, image, and table. This model is based on the LayoutLMv3 model and uses a text-image multimodal transformer to learn cross-modal features. The capture of text information and image information is achieved through three modules: Masked Language Modeling (MLM), Masked Image Modeling (MIM), and Word-Patch Alignment (WPA).
[0048] Specifically, Masked Language Modeling embeds the text, including word embedding and position embedding (layout). Use an off-the-shelf OCR toolkit to preprocess the document image to obtain the text content and the corresponding 2D position information. Initialize the word embedding using the embedding model. The position embedding includes 1D position and 2D layout position embedding. The 1D position refers to the index of the token in the text sequence, and the 2D layout position refers to the bounding box coordinates of the text sequence. Since adjacent words in the text usually express similar semantics, a paragraph of text shares a two-dimensional position vector.
[0049] Masked Image Modeling uses image embedding to divide all document images into a series of uniformly sized P×P patch blocks, obtains the image feature sequence through linear mapping, and then gets the image vector after adding a learnable one-dimensional position vector. Further, the image embedding resizes the document image to H×W and uses I∈R C×H×WDenote an image, where C, H, and W are the channel size, width, and height of the image respectively. Then, the image is segmented into a series of uniform P×P patches, linearly projected into D dimensions, and flattened into a sequence of vectors with length M = HW / P 2 。
[0050] Among them, masked language modeling randomly masks 30% of the text word vectors, but retains the corresponding two-dimensional position (layout) information. The task objective is to restore the masked words in the text based on the unmasked text and layout information; masked image modeling randomly masks approximately 40% of the image patches. The task objective is to restore the discretized ID of the masked image patches based on the unmasked text and image information; word-patch alignment learns the fine-grained alignment relationship between the language and visual modalities by explicitly predicting whether the corresponding image patch of a text word is masked.
[0051] Given an input document image and its corresponding text and two-dimensional position information (layout) information obtained by OCR, the model takes the linear projection of the patches and word tokens as inputs and encodes them into context-aware vector representations. The model is trained through the discrete token reconstruction objectives of the masked language model (MLM) and masked image model (MIM). In addition, LayoutLMv3 is pre-trained through a Word-Patch Alignment (WPA) objective to learn cross-modal alignment by predicting whether the image patch corresponding to a text word is masked.
[0052] During the process of training the layout detection model, the learning rate is set to 5e-5. At the same time, the cross-entropy loss function is selected to measure the difference between the model prediction and the actual label, and the model parameters are updated through the Adam optimizer.
[0053] (2) Train the table conversion model. To handle the characteristics of large volume and complex structure of table data, it is necessary to specially process the table data in the document. The table conversion model is used to convert the table data in the document into a more readable json format.
[0054] First, obtain the dataset for model training. Based on datasets such as FeTaQA, TAT-QA, and the self-built Chinese table dataset manually collected and annotated, divide the training set and test set.
[0055] Design the architecture of the table detection model, such as Figure 5As shown, the model is based on the standard Transformer architecture and adds a trainable prompt learning module to assist the model's generation task. Specifically, a 12-layer Transformer is stacked. In particular, a special symbol "instruction:" is added before the input table content to prompt text generation, and a trainable vector is added to the multi-head attention module of each layer to pre-train the prompt of the task, and the prompt length is set to 100. During the training process, the table is first flattened into a sequence so that it can be directly input into the model. By inserting several special markers to represent the boundaries of the table, a flattened table can be represented as:
[0056] TABLE = [HEAD], col1, col2,..coln, [ROW], 1, row1,row2,..rowm,[ROW],..
[0057] Here, [HEAD] and [ROW] are special markers, representing the header and row regions respectively. The number after [ROW] is used to represent the row index. Before inputting the table data into the model, a special symbol "instruction:" is added to assist text generation. The following described Table 1 will be converted to:
[0058] {[HEAD], ”Serial Number”, ”Time Period”, ”Content”, [ROW], ”1”, ”2014.1 - 2014.3”, ”Preliminary Research and Technical Preparation Stage: We will conduct an in-depth analysis of the characteristics of point cloud data and investigate existing convolutional neural network technologies to establish the theoretical basis and technical roadmap for the project.”, [ROW], ”2”, ”2014.10 - 2014.12”, ”Technical Optimization and Application Promotion Stage: In the final stage, we will further optimize the algorithm to solve the problems found in the experiment.”}
[0059] Table 1
[0060]
[0061] The self-built Chinese table-form data is as follows:
[0062] {
[0063] "instruction": "Please read the following table in the input and convert the table into a json representation according to the table content. In the table, [HEAD] and [ROW] are special markers, representing the header and row regions respectively. The number after [ROW] is used to represent the row index.",
[0064] "input_serialize":"[\"[HEAD]\",\"Serial Number\",\"Time Period\",\"Content\",\"[ROW]\",\"1\",\"2014.1 - 2014.3\",\"Preliminary Research and Technical Preparation Phase: We will conduct an in - depth analysis of the characteristics of point cloud data and investigate existing convolutional neural network technologies to establish the theoretical basis and technical roadmap for the project.\",\"[ROW]\",\"2\",\"2014.10 - 2014.12\",\"Technical Optimization and Application Promotion Phase: In the final stage, we will further optimize the algorithm to solve the problems found in the experiments.\"]",
[0065] "input":"{\"research_phases\": [{\"index\": 1,\"time_period\": \"2014.1 - 2014.3\",\"content\": \"Preliminary Research and Technical Preparation Phase: We will conduct an in - depth analysis of the characteristics of point cloud data and investigate existing convolutional neural network technologies to establish the theoretical basis and technical roadmap for the project.\"},{\"index\": 2,\"time_period\": \"2014.10 - 2014.12\",\"content\": \"Technical Optimization and Application Promotion Phase: In the final stage, we will further optimize the algorithm to solve the problems found in the experiments.\"}]}",
[0066] "output":"In the period from January 2014 to March 2014, we will conduct an in - depth analysis of the characteristics of point cloud data and investigate existing convolutional neural network technologies to establish the theoretical basis and technical roadmap for the project. From October 2014 to December 2014, this is the final stage of the project. We will further optimize the algorithm to solve the problems found during the experiments.",
[0067] }
[0068] When training the table conversion model, the learning rate is set to 3×5e - 5 to ensure that the model can fully learn during training and generalize to unseen data. The cross - entropy loss function is selected to measure the difference between the model prediction and the actual label, and the , , AdamW optimizer is used to improve the training stability.
[0069] (3)Construct a layout detection module and a content analysis module based on the layout detection model and the table conversion model obtained in steps (1) and (2), perform layout detection, OCR text recognition, and table conversion on the source document, and convert the recognized content into Markdown format and store it. The specific implementation process is as follows:
[0070] (3.1)Identify the type of the document and use corresponding tools such as pdfkit, python-docx, pptx, etc. to convert the document into a document image.
[0071] (3.2)Use the layout detection module to analyze the content layout of the document image and divide the document image into three types of regions: text, image, and table. Specifically, use a blue background to represent text lines, including the title, abstract, and keywords of the document, use a purple background to represent images, and each image has an image frame, and use a yellow background to represent each table. Finally, the layout detection module outputs the four boundary values of all region boxes, indicating the position of the region in the image, and each detected region will be assigned a category label, such as "TEXT", "IMAGE", "TABLE", etc.
[0072] (3.3)Use the content analysis module to extract the content of the three types of regions, namely text, image, and table, and save the extracted data in Markdown format.
[0073] For the extraction of text regions, use an OCR (Optical Character Recognition) tool to recognize and extract the text regions in the document. OCR technology analyzes the characters in the image and converts them into an editable text format. During this process, OCR can accurately recognize the paragraph structure of the document, and the extracted text will be saved according to the paragraphs of the original document to ensure the integrity of the document structure and the coherence of the content. For the formulas in the document, choose to save them in latex format. latex is a standard format for typesetting complex mathematical formulas and can accurately represent the structure and symbols of the formulas.
[0074] For the table region, flatten the recognized table and input it into the table conversion model. The model can parse the structure of the table and convert it into a JSON format representation. JSON format is a lightweight data exchange format, which is convenient for machine processing and storage and can retain the row and column structure and data content of the table. The extracted table will be embedded in the Markdown file for subsequent data storage.
[0075] For the image regions recognized in the document, the present invention chooses to directly save a copy of the picture and store it locally. The Markdown file saves the relative path of the image locally, retaining the original information of the image and providing materials for direct calling in subsequent tasks. In particular, extract important data in the document, such as file name, file theme, reference files, attachments, etc.
[0076] (4)Based on the Markdown format file obtained in step (3), input the content into the Qwen72B large language model to extract key information such as the theme, reference documents, and appendices of the document; deploy the LLM locally and design the prompt as:
[0077] prompt = "Please extract the following key information from the following document:
[0078] 1. The title of the document
[0079] 2. List of reference documents or references
[0080] 3. Content or list of the appendix section
[0081] The content of the document is as follows: [Insert document content]
[0082] Please output in the following format: Title: [Extracted title] References: [List of extracted reference documents] Appendix: [Extracted appendix content]"
[0083] Record the extracted key information and save it at the beginning of the Markdown file in (4) for constructing the knowledge graph.
[0084] (5)Construct a data storage module. Based on the Markdown format file obtained in step (3), transform the unstructured data in the file, split it into structured data chunks of different sizes, and store them in the vector database; based on the key information of the document obtained in step (4), construct a set of relationship types between documents and generate a visualized knowledge graph. The specific implementation process is as follows:
[0085] (5.1)Conduct a structured division of the Markdown document. When splitting the document, targeted processing strategies are adopted for different types of data. For text data, as Figure 4 shown, using the title as the basis for paragraph division, abstract all paragraphs of the entire document into a title text tree. For the text within paragraphs, the sliding window technique is used to divide the data chunks. Assuming each data chunk contains 100 words, then each slide may slide 50 words, that is, expand 50 words before and after, so that there is an overlap between adjacent chunks, ensuring the coherence and integrity of the text information. For table and image data, taking each entity as a unit, each table or image is regarded as an independent data chunk.
[0086] (5.2) Define the entity type set and relationship type set based on the document data extracted in steps (3.1) and (4). The entity type set is the document name and document type. The relationship type set relies on the key information such as the document theme, reference documents, and appendices extracted by the large language model in step (4) to generate the connections between documents. Import the obtained entity-relationship binary tuples into the neo4j database to form a visualized document relationship knowledge graph.
[0087] Construct a graph database with "theme" as the core to form a highly interconnected information network. In this network, the "theme" node occupies the central position and is connected to multiple "file" nodes. Each "file" node is a rich information collection, containing sub-nodes such as "reference document", "es index", and "attachment". Among them, "reference document" and "appendix" play the role of bridges between files, building a complex knowledge network, and the "es index" stores the index information about this file in elasticsearch. All nodes are connected to other nodes through various relationship types such as reference documents and appendices, forming a complex semantic network. Further, store the key information such as the reference documents and attachments of the documents obtained in step (4) to build a connection network with other relevant documents. As Figure 3 shown in the network, there are multiple themes in this network, and each theme contains multiple files. The files may have multiple attributes. These files are independent of each other but are related to other files through reference documents and appendices. Figure 3 The arrows from the theme to the file and from the file to the attribute in the figure represent the inclusion relationship, and the arrow from the attribute to the file represents the association relationship. In particular, it extends to the elasticsearch database through the es index to build a hybrid retrieval of the graph database and the vector database.
[0088] (5.3) Create an index of an association type in elasticsearch. This index stores the specific data of the document. The structure of the index is mainly composed of attribute information. The attributes include the unique id of the index, the unique file_id of the document, the theme, content, embedding, type, and level of the document. When creating the index, set the number of shards of the index to 3 and the number of replicas to 0.
[0089] The index design is as follows:
[0090] {
[0091] 'file_id': {
[0092] 'type': 'keyword'
[0093] },
[0094] 'trunk_id': {
[0095] 'type': 'keyword'
[0096] },
[0097] 'file_theme':{
[0098] 'type':'text',
[0099] 'fields':{
[0100] 'keyword': {
[0101] 'type': 'keyword'
[0102] }
[0103] },
[0104] },
[0105] 'content': {
[0106] 'type': 'text',
[0107] 'similarity': 'BM25'
[0108] },
[0109] 'embedding': {
[0110] 'type': 'dense_vector',
[0111] 'dims': EMBEDDING_DIMS,
[0112] 'index': True,
[0113] 'similarity': 'cosine'
[0114] },
[0115] 'type': {
[0116] 'type': 'text'
[0117] },
[0118] 'level': {
[0119] 'type': 'text'
[0120] }
[0121] (5.4) Based on the divided data blocks obtained in step (5.1), use the BGE-embedding vectorization model to treat all data blocks as a single token, expand it to 1024 dimensions, and then save it in the elasticsearch vector database. In the elasticsearch database, files are used as indexes. Each file contains multiple data blocks, and each data block has its own data type. Each data block contains file_id representing the unique id of the file, trunk_id representing the unique id of the data block, file_theme representing the theme of the file, content representing the initial expression of the content of the data block, embedding representing the vectorized expression of the content of the data block, type representing the initial category of the data block, such as table, image, text, etc., and level representing the level of the heading of the data block. In particular, the original version of the file data block, that is, the non-vectorized version, is also saved in the elasticsearch database to achieve hybrid retrieval. Traverse the divided data blocks obtained in step (5.1), extract the attribute values required for the elasticsearch database index, and write them into the index in sequence.
[0122] Corresponding to the foregoing embodiment of a data processing and storage method based on a graph database and a vector database, the present invention also provides an embodiment of a data processing and storage device based on a graph database and a vector database.
[0123] See Figure 6 , an embodiment of a data processing and storage device based on a graph database and a vector database provided by an embodiment of the present invention includes a memory and one or more processors. Executable code is stored in the memory. When the processor executes the executable code, it is used to implement a data processing and storage method based on a graph database and a vector database in the foregoing embodiment.
[0124] An embodiment of a data processing and storage device based on a graph database and a vector database provided by the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From a hardware perspective, as Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where a data processing and storage device based on a graph database and a vector database provided by the present invention is located. Except for Figure 6In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located generally may further include other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated herein.
[0125] The implementation processes of the functions and roles of each unit in the above device are specifically described in detail in the implementation processes of the corresponding steps in the above method, which will not be elaborated herein.
[0126] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0127] The embodiment of the present invention further provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements a data processing and storage method based on a graph database and a vector database in the above embodiment.
[0128] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by the any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.
[0129] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the data processing and storage method based on a graph database and a vector database described above.
[0130] The above embodiments are used to explain the present invention rather than limit the present invention. Any modification and change made to the present invention within the spirit and scope of the protection of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A data processing and storage method based on a graph database and a vector database, characterized in that: The method comprises the following steps: Step 1: Identify the document layout based on the LayoutLMv3 model and convert the table format based on the Transformer model, convert the identified content into Markdown format and store it; analyze the document layout by building a layout detection model. The construction steps are as follows: 1) Obtain the dataset for model training, including the Publaynet dataset and the self-built dataset with manual collection and annotation, and divide it into training set and test set; 2) Design a layout detection model architecture. Based on the LayoutLMv3 architecture, use document images to identify the structure of documents. After converting documents into document images, use the layout detection model to analyze the content layout of the document images, divide the document images into three types of areas: text, image, and table, and represent them with different colors. The layout detection model outputs four boundary values of all area boxes, indicating the location of the area in the image. Each detected area will be assigned a category label. Extract the contents of text, image and table areas, and save the extracted data in Markdown format; the layout detection model uses a text-image multimodal transformer to learn cross-modal features, and uses three modules: masked language modeling MLM, masked image modeling MIM, and word block alignment WPA to capture text information and image information; the masked language modeling randomly masks a part of the text word vector, but retains the corresponding two-dimensional position information, and the task goal is to restore the masked words in the text based on the unmasked image and layout information; the masked image modeling randomly masks a part of the image block, and the task goal is to restore the discretized ID of the masked image block based on the unmasked text and image information; the word block alignment learns the fine-grained alignment relationship between language and visual modalities by explicitly predicting whether the corresponding image block of a text word is masked; 3) Train the layout detection model, set the learning rate, measure the difference between the model prediction and the actual label through the cross entropy loss function, and update the model parameters through the Adam optimizer; Step 2: Extract the subject, reference files, and appendix key information of the document in the Markdown format file based on the large language model; Step 3: Convert the Markdown format file into structured data, divide the data blocks into blocks and store them in the vector database. Based on the key information of the document, build a set of relationship types between documents and generate a visual knowledge graph. Specifically, divide the converted Markdown format document into structured blocks and save them in the graph database neo4j and the vector database elasticsearch. The specific steps are as follows: 1) Structural division of Markdown documents; for text data, the title is used as the basis for paragraph division, and the text within the paragraph is divided into data blocks using the sliding window technology; for table and image data, each entity is taken as a unit, and each table or image is regarded as an independent data block; 2) Define entity type sets and relationship type sets, where the entity type set is the document name and document type; the relationship type set generates the connection between documents based on the subject, reference file, and appendix key information of the document; import the obtained entity relationship tuples into the neo4j database, build a graph database with "subject" as the core, and form an interconnected information network. In this network, the "subject" node occupies a central position and is connected to multiple "file" nodes. Each "file" node is a rich information set, including "reference file", "es index", and "attachment" sub-nodes; all nodes are connected to other nodes through the relationship types of reference files and appendices to form a complex semantic network and obtain a visual document relationship knowledge graph; 3) Create an associated index in elasticsearch, which stores the specific data of the document. The index structure consists of attribute information, including the unique id of the index, the unique file_id of the document, the theme, content, embedding, type and level of the document; 4) Based on the divided data blocks, all data blocks are expanded using the BGE-embedding vectorization model and saved in the elasticsearch vector database. Specifically, the data blocks are represented by vectors. For image data, after being converted into vector representations, they are stored separately in elasticsearch. The BGE-embedding vectorization model is used to vectorize the divided data blocks. Files are used as indexes in the elasticsearch database. Each file includes multiple data blocks, and each data block has its own data type. At the same time, the original version of the file data block, that is, the unvectorized version, is saved in the elasticsearch database to achieve hybrid retrieval.
2. A data processing and storage method based on a graph database and a vector database according to claim 1, characterized in that: In step 1, the table is converted into JSON format by building a table conversion model. The construction steps are as follows: 1) Obtain data sets for model training, including FeTaQA, TAT-QA data sets, and manually collected and annotated self-built Chinese table data sets, and divide them into training sets and test sets; 2) Design the architecture of the table detection model and add a trainable prompt learning module based on the Transformer architecture to assist the model in generating tasks; 3) Train the table conversion model, set the learning rate, use the cross-entropy loss function to measure the difference between the model prediction and the actual label, and use the AdamW optimizer to improve training stability.
3. A data processing and storage method based on a graph database and a vector database according to claim 2, characterized in that: The table conversion model uses a 12-layer Transformer stack, adds a special symbol "instruction:" before the input table content to prompt text generation, and adds a trainable vector in the multi-head attention module of each layer to pre-train the task prompt; During training, the table is first flattened into a sequence so that it can be directly input into the model; special markers are inserted to indicate the boundaries of the table.
4. A data processing and storage device based on a graph database and a vector database, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, it implements a data processing and storage method based on a graph database and a vector database as described in any one of claims 1 to 3.