Text retrieval method combining MBSE with RAG
By using ETL tools and word embedding technology in MBSE, combined with the weighted TF-IDF algorithm, the problem of incomplete information in the text retrieval method in the existing technology is solved, and a more accurate and comprehensive large-language model retrieval effect is achieved.
Patent Information
- Application Number
- CN202411561614.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-04
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the text search method only considers the terms to be retrieved and does not consider including terms with high correlation, resulting in incomplete information read by the model.
The text retrieval method of MBSE combined with RAG is adopted to convert the document content into a structured form through ETL tools, extract high-frequency terms and embed word, and calculate document scores using the weighted TF-IDF algorithm, considering the context information of the terms to improve retrieval accuracy.
By considering the contextual information of the term, the accuracy and rationality of text retrieval is improved, ensuring that the information read by the model is more comprehensive, and suitable for the retrieval and enhanced generation of large language models.
Smart Images

Figure CN120067281A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular, to a text retrieval method combining MBSE and RAG. Background Art
[0002] At the current stage, large language models are prone to drawbacks such as hallucinations and lack of timeliness, especially when faced with the massive data of MBSE. This can be effectively alleviated by the method of Retrieval-augmented Generation. When the model needs to generate text or answer questions, it will retrieve relevant information from a huge document collection and then use this information to guide the generation of text, thereby improving the quality and accuracy of the answer. Among them, retrieval is an important part, and retrieval methods include TF-IDF, BM25, Elasticsearch, etc.
[0003] However, the document ranking retrieved by the retrieval methods in the existing methods only considers the terms to be retrieved themselves and does not consider the terms with high relevance to them, resulting in incomplete information read by the model. Summary of the Invention
[0004] The purpose of the present invention is to provide a text retrieval method combining MBSE and RAG, aiming to solve the technical problem that the document ranking retrieved by the retrieval method in the existing technology only considers the terms to be retrieved themselves and does not consider the terms with high relevance to them, resulting in incomplete information read by the model.
[0005] To achieve the above purpose, a text retrieval method combining MBSE and RAG adopted by the present invention includes the following steps:
[0006] Step 1: Use an ETL tool to convert the content of the massive document data in MBSE into a processable structured form, provide a suitable input for knowledge retrieval, and preprocess the data;
[0007] Step 2: Extract the high-frequency terms in each section to create a vocabulary, and perform word embedding on the terms using a neural network, map the terms to the semantic space, and obtain the word vectors with their structured representations;
[0008] Step 3: Perform word embedding on the input terms to be matched, and calculate the k word vectors with the closest relationship to the terms to be matched through cosine similarity;
[0009] Step 4: Use the weighted TF-IDF algorithm for the k + 1 terms. The weight calculation is that the weight of the term to be matched is 1, and the remaining k terms use softmax to calculate the weights for the cosine similarity results. Calculate the sum of the scores of the k + 1 terms in the document to obtain the total score of the document regarding the terms to be matched and sort them;
[0010] Step 5: Select the top n documents from the sorted documents. These n documents are the most relevant documents matching the input terms, and the text retrieval of MBSE combined with RAG is completed.
[0011] Among them, the weighted TF-IDF algorithm is as follows:
[0012]
[0013] TF-IDF(t,D) = TF(t,D)·IDF(t)
[0014] Among them, λ represents the weight, f(t,D) represents the number of times the term t appears in the document D, ∑ t′∈D f(t′,D) is the total number of words in the document, N is the total number of documents, and DF(t) is the number of documents containing the word t.
[0015] Among them, the massive document data in MBSE includes document data in PDF, PPT, and docx formats.
[0016] Among them, the processing method for PDF document data is as follows:
[0017] Use the PDF content extraction tool PDFplumber to extract PDF text and tables;
[0018] For PDF documents that cannot be directly recognized, use OCR to extract the content of each page, and use the PyMuPDF library of the Python tool to extract text blocks;
[0019] Perform layout recognition on the PDF using the object recognition model, and put the extracted content into the corresponding sections for convenient subsequent query and call.
[0020] Among them, the methods for preprocessing the data include removing punctuation marks, converting to lowercase, and stemming.
[0021] Among them, data preprocessing is performed using Pandas.
[0022] Among them, when extracting the high-frequency terms in each section to create a vocabulary, the extraction method is as follows: for documents in docx, doc, ppt, pptx, and xls formats, use Spring AI and LangChain4j based on the Apache Tika framework for extraction;
[0023] Based on the fact that the content of the PDF type is unstructured, use the AI-driven OCR for data extraction.
[0024] A text retrieval method combining MBSE and RAG of the present invention structures a large number of requirement documents, design specifications, and model documents in MBSE into a structured format and creates a vocabulary, and then performs word embedding on the terms in the vocabulary. Subsequently, word embedding is performed on the terms to be matched input by the user, similar terms are found through a distance algorithm, and the document most matching the terms is found through a text retrieval algorithm. The present invention uses a high-dimensional distance algorithm for word vectors and converts them into standardized weights, which can better understand the meaning of terms in retrieval. At the same time, the weighted text retrieval algorithm used in the present invention not only considers the frequency and inverse frequency of terms in the document, but also considers the context information of the terms, and the obtained document results are more semantic, improving the accuracy and rationality of text retrieval, which is very valuable for retrieval-enhanced generation of large language models. In this way, the technical problem in the prior art that the document sorting retrieved by the retrieval method only considers the terms to be retrieved itself and does not consider the terms with high relevance to it, resulting in incomplete information read by the model, is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 is the flowchart of the steps of the text retrieval method combining MBSE and RAG of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will describe in detail the embodiments of the present invention. The examples of the embodiments are shown in the drawings. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as a limitation of the present invention.
[0028] Please refer to Figure 1 , Figure 1 is the flowchart of the steps of the text retrieval method combining MBSE and RAG of the present invention.
[0029] The present invention provides a text retrieval method combining MBSE and RAG, including the following steps,
[0030] Step 1: Use an ETL tool to convert the document content in MBSE into a processable structured form, provide suitable input for knowledge retrieval, and preprocess the data;
[0031] Step 2: Extract the high-frequency terms in each section to create a vocabulary, and perform word embedding on the terms using a neural network to map the terms to the semantic space and obtain the word vectors with their structured representations;
[0032] Step 3: Perform word embedding on the input terms to be matched, and calculate the k word vectors that are closest to the terms to be matched through cosine similarity;
[0033] Step 4: Use the weighted TF-IDF algorithm for the k + 1 terms. The weight calculation is that the weight of the term to be matched is 1, and for the remaining k terms, the softmax is used to calculate the weight based on the cosine similarity result. Calculate the sum of the scores of the k + 1 terms in the document to obtain the total score of the document regarding the term to be matched and sort them;
[0034] Step 5: Select the top n documents from the sorted documents. These n documents are the most relevant documents matching the input terms, and the text retrieval combining MBSE and RAG is completed.
[0035] The weighted TF-IDF algorithm is used as follows:
[0036]
[0037] TF-IDF(t,D) = TF(t,D)·IDF(t)
[0038] where λ represents the weight, f(t,D) represents the number of occurrences of the term in document D, ∑ t′∈D f(t′,D) is the total number of words in the document, N is the total number of documents, and DF(t) is the number of documents containing the term t.
[0039] The massive document data in MBSE includes document data in PDF, PPT, and docx formats.
[0040] The processing method for PDF document data is as follows:
[0041] Use the PDF content extraction tool PDFplumber to extract PDF text and tables;
[0042] For PDF documents that cannot be directly recognized, use OCR to extract the content of each page, and use the PyMuPDF library of Python to extract text blocks;
[0043] Perform layout recognition on the PDF using an object recognition model, and put the extracted content into the corresponding sections for convenient subsequent query and call.
[0044] The methods for preprocessing the data include removing punctuation marks, converting to lowercase, and stemming.
[0045] Data preprocessing is carried out using Pandas.
[0046] When extracting high-frequency terms in each section to create a vocabulary, the extraction method is as follows: for documents in docx, doc, ppt, pptx, and xls formats, extraction is performed using Spring AI and LangChain4j based on the Apache Tika framework.
[0047] Since the content based on PDF type is unstructured, AI-driven OCR is used for data extraction.
[0048] Using a text retrieval method combining MBSE and RAG in this embodiment, a large number of requirement documents, design specifications, and model documents in MBSE are organized into a structured format and a vocabulary is created. Then, word embedding is performed on the terms in the vocabulary. Subsequently, word embedding is performed on the term to be matched input by the user, similar terms are found through a distance algorithm, and the document most matching the term is found through a text retrieval algorithm. The present invention uses a high-dimensional distance algorithm for word vectors and converts them into standardized weights, which can better understand the meaning of terms in retrieval. At the same time, the weighted text retrieval algorithm used in the present invention not only considers the frequency and inverse frequency of terms in the document, but also considers the context information of the terms, and the obtained document results are more in line with semantics, improving the accuracy and rationality of text retrieval, which is very valuable for retrieval-enhanced generation of large language models. In this way, the technical problem in the prior art that the document ranking retrieved by the retrieval method only considers the term to be retrieved itself and does not consider terms with high relevance to it, resulting in incomplete information read by the model, is solved.
[0049] Generating a vocabulary through MBSE documents significantly improves the efficiency of retrieving vocabulary in RAG. The generated vocabulary can be adjusted according to the data and requirements of the actual task, so as to better capture the vocabulary and context of domain-specific requirement documents or design documents. And by customizing the vocabulary, more precise control can be exerted on the generated content to avoid generating irrelevant or inappropriate vocabulary. Using word embedding and cosine similarity techniques to map the term to be matched into a semantically rich semantic vector space can better capture semantic relationships. Using the weighted TF-IDF algorithm to calculate the weighted document scores can better match highly relevant documents.
[0050] The above-disclosed is only a preferred embodiment of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A text retrieval method combining MBSE and RAG, It is characterized in that The following steps are involved: Step 1: Use ETL tools to convert the massive document data in MBSE into a processable structured form, provide appropriate input for knowledge retrieval, and pre-process the data; Step 2: Extract the high-frequency words in each section to create a vocabulary, and use a neural network to embed the words, map the words to the semantic space, and obtain the word vector of its structured representation; Step 3: embed the input word to be matched, and calculate the k word vectors closest to the word to be matched through cosine similarity; Step 4: Use the weighted TF-IDF algorithm for k+1 terms. The weight calculation is that the weight of the term to be matched is 1. The weights of the remaining k terms are calculated using softmax on the cosine similarity results. The scores of the k+1 terms in the document are calculated and added to get the total score of the document for the term to be matched and sorted. Step 5: Select the first n documents from the sorted documents. These n documents are the most relevant documents matching the input term, and complete the text retrieval of MBSE combined with RAG.
2. The MBSE combined with RAG text retrieval method according to claim 1, characterized in that: The weighted TF-IDF algorithm is as follows: TF-IDF(t,D)=TF(t,D)·IDF(t) Where λ represents the weight, f(t,D) represents the number of times the term appears in document D, ∑ t′∈D f(t′,D) is the total number of words in the document, N is the total number of documents, and DF(t) is the number of documents containing word t.
3. The MBSE combined with RAG text retrieval method as claimed in claim 2, characterized in that: The massive document data in MBSE includes document data in PDF, PPT, and docx formats.
4. The MBSE combined with RAG text retrieval method as claimed in claim 3, characterized in that: The processing method for PDF document data is as follows: Use PDF content extraction tool PDFplumber to extract PDF text and tables; For PDF documents that cannot be directly recognized, use OCR to extract the content of each page, and use Python's PyMuPDF library to extract text blocks; The target recognition model is used to perform layout recognition on PDF, and the extracted content is placed in the corresponding section for subsequent query and call.
5. The MBSE combined with RAG text retrieval method as claimed in claim 4, characterized in that: The data was preprocessed by removing punctuation, converting to lowercase, and stemming.
6. The MBSE combined with RAG text retrieval method as claimed in claim 5, characterized in that: Pandas is used for data preprocessing.
7. The MBSE combined with RAG text retrieval method as claimed in claim 6, characterized in that: When extracting high-frequency terms from each section to create a vocabulary, the extraction method is as follows: For documents in docx, doc, ppt, pptx, and xls formats, Spring AI and LangChain4j based on the Apache Tika framework are used for extraction; The content of PDF files is unstructured and data extraction is performed using AI-driven OCR.
Citation Information
Patent Citations
Information retrieval method and device based on thesaurus
CN103778262A
Multisource semantic analysis based information retrieval method
CN106156272A
Word2vec-based semantic query expansion method and device
CN108491462A
Document retrieval method and device, electronic equipment and storage medium
CN116821280A
Class case retrieval system and method based on retrieval enhancement generation technology
CN118260391A
Cited By
Model cluster-based complex system intelligent design method and system
CN121189188A