Retrieval enhancement letter text analysis method and system based on knowledge graph
By constructing a letter text analysis method based on knowledge graphs, using vectorization and entity relationship recognition technology, the problems of inaccurate semantic search and high computational complexity in the existing technology are solved, and efficient and accurate answers to text analysis in professional fields are achieved.
Patent Information
- Application Number
- CN202510498869.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing search enhancement generation techniques have problems in semantic search accuracy and computational complexity. Especially in text analysis in professional fields, entity disambiguation and semantic correlation are poor, and the computing resource consumption is not proportional.
A search-enhanced letter text analysis method based on knowledge graph is constructed, and a letter knowledge graph is constructed by vectorizing text blocks and constructing a letter knowledge graph. It uses a two-way long and short-term memory network and a multi-head attention mechanism for feature extraction and entity relationship recognition, and combines a large language model for mixed search and answer generation.
It improves the accuracy and professionalism of text analysis in the professional field of large language models, reduces the amount of calculation, provides richer information support, and generates answers that meet user needs.
Smart Images

Figure CN120386874A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and specifically to a retrieval-enhanced correspondence text analysis method and system based on a knowledge graph. Background Art
[0002] Retrieval-Augmented Generation (RAG) is an artificial intelligence technology that combines information retrieval technology with language generation models. It retrieves relevant information from an external knowledge base and uses it as a prompt to input to large language models (LLMs) to enhance the model's ability to handle knowledge-intensive tasks. Although retrieval-augmented generation has played a certain role in improving the accuracy and knowledgeability of large language models, there are still problems such as inaccurate semantic search and high computational costs due to complex architectures.
[0003] In recent years, to improve the accuracy and domain adaptability of text analysis, it has been achieved by integrating the semantic reasoning ability of knowledge graphs and the context correlation characteristics of retrieval enhancement technologies. However, there are still problems at the technical integration level: First, the mainstream approach only shallowly concatenates the entity linking results of knowledge graphs with the document features of retrieval enhancement. This linear superposition method leads to an exponential increase in algorithm complexity and generates a large amount of redundant calculations. Second, the lack of technical synergy makes the improvement in accuracy not proportional to the consumption of computing resources. Especially when processing text in some professional fields, the existing methods have poor improvement effects in entity disambiguation and semantic association. The current architecture fails to establish a dynamic coupling mechanism between knowledge embedding and retrieval features, resulting in the inability of domain-specific knowledge to effectively guide the retrieval direction. Summary of the Invention
[0004] Based on this, in view of the problem that the combination of existing retrieval enhancement technology and knowledge graph technology alone causes a large amount of calculation and insignificant improvement in accuracy, it is necessary to provide a retrieval-enhanced correspondence text analysis method and system based on a knowledge graph.
[0005] To achieve the above object, the present invention adopts the following technical solutions: A retrieval-enhanced correspondence text analysis method based on a knowledge graph, comprising the following steps: Obtain sample correspondence text for preprocessing, inject spatio-temporal watermarks and invisible identifiers into the preprocessed text, perform text chunking and vectorization, obtain vector text chunks, and construct a vector database; Perform entity relationship recognition on the vector text chunks, construct a correspondence knowledge graph in the form of triples, and use the correspondence knowledge graph index to uniquely identify the vector text chunks in the vector database, establishing the association between the vector text chunks and the correspondence knowledge graph; Obtain the user's question information, retrieve it in the vector database, identify the spatio-temporal watermark and invisible identifier contained in the retrieved vector text block, perform source exclusion calculation, optimize to obtain vector retrieval information, and retrieve in the correspondence knowledge graph through the unique ID of the vector text block in the vector retrieval information to obtain knowledge graph retrieval information. Then, fuse the vector retrieval information and the knowledge graph retrieval information to obtain hybrid retrieval information; Use the hybrid retrieval information as the prompt word of the large model and input it into the large language model together with the user's question information to output the answer text information corresponding to the user's question information.
[0006] Furthermore, the specific steps for identifying the entity relationship of the vector text block and constructing the correspondence knowledge graph in the form of triples are as follows: Use a bidirectional long short-term memory network to extract features from the vector text block and input them into the multi-head attention mechanism to establish the association between the features and the vector text block; Classify the features into different labels for probability calculation, sort to obtain the labeled sequence, arrange the different labeled sequences in descending order of probability, and select the TOP1 labeled sequence to form an ontological knowledge expression and use it as new knowledge; Fuse the new knowledge, then perform knowledge evaluation, and store the qualified part in the correspondence knowledge graph in the form of <entity, relationship, entity> triples.
[0007] Furthermore, the specific steps for obtaining the user's question information, retrieving it in the vector database, identifying the spatio-temporal watermark and invisible identifier contained in the retrieved vector text block, and performing source exclusion calculation to optimize and obtain vector retrieval information are as follows: Preprocess the user's question information, calculate the similarity between the preprocessed user's question information vector and the constructed vector database, sort in reverse order, select the TOPN vector text blocks, identify the spatio-temporal watermark and invisible identifier of the TOPN vector text blocks, determine their source correspondence text, and perform source exclusion reordering to select the TOPK vector text blocks as vector retrieval information; where K < N and is a positive integer.
[0008] Furthermore, the specific steps for retrieving in the correspondence knowledge graph through the unique ID of the vector text block in the vector retrieval information to obtain knowledge graph retrieval information are as follows: Read the ID of the vector text block in the vector retrieval information, and obtain the corresponding entity representation in the correspondence knowledge graph according to the ID; Obtain the adjacent entity representation of the entity representation through multi-hop reasoning, and use it together with the entity representation as the knowledge graph retrieval information.
[0009] Further, the specific steps for obtaining sample correspondence texts for preprocessing, injecting spatio-temporal watermarks and invisible identifiers into the preprocessed texts, performing text chunking and vectorization, obtaining vector text chunks, and constructing a vector database are as follows: Obtain sample correspondence texts for data cleaning, remove irrelevant special symbols and redundant information, and inject spatio-temporal watermarks and invisible identifiers into the cleaned texts to obtain enhanced texts; use the exact word segmentation mode in the jieba word segmentation tool to divide the enhanced texts into text chunks, then use the BERT model to perform word embedding on the text chunks, and vectorize the text chunks to obtain vector text chunks.
[0010] Further, the sample correspondence texts include official documents, communication records, agreements, and reports in various industries and fields.
[0011] Further, the mixed retrieval information is used as prompts for the large model in a custom structured form, where the custom structured form includes the form of the sender unit, the sending date, the recipient unit, the correspondence requirements, the original plan of the correspondence, and the modification plan of the correspondence.
[0012] The present invention also relates to a retrieval-enhanced correspondence text analysis system based on a knowledge graph, which mainly includes a vector database construction module, a correspondence knowledge graph association module, a retrieval information mixing module, and an output module.
[0013] The vector database construction module is used to obtain sample correspondence texts for preprocessing, perform text chunking and vectorization on the preprocessed texts, obtain vector text chunks, and construct a vector database; The correspondence knowledge graph association module is used to identify entity relationships in the vector text chunks, construct a correspondence knowledge graph in the form of triples, and perform unique ID identification on the vector text chunks in the vector database through the correspondence knowledge graph index to establish the association between the vector text chunks and the correspondence knowledge graph; The retrieval information mixing module is used to obtain user question information, perform retrieval in the vector database to obtain vector retrieval information, and perform retrieval in the correspondence knowledge graph through the unique ID of the vector text chunks in the vector retrieval information to obtain knowledge graph retrieval information, and fuse the vector retrieval information and the knowledge graph retrieval information to obtain mixed retrieval information; The output module is used to input the mixed retrieval information as prompts for the large model together with the user question information into the large language model, and output the answer text information corresponding to the user question information.
[0014] Compared with the prior art, the beneficial effects of the present invention include: 1. The present invention utilizes knowledge graph technology, combines knowledge engineering with retrieval technology, constructs a correspondence knowledge graph, which can provide guidance for knowledge management, and establishes an association between the knowledge graph database and the vector database by assigning IDs to each text block vector according to the knowledge graph index. When performing hybrid retrieval, the workload is greatly reduced, and the combination of the two technologies provides richer information for the large language model; 2. The present invention utilizes retrieval-enhanced generation technology, combines retrieval technology with a language generation model, and can retrieve more context information according to the knowledge base constructed based on the collected correspondence information, making up for the deficiencies of traditional large language models in the professional field of correspondence analysis, enabling the large language model to generate answers that more meet the user's needs, and improving the accuracy and professionalism of the large model's question answering. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. Among them: Figure 1 is a flowchart of a retrieval-enhanced correspondence text analysis method based on a knowledge graph introduced in the present invention; Figure 2 is based on Figure 1 the logical framework diagram; Figure 3 is based on Figure 1 the flowchart of constructing a correspondence knowledge graph. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various structural ways and implementation ways that can be mutually replaced. Therefore, the following detailed embodiments and the accompanying drawings are only illustrative descriptions of the technical solution of the present invention, and should not be regarded as the whole of the present invention or as a limitation or restriction on the technical solution of the present invention.
[0017] Embodiment 1 Please refer to Figure 1 , this embodiment introduces a retrieval-enhanced correspondence text analysis method based on a knowledge graph, including the following steps: S1. Obtain sample correspondence text for preprocessing, inject spatio-temporal watermarks and invisible identifiers into the preprocessed text, perform text chunking and vectorization, obtain vector text chunks, and construct a vector database. Specifically as follows: Taking the massive letter texts obtained as sample letter texts, cleaning the data of the letter texts to remove irrelevant special symbols and redundant information, and injecting spatio-temporal watermarks and invisible identifiers into the cleaned texts to obtain enhanced texts. Using the jieba word segmentation tool to perform text chunking on the enhanced letter texts, adopting the accurate mode of word segmentation, and on this basis, dividing long texts; using the BERT model for word embedding, vectorizing the text chunks and constructing a vector database.
[0018] The process of collecting massive letter information involves obtaining sent letters and received letters in different fields from multiple sources. The content of these letters includes official documents, communication records, agreements, reports in various industries and fields. Relevant documents in the relevant fields are sorted out from them to reduce the workload of text analysis.
[0019] Using the BERT model to perform word embedding on the segmented text chunks to obtain a word vector matrix . In the BERT model, the Encoder part in Transformer is mainly used, which is also called Transformer Block. Transformer Block is divided into two parts: the multi-head attention mechanism and the feed-forward neural network. Finally, it is transformed into an output feature vector and output to the hidden layer. The input vector of BERT is composed of three vectors: Token Embedings obtained after splitting the word vector, Segment Embeddings obtained after splitting the sentence vector, and the position embedding vector Position Embedings that records the sequence order.
[0020] S2. Perform entity relationship recognition on the vector text chunks, construct a letter knowledge graph in the form of triples, and use the letter knowledge graph index to uniquely identify the vector text chunks in the vector database, establishing the association between the vector text chunks and the letter knowledge graph. Specifically as follows: As Figure 3 shown, use the bidirectional long short-term memory network (BiLSTM) for feature extraction; take the word vector matrix after word embedding as the input, and finally obtain the corresponding output column The expression is: In formula (1) represents the hidden state sequence output by the forward LSTM network, represents the hidden state sequence output by the reverse LSTM network; Input the complete output sequence into the multi-head attention mechanism, which is used to map the dimension of the word vector matrix in the output sequence to a low dimension and establish the association between the word vector and the whole sentence; In the multi-head attention mechanism, the calculation method of each self-attention mechanism head is as follows: In formula (2) represents the i-th self-attention mechanism head, , , represents the corresponding weight parameter, represents the process of single-scale dot product, where Q, K, and V are the query, key, and value matrices of the word vectors in the output sequence respectively; The single Attention expression is: In formula (3) represents the scaling factor, represents the activation function; Concatenate the output vectors of multiple parallel structures to obtain a word vector matrix S containing the information of the entire sentence. The vector concatenation method is: In formula (4) represents the multi-head attention formula, h represents the number of self-attention heads in the multi-layer attention mechanism, is the matrix parameter used to concatenate the sentence matrix; Use CRF (Conditional Random Field) for feature classification. The loss function expression of the CRF model is: In formula (5) represents the scoring of different words in the sentence classified into each label, represents the scores of the word vectors on different labels after feature extraction. N represents the number of word vectors. BIO is used as the dataset annotation method. B represents the start of the entity, I represents the non-start part of the entity, and O represents the non-entity. The path score expression is: In formula (6) is the transition score, representing the possibility of the label i to label i+1 change in the sequence, represents the word corresponding to the th label score; Use the maximum likelihood estimation algorithm to calculate the likelihood probability of the label, sort it, select the best annotation sequence, and form an ontological knowledge expression as new knowledge; The log-likelihood formula is: In formula (7) Denotes the probability that the input text sequence is x and the annotation sequence is y. x is the word vector matrix after word embedding. n represents the length of the input text sequence. Denotes the actual label. Denotes the set of annotation sequences that the sequence x may be transformed into. Fuse the acquired new knowledge, eliminate contradictions and ambiguities, evaluate the fused knowledge, add the qualified parts to the knowledge base and store them in the form of <entity, relationship, entity> triples in the correspondence knowledge graph (Neo4j graph database).
[0021] Determine the ID of each vector text block by the index in the correspondence knowledge graph, and each vector text block ID is associated with the entity in the knowledge graph.
[0022] It should be noted that to construct the correspondence knowledge graph, it is necessary to pre-define the entities in the correspondence according to the basic understanding of the correspondence information, specifically including: the sending unit, the sending date, the receiving unit, the correspondence requirements, the original plan of the correspondence, and the modified plan of the correspondence. Use the entity recognition model to identify the entities and relationships in the collected correspondences according to the specified entities, and finally generate the correspondence knowledge graph.
[0023] S3. Obtain the user's question information, retrieve it in the vector database, identify the spatio-temporal watermark and invisible identifier contained in the retrieved vector text block and perform source exclusion calculation, optimize to obtain the vector retrieval information, and retrieve in the correspondence knowledge graph through the unique ID of the vector text block in the vector retrieval information to obtain the knowledge graph retrieval information, and fuse the vector retrieval information and the knowledge graph retrieval information to obtain the hybrid retrieval information. Specifically as follows: Preprocess the user's question information, calculate the similarity between the preprocessed user's question information vector and the constructed vector database, sort it in reverse order, select the TOPN vector text blocks, identify the spatio-temporal watermark and invisible identifier of the TOPN vector text blocks, determine their source correspondence text, and reorder through source exclusion to select the TOPK vector text blocks as the vector retrieval information. Among them, K < N and is a positive integer.
[0024] The similarity calculation formula is: In formula (8), n represents the weight, l represents the user's question information vector, and m represents the vector in the constructed vector database. Denotes the dot product. Denotes the norm of vector 2. Read the ID of the vector text block in the vector retrieval information, and obtain the entity representation of the corresponding knowledge graph according to the ID. Obtain adjacent entities based on multi-hop reasoning and obtain relevant information; Extract a subgraph from the correspondence knowledge graph , where V is the entity set of the subgraph and R is the edge set connecting entities. Obtain the embedding vector corresponding to each entity from the knowledge graph database. According to the query q and the adjacent entities after multi-hop reasoning, define an attention score using an activation function The formula is: In formula (9) represents the attention score of the k-th hop in the i-th layer, represents the entity vector of the k-th hop in the i-th layer, is the learning parameter matrix, represents element-wise multiplication, represents vector concatenation; Calculate the weighted sum of the k-th hop entities to obtain the result of multi-hop reasoning, and jointly use the entity representation of the corresponding knowledge graph obtained by the ID as the knowledge graph retrieval information; Organize and fuse the vector retrieval information and the knowledge graph retrieval information to obtain the hybrid retrieval information.
[0025] S4. Use the hybrid retrieval information as the prompt words of the large model and input them together with the user's question information into the large language model to output the answer text information corresponding to the user's question.
[0026] The large language model is based on the Transformer architecture and completely relies on the attention mechanism to process the encoding and decoding of sequences. When generating text output, the large language model will calculate the attention score between the input and the output, obtain the weight through the softmax activation function, and then calculate the weighted sum of the input sequence to obtain the context vector and obtain the relationship between words.
[0027] The hybrid retrieval related information will be provided to the large language model in a structured manner, specifically including the form of the sending unit, sending date, receiving unit, correspondence requirements, original correspondence plan, and correspondence modification plan. These are input together with the user's query as prompt information to ensure that the large language model can reason based on the knowledge domain of the correspondence text and generate professional answers in the correspondence field.
[0028] In this embodiment, the knowledge graph technology is utilized to combine knowledge engineering with retrieval technology. Constructing a correspondence knowledge graph can provide guidance for knowledge management. An association between the knowledge graph database and the vector database is established by assigning IDs to each text block vector according to the knowledge graph index. When performing hybrid retrieval, the workload is greatly reduced. Meanwhile, the combination of the two technologies provides richer information for the large language model. By using retrieval-enhanced generation technology, the retrieval technology is combined with the language generation model. A knowledge base constructed based on the collected correspondence information can retrieve more context information, making up for the deficiencies of traditional large language models in the professional field of correspondence analysis, enabling the large language model to generate answers that better meet user needs and improving the accuracy and professionalism of the large model's question answering.
[0029] Embodiment 2 This embodiment introduces a retrieval-enhanced correspondence text analysis system based on a knowledge graph, which mainly includes a vector database construction module, a correspondence knowledge graph association module, a retrieval information mixing module, and an output module.
[0030] The vector database construction module is used to obtain sample correspondence texts for preprocessing, perform text chunking and vectorization on the preprocessed texts to obtain vector text chunks, and construct a vector database. The correspondence knowledge graph association module is used to identify entity relationships in the vector text chunks, construct a correspondence knowledge graph in the form of triples, and perform unique ID identification on the vector text chunks in the vector database through the correspondence knowledge graph index, establishing an association between the vector text chunks and the correspondence knowledge graph. The retrieval information mixing module is used to obtain user question information, perform retrieval in the vector database to obtain vector retrieval information, and perform retrieval in the correspondence knowledge graph through the unique IDs of the vector text chunks in the vector retrieval information to obtain knowledge graph retrieval information, and fuse the vector retrieval information and the knowledge graph retrieval information to obtain hybrid retrieval information. The output module is used to input the hybrid retrieval information as a prompt word for the large model and the user question information into the large language model together, and output the answer text information corresponding to the user question information.
[0031] Embodiment 3 This embodiment provides a computer terminal, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the lower limb biomechanical data estimation method based on the data augmentation technology in Embodiment 1.
[0032] When the lower limb biomechanical data estimation method based on the data augmentation technology in Embodiment 1 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program and installed on a computer terminal, which can be a computer, a smart phone, etc. It can also be designed as an embedded running program and installed on a computer terminal, such as installed on a single-chip microcomputer.
[0033] Embodiment 4 This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the lower limb biomechanical data estimation method based on the data augmentation technology in Embodiment 1 are implemented.
[0034] When the lower limb biomechanical data estimation method based on the data augmentation technology in Embodiment 1 is applied, it can be applied in the form of software. For example, it can be designed as an independently running program on a computer-readable storage medium. The computer-readable storage medium can be a USB flash drive, designed as a USB key, and the program that starts the whole method by external trigger is designed through the USB flash drive.
[0035] The technical scope of the present invention is not limited to the content in the above description. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.
Claims
1. A retrieval-enhanced correspondence text analysis method based on a knowledge graph, characterized in that It includes the following steps: Obtain the sample correspondence text for preprocessing, inject spatio-temporal watermarks and invisible identifiers into the preprocessed text, perform text chunking and vectorization, obtain vector text chunks, and construct a vector database; Perform entity relationship recognition on the vector text chunks, construct a correspondence knowledge graph in the form of triples, use the correspondence knowledge graph index to uniquely identify the vector text chunks in the vector database, and establish the association between the vector text chunks and the correspondence knowledge graph; Obtain the user's question information, retrieve it in the vector database, identify the spatio-temporal watermarks and invisible identifiers contained in the retrieved vector text chunks, perform source exclusion calculation, optimize to obtain vector retrieval information, and retrieve in the correspondence knowledge graph through the unique ID of the vector text chunks in the vector retrieval information to obtain knowledge graph retrieval information, and fuse the vector retrieval information with the knowledge graph retrieval information to obtain mixed retrieval information; Use the mixed retrieval information as the prompt of the large model and input it into the large language model together with the user's question information to output the answer text information corresponding to the user's question information.
2. The retrieval enhanced correspondence text analysis method based on a knowledge graph according to claim 1, wherein, The specific steps of performing entity relationship recognition on the vector text chunks and constructing a correspondence knowledge graph in the form of triples are as follows: Use a bidirectional long short-term memory network to extract features from the vector text chunks and input them into a multi-head attention mechanism to establish the association between the features and the vector text chunks; Classify the features into different labels for probability calculation, obtain the labeled sequence after sorting, arrange the different labeled sequences in descending order of probability, and select the TOP1 labeled sequence to form an ontological knowledge expression and use it as new knowledge; Perform knowledge fusion on the new knowledge, then perform knowledge evaluation, and store the qualified part in the correspondence knowledge graph in the form of <entity, relationship, entity> triples.
3. The retrieval enhanced correspondence text analysis method based on a knowledge graph according to claim 1, wherein The specific steps of obtaining the user's question information, retrieving it in the vector database, identifying the spatio-temporal watermarks and invisible identifiers contained in the retrieved vector text chunks, and performing source exclusion calculation to optimize and obtain vector retrieval information are as follows: Preprocess the user's question information, calculate the similarity between the preprocessed user's question information vector and the constructed vector database, sort it in reverse order, select the TOPN vector text chunks, identify the spatio-temporal watermarks and invisible identifiers of the TOPN vector text chunks, determine their source correspondence text, and perform source exclusion reordering to select the TOPK vector text chunks as vector retrieval information; where K < N and is a positive integer.
4. The retrieval-enhanced correspondence text analysis method based on a knowledge graph according to claim 1, characterized in that The specific steps of retrieving in the correspondence knowledge graph through the unique ID of the vector text chunks in the vector retrieval information to obtain knowledge graph retrieval information are as follows: Read the ID of the vector text chunks in the vector retrieval information, and obtain the corresponding entity representation in the correspondence knowledge graph according to the ID; Obtain the adjacent entity representation through multi-hop reasoning of the entity representation, and use it together with the entity representation as the knowledge graph retrieval information.
5. The retrieval enhanced correspondence text analysis method based on a knowledge graph according to claim 1, characterized in that The specific steps of obtaining the sample correspondence text for preprocessing, injecting spatio-temporal watermarks and invisible identifiers into the preprocessed text, performing text chunking and vectorization, obtaining vector text chunks, and constructing a vector database are as follows: Obtain the sample correspondence text for data cleaning, remove irrelevant special symbols and redundant information, and inject spatio-temporal watermarks and invisible identifiers into the cleaned text to obtain enhanced text; Use the exact word segmentation mode in the jieba word segmentation tool to divide the enhanced text into text blocks, and then use the BERT model to perform word embedding on the text blocks to vectorize the text blocks to obtain vector text blocks.
6. The method for retrieving and enhancing correspondence text analysis based on a knowledge graph according to claim 5, wherein, The sample correspondence text includes official documents, communication records, agreements, and reports in various industries and fields.
7. The method for retrieving and enhancing correspondence text analysis based on a knowledge graph according to claim 1, wherein Mix the retrieval information as prompt words for the large model in a customized structured form, where the customized structured form includes the form of the sender unit, the sending date, the recipient unit, the correspondence requirements, the original plan of the correspondence, and the modification plan of the correspondence.
8. A retrieval-enhanced correspondence text analysis system based on a knowledge graph, characterized in that, It includes: A vector database construction module, which is used to obtain the sample correspondence text for preprocessing, divide the preprocessed text into text blocks and vectorize them to obtain vector text blocks and construct a vector database; A correspondence knowledge graph association module, which is used to identify entity relationships in the vector text blocks and construct a correspondence knowledge graph in the form of triples, and use the correspondence knowledge graph index to uniquely identify the vector text blocks in the vector database, and establish the association between the vector text blocks and the correspondence knowledge graph; A retrieval information mixing module, which is used to obtain user question information, retrieve in the vector database to obtain vector retrieval information, and retrieve in the correspondence knowledge graph through the unique ID of the vector text blocks in the vector retrieval information to obtain knowledge graph retrieval information, and fuse the vector retrieval information and the knowledge graph retrieval information to obtain mixed retrieval information; An output module, which is used to input the mixed retrieval information as prompt words for the large model and the user question information into the large language model together, and output the answer text information corresponding to the user question information.
9. A computer terminal, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the knowledge graph-based retrieval enhanced correspondence text analysis method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the knowledge graph-based retrieval enhanced correspondence text analysis method according to any one of claims 1-7.
Citation Information
Patent Citations
RAG question and answer method and system based on knowledge graph and medium
CN118673126A
Knowledge retrieval enhancement-based characteristic agricultural product standardized file content generation method
CN118779407A
Method and system for generating enhanced knowledge questions and answers for mixed retrieval of heterogeneous database
CN119311831A
Mixing retrieval method and device based on knowledge graph and vector, equipment and storage medium
CN119357376A
Vector database watermarking method with transparent vector priority
CN119646773A
Cited By
Multi-index knowledge base and updating method and query self-healing method thereof
CN122432172A