Anaphora resolution system and method based on large model and RAG technology
Through the reference digestion system of large model and RAG technology, combined with multi-dimensional similarity evaluation and contextual understanding, the difficulty of referential digestion caused by diverse entity expressions is solved, and high-accurate entity alignment and knowledge base construction are achieved.
Patent Information
- Application Number
- CN202510780396.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-07-08
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has various entities built in different local text contexts, which leads to difficulty in deriving the reference and inaccurate calculation methods, which may lead to failure of entity alignment and inaccurate results of deriving the reference and inaccurate results.
The reference digestion system based on big model and RAG technology is adopted, including knowledge cleaning, text segmentation, vectorization, vector search and RAG semantic processing modules, and multi-dimensional evaluation is carried out in combination with entity text similarity, structural similarity and semantic similarity, and context comprehensive understanding is carried out through RAG technology.
Improve the accuracy of entity alignment, ensure that the reference dissolution results are in line with actual semantics, effectively deal with ambiguity in complex contexts, and form a high-quality knowledge base.
Smart Images

Figure CN120278130A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large model data, and particularly relates to a coreference resolution system and method based on large models and RAG technology. Background Art
[0002] Essentially, a knowledge graph is a semantic network that describes objective things in the form of a graph. The graph consists of nodes and edges. Nodes in the knowledge graph represent concepts and entities. Concepts are abstracted things, and entities are specific things. Edges represent the relationships and attributes of things. The internal characteristics of things are represented by attributes, and the external connections are represented by relationships. Entities can be people, places, organizations, concepts, etc. There are more types of relationships, such as relationships between people, relationships between people and organizations, relationships between concepts and a certain object, etc. The knowledge graph is stored in the form of triples of "entity-relationship-entity" or "entity-attribute-attribute value" to form a graph-shaped knowledge base. Among them, entities are the basic elements of the knowledge graph, referring to specific names of people, organization names, place names, dates, times, etc. Relationships are the semantic relationships between two entities, which are instances of the relationships defined in the schema layer. Attributes are the descriptions of entities, which are the mapping relationships between entities and attribute values. In the data layer of the knowledge graph, nodes represent entities, and edges represent the relationships between entities or the attributes of entities;
[0003] The current coreference resolution has the following problems. In different local text contexts, the same entity constructed by a large model usually has multiple different expressions, which brings difficulties to coreference resolution and requires the use of entity alignment methods to solve. When judging whether two entities point to the same physical object, relying solely on a single similarity calculation method (such as text similarity, structural similarity, or semantic similarity) may be inaccurate and cannot comprehensively and accurately evaluate the corresponding relationship between entities. In actual operation, entity alignment failures may occur, and the existing technology's handling strategies for such situations are not perfect enough, which may lead to inaccurate coreference resolution results or hallucination phenomena. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the present application provides a coreference resolution system and method based on large models and RAG technology.
[0005] In a first aspect, the present application proposes a coreference resolution system based on large models and RAG technology, including a knowledge cleaning module, a text segmentation module, a vectorization module, a vector retrieval module, a coreference resolution module, and a RAG semantic processing module;
[0006] The knowledge cleaning module is used to preprocess the original text data, remove special symbols and redundant spaces, and convert the text into a unified format and encoding to ensure data consistency;
[0007] The text segmentation module is used to split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and segment it into small blocks according to rules to fit the context window of the large model, and obtain text data blocks;
[0008] The vectorization module is used to convert the segmented text data blocks into a vector matrix, obtain the digital representation of the text through the embedding technology, and obtain the final vectorization result;
[0009] The vector retrieval module is used to store the final vectorization result in the vector database and establish an index, and combine various retrieval methods such as similarity retrieval and full-text retrieval to retrieve vectors similar to the query vector from the vector database;
[0010] The anaphora resolution module is used to perform retrieval reasoning on the text data blocks using the search strategy of the vector database, cooperate with the large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction results in combination with entity text similarity, structural similarity and semantic similarity, judge whether two entities point to the same physical object, and merge the same entities;
[0011] The RAG semantic processing module is used to perform comprehensive context understanding and analysis based on the RAG technology, obtain the semantic understanding result, fuse the extracted knowledge, eliminate ambiguity and duplicate information, form a knowledge base, and store the constructed knowledge graph.
[0012] In some embodiments, the knowledge cleaning module includes a regular expression processing unit, an encoding conversion unit, and a standardization processing unit;
[0013] The regular expression processing unit is used to match and remove special symbols through a predefined regular expression pattern;
[0014] The encoding conversion unit is used to uniformly convert the text into the UTF-8 encoding format;
[0015] The standardization processing unit is used to perform operations on the unification of the date format and the conversion of case.
[0016] In some embodiments, the text segmentation module includes a confidence grading unit, a dynamic segmentation unit, and a symbol integrity verification unit;
[0017] The confidence grading unit is used to divide the delimiters into three confidence levels according to the semantic segmentation ability, and the confidence levels are line break > end-of-sentence punctuation > semicolon / comma;
[0018] The dynamic segmentation unit is used to force segmentation and retain the overlapping retention length when the character traversal reaches the maximum block length threshold. When the text block length exceeds the maximum block length threshold, a text block not exceeding the maximum block length threshold is constructed by merging the temporarily stored block and the split sub-blocks, and overlap_size characters are retained as context overlap.
[0019] The symbol integrity verification unit is used to perform paired verification on paired punctuation marks to avoid cutting breaks.
[0020] In some embodiments, the calculation method of the maximum block length threshold is as follows: Set chunk_size as the maximum block length threshold, and the maximum block length threshold is determined according to the context window size of the large model. Set overlap_size as the overlapping retention length, and the overlapping retention length is set to 10%-20% of chunk_size. When the text block length exceeds chunk_size*1.2, split preferentially at semicolons / commas; when the text block length exceeds chunk_size*1.5 and there is no delimiter, force a split at chunk_size and record a warning log.
[0021] In some embodiments, the vectorization module includes a model selection unit, a chunk embedding unit, and a vector normalization unit;
[0022] The model selection unit is used to select at least two embedding models from BERT, Sentence-BERT, and E5;
[0023] The chunk embedding unit is used to independently generate vectors for each semantic chunk after text segmentation;
[0024] The vector normalization unit is used to process vectors using the L2 normalization formula.
[0025] In some embodiments, the vector retrieval module includes a hybrid retrieval unit and a semantic optimization unit;
[0026] The hybrid retrieval unit is used to fuse the cosine similarity score and the TF-IDF full-text retrieval score with a weight of 6:4;
[0027] The semantic optimization unit is used to generate query expansion terms through the RAG technology to enhance the retrieval accuracy.
[0028] In some embodiments, the coreference resolution module includes a knowledge retrieval unit, a knowledge extraction unit, a multi-dimensional similarity calculation unit, and an entity alignment unit;
[0029] The knowledge retrieval unit is used to retrieve knowledge related to entity pairs from the vector database using the RAG system;
[0030] The knowledge extraction unit is used to generate the relationships between entities based on the retrieved knowledge fragments, combined with the specific requirements of the task, through a pre-trained large language model to complete global knowledge extraction;
[0031] The multi-dimensional similarity calculation unit is used to perform text similarity calculation: using the edit distance algorithm; structural similarity calculation: neighbor set comparison based on the Jaccard coefficient; semantic similarity calculation: cosine similarity of semantic vectors generated by RAG;
[0032] The entity alignment unit is used to align the entities in the knowledge extraction results by combining entity text similarity, structural similarity, and semantic similarity.
[0033] In some embodiments, the determination condition for the entity alignment unit to perform entity alignment is: Comprehensive similarity Score = 0.4 × text similarity + 0.3 × structural similarity + 0.3 × semantic similarity. When Score ≥ 0.75, an entity merging operation is triggered.
[0034] In a second aspect, the present application proposes a method for coreference resolution based on large models and RAG technology, including the following steps;
[0035] Preprocess the original text data, remove special symbols and redundant spaces, and convert the text into a unified format and encoding;
[0036] Split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and split it into small blocks according to the rules to adapt to the context window of the large model, and obtain text data blocks;
[0037] Convert the segmented text data blocks into a vector matrix, and obtain the digital representation of the text through embedding technology to get the final vectorized result;
[0038] Store the final vectorized result in a vector database and establish an index, and use a combination of multiple retrieval methods such as similarity retrieval and full-text retrieval to retrieve vectors similar to the query vector from the vector database;
[0039] Use the search strategy of the vector database to retrieve and reason about the text data blocks, cooperate with the large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction results by combining entity text similarity, structural similarity, and semantic similarity, judge whether two entities point to the same physical object, and merge the same entities;
[0040] Based on the RAG technology, comprehensive context understanding and analysis are carried out to obtain semantic understanding results. The extracted knowledge is fused to eliminate ambiguity and duplicate information, forming a knowledge base, and the constructed knowledge graph is stored.
[0041] In a third aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0042] In a fourth aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0043] Advantages of the present invention:
[0044] Based on the separator confidence grading, text blocks are dynamically segmented while preserving context overlap, and at the same time, symbol integrity verification is used to avoid the breakage of paired punctuation marks. This mechanism not only adapts to the context window limit of large models but also maintains the semantic coherence of text blocks, providing high-quality input for subsequent anaphora resolution. Combining text similarity, structural similarity, and semantic similarity, and using weighted comprehensive scoring for entity alignment determination. This multi-dimensional fusion strategy overcomes the limitations of single similarity calculation and can more comprehensively evaluate the relevance between entities. Through the RAG technology for comprehensive context understanding and combined with the reasoning ability of large models, it can effectively handle ambiguous anaphora in complex contexts and ensure that the generation of entity relationships and knowledge fusion are more in line with the actual semantics. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a system principle block diagram of the present invention.
[0046] Figure 2 It is an overall flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be completely conveyed to those skilled in the art.
[0048] In a first aspect, the present application proposes a coreference resolution system based on large models and RAG technology, as Figure 1 shown, including a knowledge cleaning module, a text segmentation module, a vectorization module, a vector retrieval module, a coreference resolution module, and a RAG semantic processing module;
[0049] The knowledge cleaning module is used to preprocess the original text data, remove special symbols and redundant spaces, convert the text into a unified format and encoding, and ensure data consistency;
[0050] In some embodiments, the knowledge cleaning module includes a regular expression processing unit, an encoding conversion unit, and a standardization processing unit;
[0051] The regular expression processing unit is used to match and remove special symbols through predefined regular expression patterns;
[0052] The encoding conversion unit is used to uniformly convert the text into the UTF-8 encoding format;
[0053] The standardization processing unit is used for operations such as unifying the date format and converting case.
[0054] First, text data is obtained from multiple sources, including but not limited to web pages, PDF documents, database records, social media posts, etc. These data sources may contain text in different formats and encodings, so unified processing is required. During the data acquisition process, web crawler technology is used to capture web page content to ensure the integrity and accuracy of the data. For PDF and DOC documents, specialized parsing tools are used to extract the text content to avoid interference from format information.
[0055] Next, the obtained text is cleaned. The cleaning process includes removing special symbols (such as HTML tags, Markdown marks, etc.), redundant spaces, line breaks, etc., to make the text more standardized. At the same time, garbled characters and incompatible characters in the text are processed to ensure the consistency of the text between different platforms and tools. For example, for Chinese characters, they are uniformly converted to the UTF-8 encoding to avoid garbled problems. For Western characters, ensure that their encoding is compatible with Chinese characters to avoid display errors.
[0056] After cleaning, the text is converted into a unified format. Select an appropriate text encoding (such as UTF-8) to reduce the occurrence of garbled and incompatible problems. For structured data (such as tables), it is converted into JSON or CSV format for subsequent processing. For unstructured data (such as plain text), it is converted into a unified text file format to ensure the consistency of the data between different platforms and tools.
[0057] The text segmentation module is used to split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and split it into small blocks according to the rules to adapt to the context window of the large model, and obtain text data blocks;
[0058] In some embodiments, the text segmentation module includes a confidence grading unit, a dynamic segmentation unit, and a symbol integrity verification unit;
[0059] The confidence grading unit is configured to divide delimiters into three confidence levels according to their semantic segmentation capabilities, where the confidence levels are newline characters > end-of-sentence punctuation marks > semicolons / commas;
[0060] The dynamic segmentation unit is configured to force segmentation and retain an overlapping retention length when the character traversal reaches the maximum block length threshold. When the text block length exceeds the maximum block length threshold, the text block not exceeding the maximum block length threshold is constructed by merging the temporarily stored block and the split sub-blocks, and overlap_size characters are retained as context overlap;
[0061] The symbol integrity verification unit is configured to perform paired verification on paired punctuation marks to avoid cutting breaks.
[0062] In some embodiments, the maximum block length threshold is calculated as follows: Set chunk_size as the maximum block length threshold, which is determined according to the context window size of the large model. Set overlap_size as the overlapping retention length, and the overlapping retention length is set to 10%-20% of chunk_size. When the text block length exceeds 1.2 times chunk_size, split preferably at semicolons / commas; when the text block length exceeds 1.5 times chunk_size and there is no delimiter, force splitting at chunk_size and record a warning log.
[0063] Text segmentation is one of the key steps in the knowledge cleaning stage. Its goal is to segment the original text into semantic blocks suitable for large model processing while retaining context information to improve the accuracy and efficiency of knowledge retrieval. The specific implementation steps are as follows:
[0064] In the Chinese scenario, during the process of constructing text blocks, we perform text splitting based on the confidence of delimiters (newline characters > full stops / question marks / exclamation marks > semicolons / commas). First, consider newline characters, full stops, question marks, and exclamation marks, and then semicolons and commas. Newline characters represent the semantic pause of a complete paragraph, full stops / question marks / exclamation marks represent the pause of an entire sentence, and semicolons and commas represent the semantic pause within an entire sentence. This priority setting helps to better understand the structure and semantics of the text.
[0065] The `chunk_size` is used to represent the preset block size. During the process of constructing text blocks, blocks not exceeding `chunk_size` need to be temporarily stored in sequence. When encountering a block larger than `chunk_size`, the temporarily stored set and the larger block will be disassembled and merged. The premise for splitting the largest block is the availability of a delimiter. Without a delimiter, even if the text block is extremely long, it cannot be split and can only be directly added to the final set. In the case of having a delimiter, we need to split the text block into small blocks and then merge and write them into the final set. At the same time, to improve the continuity and integrity of the text blocks, a part of the rightmost text needs to be retained to merge with the subsequent blocks, forming an overlapping part, which can better retain the context and prevent inappropriate short sentence segmentation from splitting the semantics.
[0066] According to the above rules and the characteristics of Chinese text, the following are the refined text chunking steps:
[0067] 1. Initialize parameters;
[0068] Set `chunk_size` (the maximum block length threshold);
[0069] Set `overlap_size` (the overlapping retention length, recommended to be 10 - 20% of `chunk_size`);
[0070] Create a temporarily stored block (the text block currently being constructed);
[0071] Create a final set (to store the text after chunking is completed);
[0072] Traverse the text character by character;
[0073] Read each character in sequence, add the character to the temporarily stored block, and check in real-time whether the length of the temporarily stored block exceeds `chunk_size`, and handle the delimiter priority
[0074] a. When encountering a newline character (\n):
[0075] Immediately add the temporarily stored block (including the newline character) to the final set;
[0076] Empty the temporarily stored block, and retain the last `overlap_size` characters as the starting content of the new temporarily stored block;
[0077] b. When not encountering a newline character, check for end-of-sentence punctuation, which includes a period, a question mark, or an exclamation point:
[0078] When the length of the temporarily stored block ≥ `chunk_size`:
[0079] 1. Search forward for the nearest end-of-sentence punctuation;
[0080] 2. If found, split after the punctuation;
[0081] 3. Add the first part to the final set and keep the second part as a new temporary block;
[0082] 4. Keep overlap_size characters as the overlap;
[0083] c. When the end punctuation of the sentence is not found, check for semicolons / commas (;,):
[0084] When the length of the temporary block ≥ chunk_size * 1.2 (allowing moderate overlength):
[0085] 1. Look forward to find the nearest semicolon / comma;
[0086] 2. If found, split after the punctuation;
[0087] 3. Add the first part to the final set and keep the second part as a new temporary block;
[0088] 4. Keep overlap_size characters as the overlap;
[0089] d. Handling of overlong blocks:
[0090] When the length of the temporary block ≥ chunk_size * 1.5 and there is no separator:
[0091] Force to split at chunk_size, add the first part to the final set, and keep the second part as a new temporary block (do not keep the overlap). After the traversal, if the temporary block is not empty, directly add it to the final set. If the length exceeds chunk_size * 2, a warning log needs to be recorded.
[0092] Suppose there is an original text: "The quick brown fox jumps over the lazy dog. The dog, however, was not amused. 2023-10-01". First, set the chunk_size to 30 characters and the overlap_size to 5 characters. Traverse the text character by character. When a period (a high-confidence separator) is encountered, perform segmentation to obtain the first text chunk: "The quick brown fox jumps over the lazy dog.". Then, continue traversing. When the text length reaches the chunk_size, force segmentation and retain the overlap_size to obtain the second text chunk: "The dog, however, was not amused. 2023-10-01". Finally, use a dynamic semantic segmentation algorithm to optimize the segmentation results to ensure that each text chunk has a complete semantics. Through this series of operations, the original text is segmented into semantic chunks suitable for large model processing while retaining the context information.
[0093] The vectorization module is used to convert the segmented text data chunks into a vector matrix, obtain the digital representation of the text through embedding technology, and get the final vectorization result;
[0094] In some embodiments, the vectorization module includes a model selection unit, a chunk embedding unit, and a vector normalization unit;
[0095] The model selection unit is used to select at least two embedding models from BERT, Sentence-BERT, and E5;
[0096] The chunk embedding unit is used to independently generate vectors for each semantic chunk after text segmentation;
[0097] The vector normalization unit is used to process the vectors using the L2 normalization formula.
[0098] Vectorization is the process of converting text data into a vector matrix. Its core goal is to obtain the digital representation of a piece of text through embedding technology, so that texts with similar semantics have similar vector representations in the vector space. The following are the detailed implementation steps of vectorization:
[0099] Text embedding model selection:
[0100] Select a suitable text embedding model according to the task requirements. Commonly used embedding models include BERT, Sentence-BERT, E5, etc. These models can map text into a high-dimensional vector space and capture the semantic information of the text.
[0101] For specific scenarios, an open-source embedding model can be selected for fine-tuning, or an embedding model suitable for the own scenario can be directly trained. When fine-tuning, a domain-specific dataset is used for training to improve the model's performance on specific tasks.
[0102] The output of the embedding model is a vector matrix, representing the semantic information of the text. The formula is as follows:
[0103]
[0104] Among them, represents the text, represents the embedding model, represents the vector representation of the text.
[0105] Text chunking and embedding:
[0106] The text is segmented into smaller, semantically meaningful chunks to better capture local semantic information. When chunking, ensure the semantic integrity of each chunk to avoid semantic breaks caused by segmentation.
[0107] Each text chunk is embedded to obtain its vector representation. The embedded vector matrix can be used for subsequent vector retrieval and similarity calculation.
[0108] The chunking and embedding formula is as follows:
[0109]
[0110] Among them, represents the th text chunk, represents the embedding model, represents the vector representation of the text chunk.
[0111] Vector normalization:
[0112] The embedded vectors are normalized to eliminate the influence of vector length on similarity calculation. The normalized vectors have unit length, facilitating the calculation of cosine similarity.
[0113] The normalization formula is as follows:
[0114]
[0115] Among them, represents the original vector, represents the length of the vector, represents the normalized vector, that is, the final vectorization result.
[0116] Suppose there is a piece of original text:
[0117] "The quick brown fox jumps over the lazy dog. The dog, however, was not amused.". First, split the text into two text blocks: "The quick brown fox jumps over the lazy dog." and "The dog, however, was not amused.". Then, select Sentence - BERT as the embedding model to embed each text block, obtaining two vector representations. Next, normalize these two vectors to make them have unit length. Store the normalized vectors in the FAISS vector database and build an HNSW index. Finally, calculate the similarity between the two vectors using cosine similarity, and the similarity result is 0.85, indicating that these two text blocks are semantically highly similar. Through this series of operations, the original text is transformed into a vector matrix, facilitating subsequent retrieval and similarity calculation.
[0118] The vector retrieval module is used to store the final vectorized result in the vector database and build an index, and combines various retrieval methods such as similarity retrieval and full - text retrieval to retrieve vectors similar to the query vector from the vector database;
[0119] In some embodiments, the vector retrieval module includes a hybrid retrieval unit and a semantic optimization unit;
[0120] The hybrid retrieval unit is used to fuse the cosine similarity score and the TF - IDF full - text retrieval score with a weight of 6:4;
[0121] The semantic optimization unit is used to generate query expansion terms through the RAG technology to enhance the retrieval accuracy.
[0122] Vector retrieval is a key step in knowledge graph construction, and its goal is to quickly find the vector most similar to the query vector from the vector database through an efficient retrieval method. The following are the detailed implementation steps of vector retrieval:
[0123] Vector index establishment:
[0124] Build an index for the vectorized data for quick retrieval. Common index methods include inverted index, HNSW, etc. HNSW is a graph - based index method that can efficiently process high - dimensional vector data.
[0125] The index establishment formula is as follows:
[0126]
[0127] Where, represents the vector set, Represents an indexing method, represents the established vector index.
[0128] Similarity search:
[0129] The cosine similarity is used to calculate the similarity between the query vector and the vectors in the database. Cosine similarity measures the similarity degree of two vectors in direction, with a value range of [-1, 1]. The larger the value, the higher the similarity.
[0130] The cosine similarity formula is as follows:
[0131]
[0132] Among them, and represent two vectors, represents the dot product of vectors, represents the length of the vector.
[0133] To improve the retrieval efficiency, the approximate nearest neighbor search algorithm can be adopted, which can significantly improve the retrieval speed on the premise of ensuring the retrieval accuracy.
[0134] Full-text search:
[0135] Based on vector retrieval, combined with the full-text search method, the recall rate is further improved. Full-text search finds the documents related to the query from the text data by means of keyword matching.
[0136] The full-text search formula is as follows:
[0137]
[0138] Among them, represents the document, represents the query, represents the term in the query, represents the term in the document TF-IDF value.
[0139] Hybrid search:
[0140] Combining similarity search and full-text search, the hybrid search method is adopted to improve the retrieval precision and recall rate. Hybrid search obtains the final retrieval result by weighted fusion of the results of the two search methods.
[0141] The hybrid search formula is as follows:
[0142]
[0143] Among them, Represents the weight for similarity retrieval, Represents the weight for full-text retrieval, Represents the score for similarity retrieval, Represents the score for full-text retrieval.
[0144] Sorting of retrieval results:
[0145] Sort the retrieval results so that users can quickly find the most relevant results. The sorting method can be comprehensively sorted according to various factors such as retrieval scores, timestamps, user preferences, etc.
[0146] The sorting formula is as follows:
[0147]
[0148] Among them, Represents the weight of the document which can be adjusted according to factors such as the importance of the document and user preferences.
[0149] There is a query vector , and we need to retrieve the most similar vector to it from the vector database. First, use the HNSW indexing method to establish a vector index. Then, calculate the similarity between the query vector and each vector in the database through cosine similarity to obtain the similarity score . At the same time, use the full-text retrieval method to calculate the relevance between the query keywords and the document through TF-IDF to obtain the full-text retrieval score . Then, perform weighted fusion on the results of similarity retrieval and full-text retrieval to obtain the final retrieval score . Finally, sort the results according to the retrieval score and output the most relevant retrieval results. Through this series of operations, the most similar vector to the query can be retrieved efficiently from the vector database, improving the precision and recall of the retrieval.
[0150] The anaphora resolution module is used to perform retrieval reasoning on the text data block using the search strategy of the vector database, cooperate with the large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction results by combining entity text similarity, structural similarity, and semantic similarity, determine whether two entities refer to the same physical object, and merge the same entities;
[0151] In some embodiments, the anaphora resolution module includes a knowledge retrieval unit, a knowledge extraction unit, a multi-dimensional similarity calculation unit, and an entity alignment unit;
[0152] The knowledge retrieval unit is used to retrieve knowledge related to entity pairs from the vector database using the RAG system;
[0153] The knowledge extraction unit is used to generate the relationships between entities based on the retrieved knowledge fragments and in combination with the specific requirements of the task, and complete the global knowledge extraction through a pre-trained large language model;
[0154] The multi-dimensional similarity calculation unit is used to perform text similarity calculation: using the edit distance algorithm; structural similarity calculation: neighbor set comparison based on the Jaccard coefficient; semantic similarity calculation: cosine similarity of semantic vectors generated by RAG;
[0155] The entity alignment unit is used to align the entities in the knowledge extraction results by combining entity text similarity, structural similarity, and semantic similarity.
[0156] In some embodiments, the determination condition for the entity alignment unit to perform entity alignment is: comprehensive similarity Score = 0.4 × text similarity + 0.3 × structural similarity + 0.3 × semantic similarity. When Score ≥ 0.75, the entity merging operation is triggered.
[0157] In the process of constructing a knowledge graph, entity alignment is a key step, aiming to determine whether the entities in different texts refer to the same physical object and merge the same entities. Entity alignment usually requires a comprehensive judgment by combining multi-dimensional information such as entity text similarity, structural similarity, and semantic similarity.
[0158] Text similarity calculation:
[0159] Text similarity is used to measure the similarity degree of the text expressions of two entities. Common calculation methods include the edit distance, which measures the similarity by calculating the minimum number of edit operations required to convert one string into another. The smaller the edit distance, the more similar the text expressions of the two entities.
[0160] The edit distance formula is as follows:
[0161]
[0162] Where and represent the text of two entities, 、 and represent the number of insertion, deletion, and replacement operations respectively.
[0163] Structural similarity calculation:
[0164] Structural similarity is used to measure whether the structural relationships of two entities in the knowledge graph are similar. Common calculation methods include the Jaccard coefficient, which measures structural similarity by calculating the ratio of the intersection to the union of the neighbor sets of the two entities.
[0165] The Jaccard coefficient formula is as follows:
[0166]
[0167] Where and represent the neighbor sets of the two entities respectively.
[0168] Semantic similarity calculation:
[0169] Semantic similarity is used to measure the semantic similarity degree between two entities. Common calculation methods include the knowledge extraction method based on RAG, which retrieves relevant information from an external knowledge base and combines it with a pre-trained large language model to generate semantic representations and calculate the semantic similarity of the two entities.
[0170] The semantic similarity formula is as follows:
[0171]
[0172] Where and represent the two entities, represents the entity semantic representation generated by the RAG technology, represents the cosine similarity.
[0173] Multi-dimensional similarity fusion:
[0174] Weightedly fuse text similarity, structural similarity, and semantic similarity to obtain a comprehensive similarity score for judging whether two entities point to the same physical object.
[0175] The comprehensive similarity formula is as follows:
[0176] Where , and represent the weights of text similarity, structural similarity, and semantic similarity respectively, which are usually adjusted according to specific application scenarios.
[0177] Suppose there are two entities and , and it is necessary to judge whether they point to the same physical object. First, calculate their text similarity using the edit distance formula , and obtain the text similarity score. Then, calculate their structural similarity using the Jaccard coefficient formula , the structural similarity score is obtained. Then, their semantic similarity is calculated. The knowledge extraction method based on RAG is used to generate semantic representations, and through the cosine similarity formula the semantic similarity score is calculated. Finally, the text similarity, structural similarity, and semantic similarity are weighted and fused to obtain the comprehensive similarity score , if the comprehensive similarity score exceeds the preset threshold, it is considered that and point to the same physical object.
[0178] The RAG semantic processing module is used to perform comprehensive context understanding and analysis based on RAG technology, obtain semantic understanding results, fuse the extracted knowledge, eliminate ambiguity and duplicate information, form a knowledge base, and store the constructed knowledge graph.
[0179] Among them, the RAG technology includes two main stages: the retrieval stage and the generation stage.
[0180] Retrieval stage: The RAG system uses an encoding model (such as BM25, SentenceBERT, ColBERT, etc.) to retrieve relevant information from the knowledge base according to the task requirements. The purpose of this step is to find the knowledge fragments most relevant to the user's question and provide a basis for subsequent text generation.
[0181] Generation stage: Based on the retrieved information and combined with the specific requirements of the task, the system generates the required text through a pre-trained large language model. This step uses the retrieved knowledge fragments to enhance the process of generating answers to questions, thereby producing more accurate and relevant text.
[0182] Visualize the identified entities and their relationships using the Neo4j graph database, and construct and store the knowledge graph.
[0183] In the second aspect, the present application proposes a coreference resolution method based on a large model and RAG technology, as Figure 2 shown, including the following steps;
[0184] S100: Preprocess the original text data, remove special symbols and extra spaces, and convert the text into a unified format and encoding;
[0185] S200: Split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and split it into small blocks according to the rules to adapt to the context window of the large model, and obtain text data blocks;
[0186] S300: Convert the segmented text data blocks into a vector matrix, obtain the digital representation of the text through embedding technology, and get the final vectorization result;
[0187] S400: Store the final vectorization result in a vector database and build an index. Combine multiple retrieval methods such as similarity retrieval and full-text retrieval to retrieve vectors similar to the query vector from the vector database;
[0188] S500: Use the search strategy of the vector database to retrieve and reason about the text data blocks, cooperate with a large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction results by combining entity text similarity, structural similarity, and semantic similarity, determine whether two entities point to the same physical object, and merge the same entities;
[0189] S600: Conduct comprehensive context understanding and analysis based on the RAG technology to obtain the semantic understanding result, fuse the extracted knowledge, eliminate ambiguity and duplicate information, form a knowledge base, and store the constructed knowledge graph.
[0190] In a third aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0191] In a fourth aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0192] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0193] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not elaborated or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0194] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this disclosure.
[0195] In the embodiments provided in this disclosure, it should be understood that the disclosed device / computer equipment and method can be implemented in other ways. For example, the device / computer equipment embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0196] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0197] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0198] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0199] The above are only the preferred embodiments of the present invention. It should be pointed out that for those skilled in the art, without departing from the premise of the technical solution of the present invention, several deformed and improved technical solutions should be equally regarded as falling within the scope protected by this solution.
Claims
1. A pronoun resolution system based on large models and RAG technology, characterized in that: It includes a knowledge cleaning module, a text segmentation module, a vectorization module, a vector retrieval module, a coreference resolution module, and a RAG semantic processing module; The knowledge cleaning module is used to preprocess the original text data, remove special symbols and redundant spaces, convert the text into a unified format and encoding, and ensure data consistency; The text segmentation module is used to split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and segment it into small blocks according to the rules to adapt to the context window of the large model, and obtain text data blocks; The vectorization module is used to convert the segmented text data blocks into a vector matrix, obtain the digital representation of the text through embedding technology, and obtain the final vectorization result; The vector retrieval module is used to store the final vectorization result in a vector database and establish an index, and combine multiple retrieval methods such as similarity retrieval and full-text retrieval to retrieve vectors similar to the query vector from the vector database; The coreference resolution module is used to retrieve and reason about the text data blocks using the search strategy of the vector database, cooperate with the large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction result by combining entity text similarity, structural similarity, and semantic similarity, judge whether two entities point to the same physical object, and merge the same entities; The RAG semantic processing module is used to perform comprehensive context understanding and analysis based on RAG technology, obtain semantic understanding results, fuse the extracted knowledge, eliminate ambiguity and duplicate information, form a knowledge base, and store the constructed knowledge graph; 2. The system according to claim 1, wherein: The knowledge cleaning module includes a regular expression processing unit, an encoding conversion unit, and a standardization processing unit; The regular expression processing unit is used to match and remove special symbols through predefined regular expression patterns; The encoding conversion unit is used to uniformly convert the text into the UTF-8 encoding format; The standardization processing unit is used for operations such as unifying the date format and converting the case; 3. The system according to claim 2, characterized in that: The text segmentation module includes a confidence grading unit, a dynamic segmentation unit, and a symbol integrity verification unit; The confidence grading unit is used to divide the delimiters into three confidence levels according to their semantic segmentation capabilities, and the confidence levels are line break > end-of-sentence punctuation > semicolon / comma; The dynamic segmentation unit is used to force segmentation and retain the overlapping retention length when the character traversal reaches the maximum block length threshold. When the text block length exceeds the maximum block length threshold, a text block not exceeding the maximum block length threshold is constructed by merging the temporarily stored block and the split sub-blocks, and overlap_size characters are retained as context overlap; The symbol integrity verification unit is used to perform paired verification on paired punctuation marks to avoid cutting breaks.
4. The system according to claim 3, wherein: The calculation method of the maximum block length threshold is as follows: Set chunk_size as the maximum block length threshold, which is determined according to the context window size of the large model. Set overlap_size as the overlapping retention length, and the overlapping retention length is set to 10%-20% of chunk_size. When the text block length exceeds 1.2 * chunk_size, it is preferentially split at semicolons / commas; when the text block length exceeds 1.5 * chunk_size and there is no delimiter, it is forcibly split at chunk_size and a warning log is recorded.
5. The system according to claim 4, wherein: The vectorization module includes a model selection unit, a chunk embedding unit, and a vector normalization unit; The model selection unit is used to select at least two embedding models from BERT, Sentence-BERT, and E5; The chunk embedding unit is used to independently generate vectors for each semantic chunk after text segmentation; The vector normalization unit is used to process vectors using the L2 normalization formula.
6. The system according to claim 5, wherein: The vector retrieval module includes a hybrid retrieval unit and a semantic optimization unit; The hybrid retrieval unit is used to fuse the cosine similarity score and the TF-IDF full-text retrieval score with a weight of 6:4; The semantic optimization unit is used to generate query expansion terms through the RAG technology to enhance the retrieval accuracy.
7. The system according to claim 6, wherein: The anaphora resolution module includes a knowledge retrieval unit, a knowledge extraction unit, a multi-dimensional similarity calculation unit, and an entity alignment unit; The knowledge retrieval unit is used to retrieve knowledge related to entity pairs from the vector database using the RAG system; The knowledge extraction unit is used to generate the relationships between entities based on the retrieved knowledge fragments and in combination with the specific requirements of the task, and complete the global knowledge extraction through a pre-trained large language model; The multi-dimensional similarity calculation unit is used to perform text similarity calculation: using the edit distance algorithm; structure similarity calculation: neighbor set comparison based on the Jaccard coefficient; Semantic similarity calculation: cosine similarity of semantic vectors generated by RAG; The entity alignment unit is used to align the entities in the knowledge extraction results by combining entity text similarity, structure similarity, and semantic similarity.
8. The system according to claim 7, characterized in that: The determination condition for the entity alignment unit to perform entity alignment is: Comprehensive similarity Score = 0.4 × text similarity + 0.3 × structure similarity + 0.3 × semantic similarity. When Score ≥ 0.75, the entity merging operation is triggered.
9. A method for anaphora resolution based on large models and RAG technology, characterized in that: Including the following steps; Preprocess the original text data, remove special symbols and extra spaces, and convert the text to a unified format and encoding; Split the cleaned text according to the confidence of the delimiter, set the maximum block length threshold and the overlapping retention length, traverse the text character by character and split it into small chunks to fit the context window of the large model, and obtain text data chunks; Convert the segmented text data chunks into a vector matrix, and obtain the digital representation of the text through the embedding technology to get the final vectorization result; Store the final vectorized results in a vector database and build an index, and use a combination of various retrieval methods such as similarity retrieval and full-text retrieval to retrieve vectors similar to the query vector from the vector database; Use the search strategy of the vector database to retrieve and reason about the text data block, cooperate with the large model to generate the relationships between entities, complete knowledge extraction, align the entities in the knowledge extraction results by combining entity text similarity, structural similarity and semantic similarity, judge whether two entities point to the same physical object, and merge the same entities; Conduct comprehensive context understanding and analysis based on the RAG technology, obtain the semantic understanding results, fuse the extracted knowledge, eliminate ambiguity and duplicate information, form a knowledge base, and store the constructed knowledge graph.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in claim 9 are implemented.
Citation Information
Patent Citations
Intelligent document retrieval generation method and system based on RAG technology
CN118332072A
Enhanced document generation and retrieval method based on knowledge graph
CN119646178A