Knowledge construction method and system based on large model and RAG technology

Through the combination of large models and RAG technology, the problem of word segmentation accuracy and cross-text block association in traditional knowledge graph construction is solved, and efficient and accurate knowledge graph construction and entity alignment are achieved to meet the needs of dynamic fields.

CN120296111AInactive Publication Date: 2025-07-11COMMUNICATION UNIVERSITY OF CHINA

Patent Information

Application Number
CN202510780524.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional knowledge graph construction methods are susceptible to word segmentation accuracy, lack the ability to globally correlate across text blocks, cannot dynamically capture potential associations across text blocks, and rely on manual intervention to adapt to the needs of dynamic fields.

Method used

Using large model and RAG technology, through data cleaning, text segmentation, vectorization, index establishment, local knowledge extraction and global knowledge extraction, combined with multi-dimensional reference digestion, mutual information calculation and semantic segmentation across text blocks are realized, and knowledge graphs are dynamically constructed.

Benefits of technology

It improves the accuracy of entity recognition and the integrity of cross-text block associations, reduces the error digestion rate, improves the recall rate and entity alignment accuracy of the knowledge graph, and adapts to the needs of the dynamic domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296111A_ABST
    Figure CN120296111A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge construction method and system based on a large model and an RAG technology, and belongs to the technical field of large models.According to the knowledge construction method and system based on the large model and the RAG technology, through cross-text-block mutual information calculation and RAG enhanced reasoning, the system can quantify statistical correlation between entities, and in combination with context information of an external knowledge base, potential incidence relations of cross-paragraphs or documents are mined, so that the knowledge construction efficiency is improved. A dynamic semantic segmentation strategy is adopted to ensure that semantics in text blocks are consistent, context continuity is maintained by retaining overlapped parts, multiple expressions of the same entity are comprehensively judged through a multi-dimensional anaphora resolution mechanism, the error resolution rate is reduced, dynamic segmentation and mixed retrieval are combined, and the problem of large model input length limitation is solved. And through semantic retrieval and keyword matching complementation, the recall rate is improved, and the key problems of semantic fracture, cross-text association missing, low entity alignment precision and the like in traditional knowledge construction are remarkably solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models, and in particular relates to a knowledge construction method and system based on large models and RAG technology. Background Art

[0002] Knowledge graphs organize the knowledge of the human world in an orderly manner in the form of triples to achieve effective organization of information. However, since many sources of knowledge are mostly unstructured, there are many problems in automatically constructing knowledge graphs, so they need to be optimized.

[0003] First of all, data cleaning and preprocessing are crucial to ensure data quality and accuracy. In addition, in technical fields such as entity recognition, relationship extraction, and attribute extraction, researchers have adopted a variety of methods, such as entity recognition based on text mining, association relationship extraction based on graph neural networks, etc. Through technologies such as entity disambiguation, relationship update, and graph fusion, the ambiguity between referents such as entities, relationships, and attributes and factual objects can be eliminated to form a high-quality knowledge base. However, the complexity and dynamism of domain graphs have brought many challenges to existing technologies. The construction of knowledge graphs involves structured, semi-structured, and unstructured data. A series of automated and semi-automated technical means are used to extract knowledge facts from original databases and third-party databases and store them in knowledge bases. During the entire knowledge graph construction process, the knowledge graph needs to be continuously optimized and updated to maintain its timeliness and accuracy. Therefore, the traditional method of constructing knowledge graphs has the following defects:

[0004] Traditional methods rely on Chinese word segmentation, which is easily affected by the accuracy of word segmentation, resulting in entity recognition errors; and only focus on entities within the current sentence, lacking the ability to globally associate across text blocks; relationship extraction is limited to predefined relationship tables or local sentences, and cannot dynamically capture potential associations across text blocks, especially in long documents, which is prone to information omissions due to semantic mixing. When using a fixed window to segment text, mechanical segmentation may truncate complete semantic units (such as sentences or paragraphs), destroy contextual coherence, and reduce retrieval accuracy. Traditional methods rely on a single similarity indicator, which makes it difficult to integrate text, structure, and semantic features, resulting in limited accuracy in reference resolution. In addition, knowledge graph updates rely on manual intervention, lack an automated mechanism to handle complex relationships and real-time data, and are difficult to adapt to dynamic domain needs. Summary of the invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present application provides a knowledge construction method and system based on large models and RAG technology.

[0006] In the first aspect, the present application proposes a knowledge construction method based on a large model and RAG technology, including knowledge cleaning and knowledge construction stages:

[0007] The knowledge cleaning stage includes:

[0008] Obtain text data from multiple sources, clean, format-convert, and verify the text data, and then store it in a text database;

[0009] Set parameters and split large segments of text in the text database into text data blocks with semantic meanings according to the delimiter priority;

[0010] Select an embedded model to convert the text data blocks into text vectors, perform normalization processing, and then store them in a vector database;

[0011] Build a vector index, combine similarity retrieval and full-text retrieval, and optimize the retrieval strategy through hybrid search;

[0012] The knowledge construction stage includes:

[0013] Use a large model to perform local knowledge extraction, and extract entities, attributes, and relationships from the text data blocks;

[0014] Adopt the mutual information method to construct relevant entity pairs across the text data blocks, and combine the RAG technology and the retrieval strategy to perform global knowledge extraction to generate entity relationships;

[0015] Judge whether two entities refer to the same physical object through the multi-dimensional anaphora resolution method, merge the same entities to form a knowledge base, and store the constructed knowledge.

[0016] In some embodiments, storing the text data in the text database after performing cleaning operations, format conversion operations, and verification operations on the text data includes:

[0017] The cleaning operation is: remove HTML tags, emojis, garbled characters, and extra spaces, while retaining technical terms, numbers, and date information;

[0018] The format conversion operation is: uniformly encode the cleaned text into UTF-8 or GBK format, and convert it into plain text or rich text format;

[0019] The verification operation is: after cleaning and conversion are completed, verify the data. The verification content includes text length, whether key information is retained, and whether the encoding is correct. If data anomalies are found, re-perform the cleaning operation and format conversion operation.

[0020] In some embodiments, setting parameters and splitting large segments of text in the text database into text data blocks with semantic meanings according to the delimiter priority includes:

[0021] Set the maximum block length threshold: chunk_size and the overlapping retention length: overlap_size, and create a staging block and a final set to store the segmented text data blocks;

[0022] Read each character in sequence, add the character to the staging block, and check in real-time whether the length of the staging block exceeds chunk_size. Process according to the separator priority, and the separator priority is: line break > end-of-sentence punctuation > semicolon / comma;

[0023] When a line break is encountered, immediately add the current staging block to the final set, clear the current staging block, and retain the last overlap_size characters as the starting content of the new staging block;

[0024] When the length of the staging block ≥ chunk_size, search forward for the nearest end-of-sentence punctuation. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new staging block, and retain overlap_size characters as the overlap;

[0025] When the length of the staging block ≥ 1.2 * chunk_size, search forward for the nearest semicolon / comma. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new staging block, and retain overlap_size characters as the overlap;

[0026] When the length of the staging block ≥ 1.5 * chunk_size and there is no separator, force a split at chunk_size, add the first segment to the final set, and retain the second segment as the new staging block;

[0027] After the traversal ends, if the current staging block is not empty, directly add it to the final set. If the length exceeds 2 * chunk_size, record a warning log.

[0028] In some embodiments, the selected embedded model converts the text data blocks into text vectors, normalizes them, and stores them in a vector database, including:

[0029] Use the segmented text data blocks as input, and each text data block is represented as

[0030] , where represents the th word or character;

[0031] For each word , generate its corresponding word vector through a word embedding model , and the word embedding model maps words to a low-dimensional continuous vector space, so that words with similar semantics are close in the vector space;

[0032] For the entire text data block Generate text vectors by means of text vectorization ;

[0033] Normalize the generated text vectors and store them in a vector database after normalization.

[0034] In some embodiments, for establishing the vector index, a combined similarity retrieval and full-text retrieval is adopted and the retrieval strategy is optimized through hybrid search, including:

[0035] Store the vectorized data in a vector database and establish an index for fast retrieval. The vector database is used to cooperate with knowledge construction by adopting similarity retrieval, full-text retrieval or hybrid search strategy. The process of establishing the index is as follows:

[0036] Vector storage: Store the vectors of each text data block in the database and add metadata to it, including text content, source and timestamp;

[0037] Index construction: Select the corresponding index structure according to the characteristics of the vector database;

[0038] The similarity retrieval is: By calculating the similarity scores between the query vector and all vectors in the database, return the record with the highest score;

[0039] The full-text retrieval is: Establish an inverted index through keywords, and perform full-text retrieval through keywords during retrieval to find the corresponding records;

[0040] The hybrid search strategy is: Define multiple retrievers, query respectively using similarity retrieval and full-text retrieval to obtain their respective retrieval results, and use the RRF algorithm to re-rank the retrieval results to obtain the final records.

[0041] In some embodiments, for extracting local knowledge using a large model, entities, attributes and relationships are extracted from the text data block, including:

[0042] Set prompt examples in the prompt words, and guide the large model to extract entities, attributes and relationships from the text data block through the prompt examples. The setting elements of the prompt examples include role definition instructions, input-output specifications, term extraction rules and relationship generation constraints.

[0043] In some embodiments, for constructing relevant entity pairs across the text data block using the mutual information method and combining with RAG technology and retrieval strategy for global knowledge extraction to generate entity relationships, including:

[0044] Mutual information calculation: Calculate the mutual information between entities in different text blocks. The formula is:

[0045]

[0046] Among them, and represent the entity sets in two text blocks, represents the entity and the probability of simultaneous occurrence, and respectively represent the probabilities of the entities and appearing alone. By calculating the mutual information, it is judged whether the entities in different text blocks are related;

[0047] Use the RAG system to retrieve the context knowledge related to the entity pair from the vector database;

[0048] Based on the retrieved knowledge fragments, combined with the specific requirements of the task, the relationship between entities is generated through a pre-trained large language model to complete global knowledge extraction.

[0049] In some embodiments, the method for judging whether two entities refer to the same physical object through the multi-dimensional anaphora resolution method, and merging the same entities to form a knowledge base and storing the constructed knowledge includes:

[0050] Text similarity calculation: Calculate the text similarity of entities in different text blocks through cosine similarity or edit distance;

[0051] Structural similarity calculation: Analyze the structural relationship of entities in the knowledge graph to judge whether the entities have similar structures.

[0052] Semantic similarity calculation: Calculate the semantic similarity of entities through a large model to judge whether the entities have the same semantics.

[0053] Entity alignment and merging: According to text, structure, and semantic similarities, judge whether two entities refer to the same physical object. If so, merge them to obtain the entity alignment result.

[0054] In a second aspect, the present application proposes a knowledge construction system based on a large model and RAG technology, including a data cleaning module, a text segmentation module, a vector conversion module, a vector indexing module, a local knowledge extraction module, a global knowledge extraction module, and an anaphora resolution module;

[0055] The data cleaning module is used to obtain text data from multiple sources, clean, format-convert, and verify the text data, and then store it in the text database;

[0056] The text segmentation module is used to set parameters and split large chunks of text in the text database into text data blocks with semantic meanings according to the separator priority;

[0057] The vector conversion module is used to select an embedding model to convert the text data blocks into text vectors, perform normalization processing, and store them in the vector database;

[0058] The vector indexing module is used to establish a vector index, combine similarity retrieval and full-text retrieval, and optimize the retrieval strategy through hybrid search;

[0059] The local knowledge extraction module is used to perform local knowledge extraction using a large model, and extract entities, attributes, and relationships from the text data blocks;

[0060] The global knowledge extraction module is used to construct relevant entity pairs across the text data blocks using the mutual information method, and perform global knowledge extraction in combination with the RAG technology and the retrieval strategy to generate entity relationships;

[0061] The coreference resolution module is used to determine whether two entities refer to the same physical object through a multi-dimensional coreference resolution method, merge the same entities to form a knowledge base, and store the constructed knowledge.

[0062] In a third aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0063] In a fourth aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0064] Advantages of the present invention:

[0065] Through cross-text block mutual information calculation and RAG-enhanced reasoning, the system can quantify the statistical correlation between entities, combine the context information of the external knowledge base, mine potential association relationships across paragraphs or documents, adopt a dynamic semantic segmentation strategy to ensure the internal semantic consistency of the text blocks, maintain context continuity by retaining overlapping parts, and through a multi-dimensional coreference resolution mechanism, comprehensively determine multiple expressions of the same entity, reduce the error resolution rate, combine dynamic segmentation and hybrid retrieval to solve the problem of the input length limit of the large model, and improve the recall rate through the complementarity of semantic retrieval and keyword matching. The present invention significantly solves key problems such as semantic breakage, lack of cross-text association, and low entity alignment accuracy in traditional knowledge construction. Description of the Drawings

[0066] Figure 1This is the overall flowchart of the present invention.

[0067] Figure 2 This is the system principle block diagram of the present invention. Detailed implementation manners

[0068] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein; on the contrary, these embodiments are provided so that the present invention can be understood more thoroughly and the scope of the present invention can be fully communicated to those skilled in the art.

[0069] In a first aspect, the present application proposes a knowledge construction method based on large models and RAG technology, as Figure 1 shown, including a knowledge cleaning and a knowledge construction stage:

[0070] The knowledge cleaning stage includes:

[0071] S100: Obtain text data from multiple sources, clean, format-convert, and verify the text data, and then store it in a text database;

[0072] In some embodiments, storing the text data in the text database after performing the cleaning operation, format conversion operation, and verification operation on the text data includes:

[0073] The cleaning operation is: removing HTML tags, emojis, garbled characters, and extra spaces, while retaining technical terms, numbers, and date information;

[0074] The format conversion operation is: uniformly encoding the cleaned text into UTF-8 or GBK format, and converting it into plain text or rich text format;

[0075] The verification operation is: after cleaning and conversion are completed, verifying the data, and the verification content includes text length, whether key information is retained, and whether the encoding is correct. If data anomalies are found, the cleaning operation and format conversion operation are performed again.

[0076] Data preprocessing is the first step in knowledge construction. Its purpose is to obtain the required text data from multiple sources (such as web pages crawled by web crawlers, emails, social media posts, etc.), and clean and standardize these data to ensure the integrity and accuracy of the data. The specific steps are as follows:

[0077] Obtain text data from multiple sources, including but not limited to web pages, emails, social media posts, news reports, contract texts, etc. These data may contain different formats and encodings, so they need to be uniformly processed;

[0078] Convert the cleaned text into a unified format and encoding. Common text formats include plain text (.txt), rich text (.rtf), etc., and encoding methods can choose UTF-8, GBK, etc. Select an appropriate text encoding to reduce garbled characters and incompatibility issues, facilitating subsequent data analysis and mining;

[0079] After cleaning and conversion, verify the data to ensure its integrity and accuracy. The verification content includes text length, whether key information is retained, whether the encoding is correct, etc. If data anomalies are found, cleaning and conversion need to be performed again.

[0080] Clean the obtained text, removing special symbols (such as HTML tags, emojis, garbled characters, etc.), extra spaces, line breaks, etc., to make the text more standardized. During the cleaning process, key information in the text, such as technical terms, numbers, dates, etc., needs to be retained to ensure that these contents are not accidentally deleted or split;

[0081] Store the processed text data in a database for subsequent knowledge extraction and knowledge graph construction. Common databases include relational databases (such as MySQL, PostgreSQL) and non-relational databases (such as MongoDB, Elasticsearch).

[0082] S200: Set parameters and split the large text in the text database into text data blocks with semantic meaning according to the delimiter priority;

[0083] In some embodiments, the setting parameters and splitting the large text in the text database into text data blocks with semantic meaning according to the delimiter priority includes:

[0084] Set the maximum block length threshold: chunk_size and the overlapping retention length: overlap_size, and create a staging block and a final set for storing the split text data blocks;

[0085] Read each character in sequence, add the character to the staging block, and check in real time whether the length of the staging block exceeds chunk_size, and process according to the delimiter priority, where the delimiter priority is: line break > end punctuation > semicolon / comma;

[0086] When a line break is encountered, immediately add the current staging block to the final set, clear the current staging block, and retain the last overlap_size characters as the starting content of the new staging block;

[0087] When the length of the temporary block ≥ chunk_size, search forward for the nearest end-of-sentence punctuation. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new temporary block, and retain overlap_size characters as the overlap;

[0088] When the length of the temporary block ≥ 1.2 * chunk_size, search forward for the nearest semicolon / comma. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new temporary block, and retain overlap_size characters as the overlap;

[0089] When the length of the temporary block ≥ 1.5 * chunk_size and there are no delimiters, force a split at chunk_size, add the first segment to the final set, and retain the second segment as the new temporary block;

[0090] After the traversal, if the current temporary block is not empty, add it directly to the final set. If the length exceeds 2 * chunk_size, record a warning log.

[0091] Text segmentation is the process of splitting a large text into smaller, semantically meaningful chunks for subsequent vectorization and knowledge extraction. The specific steps are as follows:

[0092] Set chunk_size (the maximum block length threshold) and overlap_size (the length of the retained overlap, recommended to be 10 - 20% of chunk_size). Create a temporary block and a final set to store the segmented text chunks;

[0093] Read each character in order, add the character to the temporary block, and check in real-time whether the length of the temporary block exceeds chunk_size. Process according to the delimiter priority;

[0094] When a newline character is encountered, immediately add the temporary block (including the newline character) to the final set, clear the temporary block, and retain the last overlap_size characters as the starting content of the new temporary block;

[0095] When the length of the temporary block ≥ chunk_size, search forward for the nearest end-of-sentence punctuation. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new temporary block, and retain overlap_size characters as the overlap;

[0096] When the length of the temporary block ≥ 1.2 * chunk_size (allowing moderate overlength), search forward for the nearest semicolon / comma. If found, split after the punctuation, add the first segment to the final set, retain the second segment as the new temporary block, and retain overlap_size characters as the overlap;

[0097] When the length of the temporary storage block ≥ 1.5 * chunk_size and there are no delimiters, it is forced to split at chunk_size. The first segment is added to the final set, and the second segment is retained as a new temporary storage block (no overlap is retained);

[0098] After the traversal ends, if the temporary storage block is not empty, it is directly added to the final set. If the length exceeds 2 * chunk_size, a warning log needs to be recorded;

[0099] For paired symbols such as quotation marks / brackets, integrity must be maintained. Content that cannot be split, such as technical terms, numbers, and dates, must be retained as a whole. When Chinese and Western languages are mixed, the length is uniformly calculated according to Chinese characters.

[0100] Suppose there is a text containing enterprise contract terms with a length of 5000 characters. First, set chunk_size to 1000 characters and overlap_size to 150 characters. Then, traverse the text character by character, read each character in sequence, add the character to the temporary storage block, and check in real-time whether the length of the temporary storage block exceeds 1000 characters. When a newline character is encountered, the temporary storage block (including the newline character) is immediately added to the final set, the temporary storage block is cleared, and the last 150 characters are retained as the starting content of the new temporary storage block. When the length of the temporary storage block ≥ 1000 characters, search forward for the nearest end-of-sentence punctuation. If found, split after the punctuation. The first segment is added to the final set, and the second segment is retained as a new temporary storage block, and 150 characters are retained as the overlap. If the end-of-sentence punctuation is not found, then check for semicolons / commas. When the length of the temporary storage block ≥ 1200 characters, search forward for the nearest semicolon / comma. If found, split after the punctuation. The first segment is added to the final set, and the second segment is retained as a new temporary storage block, and 150 characters are retained as the overlap. When the length of the temporary storage block ≥ 1500 characters and there are no delimiters, it is forced to split at 1000 characters. The first segment is added to the final set, and the second segment is retained as a new temporary storage block (no overlap is retained). After the traversal ends, if the temporary storage block is not empty, it is directly added to the final set. If the length exceeds 2000 characters, a warning log needs to be recorded. Through the above steps, we have successfully split the 5000-character contract terms text into multiple semantically meaningful blocks for subsequent vectorization and knowledge extraction.

[0101] S300: Select an embedded model to convert the text data block into a text vector, perform normalization processing, and store it in a vector database;

[0102] In some embodiments, the selecting the embedded model to convert the text data block into a text vector, perform normalization processing, and store it in a vector database includes:

[0103] Taking the segmented text data block as an input, each text data block is represented as

[0104] , where Indicates the th word or character;

[0105] For each word , generate its corresponding word vector through a word embedding model . The word embedding model maps words to a low-dimensional continuous vector space, such that words with similar semantics are close in the vector space;

[0106] For the entire text data block , generate a text vector through text vectorization ;

[0107] Normalize the generated text vector and store it in a vector database.

[0108] Select a suitable embedding model according to the specific application scenario. You can refer to the MTEB (Massive Text Embedding Benchmark) leaderboard and choose a model with excellent performance, such as bge-large or the E5 series. For a specific domain, you can choose an open-source model for fine-tuning or training from scratch to adapt to the domain-specific semantic features.

[0109] Text preprocessing: Before vectorization, further preprocess the text, including removing stop words, stemming, lemmatization, etc. For Chinese text, word segmentation is also required. The purpose of preprocessing is to reduce noise and improve the accuracy of vectorization.

[0110] Text embedding: Input the preprocessed text into an embedding model to generate the corresponding vector representation. The specific process is as follows:

[0111] Input text: Use the segmented text block as input. Each text block can be represented as , where indicates the th word or character.

[0112] Word embedding: For each word , generate its corresponding word vector through a word embedding model (such as Word2Vec, GloVe, BERT, etc.) . The word embedding model maps words to a low-dimensional continuous vector space, such that words with similar semantics are close in the vector space.

[0113] Text vectorization: For the entire text block , its vector representation can be generated in multiple ways . Common methods include:

[0114] Average pooling: Take the average of the word vectors of all words in the text to obtain the text vector .

[0115] Max pooling: Take the maximum value for each dimension to obtain the text vector .

[0116] Pooling based on attention mechanism: Calculate the weight of each word through the attention mechanism, and then perform weighted summation to obtain the text vector , where represents the attention weight of the

[0117] Vector normalization: Perform normalization on the generated text vector so that the vector is distributed on the unit sphere for subsequent similarity calculation. The normalization formula is: , where represents the vector Euclidean norm of

[0118] Vector storage: Store the generated text vector in a vector database for subsequent retrieval and comparison. Commonly used vector databases include FAISS, Chromadb, Elasticsearch, Milvus, etc. When storing, metadata such as the content, source, and timestamp of the text block can be added to the vector for subsequent query and analysis.

[0119] S400: Establish a vector index, adopt a combination of similarity retrieval and full-text retrieval, and optimize the retrieval strategy through hybrid search;

[0120] In some embodiments, the establishing of the vector index, adopting a combination of similarity retrieval and full-text retrieval and optimizing the retrieval strategy through hybrid search, includes:

[0121] Store the vectorized data in a vector database and establish an index for quick retrieval. The vector database is used to cooperate with knowledge construction using similarity retrieval, full-text retrieval, or hybrid search strategies. The process of establishing the index is as follows:

[0122] Vector storage: Store the vector of each text data block in the database and add metadata to it, including text content, source, and timestamp;

[0123] Index construction: Select the corresponding index structure according to the characteristics of the vector database;

[0124] The similarity retrieval is: By calculating the similarity score between the query vector and all vectors in the database, return the record with the highest score;

[0125] The full - text retrieval is as follows: An inverted index is established through keywords. During retrieval, full - text retrieval is performed through keywords to find the corresponding records;

[0126] The hybrid search strategy is as follows: Define multiple retrievers, use similarity retrieval and full - text retrieval respectively for querying to obtain their respective retrieval results, and use the RRF algorithm to re - rank the retrieval results to obtain the final records.

[0127] Vector retrieval is a key step in knowledge graph construction. Its goal is to quickly find the vector most similar to the query vector from the vector database through an efficient retrieval method. The following are the detailed implementation steps of vector retrieval:

[0128] Vector index establishment:

[0129] Index the vectorized data for quick retrieval. Common index methods include inverted index, HNSW, etc. HNSW is a graph - based index method that can efficiently process high - dimensional vector data.

[0130] The index establishment formula is as follows:

[0131]

[0132] Among them, represents the vector set, represents the index method, represents the established vector index.

[0133] Similarity retrieval:

[0134] Use cosine similarity to calculate the similarity between the query vector and the vectors in the database. Cosine similarity measures the similarity degree of two vectors in direction, and its value range is [-1, 1]. The larger the value, the higher the similarity.

[0135] The cosine similarity formula is as follows:

[0136]

[0137] Among them, and represent two vectors, represents the vector dot product, represents the length of the vector.

[0138] To improve the retrieval efficiency, an approximate nearest neighbor search algorithm can be adopted, which can greatly improve the retrieval speed on the premise of ensuring the retrieval accuracy.

[0139] Full - text retrieval:

[0140] Based on vector retrieval, combined with the full-text retrieval method, the recall rate is further improved. Full-text retrieval finds documents relevant to the query from text data through keyword matching.

[0141] The full-text retrieval formula is as follows:

[0142]

[0143] Where, represents the document, represents the query, represents the term in the query, represents the term in the document 's TF-IDF value.

[0144] Hybrid retrieval:

[0145] Combine similarity retrieval and full-text retrieval, and adopt the hybrid retrieval method to improve the precision and recall rate of retrieval. Hybrid retrieval obtains the final retrieval result by weighted fusion of the results of the two retrieval methods.

[0146] The hybrid retrieval formula is as follows:

[0147]

[0148] Where, represents the weight of similarity retrieval, represents the weight of full-text retrieval, represents the score of similarity retrieval, represents the score of full-text retrieval.

[0149] Retrieval result sorting:

[0150] Sort the retrieval results so that users can quickly find the most relevant results. The sorting method can be comprehensively sorted according to various factors such as retrieval score, timestamp, user preference, etc.

[0151] The sorting formula is as follows:

[0152]

[0153] Where, represents the weight of the document and can be adjusted according to factors such as the importance of the document and user preference.

[0154] Suppose there is a query vector , and we need to retrieve the vector most similar to it from the vector database. First, use the HNSW indexing method to build a vector index. Then, calculate the query vector The similarity with each vector in the database is obtained to get a similarity score . At the same time, using the full-text retrieval method, the relevance between the query keywords and the document is calculated by TF-IDF to obtain a full-text retrieval score . Then, the results of similarity retrieval and full-text retrieval are weighted and fused to obtain a final retrieval score . Finally, the results are sorted according to the retrieval score, and the most relevant retrieval results are output. Through this series of operations, the vector most similar to the query can be retrieved from the vector database efficiently, improving the precision and recall rate of the retrieval.

[0155] The knowledge construction stage includes:

[0156] S500: Use a large model for local knowledge extraction to extract entities, attributes, and relationships from the text data block;

[0157] In some embodiments, the use of a large model for local knowledge extraction to extract entities, attributes, and relationships from the text data block includes:

[0158] Set prompt examples in the prompt words, and guide the large model to extract entities, attributes, and relationships from the text data block through the prompt examples. The setting elements of the prompt examples include role definition instructions, input-output specifications, term extraction rules, and relationship generation constraints.

[0159] By setting some examples in the prompt words, the large model can better understand which information to extract, or guide the large model to perform step-by-step reasoning during extraction, which also helps to improve the accuracy of the extraction results. The following gives prompt examples:

[0160] "You are a network diagram maker, extracting terms and their relationships from the given context."

[0161] "A context block is provided for you (separated by \n), and your task is to extract the ontology."

[0162] "Terms mentioned in the given context. These terms should represent key concepts according to the context.\n"

[0163] "Step 1: When traversing each sentence, think about the key terms mentioned in it.\n"

[0164] "\tTerms may include objects, entities, locations, organizations, people,\n"

[0165] "\tconditions, acronyms, documents, services, concepts, etc.\n"

[0166] "The terms should be as atomic as possible.\n\n"

[0167] "Step 2: Consider how these terms can establish one-to-one relationships with other terms.\n"

[0168] "\tTerms mentioned in the same sentence or paragraph are usually related to each other.\n"

[0169] "\tA term can be related to many other terms\n\n"

[0170] "Step 3: Identify the relationships between each pair of related terms.\n\n"

[0171] "Format the output as a json list, where each element in the list contains a pair of terms."

[0172] "And the relationship between them is as follows:\n"

[0173] "[\n"

[0174] "{\n"

[0175] '"node_1": "The concept extracted from the extracted ontology",\n'

[0176] '"node_2": "The related concept extracted from the extracted ontology",\n'

[0177] '"edge": "The relationship between the two concepts of node_1 and node_2 in one or two sentences."\n'

[0178] "}, {...}\n"

[0179] "]"

[0180] S600: Use the mutual information method to construct related entity pairs across the text data blocks, and combine RAG technology and retrieval strategies to perform global knowledge extraction to generate entity relationships;

[0181] In some embodiments, the using the mutual information method to construct related entity pairs across the text data blocks and combining RAG technology and retrieval strategies to perform global knowledge extraction to generate entity relationships includes:

[0182] Mutual information calculation: Calculate the mutual information between entities in different text blocks, and the formula is:

[0183]

[0184] Where and represent the entity sets in two text blocks, represents the entity and the probability of simultaneous occurrence, and respectively represent the entities and the probabilities of separate occurrences, and by calculating the mutual information, it is determined whether the entities in different text blocks are related;

[0185] Use the RAG system to retrieve the context knowledge related to the entity pair from the vector database;

[0186] Based on the retrieved knowledge fragments, combined with the specific requirements of the task, the relationships between entities are generated through a pre-trained large language model to complete global knowledge extraction.

[0187] Mutual information is an index in information theory used to measure the degree of mutual dependence between two variables. That is, the larger the value of mutual information, the stronger the correlation between the two variables. Suppose an example of weather and whether to bring an umbrella is selected. Suppose there are two random variables X and Y, where X represents the weather (0 represents sunny, 1 represents rainy), and Y represents whether to bring an umbrella (0 represents not bringing, 1 represents bringing). Suppose 10 days of data are observed, and the following joint frequency distribution is obtained:

[0188] When X = 0 (sunny), Y = 0 for 4 days and Y = 1 for 1 day;

[0189] When X = 1 (rainy), Y = 0 for 1 day and Y = 1 for 4 days.

[0190] In this way, there are a total of 10 days. Then the joint probability distribution can be calculated as:

[0191] P(X = 0, Y = 0) = 4 / 10 = 0.4;

[0192] P(X = 0, Y = 1) = 1 / 10 = 0.1;

[0193] P(X = 1, Y = 0) = 1 / 10 = 0.1;

[0194] P(X = 1, Y = 1) = 4 / 10 = 0.4;

[0195] Next, calculate the marginal probabilities:

[0196] P(X = 0) = 0.4 + 0.1 = 0.5;

[0197] P(X = 1) = 0.1 + 0.4 = 0.5;

[0198] P(Y = 0) = 0.4 + 0.1 = 0.5;

[0199] P(Y = 1) = 0.1 + 0.4 = 0.5;

[0200] Then, calculate the terms corresponding to each (x,y) according to the mutual information formula and then sum them up.

[0201] Now, calculate each term:

[0202] For X = 0, Y = 0:

[0203] P(x,y) = 0.4, P(x)P(y) = 0.5 * 0.5 = 0.25;

[0204] log2(0.4 / 0.25) = log2(1.6) ≈ 0.678;

[0205] So the term is: 0.4 * 0.678 ≈ 0.2712; For X = 0, Y = 1:

[0206] P(x,y) = 0.1, P(x)P(y) = 0.25;

[0207] 0.1 / 0.25 = 0.4, log2(0.4) = -1.3219;

[0208] The term is: 0.1 * (-1.3219) ≈ -0.1322; For X = 1, Y = 0:

[0209] Similarly, P(x,y) = 0.1, P(x)P(y) = 0.25, so like above, the term is: -0.1322; For X = 1, Y = 1:

[0210] P(x,y) = 0.4, P(x)P(y) = 0.25, so like the case of X = 0, Y = 0, the term is 0.2712,

[0211] Sum these up: 0.2712 - 0.1322 - 0.1322 + 0.2712 ≈ 0.2712 * 2 + (-0.1322) * 2 = (0.5424) - (0.2644) = 0.278

[0212] So the mutual information in the above example is approximately 0.278 bits. This indicates that there is a certain correlation between X and Y. In this example, the probability is higher when X and Y are the same, so the mutual information should be a positive value.

[0213] I. RAG Enhanced Inference Process: Combining Retrieval and Generation for Relationship Prediction, Process Design:

[0214] Step 1: Cross - text - block Entity Retrieval:

[0215] Filter high - correlation entity pairs through mutual information (even if they are not in the same text block). For example, calculate the MI value of entity A and entity B. If it exceeds the threshold, trigger the retrieval.

[0216] Step 2: Context Enhancement Retrieval:

[0217] For entity pairs with high MI values, retrieve all relevant text chunks (such as paragraphs, documents) containing these entities from the knowledge base to form enhanced context.

[0218] Step 3: Prompt Design and Generation Constraints, Prompt Template:

[0219] Based on the following context, analyze the potential relationship between entity [Entity A] and [Entity B]:

[0220] [Retrieved Text Chunk 1]

[0221] [Retrieved Text Chunk 2]

[0222] Possible association types include: investment, cooperation, competition, kinship, etc.

[0223] Output Format: {"relation": "<type>", "evidence": "<key sentence>"}

[0224] Generation Constraints:

[0225] Restrict the LLM to only output in JSON format, and the relation type is selected from a predefined list;

[0226] Combine RAG technology for knowledge extraction, demonstrated with examples:

[0227] Input:

[0228] Text Chunk 1 (News A): "Company X announced its entry into the new energy vehicle field."

[0229] Text Chunk 2 (Earnings Report B): "The CEO of Company Y once led a battery technology R & D project."

[0230] Process:

[0231] 1. Calculate the MI value of "Company X" and "Company Y", assuming it exceeds the threshold.

[0232] 2. Retrieve the context containing both, and find that there may be a supply chain relationship between "Company Y's battery technology" and "Company X's new energy vehicles".

[0233] 3. LLM generates the result: {"relation": "Supply Chain Cooperation", "evidence": "Company X's new energy vehicles may rely on Company Y's battery technology"};

[0234] Through the above steps, we can not only extract the relationships between entities from multiple text blocks, but also combine the information in the external knowledge base to generate more accurate and rich relationships, thereby enhancing the content and depth of the knowledge graph and completing global knowledge extraction.

[0235] Mutual information is used to construct entity pairs across text blocks, and entity pairs that meet the set threshold are semantically inferred using the RAG technique. At the same time, large models have powerful reasoning capabilities and can infer and deduce by analyzing the relationships between entities. This reasoning ability can help uncover more associative relationships and implicit knowledge, enriching the content and depth of the knowledge graph. These advantages enable large models to play an important role in the construction and application of knowledge graphs.

[0236] S700: Determine whether two entities refer to the same physical object through a multi-dimensional coreference resolution method, merge the same entities to form a knowledge base, and store the constructed knowledge.

[0237] In some embodiments, the determining whether two entities refer to the same physical object through a multi-dimensional coreference resolution method, merging the same entities to form a knowledge base, and storing the constructed knowledge includes:

[0238] Text similarity calculation: Calculate the text similarity of entities in different text blocks through cosine similarity or edit distance;

[0239] Structural similarity calculation: Analyze the structural relationships of entities in the knowledge graph to determine whether the entities have similar structures.

[0240] Semantic similarity calculation: Calculate the semantic similarity of entities through a large model to determine whether the entities have the same semantics.

[0241] Entity alignment and merging: Based on text, structural, and semantic similarities, determine whether two entities refer to the same physical object. If so, merge them to obtain the entity alignment result.

[0242] Text similarity calculation:

[0243] Text similarity is used to measure the similarity degree of the text expressions of two entities. Common calculation methods include edit distance, which measures similarity by calculating the minimum number of edit operations required to convert one string into another. The smaller the edit distance, the more similar the text expressions of the two entities.

[0244] The edit distance formula is as follows:

[0245]

[0246] Where, and Represents two entity texts, , and represent the number of insert, delete, and replace operations respectively.

[0247] Structural similarity calculation:

[0248] Structural similarity is used to measure whether the structural relationships of two entities in the knowledge graph are similar. Common calculation methods include the Jaccard coefficient, which measures structural similarity by calculating the ratio of the intersection to the union of the neighbor sets of two entities.

[0249] The Jaccard coefficient formula is as follows:

[0250]

[0251] where and represent the neighbor sets of two entities respectively.

[0252] Semantic similarity calculation:

[0253] Semantic similarity is used to measure the semantic similarity degree of two entities. Common calculation methods include the knowledge extraction method based on RAG, which calculates the semantic similarity of two entities by retrieving relevant information from an external knowledge base and combining the pre-trained large language model to generate semantic representations.

[0254] The semantic similarity formula is as follows:

[0255]

[0256] where and represent two entities, represents the entity semantic representation generated by the RAG technology, represents the cosine similarity.

[0257] Multi-dimensional similarity fusion:

[0258] Weightedly fuse text similarity, structural similarity, and semantic similarity to obtain a comprehensive similarity score, which is used to judge whether two entities point to the same physical object.

[0259] The comprehensive similarity formula is as follows:

[0260] where , and represent the weights of text similarity, structural similarity, and semantic similarity respectively, and are usually adjusted according to specific application scenarios.

[0261] There are two entities and , and it is necessary to determine whether they point to the same physical object. First, calculate their text similarity using the edit distance formula to obtain the text similarity score. Then, calculate their structural similarity using the Jaccard coefficient formula to obtain the structural similarity score. Next, calculate their semantic similarity. Use the RAG-based knowledge extraction method to generate semantic representations and calculate the semantic similarity score through the cosine similarity formula . Finally, perform weighted fusion on the text similarity, structural similarity, and semantic similarity to obtain the comprehensive similarity score . If the comprehensive similarity score exceeds the preset threshold, it is considered that and point to the same physical object.

[0262] In a second aspect, the present application proposes a knowledge construction system based on large models and RAG technology. As Figure 2 shown, it includes a data cleaning module, a text segmentation module, a vector conversion module, a vector index module, a local knowledge extraction module, a global knowledge extraction module, and a coreference resolution module;

[0263] The data cleaning module is used to obtain text data from multiple sources, clean, format-convert, and verify the text data, and then store it in the text database;

[0264] The text segmentation module is used to set parameters and split the large text in the text database into text data blocks with semantic meanings according to the delimiter priority;

[0265] The vector conversion module is used to select an embedding model to convert the text data blocks into text vectors, perform normalization processing, and then store them in the vector database;

[0266] The vector index module is used to establish a vector index, combine similarity retrieval and full-text retrieval, and optimize the retrieval strategy through hybrid search;

[0267] The local knowledge extraction module is used to perform local knowledge extraction using a large model, and extract entities, attributes, and relationships from the text data blocks;

[0268] The global knowledge extraction module is used to construct relevant entity pairs across the text data blocks using the mutual information method, and combine RAG technology and retrieval strategies to perform global knowledge extraction and generate entity relationships;

[0269] The anaphora resolution module is used to determine whether two entities refer to the same physical object through a multi-dimensional anaphora resolution method, merge the same entities to form a knowledge base, and store the constructed knowledge.

[0270] In a third aspect, the present application proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.

[0271] In a fourth aspect, the present application proposes a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0272] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be repeated here.

[0273] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0274] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0275] In the embodiments provided in the present disclosure, it should be understood that the disclosed apparatus / computer device and method can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the apparatus or unit can be in electrical, mechanical or other forms.

[0276] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0277] In addition, in each embodiment of the present disclosure, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0278] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0279] The above are only the preferred embodiments of the present invention. It should be noted that for those skilled in the art, without departing from the technical solution of the present invention, several modified and improved technical solutions should also be regarded as falling within the scope protected by this solution.

Claims

1. A knowledge construction method based on large models and RAG technology, characterized in that: It includes a knowledge cleaning and a knowledge construction phase: The knowledge cleaning phase includes: Obtaining text data from multiple sources, cleaning, format-converting, and validating the text data, and then storing it in a text database; Setting parameters and splitting large segments of text in the text database into text data chunks with semantic meanings according to the separator priority; Selecting an embedding model to convert the text data chunks into text vectors, performing normalization processing, and then storing them in a vector database; Establishing a vector index, combining similarity retrieval and full-text retrieval, and optimizing the retrieval strategy through hybrid search; The knowledge construction phase includes: Using a large model to perform local knowledge extraction, extracting entities, attributes, and relationships from the text data chunks; Adopting the mutual information method to construct relevant entity pairs across the text data chunks, and combining with the RAG technology and retrieval strategy to perform global knowledge extraction to generate entity relationships; Judging whether two entities refer to the same physical object through the multi-dimensional anaphora resolution method, merging the same entities to form a knowledge base, and storing the constructed knowledge.

2. The method according to claim 1, characterized in that: The storing the text data in the text database after performing the cleaning operation, format conversion operation, and validation operation on the text data includes: The cleaning operation is: removing HTML tags, emojis, garbled characters, and extra spaces, while retaining technical terms, numbers, and date information; The format conversion operation is: uniformly encoding the cleaned text into UTF-8 or GBK format, and converting it into plain text or rich text format; The validation operation is: after cleaning and conversion are completed, validating the data, and the validation content includes text length, whether key information is retained, and whether the encoding is correct. If data anomalies are found, the cleaning operation and format conversion operation are performed again.

3. The method according to claim 2, wherein: The setting parameters and splitting large segments of text in the text database into text data chunks with semantic meanings according to the separator priority includes: Setting the maximum chunk length threshold: chunk_size and the overlapping retention length: overlap_size, and creating a temporary chunk and a final set for storing the split text data chunks; Reading each character in order, adding the character to the temporary chunk, and checking in real time whether the length of the temporary chunk exceeds chunk_size, and processing according to the separator priority, where the separator priority is: line break > end-of-sentence punctuation > semicolon / comma; When a line break is encountered, immediately add the current temporary chunk to the final set, clear the current temporary chunk, and retain the last overlap_size characters as the starting content of the new temporary chunk; When the length of the temporary chunk ≥ chunk_size, look forward to find the nearest end-of-sentence punctuation. If found, split after the punctuation, add the previous segment to the final set, retain the latter segment as the new temporary chunk, and retain overlap_size characters as the overlap; When the length of the temporary chunk ≥ 1.2 * chunk_size, look forward to find the nearest semicolon / comma. If found, split after the punctuation, add the previous segment to the final set, retain the latter segment as the new temporary chunk, and retain overlap_size characters as the overlap; When the length of the temporary storage block is ≥ 1.5 * chunk_size and there are no delimiters, it is forced to split at chunk_size. The first segment is added to the final set, and the second segment is retained as a new temporary storage block; After the traversal is completed, if the current temporary storage block is not empty, it is directly added to the final set. If the length exceeds 2 * chunk_size, a warning log is recorded.

4. The method according to claim 3, characterized in that: The selected embedded model converts text data blocks into text vectors, normalizes them, and stores them in the vector database, including: Using the segmented text data blocks as input, each text data block is represented as , where represents the th word or character; For each word , generate its corresponding word vector through a word embedding model , the word embedding model maps words to a low-dimensional continuous vector space, such that words with similar semantics are close in the vector space; For the entire text data block , generate text vectors through text vectorization ; The generated text vector is normalized and then stored in the vector database.

5. The method according to claim 4, characterized in that: The establishment of the vector index adopts a combination of similarity retrieval and full-text retrieval and optimizes the retrieval strategy through hybrid search, including: The vectorized data is stored in the vector database, and an index is established for quick retrieval. The vector database is used to cooperate with knowledge construction using similarity retrieval, full-text retrieval, or hybrid search strategies. The process of establishing the index is as follows: Vector storage: Store the vectors of each text data block in the database and add metadata to it, including text content, source, and timestamp; Index construction: Select the corresponding index structure according to the characteristics of the vector database; The similarity retrieval is: By calculating the similarity scores between the query vector and all vectors in the database, return the record with the highest score; The full-text retrieval is: Establish an inverted index through keywords and perform full-text retrieval through keywords during retrieval to find the corresponding records; The hybrid search strategy is: Define multiple retrievers, use similarity retrieval and full-text retrieval to query respectively, obtain their respective retrieval results, and use the RRF algorithm to re-rank the retrieval results to obtain the final records.

6. The method according to claim 5, wherein: The extraction of local knowledge using the large model extracts entities, attributes, and relationships from the text data blocks, including: Set prompt examples in the prompt words, and guide the large model to extract entities, attributes, and relationships from the text data blocks through the prompt examples. The setting elements of the prompt examples include role definition instructions, input-output specifications, term extraction rules, and relationship generation constraints.

7. The method according to claim 6, wherein: The construction of relevant entity pairs across the text data blocks using the mutual information method and the generation of entity relationships through global knowledge extraction in combination with the RAG technology and retrieval strategy, including: Mutual information calculation: Calculate the mutual information between entities in different text blocks. The formula is: Among them, and represent the entity sets in two text blocks, represents the entity and the probability of simultaneous occurrence, and respectively represent the probabilities of the entities and appearing alone. By calculating the mutual information, it is determined whether the entities in different text blocks are relevant; Use the RAG system to retrieve context knowledge related to the entity pair from the vector database; Based on the retrieved knowledge fragments, combined with the specific requirements of the task, generate the relationships between entities through a pre-trained large language model to complete global knowledge extraction.

8. The method according to claim 7, wherein: The judgment of whether two entities refer to the same physical object through the multi-dimensional anaphora resolution method, merging the same entities to form a knowledge base, and storing the constructed knowledge, including: Text similarity calculation: Calculate the text similarity of entities in different text blocks through cosine similarity or edit distance; Structural similarity calculation: Analyze the structural relationships of entities in the knowledge graph to judge whether the entities have similar structures; Semantic similarity calculation: Calculate the semantic similarity of entities through a large model to determine whether the entities have the same semantics; Entity alignment and merging: Based on text, structure, and semantic similarity, determine whether two entities refer to the same physical object. If so, merge them to obtain the entity alignment result.

9. A knowledge construction system based on large models and RAG technology, characterized in that: It includes a data cleaning module, a text segmentation module, a vector conversion module, a vector indexing module, a local knowledge extraction module, a global knowledge extraction module, and a coreference resolution module; The data cleaning module is used to obtain text data from multiple sources, clean, format-convert, and verify the text data, and then store it in a text database; The text segmentation module is used to set parameters and split the large text in the text database into text data blocks with semantic meanings according to the delimiter priority; The vector conversion module is used to select an embedding model to convert the text data blocks into text vectors, perform normalization processing, and then store them in a vector database; The vector indexing module is used to establish a vector index, combine similarity retrieval and full-text retrieval, and optimize the retrieval strategy through hybrid search; The local knowledge extraction module is used to extract local knowledge using a large model, and extract entities, attributes, and relationships from the text data blocks; The global knowledge extraction module is used to construct relevant entity pairs across the text data blocks using the mutual information method, and combine the RAG technology and the retrieval strategy to perform global knowledge extraction to generate entity relationships; The coreference resolution module is used to determine whether two entities refer to the same physical object through a multi-dimensional coreference resolution method, merge the same entities to form a knowledge base, and store the constructed knowledge.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Intelligent document retrieval generation method and system based on RAG technology

    CN118332072A

  • Entity relationship extraction method and system for construction of knowledge graph in financial field

    CN119378494A

  • Knowledge management system and method for constructing large voice model

    CN119476458A

  • Multi-source cross-domain data query method and system

    CN119691003A

  • Multi-modal document retrieval enhancement generation method based on large model

    CN119988588A

Cited By

  • Method and system for constructing guarantee knowledge think tank based on LLM and RAG technologies

    CN120706525A

  • Method and system for constructing a guarantee knowledge library based on LLM and RAG technologies

    CN120706525B

  • Knowledge question-answering method based on improved RAG and agent workflow

    CN120929577A

  • Tobacco certificate handling business text block dynamic segmentation method based on near-end strategy optimization

    CN121031539A

  • Tobacco licensing business text block dynamic segmentation method based on proximal policy optimization

    CN121031539B