RAG database construction method based on information compression and pruning

By compressing and pruning the text information in the RAG database, the problems of large storage overhead and poor adaptability of existing RAG databases are solved, and more efficient information storage and better answer quality are achieved.

CN120045634AActive Publication Date: 2025-05-27XIDIAN UNIV

Patent Information

Application Number
CN202510017185.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-27
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The existing RAG database has a high storage overhead and poor adaptability of text information to specific large language models, resulting in a low quality of generated answers.

Method used

The RAG database is constructed based on information compression and pruning. By compressing the sentence unit text information in the child nodes covered by its parent node in the cluster tree, and pruning the cluster tree after the text information of some node objects is compressed based on the QR decomposition method, filtering and storing the most valuable external information.

Benefits of technology

It effectively reduces the storage overhead of redundant information in the RAG database, improves the storage efficiency of the database and the retrieval efficiency of the large language model, and ensures the quality of the generated answers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045634A_ABST
    Figure CN120045634A_ABST
Patent Text Reader

Abstract

The invention provides an information compression and pruning RAG database construction method. The method comprises the implementation steps that node objects and a hierarchical clustering tree are constructed; compressing the node object text information in the clustering tree; and pruning the clustering tree after the partial node object text information is compressed based on a QR-like decomposition method. According to the method, the sentence unit text information in the sub-nodes in the clustering tree is compressed, the embedded vectors of the compressed sub-nodes are updated, and the most valuable external information can be screened and stored according to the actual requirements of the large language model, so that invalid data storage in an RAG database is reduced, and the reliability of the RAG database is improved. According to the method, text information stored in the RAG database is guaranteed to have gain all the time, meanwhile, pruning is carried out on a clustering tree obtained after partial node object text information is compressed based on a QR-like decomposition method, redundant child node objects with high semantic similarity are effectively recognized and deleted, and the storage overhead of the node objects in the RAG database is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence large language models, and relates to a method for constructing a retrieval-enhanced RAG database, and specifically to a method for constructing a retrieval-enhanced RAG database based on information compression and pruning, which can be used in the fields of natural language processing, information retrieval, and external knowledge base management of large language models. Background Art

[0002] With the rapid development of Large Language Models (LLMs), the technology based on Retrieval-Augmented Generation (RAG) has gradually become an important means to improve the reasoning ability of models. In the RAG system, the external knowledge base is usually stored as a vector database, and the relevant information is retrieved by similarity to provide additional context for the language model, so that it can refer to more information when generating answers, especially when the large model cannot directly obtain the answer from its parameters. RAG technology improves the performance of the model by combining external knowledge sources. Among them, how to efficiently build a RAG database is one of the core issues in the RAG technology of large language models.

[0003] The basic process of constructing the RAG database: (1) Text segmentation and vectorization: First, the text to be retrieved is segmented into several small sentence units, usually paragraphs or fixed-length sentence fragments, and then these sentence units are vectorized into fixed-length embedding vectors through the embedding model for subsequent retrieval and matching. These vectors are used to construct the RAG database, and each vector represents the semantic information of a sentence unit. (2) Database storage of sentence units and embedding vectors: The above-mentioned segmented sentence units and embedding vectors are stored in the database according to certain relationship rules or data structures to form a RAG database. Common RAG databases combine a single sentence unit and its corresponding embedding vector into the minimum query unit and store them directly in sequence in the database.

[0004] In order to improve the retrieval efficiency in large-scale text databases, for example, P Sarthi et al. proposed a RAG database construction method based on hierarchical iterative clustering tree RAPTOR in the paper "Recursive Abstractive Processing for Tree-organized Retrieval" in 2024. This method first divides the text into several sentence units, vectorizes each sentence unit into an embedding vector through an embedding model and combines it with the sentence unit into a node object; clustering algorithms are used to aggregate similar sentence units together, and these clusters are semantically summarized using a large language model to generate higher-level parent node sentence units. This process is iterated to gradually build a multi-level sentence unit tree structure. This method iteratively clusters node objects, so that the RAG database has a better data management structure and retrieval efficiency. However, when constructing the sentence unit tree structure, there is redundant information between the node sentence units, resulting in a large amount of similar or repeated text content in the database, which not only increases the storage overhead of the RAG database, but also affects the efficiency of the large language model when retrieving the RAG database and the quality of generating answers based on the retrieval results. Moreover, this method directly stores the sentence units segmented from the original text as node objects, resulting in poor adaptability of the text information in the RAG database to specific large language models. For certain specific fields of knowledge, the text retrieved from the RAG database contains less effective information required by the large language model, resulting in lower quality of generated answers. Summary of the invention

[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and propose a RAG database construction method based on information compression and pruning to solve the technical problems in the prior art of large storage overhead of RAG database and poor adaptability of text information to specific large language models.

[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0007] (1) Construct a node object:

[0008] Obtain N text data, each containing multiple sentences, and split each text data according to the separator to obtain M sentence units of the N text data, then vectorize each sentence unit through the embedding model, and then combine the embedding vector of each sentence unit after vectorization with the sentence unit into a node object, to obtain M node objects, where N ≥ 1000;

[0009] (2) Constructing a hierarchical clustering tree:

[0010] Perform S iterations of semantic clustering on the M node objects, and construct an S+1-layer clustering tree with the M node objects as leaf nodes and the G node objects composed of the summary sentence units of multiple grouped child nodes obtained by each iteration of semantic clustering and their corresponding embedding vectors as parent nodes;

[0011] (3) Compress the text information of node objects in the clustering tree:

[0012] Compress the sentence unit text information in the child nodes covered by their parent nodes in the clustering tree, and update the embedding vectors of the child nodes corresponding to the compressed sentence unit text information based on the embedding model to obtain a clustering tree after the text information of some node objects is compressed;

[0013] (4) Prune the clustering tree after compressing the text information of some node objects based on the QR-like decomposition method:

[0014] The multiple child node objects contained in each parent node in the clustering tree after the text information of some node objects is compressed are sorted in descending order of importance metric scores, and based on the quasi-orthogonal triangular QR decomposition method, the linear independence coefficient of the embedding vector of each sorted child node object contained in each parent node is calculated, and then the child node objects whose independence coefficient is lower than the independence threshold ∈ are pruned to obtain the RAG database containing the structure and information of the pruned clustering tree.

[0015] Compared with the prior art, the present invention has the following advantages:

[0016] 1. The present invention compresses the sentence unit text information in the child nodes covered by their parent nodes in the clustering tree, and updates the embedding vectors of the child nodes corresponding to the compressed sentence unit text information based on the embedding model, so as to screen and store the most valuable external information, avoid the problem of excessive invalid data storage in the database in the prior art, ensure that the text information stored in the RAG database is always beneficial, thereby improving the reasoning effect of the overall RAG system.

[0017] 2. The present invention prunes the clustering tree after compressing the text information of some node objects based on a QR-like decomposition method, effectively identifies and deletes redundant sub-node objects with high semantic similarity, reduces the storage overhead of node objects in the RAG database, and avoids the impact of the prior art on storage space occupancy and data management complexity due to the large amount of redundant information in the database. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is an implementation flow chart of the present invention. DETAILED DESCRIPTION

[0019] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0020] Reference Figure 1 , the present invention comprises the following steps:

[0021] Step 1) Build a node object:

[0022] (1a) Obtain 1,000 text data, each containing multiple sentences, and perform preliminary segmentation based on common delimiters including period ".", exclamation mark "!", question mark "?", and line break, ensuring that each sentence is an independent unit. This operation is completed using regular expressions, which can effectively handle different types of text delimiters; each sentence after segmentation will be encoded by the tokenizer of the embedding model BERT, and the number of block tokens encoded for each sentence will be calculated. The purpose is to determine the number of tokens occupied by each sentence, so as to facilitate the subsequent control of the maximum number of tokens in the text block;

[0023] (1b) Based on the maximum allowed number of tokens of 100, multiple sentences are gradually merged into one text block. If the number of tokens in the current text block does not exceed the upper limit after adding a certain sentence, the sentence will be added to the current text block; if adding the sentence will exceed the upper limit, the current text block will be saved and a new text block will be built. When the number of tokens in a sentence exceeds the maximum token limit, the sentence will be further split according to punctuation marks such as commas, semicolons, colons, etc. to ensure that the length of each part is within the allowed range. This process builds a suitable text block by gradually adding smaller clauses;

[0024] (1c) When constructing adjacent text blocks, the overlap parameter is set to retain part of the content at the end of the previous text block and copy this part to the beginning of the next text block. This can form a certain contextual association between text blocks and prevent the loss of important contextual information during the segmentation process. Finally, a list of 16,000 sentence units is returned. The length of each sentence unit is controlled within the specified maximum number of tokens, and the overlap between text blocks is set according to requirements;

[0025] (1d) Obtain the embedding model nomic-embed-text, which includes a cascaded tokenizer, an embedding layer, an encoder composed of multiple cascaded Transformer blocks, and an output layer. The Transformer block contains a cascaded self-attention network and two perceptron MLP feed-forward networks. The embedding layer consists of a word embedding matrix of size 30000×768, where 30000 represents the number of all subwords or word identifiers; the encoder consists of 12 cascaded Transformer blocks, and the output layer consists of one Transformer block;

[0026] (1e)nomic-embed-text The tokenizer of the input layer breaks down each sentence unit into 100 subwords or word identifiers, and then the embedding layer maps the subwords or word identifiers into 100 vectors with a fixed dimension of 768 through the word embedding matrix;

[0027] The multi-layer Transformer of the (1f)nomic-embed-text encoder layer uses the self-attention mechanism to perform context modeling on the 100 vectors of each sentence unit to obtain 100 context-related vectors. Then the input-output layer uses the feature dimensionality reduction technology to reduce the dimension of the context-related vectors to obtain an embedding vector with a fixed dimension of 768 dimensions.

[0028] (1g) Combining the vectorized embedding vector of each sentence unit with the sentence unit into a node object, obtaining 16,000 node objects;

[0029] Step 2) Construct a hierarchical clustering tree:

[0030] (2a) Initialize the number of iterations of semantic clustering to s, and set s = 1, and initialize the node object O 1 The 16,000 node objects obtained in step 1) are used as leaf nodes of the clustering tree;

[0031] Write a prompt template prompt_sum for the large language model to summarize the sentence unit:

[0032] "Write a summary of the following, including as many key details aspossible:{sentences}:"

[0033] Where {sentences} represents the text content set of sentence units of all child nodes under the cluster grouping.

[0034] (2b) Using uniform manifold approximation and projection UMAP to map the node object O 1The embedding vector of is reduced to 10 dimensions to alleviate the problem that the distance metric may perform poorly when measuring similarity in the high-dimensional space of the Gaussian mixture model GMM; then the Bayesian Information Criterion BIC is used for Gaussian model selection to determine the optimal number of clusters. BIC not only penalizes model complexity but also rewards goodness of fit. The BIC of the given GMM is Where T is the number of text segments (or data points), p is the number of model parameters, is the maximum value of the model likelihood function. In GMM, the number of parameters p is a function of the input vector dimension and the number of clusters. The node object O is clustered according to the optimal number of clusters and GMM clustering. 1 Clustered as G s Semantic groupings;

[0035] (2c) The prompt method of the large language model Qwen2 based on the summary sentence unit template summarizes all the sentence units in each group to obtain the summary sentence unit: The obtained summary sentence unit Sum:

[0036] Sum=RAG_LLM(prompt_sum,sentences)

[0037] Then, the same vectorization operation as in step 1) is performed to vectorize each summary sentence unit to obtain a summary embedding vector. Then, each summary sentence unit and its corresponding embedding vector are combined into a parent node object to obtain G. s parent node objects, and the obtained G s Node object O s As the s-th layer of the clustering tree.

[0038] (2d) Iterative semantic clustering finally obtains a five-layer clustering tree consisting of 16,000 node objects as leaf nodes, 695 parent nodes in the first layer, 156 parent nodes in the second layer, 23 parent nodes in the third layer, and 5 parent nodes in the fourth layer.

[0039] Step 3) Compress the node object text information in the clustering tree:

[0040] (3a) Use the large model to quantify the additional information provided in the child node text beyond the parent node summary text, using a score of 1-6, and write a prompt template based on the large language model to calculate the incremental score of one sentence unit text information relative to another sentence unit text information, prompt_add:

[0041] "Here is a summary of a long text and an excerpt from that text. Please evaluate whether the excerpt provides important information that is not included in the summary. Use an integer score from 1to 6to measure the amount of new information provided by the excerpt.

[0042] Summary:

[0043] {sentence2}

[0044] Excerpt:

[0045] {sentence1}

[0046] Please answer using Arabic numerals only 1,2,3,4,5,6:”

[0047] Where {sentence2} represents the content of the summary text in the parent node, and {sentence1} represents the text content of the child node; then write the prompt template prompt_com that compresses the text information:

[0048] "Here is a summary of a long text and an excerpt from that text. Carefully read both and identify the important additional information in the excerpt that is not mentioned in the summary. Concisely summarize this additional information.

[0049] Summary:

[0050] {sentence2}

[0051] Excerpt:

[0052] {sentence1}:”

[0053] (3b) The prompt method of the large language model Qwen2 based on the incremental score calculation prompt template calculates the incremental score of the sentence unit text information of each child node in the first 4 layers of the 5-layer clustering tree relative to the sentence unit text information of its parent node. The incremental score A of the kth child node contained in the jth parent node of the sth layer is sjk The calculation formula is:

[0054] A sjk =RAG_LLM(prompt_add,sentence sjk ,sentence sj )

[0055] Among them, sentence 1 Represents the child node sentence unit, sentence 2 Represents the parent node sentence unit;

[0056] (3c) The text of the parent node is the semantic summary of all child nodes. Therefore, by comparing the text of the child node with that of the parent node, it is possible to evaluate whether the child node provides additional information beyond the summary of the parent node. If most of the content of the child node has been completely covered by the parent node, it means that the text redundancy of the child node is high and is suitable for compression. Determine whether the incremental score of each child node is less than the incremental score threshold of 4. If so, the prompt method of the large language model Qwen2 based on the text information compression prompt template compresses the sentence unit text of the child node:

[0057] sentence sjk =RAG_LLM(prompt_com,sentence sjk ,sentence sj )

[0058] Otherwise, the sentence unit text of the node remains unchanged.

[0059] (3d) Perform the same vectorization operation as in step 1) and vectorize the sentence unit of the child node corresponding to the compressed sentence unit text information based on the embedding model nomic-embed-text, and replace the original embedding vector of the child node with the vectorized embedding vector with a dimension of 768.

[0060] (3e) The original embedding vector of the child node after text information compression is replaced with a compressed embedding vector of 768 dimensions to obtain a clustering tree after the text information of some node objects is compressed.

[0061] Step 4) Prune the clustering tree after compressing the text information of some node objects based on the QR-like decomposition method:

[0062] (4a) Use Qwen2 to measure the amount of information in the text of each child node and calculate the amount of new information that each child node can provide to the Qwen2 knowledge base. This information measurement model can identify unique content in child nodes, especially content that is not mentioned in the parent node summary, and knowledge that Qwen2 has not yet learned from other data. The more informative a child node is, the more important the text is to the retrieval generation task. The large model prompt method is used to achieve quantitative measurement of information (scoring 1-6 points). Write a prompt template prompt_imp based on the large language model prompt method to calculate the incremental score of the sentence unit text information relative to the large language model Qwen2 internal parameter knowledge base:

[0063] "I am providing you with a summary of a long text and a specific excerpt from that text. What is the article title of the original text? If you know the title of the article from which this excerpt is taken, meaning you are familiar with the content of the original article, please assess how much important information this excerpt provides that you did not already know from the original.

[0064] Use an integer score from 1 to 6 to measure the amount of new information provided by the excerpt.

[0065] If you do not know the title of the article,answer'6'.If this excerptprovides no important information to you,answer'1'.

[0066] Summary:

[0067] {sentence2}

[0068] Excerpt:

[0069] {sentence1}

[0070] Please answer using Arabic numerals only 1,2,3,4,5,6:”

[0071] (4b) The prompt method of the large language model Qwen2 based on the incremental score calculation prompt template calculates the incremental score of the sentence unit text information of the child nodes in the first 4 layers of the 5-layer clustering tree after the compression of the text information of some node objects relative to the internal parameter knowledge base of the large language model Qwen2, where the importance measurement score Z of the kth child node contained in the jth parent node of the sth layer is sjk The calculation formula is:

[0072] Z sjk =RAG_LLM(prompt_imp,sentence sjk ,sentence sj )

[0073] (4c) In order to make it easier to delete nodes containing less information, the sub-node objects under the same cluster group are arranged from large to small, and a clustering tree is obtained after the importance of the grouped sub-nodes is sorted.

[0074] (4d) Combine the embedding vectors of all the child node objects sorted by importance metric contained in each parent node in the clustering tree into the embedding vector matrix E under the parent node. sj =[r sj1 ,r sj2 ,...,r sjk ,...,r sjK ], among which, E sj represents the embedding vector matrix of the j-th parent node in the s-th layer in the sorted clustering tree, s∈[2,3,...,5], j∈[1,2,...,J], J represents the number of parent nodes in the s-th layer, k∈[1,2,...,K], K represents the number of child nodes contained in the j-th parent node in the s-th layer; r sjk represents the embedding vector of the kth child node contained in the jth parent node of the sth layer in the sorted clustering tree, so the dimension of the embedding matrix is ​​768×K, 768 is the dimension of each embedding vector; then the UMAP dimensionality reduction algorithm is used to reduce the dimension of the embedding vector in each embedding matrix to 10 dimensions, and the embedding vector matrix E with a dimension of 10×K is obtained. sj:

[0075] E sj =UMAP(E sj ,n_components=10)

[0076] Where n_components is the target dimension of the UMAP algorithm. If the dimension of the embedding vector is small, there is no need to reduce the dimension and the original embedding matrix can be used directly.

[0077] (4e) Initialize the orthogonal vector matrix Q of the j-th parent node in the s-th layer of the clustering tree sj =[q sj1 ,q sj2 ,...,q sjk ,...,q sjK ], let the first column of the Q matrix be equal to the first column of its embedding vector matrix r 1 ,q sj1 =r sj1 ; The linear independence coefficient of this vector is set to 1, L sj1 =1, where L sjk The linear independence coefficient of the embedding vector of the i-th child node of the j-th parent node in the s-th layer of the clustering tree is also the linear independence score of the node object;

[0078] (4f) Calculate the embedding vector r of the child node sjk The orthogonalized matrix Q at its parent node sj The projection vector on the first K' column vectors of , where K' represents the number of child nodes in the parent node that precede the child node, and then subtract the projection vector to obtain the residual vector q of the child node sjk :

[0079]

[0080] · represents the inner product operation, ||·|| represents the L1 norm, and ∑ represents the sum operation;

[0081] (4g) Calculate the ratio of the norm of the residual vector of the child node object to the norm of the embedding vector to obtain the linear independence coefficient of the embedding vector of the child node:

[0082]

[0083] The closer this ratio is to 1, the more linearly independent the current node object is from the previous node object in the vector space; the closer the ratio is to 0, the more redundant the content of the node object is from the previous node object. In the QR-like decomposition of vector linear independence, i.e., node redundancy evaluation, the vectors with the lower rankings map most of their own vector components to the projection of the previous vectors, and usually get a smaller L. Therefore, the nodes with less information metrics in step 4) are placed at the back to make them easier to be deleted in the subsequent pruning step.

[0084] (4h) The linear independence of each node object is determined according to the preset threshold value of 0.01. If the linear independence score of the node object is lower than 0.01, the node object is removed, otherwise the node object is retained. Then the RAG database containing the structure and information of the pruned clustering tree is obtained.

[0085] The effect of the present invention is further described below in conjunction with simulation experiments.

[0086] 1. Simulation conditions:

[0087] The hardware platform for the simulation experiment is: Intel(R)Core(TM)i5-8400 CPU@2.80GHz

[0088] 2.81GHz CPU, 32GB memory, NVIDIA GeForce RTX2080 graphics processor, software platform: Ubuntu 22.04 operating system, python3.12, pytorch2.4.1.

[0089] 2. Simulation content and result analysis:

[0090] The size and question-answering accuracy of the RAG database constructed by the present invention and the prior art on two typical RAG datasets, QuALITY dataset and QASPER dataset, are compared and simulated, and the results are shown in Table 1. Among them, the question-answering accuracy is the ratio of the number of correctly answered questions to the total number of questions. If and only if the answer output by the large language model is exactly the same as the true answer, the answer is recorded as correct. The question-answering accuracy can reflect the effectiveness of the text information stored in the RAG database. The database compression rate represents the ratio of the text length, the number of tokens, and the number of nodes stored in the RAG database reduced by the present invention compared with the prior art.

[0091] Table 1

[0092]

[0093] As can be seen from Table 1, the RAG database obtained by the present invention can effectively reduce the memory requirement of the database while ensuring the validity of the stored text information.

Claims

1. A method for constructing a RAG database with information compression and pruning, characterized in that: The steps include: (1) Construct a node object: Obtain N text data, each containing multiple sentences, and split each text data according to the separator to obtain M sentence units of the N text data, then vectorize each sentence unit through the embedding model, and then combine the embedding vector of each sentence unit after vectorization with the sentence unit into a node object, to obtain M node objects, where N ≥ 1000; (2) Constructing a hierarchical clustering tree: Perform S iterations of semantic clustering on the M node objects, and construct an S+1-layer clustering tree with the M node objects as leaf nodes and the G node objects composed of the summary sentence units of multiple grouped child nodes obtained by each iteration of semantic clustering and their corresponding embedding vectors as parent nodes; (3) Compress the text information of node objects in the clustering tree: Compress the sentence unit text information in the child nodes covered by their parent nodes in the clustering tree, and update the embedding vectors of the child nodes corresponding to the compressed sentence unit text information based on the embedding model to obtain a clustering tree after the text information of some node objects is compressed; (4) Prune the clustering tree after compressing the text information of some node objects based on the QR-like decomposition method: The multiple child node objects contained in each parent node in the clustering tree after the text information of some node objects is compressed are sorted in descending order of importance metric scores, and based on the quasi-orthogonal triangular QR decomposition method, the linear independence coefficient of the embedding vector of each sorted child node object contained in each parent node is calculated, and then the child node objects whose independence coefficient is lower than the independence threshold ∈ are pruned to obtain the RAG database containing the structure and information of the pruned clustering tree.

2. The method according to claim 1, characterized in that The embedding model described in step (1) includes a cascaded tokenizer, an embedding layer, an encoder formed by cascading multiple Transformer blocks, and an output layer, wherein the Transformer block includes a cascaded self-attention network and two perceptron MLP feed-forward networks.

3. The method according to claim 2, characterized in that In step (1), each sentence unit is vectorized through the embedding model, and the implementation steps are as follows: (1a) The tokenizer of the input layer breaks down each sentence unit into token_len subwords or word identifiers; the embedding layer maps each subword or word identifier to a vector of dimension D1 through the word embedding matrix; (1b) The multi-layer Transformer of the encoder layer uses the self-attention network to model the context of each vector and obtain token_len context-related vectors. The output layer reduces the dimension of the token_len context-related vectors obtained by modeling and obtains an embedding vector with a dimension of D2 corresponding to each sentence unit.

4. The method according to claim 1, characterized in that The implementation steps of performing S iterations of semantic clustering on the M node objects in step (2) are as follows: (2a) Initialize the number of iterations of semantic clustering to s, and set s=1, and initialize the node object O1 to M leaf node objects; (2b) Perform Gaussian mixture semantic clustering on node object O1 to obtain G s semantic groups, and summarize all sentence units in each semantic group based on the prompt method of the large language model RAG_LLM. The resulting summary sentence unit Sum is: Sum=RAG_LLM(prompt_sum,sentences) Among them, prompt_sum represents the summary prompt template, and sentences represents the text content of all sentence units in each semantic grouping; then each summary sentence unit obtained by the summary is vectorized, and then each summary sentence unit and its vectorized summary embedding vector are combined into a node object, and the obtained G s Node object O s As the sth layer of the clustering tree; (2c) Determine whether S = s. If so, use the G node objects obtained by S iterations of semantic clustering as the S+1 layer of the clustering tree. Otherwise, let s = s+1, O1 = O s , and execute step (2b), where ∑ represents the sum operation.

5. The method according to claim 1, characterized in that The steps of compressing the sentence unit text information in the child nodes covered by their parent nodes in the clustering tree described in step (3) are as follows: (3a) Based on the prompt method of the large language model RAG_LLM, the incremental score of the sentence unit text information of each child node covered by its parent node in the clustering tree relative to the sentence unit text information summarized by its parent node is calculated, where the incremental score A of the kth child node contained in the jth parent node of the sth layer is sjk The calculation formula is: A sjk =RAG_LLM(prompt_add,sentence sjk ,sentence sj ) Among them, prompt_add represents the prompt template that calculates the incremental score of the text information of a sentence unit relative to the text information of another sentence unit, sentence sjk Represents the kth child node sentence unit contained in the jth parent node of the sth layer, sentence sj Represents the j-th parent node sentence unit of the s-th layer; (3b) Determine whether A and a preset incremental score threshold C satisfy A<C. If so, compress the sentence unit text information of the child node based on the prompt method of the large language model RAG_LLM: sentence sjk =RAG_LLM(prompt_com,sentence sjk ,sentence sj ) Among them, prompt_com represents the prompt template that compresses the sentence unit text content, otherwise the sentence unit text information of the child node is retained unchanged.

6. The method according to claim 5, characterized in that The embedding model described in step (3) is used to update the embedding vector of the sub-node corresponding to the compressed sentence unit text information. The updating method is: Based on the embedding model, the sentence unit of the sub-node corresponding to the compressed sentence unit text information is vectorized, and the original embedding vector of the sub-node is replaced by the vectorized embedding vector with a dimension of D2.

7. The method according to claim 5, characterized in that In step (4), the multiple child node objects contained in each parent node in the clustering tree after the text information of the partial node objects is compressed are sorted in descending order according to the importance measurement scores, where the importance measurement score Z of the kth child node contained in the jth parent node of the sth layer is sjk The calculation formula is: Z sjk =RAG_LLM(prompt_imp,sentence sjk ,sentence sj ) Where prompt_imp represents the prompt template that calculates the incremental score of the sub-node sentence unit text information relative to the internal parameter knowledge base of the large language model RAG_LLM.

8. The method according to claim 7, characterized in that The linear independence coefficient of the embedding vector of each sorted child node object contained in each parent node is calculated based on the quasi-orthogonal triangular QR decomposition method described in step (4), where the linear independence coefficient L of the kth child node contained in the jth parent node of the sth layer is sjk The calculation formula is: Where s∈[2,3,...,S+1], j∈[1,2,...,J], J represents the number of parent nodes in the s-th layer, k∈[1,2,...,K], K represents the number of child nodes contained in the j-th parent node in the s-th layer; r sjk represents the embedding vector of the kth child node contained in the jth parent node of the sth layer in the sorted clustering tree, q sjk Represents the residual vector of the embedding vector of the child node after orthogonal projection, q sjk' represents the residual vector of the embedding vector of the k'th child node before this child node after orthogonal projection, K' represents the number of child nodes before this child node; · represents the inner product operation, ||·|| represents the L1 norm, and Σ represents the sum operation.

Citation Information

Patent Citations

  • Hierarchical data structure of documents

    US20150006528A1

Cited By

  • Image-text mixed output large model RAG retrieval method and system

    CN120763309A

  • Multi-model compression method and device, task processing method and equipment and storage medium

    CN120952083A