RAG text partitioning method based on graph embedding
By blocking text based on graph embedding, the problem of inaccurate text chunking in the prior art is solved, and the retrieval accuracy of the RAG system and the quality of large-model generation answers are improved.
Patent Information
- Application Number
- CN202510229488.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-13
AI Technical Summary
The existing text chunking methods have problems such as sentences or paragraphs being interrupted, context information being lost, and inaccurate chunking, which affects the search accuracy of the RAG system and the quality of the big model generating answers.
The RAG text blocking method based on graph embedding is adopted, and the text is segmented, sentence feature vectors are obtained, sentence similarity matrix is established, sentence relationship diagram is constructed, nodes are updated using neighborhood aggregation, and text blocks are finally formed.
It improves the accuracy of text chunking, enhances the search accuracy of the RAG system and the quality of large-scale answers generated, reduces redundant calculations, and is suitable for processing large-scale text.
Smart Images

Figure CN120146009A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and in particular to a RAG text chunking method based on graph embedding. Background Art
[0002] With the rapid development of Large Language Models, technologies based on Retrieval-Augmented Generation (RAG) have shown great potential in various application scenarios. The RAG system can provide more accurate and rich answers by retrieving relevant information from an external knowledge base and combining the generation capabilities of a large model. In this process, building a high-quality vector database is a key step. To build a high-quality vector database, it is necessary to vectorize a large amount of text. However, directly vectorizing long text poses many challenges, such as the dilution of semantic information, the consumption of computing resources, and the reduction of retrieval efficiency. Therefore, reasonably chunking long text becomes a prerequisite for building an efficient vector database. A suitable text chunking strategy can improve the retrieval accuracy, ensure that the retrieved content is more in line with the user's query intent, and thus improve the quality of the answers generated by the large model.
[0003] Among the existing text chunking methods, the fixed-size text chunking method is the simplest and most intuitive one. It divides the text into several chunks according to a pre-set fixed length. The NTLK-based chunking method is a widely used Python natural language processing library that provides rich text processing functions. The chunking method based on the Bert model designs a binary classification task during its pre-training process to enable the BERT model to learn the relationship between two sentences. Two sentences are input into BERT simultaneously to predict whether the second sentence is the next sentence of the first sentence.
[0004] However, the fixed-size text chunking method will break sentences or paragraphs, resulting in the loss of context information and affecting the subsequent text vectorization effect and semantic understanding. The NTLK-based chunking method does not provide pre-trained weights for the Chinese sentence segmentation model, and users need to train it themselves. For the BERT-based text segmentation method, when judging the segmentation point, only the two adjacent sentences before and after are considered, and the text information at a longer distance is not utilized, which may lead to inaccurate segmentation. The text chunks may contain incomplete sentences or ideas, affecting the matching accuracy in the subsequent retrieval stage of the RAG system and the quality of the answers generated by the large model for the text generation answer task. Summary of the Invention
[0005] To overcome the deficiencies in the above-mentioned prior art, such as inaccurate text segmentation, which affects the matching accuracy in the subsequent retrieval stage of the RAG system and the quality of the answers generated by the large model, the present invention proposes a RAG text chunking method based on graph embedding.
[0006] To achieve the above object, the present invention adopts the following technical solutions. A RAG text chunking method based on graph embedding includes:
[0007] S1: Segment the text to obtain each sentence after segmentation;
[0008] S2: Obtain the feature vectors of each sentence;
[0009] S3: Based on the feature vectors of each sentence, obtain the feature similarity between sentences and establish a similarity matrix between sentences;
[0010] S4: Based on the similarity matrix between sentences, construct a sentence relationship graph;
[0011] S5: Based on the sentence relationship graph, update the nodes by means of neighborhood aggregation to obtain the updated sentence relationship graph;
[0012] S6: Based on the updated sentence relationship graph, combine the sentences corresponding to the nodes to form each text chunk.
[0013] Preferably, in step S6, based on the updated sentence relationship graph, combining the sentences corresponding to the nodes to form each text chunk includes:
[0014] S61: Based on the updated sentence relationship graph, obtain the position of each sentence in the original text as the index of the node in the sentence relationship graph;
[0015] S62: Calculate the feature similarity between adjacent nodes;
[0016] S63: Take each node and the top t nodes with the greatest feature similarity to it as a group of combined nodes;
[0017] S64: Connect and combine the sentences corresponding to each group of combined nodes according to the node index, thereby forming each text chunk.
[0018] Preferably, in step S1, segmenting the text to obtain each sentence after segmentation includes: using the punctuation marks in the text as segmentation points, and segmenting the text based on the segmentation points to obtain each sentence after segmentation;
[0019] The regular expression used for segmentation is [a1, a2,..., at,..., aT];
[0020] Among them, the square brackets [] represent the set of punctuation characters, at is the t-th punctuation character to be matched, t is the punctuation character number, and T is the number of punctuation characters.
[0021] Preferably, the punctuation includes a period, a semicolon, an exclamation mark, and a question mark.
[0022] Preferably, in step S2, the feature vectors of each sentence are obtained, that is, each segmented sentence is input into the text embedding model, and the feature vectors of each sentence are output.
[0023] Preferably, in step S3, based on the feature vectors of each sentence, the feature similarity between sentences is obtained, and a sentence similarity matrix is established, including:
[0024] S31: Calculate the feature similarity between all sentences using the cosine similarity function;
[0025] S32: Use the feature similarity between all sentences as the element value of the matrix to establish a sentence similarity matrix S;
[0026] The construction formula of the sentence similarity matrix S is:
[0027]
[0028] Among them, x i represents the feature vector of the i-th sentence, x j represents the feature vector of the j-th sentence, S ∈ R n×n ; n is the total number of sentences, S ij represents the element value in the i-th row and j-th column of the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
[0029] Preferably, in step S4, based on the sentence similarity matrix, a sentence relationship graph is constructed using the top-k method, including:
[0030] Based on the sentence similarity matrix S, each sentence is regarded as a node, and the k nodes with the greatest similarity to the current node are found to construct edges, thereby constructing a sentence relationship graph and the corresponding sentence relationship adjacency matrix A;
[0031] The construction formula of the sentence relationship adjacency matrix A is:
[0032]
[0033] Among them, A ij represents the element value in the i-th row and j-th column of the adjacency matrix A, A ij= 1 indicates that the element value at the \(i\)-th row and \(j\)-th column in the adjacency matrix \(A\) is 1, that is, there is an edge between the \(i\)-th node and the \(j\)-th node in the sentence relationship graph, \(A\) ij = 0 indicates that the element value at the \(i\)-th row and \(j\)-th column in the adjacency matrix \(A\) is 0, that is, there is no edge between the \(i\)-th node and the \(j\)-th node in the sentence relationship graph; top - k(S[i:]) represents selecting the top \(k\) largest elements in the \(i\)-th row of the sentence similarity matrix \(S\), that is, selecting the \(k\) sentences with the highest feature similarity to the \(i\)-th sentence; \(S\) ij represents the element value at the \(i\)-th row and \(j\)-th column in the sentence similarity matrix \(S\), that is, the feature similarity between the \(i\)-th sentence and the \(j\)-th sentence.
[0034] Preferably, in step S4, based on the sentence similarity matrix, a sentence relationship graph is constructed, and a threshold - based method is adopted, including:
[0035] Based on the sentence similarity matrix \(S\), both the rows and columns in the sentence similarity matrix are regarded as nodes;
[0036] A similarity threshold is set, and the elements in the sentence similarity matrix whose values are greater than or equal to the similarity threshold are found as associated elements;
[0037] Edges are constructed between the nodes corresponding to the rows and columns of the associated elements, thereby forming a sentence relationship graph and the sentence relationship adjacency matrix \(A\) corresponding to the sentence relationship graph;
[0038] The construction formula for the sentence relationship adjacency matrix \(A\) is:
[0039]
[0040] A ij represents the element value at the \(i\)-th row and \(j\)-th column in the adjacency matrix \(A\), \(a\) is the similarity threshold, \(A\) ij = 1 indicates that the element value at the \(i\)-th row and \(j\)-th column in the adjacency matrix \(A\) is 1, that is, there is an edge between the \(i\)-th node and the \(j\)-th node in the sentence relationship graph, \(A\) ij = 0 indicates that the element value at the \(i\)-th row and \(j\)-th column in the adjacency matrix \(A\) is 0, that is, there is no edge between the \(i\)-th node and the \(j\)-th node in the sentence relationship graph; \(S\) ij represents the element value at the \(i\)-th row and \(j\)-th column in the sentence similarity matrix \(S\), that is, the feature similarity between the \(i\)-th sentence and the \(j\)-th sentence.
[0041] Preferably, in step S5, based on the sentence relationship graph, the nodes are updated by means of neighborhood aggregation to obtain the updated sentence relationship graph, including:
[0042] For each node in the sentence relationship graph, add the sentence feature vector corresponding to the current node to the sentence feature vectors corresponding to its neighbor nodes, and use the result as the node embedding representation corresponding to the current node, thereby updating all nodes to obtain an updated sentence relationship graph; the neighbor nodes refer to the nodes directly connected to the current node.
[0043] Preferably, the node embedding representation H corresponding to the current node i is calculated as follows:
[0044]
[0045] where i is the sentence number, i.e., the node number; N i represents the set of neighbor nodes of node i.
[0046] The advantages of the present invention are as follows:
[0047] (1) In the method of the present invention, a sentence relationship graph and a sentence relationship adjacency matrix are constructed, and the nodes are updated by means of neighborhood aggregation to obtain an updated sentence relationship graph; based on the updated sentence relationship graph, the sentences corresponding to the nodes are combined to form each text block, which increases the accuracy of text chunking, improves the matching accuracy in the subsequent retrieval stage of the RAG system, and the quality of the answers generated by the large model.
[0048] (2) The present invention obtains the feature similarity between sentences through the feature vectors of each sentence, and establishes a sentence similarity matrix; based on the sentence similarity matrix, a sentence relationship graph is constructed, and the sentence relationship graph is constructed by using the semantic similarity between sentences, ensuring that sentences with similar semantics are assigned to the same text block, and improving the semantic consistency of chunking.
[0049] (3) The present invention updates the nodes by means of neighborhood aggregation, reduces redundant calculations, improves the chunking efficiency, and is especially suitable for processing large-scale texts.
[0050] (4) The present invention supports two methods, top-k and threshold-based, to construct the sentence relationship graph, adapts to different scenario requirements, and flexibly adjusts the chunking granularity.
[0051] (5) The present invention automatically completes the whole process from text segmentation to chunking, reduces manual intervention, and improves the processing efficiency.
[0052] (6) The present invention constructs a sentence relationship graph through a similarity matrix, avoids only focusing on adjacent sentences during chunking, and significantly improves the accuracy, flexibility, and efficiency of text chunking by combining graph embedding and neighborhood aggregation techniques, and has broad application prospects. Description of the Drawings
[0053] Figure 1 It is the structural diagram of the method of the present invention;
[0054] Figure 2 This is the flowchart for constructing sentence relationships based on top-k in the present invention;
[0055] Figure 3 This is the flowchart for constructing sentence relationships based on similarity thresholds in the present invention;
[0056] Figure 4 This is the flowchart of the method steps in the present invention. Detailed implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0058] As Figures 1-4 shown, the present invention proposes a text chunking method based on graph embedding, including:
[0059] S1: Segment the text to obtain each sentence after segmentation;
[0060] For text segmentation, punctuation marks in the text are used as segmentation points, and the text is segmented based on the segmentation points to obtain each sentence after segmentation. The regular expression is [a1, a2,..., at,..., aT]. For a text file, the text content therein usually ends a sentence with a period, a semicolon, an exclamation mark, and a question mark. In this embodiment, the regular expression [。;!?] of these four punctuation marks is used to segment the text; where, the square brackets [] represent a character set, at is the t-th punctuation mark character to be matched, t is the punctuation mark character number, T is the number of punctuation mark characters, and the period, semicolon, exclamation mark, and question mark are the characters to be matched.
[0061] S2: Obtain the feature vectors of each sentence;
[0062] For each sentence after segmentation, obtain the feature vectors of each sentence, that is, input each sentence after segmentation into the text embedding model and output the feature vectors of each sentence;
[0063] The text embedding model extracts sentences and converts text (such as words, phrases, sentences, or entire documents) into fixed-length dense vector representations. These vectors are usually continuous, low-dimensional, and can capture the semantic information of the text.
[0064] In this embodiment, the text embedding model refers to the BGE text embedding model.
[0065] S3: Based on the feature vectors of each sentence, obtain the feature similarity between sentences and establish a sentence similarity matrix. It includes: First, use the cosine similarity function to calculate the feature similarity between all sentences;
[0066] Using the cosine similarity function to calculate the feature similarity between all sentences, compared with the traditional method that only calculates the similarity between adjacent sentences, the method of the present invention calculates the feature similarity between each sentence and all other sentences, and takes the feature similarity between each sentence and all other sentences as the element value of the matrix, so as to establish a sentence similarity matrix S. The construction formula of the sentence similarity matrix S is:
[0067]
[0068] where, x i represents the feature vector of the i-th sentence, x j represents the feature vector of the j-th sentence, S ∈ R n×n ; n is the total number of sentences, S ij represents the element value of the i-th row and j-th column in the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
[0069] S4: Based on the sentence similarity matrix, construct a sentence relationship graph using the top-k method, including:
[0070] First, based on the sentence similarity matrix S, regard each sentence as a node, and find the k nodes with the greatest similarity to the current node to construct edges, so as to construct a sentence relationship graph; the construction formula of the sentence relationship adjacency matrix A corresponding to the sentence relationship graph is:
[0071]
[0072] where, A ij represents the element value of the i-th row and j-th column in the adjacency matrix A, A ij = 1 means that the element value of the i-th row and j-th column in the adjacency matrix A is 1, that is, there is an edge between the i-th node and the j-th node in the sentence relationship graph, A ij = 0 means that the element value of the i-th row and j-th column in the adjacency matrix A is 0, that is, there is no edge between the i-th node and the j-th node in the sentence relationship graph; top-k(S[i:]) represents selecting the top k largest elements in the i-th row of the sentence similarity matrix S, that is, selecting the k sentences with the greatest feature similarity to the i-th sentence; S ij represents the element value of the i-th row and j-th column in the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
[0073] As the second embodiment of the present invention, based on the sentence similarity matrix, a method based on a threshold is used to construct a sentence relationship graph, including:
[0074] Based on the sentence similarity matrix S, both the rows and columns in the sentence similarity matrix are regarded as nodes;
[0075] Set a similarity threshold, and find the elements in the sentence similarity matrix whose element values are greater than or equal to the similarity threshold as associated elements;
[0076] Construct edges between the nodes corresponding to the rows and columns of the associated elements to form a sentence relationship graph; the construction formula for the sentence relationship adjacency matrix A corresponding to the sentence relationship graph is:
[0077]
[0078] A ij represents the element value in the i-th row and j-th column of the adjacency matrix A, a is the similarity threshold, A ij = 1 means that the element value in the i-th row and j-th column of the adjacency matrix A is 1, that is, there is an edge between the i-th node and the j-th node in the sentence relationship graph, A ij = 0 means that the element value in the i-th row and j-th column of the adjacency matrix A is 0, that is, there is no edge between the i-th node and the j-th node in the sentence relationship graph.
[0079] S5: Based on the sentence relationship graph, use the neighborhood aggregation method to update the nodes to obtain the updated sentence relationship graph, including:
[0080] For each node feature in the sentence relationship graph, add the sentence feature vector corresponding to the current node to the sentence feature vector corresponding to the neighbor node as the node embedding representation corresponding to the current node, so as to update all nodes and obtain the updated sentence relationship graph; the neighbor node refers to the node directly connected to the current node.
[0081] Through the above update operation, the neighborhood information of each node is incorporated into the current node information, and the topological structure of the node is considered at the same time, so that the updated sentence relationship graph can more accurately and objectively reflect the true connection relationship between sentences.
[0082] The node embedding representation H corresponding to the current node i The calculation formula is as follows:
[0083]
[0084] i is the sentence number, that is, the node number; N i represents the set of neighbor nodes of node i.
[0085] S6: Based on the updated sentence relationship graph, combine the sentences corresponding to the nodes to form each text block. This includes:
[0086] Based on the updated sentence relationship graph, obtain the position of each sentence in the original text as the index of the node in the sentence relationship graph;
[0087] Calculate the feature similarity between adjacent nodes;
[0088] Take each node and the top t nodes with the highest feature similarity to it as a group of combined nodes;
[0089] Connect and combine the sentences corresponding to each group of combined nodes according to the node index, so as to form each text block.
[0090] In this embodiment, first, the user imports a txt text file, searches for four symbols (period, exclamation mark, semicolon, question mark) in the text file, divides the text according to the four symbols to obtain 1000 sentences, and saves them in the form of a list; then inputs the divided sentences in the list into the BCE text embedding model to obtain the feature vectors of 1000 sentences; based on the feature vectors of 1000 sentences, use the cosine similarity function to calculate the feature similarity between all sentences, and take the feature similarity between each sentence and all other sentences as the element value of the matrix, so as to establish a sentence similarity matrix S with a dimension of 1000*1000; based on the sentence similarity matrix S, regard each sentence as a node, search for the 5 nodes with the highest similarity to the current node to construct edges, so as to construct a sentence relationship graph and the corresponding sentence relationship adjacency matrix A of the sentence relationship graph; then, for each node in the sentence relationship graph, add the sentence feature vector corresponding to the current node and the sentence feature vector corresponding to the neighbor node as the node embedding representation corresponding to the current node, so as to update all nodes and obtain the updated sentence relationship graph; finally, based on the updated sentence relationship graph, obtain the positions of these 1000 sentences in the original txt text file as the indexes of the nodes in the sentence relationship graph; then calculate the feature similarity between adjacent nodes; take each node and the top 9 nodes with the highest feature similarity to it as a group of combined nodes; connect and combine the sentences corresponding to each group of combined nodes, that is, each group of 10 nodes, according to the node index, and finally form 100 text blocks.
[0091] Of course, for those skilled in the art, the present invention is not limited to the details of the above-described exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0092] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
[0093] The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.
Claims
1. A RAG text segmentation method based on graph embedding, characterized in that: include: S1: Segment the text and obtain the segmented sentences; S2: Get the feature vector of each sentence; S3: Based on the feature vectors of each sentence, obtain the feature similarity between sentences and establish a similarity matrix between sentences; S4: Construct sentence relationship graph based on sentence similarity matrix; S5: Based on the sentence relationship graph, the nodes are updated by using the neighborhood aggregation method to obtain an updated sentence relationship graph; S6: Based on the updated sentence relationship graph, the sentences corresponding to the nodes are combined to form text blocks.
2. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S6, based on the updated sentence relationship graph, the sentences corresponding to the nodes are combined to form text blocks including: S61: Based on the updated sentence relationship graph, obtain the position of each sentence in the original text as the index of the node in the sentence relationship graph; S62: Calculate feature similarity between adjacent nodes; S63: taking each node and the first t nodes with the largest feature similarity as a group of combined nodes; S64: Sentences corresponding to each group of combined nodes are connected and combined according to the node indexes to form various text blocks.
3. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S1, the text is segmented to obtain each segmented sentence, including: using punctuation marks in the text as segmentation points, segmenting the text based on the segmentation points, and obtaining each segmented sentence; The regular expression used for segmentation is [a1,a2,...,at,...,aT]; The square brackets [] represent a set of punctuation characters, at is the tth punctuation character to be matched, t is the punctuation character number, and T is the number of punctuation characters.
4. A RAG text segmentation method based on graph embedding as claimed in claim 3, characterized in that: The punctuation marks include period, semicolon, exclamation mark and question mark.
5. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S2, the feature vector of each sentence is obtained, that is, each segmented sentence is input into the text embedding model, and the feature vector of each sentence is output.
6. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S3, based on the feature vectors of each sentence, the feature similarity between sentences is obtained, and a sentence similarity matrix is established, including: S31: The cosine similarity function is used to calculate the feature similarity between all sentences; S32: Taking the feature similarities between all sentences as the element values of the matrix, a sentence similarity matrix S is established; The construction formula of the sentence similarity matrix S is: Among them, x i represents the feature vector of the i-th sentence, x j Represents the feature vector of the jth sentence, S∈R n×n ; n is the total number of sentences, S ij It represents the element value of the i-th row and j-th column in the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
7. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S4, a sentence relationship graph is constructed based on the sentence similarity matrix, using a top-k method, including: Based on the sentence similarity matrix S, each sentence is regarded as a node, and the k nodes with the greatest similarity to the current node are found to construct edges, thereby constructing a sentence relationship graph and a sentence relationship adjacency matrix A corresponding to the sentence relationship graph; The construction formula of sentence relationship adjacency matrix A is: Among them, A ij Represents the element value of the i-th row and j-th column in the adjacency matrix A. ij =1 means that the value of the element in the i-th row and j-th column of the adjacency matrix A is 1, that is, there is an edge between the i-th node and the j-th node in the sentence relationship graph. ij =0 means that the value of the element in the i-th row and j-th column of the adjacency matrix A is 0, that is, there is no edge between the i-th node and the j-th node in the sentence relationship graph; top-k(S[i:]) means selecting the largest first k elements on the i-th row of the sentence similarity matrix S, that is, selecting the k sentences with the greatest feature similarity to the i-th sentence; S ij It represents the element value of the i-th row and j-th column in the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
8. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S4, based on the sentence similarity matrix, a sentence relationship graph is constructed using a threshold-based method, including: Based on the sentence similarity matrix S, the rows and columns in the sentence similarity matrix are regarded as nodes; Set a similarity threshold, and find the elements whose element values in the sentence similarity matrix are greater than or equal to the similarity threshold as the associated elements; Construct edges between nodes in rows and columns corresponding to associated elements, thereby forming a sentence relationship graph and a sentence relationship adjacency matrix A corresponding to the sentence relationship graph; The construction formula of sentence relationship adjacency matrix A is: A ij represents the element value of the i-th row and j-th column in the adjacency matrix A, a is the similarity threshold, A ij =1 means that the value of the element in the i-th row and j-th column of the adjacency matrix A is 1, that is, there is an edge between the i-th node and the j-th node in the sentence relationship graph. ij =0 means that the value of the element in the i-th row and j-th column of the adjacency matrix A is 0, that is, there is no edge between the i-th node and the j-th node in the sentence relationship graph; S ij It represents the element value of the i-th row and j-th column in the sentence similarity matrix S, that is, the feature similarity between the i-th sentence and the j-th sentence.
9. A RAG text segmentation method based on graph embedding as claimed in claim 1, characterized in that: In step S5, based on the sentence relationship graph, the nodes are updated by using a neighborhood aggregation method to obtain an updated sentence relationship graph, including: For each node in the sentence relationship graph, the sentence feature vector corresponding to the current node is added to the sentence feature vector corresponding to the neighbor node as the node embedding representation corresponding to the current node, thereby updating all nodes and obtaining an updated sentence relationship graph; the neighbor node refers to a node directly connected to the current node.
10. A RAG text segmentation method based on graph embedding as claimed in claim 9, characterized in that: The node embedding representation H corresponding to the current node i The calculation formula is as follows: Among them, i is the sentence number, that is, the node number; N i Represents the set of neighbor nodes of node i.