A text set similarity calculation method and device for retrieval-enhanced generation

By calculating the cosine similarity and BM25 scores of any two texts in the text set, and using the Kruskal algorithm to build a minimum spanning tree, the problems of high complexity and inaccurate similarity description in large-scale text set similarity calculations are solved, and efficient and accurate text similarity evaluation is achieved.

CN119311856BActive Publication Date: 2025-05-16HANGZHOU CANGHAI GUANZHI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411803529.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-05-16
Estimated Expiration
2044-12-10

AI Technical Summary

Technical Problem

When processing large-scale text sets, the prior art has high computational complexity and is difficult to capture the global structural similarity between text sets, resulting in large time overhead and inaccurate similarity description.

Method used

By calculating the cosine similarity and BM25 scores between any two texts in the text set, a sub-graph of semantic similarity and text similarity is constructed, and the real edge weight is determined through multiple rounds of clustering. Finally, the Kruskal algorithm is used to construct a minimum spanning tree to determine the similarity of text in the text set.

Benefits of technology

It realizes that the text similarity in the text set is accurately described while ensuring processing efficiency, reducing the distortion problem in the reconciliation process, and making the similarity evaluation results more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311856B_ABST
    Figure CN119311856B_ABST
Patent Text Reader

Abstract

The invention discloses a text set similarity calculation method and device for retrieval enhanced generation, comprising: after preprocessing the acquired text set, calculating the cosine similarity and BM25 score between any two texts in the text set; using the cosine similarity and BM25 score as a semantic similarity weight and a text similarity weight respectively, and constructing two subgraphs; performing multiple rounds of clustering on the two subgraphs respectively and recording the survival time of the edge weight of each node in the subgraphs, determining the real edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with the text as a node based on the real edge weight; using the Kruskal algorithm to calculate the comprehensive graph to obtain a minimum spanning tree, and determining the similarity of the texts in the text set based on the minimum spanning tree, so that the text similarity in a text set can be accurately described while ensuring processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a text set similarity calculation method and device for retrieval-enhanced generation. Background Art

[0002] Terminology explanation:

[0003] Retrieval-Augmented Generation: Retrieval-Augmented Generation (RAG) is a new technical framework that combines information retrieval and natural language generation. It is widely used in question-answering systems, dialogue generation, and knowledge enhancement tasks. It integrates external knowledge bases with generative models to introduce the latest and most relevant information when generating natural language output, thereby overcoming the knowledge limitations of pure generative models.

[0004] Graph: A graph data structure consists of a finite (possibly mutable) set of nodes and a set of unordered pairs (for undirected graphs) or ordered pairs (for directed graphs) as edges. Nodes can be part of the graph structure or external entities denoted by integer subscripts or references.

[0005] Tree: A tree is a special type of graph in data structure, which has the following characteristics: each node has only a finite number of child nodes or no child nodes; a node without a parent node is called a root node; except for the root node, each child node can be divided into multiple non-intersecting subtrees; there is no self-loop;

[0006] Minimum spanning tree: a spanning tree with the minimum weight in a connected weighted undirected graph;

[0007] Kruskal algorithm: an algorithm used to find the minimum spanning tree;

[0008] BM25 algorithm: BM25 (Best Matching 25) algorithm is an algorithm used for information retrieval, often used to calculate the relevance score between documents and queries. It is based on the concepts of term frequency and inverse document frequency (TF-IDF), but introduces some adjustment parameters to better handle documents of different lengths and the impact of term frequency.

[0009] Union-find: is a data structure used to handle the merging and querying of disjoint sets. Union-find supports the following operations: query, merge, and add.

[0010] In order to make full use of the powerful capabilities of the RAG system, the academic community generally believes that the lower the correlation between the filled data, that is, when filling in a set of unrelated data, its performance is better. Therefore, an advanced method is needed to efficiently determine the similarity within a set of data. Traditional text similarity calculation methods are mostly based on technologies such as word frequency and word vectors, but these methods often have high computational complexity when processing large-scale text sets and are difficult to capture the global structural similarity between text sets. Kruskal algorithm is a classic minimum spanning tree algorithm that is applicable to weighted undirected graphs in graph theory and can effectively find the minimum spanning tree in the graph. Introducing the Kruskal algorithm into the field of text similarity calculation can effectively solve the global optimal problem in text set similarity calculation.

[0011] The computational complexity of the Kruskal algorithm is high on large-scale text sets. When the number of texts increases, calculating the similarity between each pair of texts will cause the time complexity and space complexity to rise sharply. Especially when processing large-scale texts, there are problems of high time overhead and inaccurate similarity description. At the same time, the cosine similarity calculation is based on the vector angle obtained by the inner and outer products between vectors, and the BM25 score is based on the calculation of word frequency and hyperparameters. The reconciliation of the two will lead to serious distortion problems, so they cannot be directly used as weights to participate in the calculation of the Kruskal algorithm. Summary of the invention

[0012] In view of the above, the purpose of the present invention is to provide a text set similarity calculation method and device for retrieval enhancement generation, which can accurately describe the text similarity within a text set while ensuring processing efficiency, so as to assist the retrieval enhancement generation system and ensure the quality of the filled-in data.

[0013] To achieve the above-mentioned purpose of the invention, an embodiment provides a text set similarity calculation method for retrieval enhanced generation, comprising the following steps:

[0014] After preprocessing the acquired text set, the cosine similarity and BM25 score between any two texts in the text set are calculated;

[0015] The cosine similarity is used as the semantic similarity weight, and a subgraph with text as node and edge weight as semantic similarity weight is constructed; the BM25 score is used as the text similarity weight, and a subgraph with text as node and edge weight as text similarity weight is constructed; multiple rounds of clustering are performed on the two subgraphs respectively and the survival time of the edge weight of each node in the subgraph is recorded, and the real edge weight of each text is determined based on the survival time of the edge weight of the same node in the two subgraphs, and a comprehensive graph with text as node is obtained based on the real edge weight;

[0016] The Kruskal algorithm is used to calculate the minimum spanning tree of the comprehensive graph, and the similarity of the texts in the text set is determined based on the minimum spanning tree.

[0017] Preferably, the acquired text set is preprocessed, including:

[0018] The text is cleaned, uniformly encoded, formatted, and then segmented.

[0019] Preferably, calculating the cosine similarity between any two texts in the text set includes:

[0020] The preprocessed text is vectorized, the text is converted into a semantic vector, and the cosine similarity is calculated based on the semantic vector.

[0021] Preferably, calculating the BM25 score between any two texts in the text set includes:

[0022] For the preprocessed text, the word frequency and inverse word order of each word are calculated, and the BM25 algorithm is used to calculate the BM25 score based on the word frequency and inverse word order.

[0023] Preferably, multiple rounds of clustering are performed on the two subgraphs respectively and the survival time of the edge weight of each node in the subgraph is recorded, including:

[0024] Cluster each node in the subgraph, delete the nodes that do not meet any centroid conditions, and record the deleted nodes and their edge weights, as well as the survival time, where the survival time is the time from participating in clustering to being deleted, or the rounds from participating in clustering to being deleted.

[0025] Preferably, the real edge weight of each text is determined based on the survival time of the edge weights of the same node in the two subgraphs, including:

[0026] The survival time of the edge weights of the same node in the two subgraphs is added together as the true edge weight of each text.

[0027] Preferably, the Kruskal algorithm is used to calculate the minimum spanning tree on the comprehensive graph, including:

[0028] All the real edge weights in the comprehensive graph are pushed into the priority queue. The edge with the smallest real weight is taken out from the priority queue each time, and the two nodes connected by the edge with the smallest real weight are determined. The union-find set is used to maintain the connectivity of the texts corresponding to the two nodes. If the two texts are already connected in the same union-find set through other paths, there is no need to introduce new edges. Otherwise, a new edge will be added between the two texts and added to the spanning tree, and the union-find set where the two nodes connected by the edge with the smallest real weight are located will be merged.

[0029] To achieve the above-mentioned object of the invention, an embodiment of the present invention further provides a text set similarity calculation device, comprising:

[0030] A preprocessing module is used to calculate the cosine similarity and BM25 score between any two texts in the text set after preprocessing the acquired text set;

[0031] A true edge weight determination module is used to use cosine similarity as a semantic similarity weight, and construct a subgraph with text as a node and edge weight as a semantic similarity weight; use BM25 score as a text similarity weight, and construct a subgraph with text as a node and edge weight as a text similarity weight; perform multiple rounds of clustering on the two subgraphs respectively and record the survival time of the edge weight of each node in the subgraphs, determine the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtain a comprehensive graph with text as a node based on the true edge weight;

[0032] The minimum spanning tree module is used to calculate the minimum spanning tree on the comprehensive graph using the Kruskal algorithm, and determine the similarity of texts in the text set based on the minimum spanning tree.

[0033] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computing device, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned text set similarity calculation method for retrieval-enhanced generation.

[0034] To achieve the above-mentioned purpose of the invention, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned text set similarity calculation method for retrieval-enhanced generation is implemented.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The present invention transforms the semantics and text similarity within a group of texts into the average semantics and text similarity edge weights of the minimum spanning tree corresponding to the group of texts by transforming the problem. In this process, the survival time of the edge weight of each node is determined by clustering, and the semantic similarity and text similarity are combined without reconciling the two weights, which significantly reduces the distortion problem in the reconciliation process, making the similarity evaluation result more accurate and reliable. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0038] Figure 1 is a flowchart of a text set similarity calculation method for retrieval-enhanced generation provided by an embodiment;

[0039] Figure 2 is a flowchart of preprocessing provided by an embodiment;

[0040] Figure 3 is a flow chart of the real edge weight provided by the embodiment;

[0041] Figure 4 This is the minimum spanning tree construction process provided by the embodiment;

[0042] Figure 5 It is a structural schematic diagram of a text set similarity calculation device provided in an embodiment. DETAILED DESCRIPTION

[0043] To make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific implementation methods described herein are only used to explain the present invention and do not limit the scope of protection of the present invention.

[0044] The inventive concept of the present invention is: to address the technical problem that the Kruskal algorithm in the prior art cannot directly combine cosine similarity calculation and BM25 to accurately and simultaneously describe text similarity and semantic similarity, resulting in the inability to accurately describe the text relevance within a document set while ensuring processing efficiency, an embodiment of the present invention provides a text set similarity calculation solution.

[0045] like Figure 1 As shown, the embodiment provides a text set similarity calculation method for retrieval enhanced generation, comprising the following steps:

[0046] S1, after preprocessing the acquired text set, calculate the cosine similarity and BM25 score between any two texts in the text set.

[0047] In the embodiment, Figure 2As shown, the obtained text set is preprocessed, including: cleaning, uniformly encoding, formatting, and then performing text segmentation on the text in the text set. In specific implementation, a batch of text sets are obtained from various data sources such as web pages, databases, and file systems. In order to ensure data quality and avoid interference from abnormal fragments, these text sets need to be cleaned, including removing HTML tags, special characters, stop words, etc., and then uniformly encoding and formatting the cleaned text; then the JIEBA word segmenter is used to segment the cleaned text. The JIEBA word segmenter is an efficient and accurate Chinese word segmentation tool that can segment text into several independent words or phrases, which are called text fragments.

[0048] In the embodiment, the process of calculating the cosine similarity between any two texts in the text set is: first, the preprocessed text is vectorized to convert the text into a semantic vector, and then a pairwise matching calculation is performed, that is, the cosine value of the angle between the two vectors is obtained based on the inner product and outer product calculation of the semantic vector as the cosine similarity.

[0049] In the embodiment, the BM25 score between any two texts in the text set is calculated, including: first, for the preprocessed text, the word frequency and the inverse word order of each word are calculated to form a BM25 object that can be directly used for calculation, and then the BM25 algorithm is used to calculate the BM25 score based on the word frequency and the inverse word order.

[0050] S2, taking cosine similarity as semantic similarity weight, and constructing a subgraph with text as node and edge weight as semantic similarity weight; taking BM25 score as text similarity weight, and constructing a subgraph with text as node and edge weight as text similarity weight; performing multiple rounds of clustering on the two subgraphs respectively and recording the survival time of the edge weight of each node in the subgraph, determining the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with text as node based on the true edge weight.

[0051] The determination of the true edge weight is the core breakthrough of the present invention. In demand analysis, the semantic similarity and text similarity that need to be taken into account cannot be directly reconciled. To solve this technical problem, subgraphs are constructed with semantic similarity weights and text similarity weights respectively. Clustering algorithms such as k-means are used to repeatedly determine the centroid, cluster, exclude irrelevant nodes, and record the survival time of each node. The true edge weight is obtained by converting the weight calculation into the survival time.

[0052] Specifically, when constructing two subgraphs, cosine similarity is used as the semantic similarity weight, and a subgraph is constructed with text as the node and edge weight as the semantic similarity weight. Similarly, BM25 score is used as the text similarity weight, and a subgraph is constructed with text as the node and edge weight as the text similarity weight. Then, multiple rounds of clustering are performed on the two subgraphs respectively and the survival time of the edge weight of each node in the subgraph is recorded. The specific process is as follows:

[0053] like Figure 3 As shown, based on the k-means++ algorithm and its probability allocation algorithm, the entire subgraph is initialized to determine the current centroid group; in each round of clustering, the binary demarcation edge conditions make the nodes that do not meet any centroid conditions become free nodes and need to be deleted, and the deleted nodes and their edge weights, as well as the survival time are recorded, where the survival time is the time from participating in clustering to being deleted, or the rounds from participating in clustering to being deleted.

[0054] Through cluster analysis, or the survival time of the edge weight of the node corresponding to each text in the two subgraphs, the survival time of the edge weight of the same node in the two subgraphs is added as the true edge weight of each text, and then based on the true edge weight, a comprehensive graph with the text as the node and the true edge weight as the edge weight is obtained.

[0055] S3, the Kruskal algorithm is used to calculate the comprehensive graph to obtain the minimum spanning tree, and the similarity of the texts in the text set is determined based on the minimum spanning tree.

[0056] The core concept of the Kruskal algorithm is that, for a given graph, if there is an edge whose weight is less than the currently selected edge, and adding it and searching the set will not form a loop, then abandoning this edge will not be able to construct a spanning tree with a smaller weight. Based on this principle, the present invention cleverly regards the screened text pairs as edges and each text as a node, thereby realizing the conversion from a text set to a text graph.

[0057] The Kruskal algorithm is used to calculate the minimum spanning tree of the comprehensive graph, including:

[0058] like Figure 4 As shown, all the real edge weights in the comprehensive graph are pushed into the priority queue, and the edge with the smallest real weight is taken out from the priority queue each time, and the two nodes connected by the edge with the smallest real weight are determined. The union-find set is used to maintain the connectivity of the texts corresponding to the two nodes. If the two texts are already connected in the same union-find set through other paths, there is no need to introduce new edges. Otherwise, a new edge will be added between the two texts and added to the spanning tree, and the union-find set where the two nodes connected by the edge with the smallest real weight are located will be merged.

[0059] When determining the similarity of texts in a text set based on the generated minimum spanning tree, it is considered that the similarity of these texts belonging to the same minimum spanning tree is relatively low overall, based on which the similarity of the text set can be evaluated.

[0060] like Figure 5 As shown, the embodiment also provides a text set similarity calculation device 50, including a preprocessing module 51, a real edge weight determination module 52, and a minimum spanning tree module 53, wherein the preprocessing module 51 is used to calculate the cosine similarity and BM25 score between any two texts in the text set after preprocessing the acquired text set; the real edge weight determination module 52 is used to use the cosine similarity as the semantic similarity weight, and construct a subgraph with the text as the node and the edge weight as the semantic similarity weight; use the BM25 score as the text similarity weight, and construct a subgraph with the text as the node and the edge weight as the text similarity weight; perform multiple rounds of clustering on the two subgraphs respectively and record the survival time of the edge weight of each node in the subgraphs, determine the real edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtain a comprehensive graph with the text as the node based on the real edge weight; the minimum spanning tree module 53 is used to use the Kruskal algorithm to calculate the comprehensive graph to obtain a minimum spanning tree, and determine the similarity of the texts in the text set based on the minimum spanning tree.

[0061] It should be noted that the text set similarity calculation device for retrieval-enhanced generation provided in the above embodiment should be illustrated by the division of the above functional modules when performing text similarity calculation. The above functions can be assigned to different functional modules as needed, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the functions described above. In addition, the text set similarity calculation device for retrieval-enhanced generation provided in the above embodiment and the text set similarity calculation method embodiment for retrieval-enhanced generation belong to the same concept. The specific implementation process is detailed in the text set similarity calculation method embodiment for retrieval-enhanced generation, which will not be repeated here.

[0062] The embodiments of the present invention provide a method and device for calculating the similarity of a text set for retrieval-enhanced generation. The method obtains the survival time in conditional continuous clustering as another representation of cosine similarity and BM25 score to construct a minimum spanning tree. The minimum spanning tree corresponds to the connotation of the actual edge weight (semantic similarity represented by cosine similarity and text similarity represented by BM25 score) to describe the semantic and grammatical similarity of a group of texts. This can significantly reduce the distortion problem in the reconciliation process and make the obtained similarity evaluation result more accurate and reliable.

[0063] Based on the same inventive concept, an embodiment further provides a computing device, including a memory and one or more processors, wherein executable code is stored in the memory, and when the one or more processors execute the executable code, the method for calculating the similarity of a text set for retrieval enhancement generation is implemented, specifically including the following steps:

[0064] S1, after preprocessing the acquired text set, calculate the cosine similarity and BM25 score between any two texts in the text set;

[0065] S2, using cosine similarity as semantic similarity weight, and constructing a subgraph with text as node and edge weight as semantic similarity weight; using BM25 score as text similarity weight, and constructing a subgraph with text as node and edge weight as text similarity weight; performing multiple rounds of clustering on the two subgraphs respectively and recording the survival time of the edge weight of each node in the subgraphs, determining the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with text as node based on the true edge weight;

[0066] S3, the Kruskal algorithm is used to calculate the comprehensive graph to obtain the minimum spanning tree, and the similarity of the texts in the text set is determined based on the minimum spanning tree.

[0067] The computing device provided in the embodiment, in addition to the processor and memory, also includes hardware required for other services such as internal bus, network interface, memory, etc. at the hardware level. The memory is a non-volatile memory, and the processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the text set similarity calculation method for retrieval-enhanced generation described in S1-S3 above. Of course, in addition to software implementation, the present invention does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0068] Based on the same inventive concept, an embodiment further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the above-mentioned text set similarity calculation method for retrieval-enhanced generation is implemented, which specifically includes the following steps:

[0069] S1, after preprocessing the acquired text set, calculate the cosine similarity and BM25 score between any two texts in the text set;

[0070] S2, using cosine similarity as semantic similarity weight, and constructing a subgraph with text as node and edge weight as semantic similarity weight; using BM25 score as text similarity weight, and constructing a subgraph with text as node and edge weight as text similarity weight; performing multiple rounds of clustering on the two subgraphs respectively and recording the survival time of the edge weight of each node in the subgraphs, determining the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with text as node based on the true edge weight;

[0071] S3, the Kruskal algorithm is used to calculate the comprehensive graph to obtain the minimum spanning tree, and the similarity of the texts in the text set is determined based on the minimum spanning tree.

[0072] In the embodiment, computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data.

[0073] The specific implementation methods described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A text set similarity calculation method for retrieval-enhanced generation, characterized in that: The following steps are involved: After preprocessing the acquired text set, the cosine similarity and BM25 score between any two texts in the text set are calculated; Use cosine similarity as the semantic similarity weight, and construct a subgraph with text as node and edge weight as semantic similarity weight; use BM25 score as text similarity weight, and construct a subgraph with text as node and edge weight as text similarity weight; Perform multiple rounds of clustering on the two subgraphs respectively and record the survival time of the edge weight of each node in the subgraphs, including: clustering each node in the subgraphs, deleting the nodes that do not meet any centroid conditions and recording the deleted nodes and their edge weights, as well as the survival time, wherein the survival time is the time from participating in clustering to being deleted, or the number of rounds from participating in clustering to being deleted; determining the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with the text as a node based on the true edge weight; The Kruskal algorithm is used to calculate the minimum spanning tree of the comprehensive graph, and the similarity of the texts in the text set is determined based on the minimum spanning tree.

2. The text set similarity calculation method for retrieval-enhanced generation according to claim 1 is characterized in that: Preprocess the acquired text set, including: The text is cleaned, uniformly encoded, formatted, and then segmented.

3. The text set similarity calculation method for retrieval-enhanced generation according to claim 1 is characterized in that: Calculate the cosine similarity between any two texts in the text set, including: The preprocessed text is vectorized, the text is converted into a semantic vector, and the cosine similarity is calculated based on the semantic vector.

4. The text set similarity calculation method for retrieval-enhanced generation according to claim 1 is characterized in that: Calculate the BM25 score between any two texts in the text set, including: For the preprocessed text, the word frequency and inverse word order of each word are calculated, and the BM25 algorithm is used to calculate the BM25 score based on the word frequency and inverse word order.

5. The text set similarity calculation method for retrieval-enhanced generation according to claim 1 is characterized in that: The real edge weight of each text is determined based on the survival time of the edge weights of the same node in the two subgraphs, including: The survival time of the edge weights of the same node in the two subgraphs is added together as the true edge weight of each text.

6. The text set similarity calculation method for retrieval-enhanced generation according to claim 1 is characterized in that: The Kruskal algorithm is used to calculate the minimum spanning tree of the comprehensive graph, including: All the real edge weights in the comprehensive graph are pushed into the priority queue. The edge with the smallest real weight is taken out from the priority queue each time, and the two nodes connected by the edge with the smallest real weight are determined. The union-find set is used to maintain the connectivity of the texts corresponding to the two nodes. If the two texts are already connected in the same union-find set through other paths, there is no need to introduce new edges. Otherwise, a new edge will be added between the two texts and added to the spanning tree, and the union-find set where the two nodes connected by the edge with the smallest real weight are located will be merged.

7. A text set similarity calculation device for retrieval-enhanced generation, characterized in that: include: A preprocessing module is used to calculate the cosine similarity and BM25 score between any two texts in the text set after preprocessing the acquired text set; A true edge weight determination module, which is used to use cosine similarity as a semantic similarity weight and construct a subgraph with text as a node and edge weight as a semantic similarity weight; use BM25 score as a text similarity weight and construct a subgraph with text as a node and edge weight as a text similarity weight; Perform multiple rounds of clustering on the two subgraphs respectively and record the survival time of the edge weight of each node in the subgraphs, including: clustering each node in the subgraphs, deleting the nodes that do not meet any centroid conditions and recording the deleted nodes and their edge weights, as well as the survival time, wherein the survival time is the time from participating in clustering to being deleted, or the number of rounds from participating in clustering to being deleted; determining the true edge weight of each text based on the survival time of the edge weight of the same node in the two subgraphs, and obtaining a comprehensive graph with the text as a node based on the true edge weight; The minimum spanning tree module is used to calculate the minimum spanning tree on the comprehensive graph using the Kruskal algorithm, and determine the similarity of texts in the text set based on the minimum spanning tree.

8. A computing device comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the one or more processors execute the executable code, they are used to implement the text set similarity calculation method for retrieval-enhanced generation according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the text set similarity calculation method for retrieval-enhanced generation described in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Large-scale biological data clustering method and system based on spanning tree

    CN114420215A

  • Similar asset fingerprint extraction method based on LLM model

    CN118069778A