A Multi-Level Retrieval Enhancement Method for Large Language Models

By initially parsing and multi-level searching of the query input by users, identifying and filtering redundant and missing information documents, the problem of insufficient information coverage in the prior art is solved, and higher quality and relevance search results are achieved.

CN119311806BActive Publication Date: 2025-06-13无锡锡商银行股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411862661.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-06-13
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

The existing search system based on large language models is difficult to fully capture the user's query intention when processing complex queries, resulting in insufficient information coverage and inaccurate search results.

Method used

By initially parsing the query input by the user and generating core query statements, performing primary searches, and clustering analysis of the primary search document sets to identify the risk of redundancy and information loss, performing multi-level search and document filtering, and finally generating document priority recommendation indexes for sorting and displaying.

Benefits of technology

Effectively filter redundant and low-confidence documents, improve the quality and relevance of search results, and ensure that users can obtain high-quality information that meets their needs faster and more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119311806B_ABST
    Figure CN119311806B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-level retrieval enhancement method for large language models, specifically related to the field of multi-level retrieval technology. By initially parsing the query input by the user and generating a core query statement and a primary retrieval document set, it can more accurately understand the user's query intention. Cluster analysis is performed on the primary retrieval document set to identify redundant risk document sets, retained document sets, and information missing risk document sets. The redundant detection information and credibility evaluation information of the redundant risk document sets are obtained, and the redundant risk document sets are initially screened to effectively filter redundant and less credible retrieval documents. Supplementary query statements are generated based on the information missing risk document sets for secondary retrieval in the knowledge base to more comprehensively cover the user's multi-angle needs and potential fuzzy intentions. Through the comprehensive calculation of the document quality coefficient and the document freshness coefficient, a document priority recommendation index is generated to ensure that documents with high quality and strong timeliness are preferentially displayed to the user, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-level retrieval, and more specifically, to a method for enhancing multi-level retrieval of large language models. Background Art

[0002] With the rapid development of artificial intelligence technology, large language models have made remarkable progress in the field of natural language processing and are widely used in scenarios such as text generation, question answering systems, and intelligent search engines. However, existing retrieval systems based on large language models still face many challenges when dealing with complex queries. Traditional single-query retrieval methods usually rely on a single query input by the user to retrieve relevant documents. This method is prone to insufficient information coverage and difficult to fully capture the user's query intent, especially when facing multi-angle or ambiguous requirements. In addition, since the results of information retrieval directly depend on the quality of the initial query, if the initial query fails to accurately reflect the user's intent, the retrieval results will be incomplete or inaccurate. Summary of the Invention

[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for enhancing multi-level retrieval of large language models to solve the problems raised in the above background art.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A method for enhancing multi-level retrieval of large language models, characterized by including the following steps:

[0006] Step S1, perform initial parsing on the query input by the user, generate a core query statement according to the initial parsing result, and perform primary retrieval in the knowledge base according to the core query statement to obtain a primary retrieval document set;

[0007] Step S2, perform clustering analysis on the primary retrieval document set and identify a redundant risk document set, a retained document set, and an information missing risk document set;

[0008] Step S3, obtain the redundancy detection information and credibility evaluation information of the redundant risk document set, and perform a primary screening on the redundant risk document set to obtain a redundant document screening set;

[0009] Step S4, generate a supplementary query statement according to the information missing risk document set, and perform secondary retrieval in the knowledge base according to the supplementary query statement to obtain an information supplementary document set;

[0010] Step S5: Merge the retained document set, redundant document screening set, and information supplement document set to obtain the final document set. Obtain the document retrieval quality information and document freshness information of the final document set, generate a document priority recommendation index, sort the documents in the final document set from largest to smallest according to the document priority recommendation index to obtain a document sorting table, and display the documents to the user in the document sorting order.

[0011] In a preferred embodiment, in step S2, perform clustering analysis on the primary retrieval document set and identify the redundant risk document set, retained document set, and information missing risk document set as follows:

[0012] Step S21: Preprocess the primary retrieval document set, including removing noise and standardizing the text format, and convert each preprocessed primary retrieval document into a vector representation;

[0013] Step S22: Use the elbow method to determine the initial number of clustering centroids K;

[0014] Step S23: According to the determined initial number of clustering centroids K, randomly select K primary retrieval documents from the primary retrieval document set as the initial clustering cluster centroids;

[0015] Step S24: Calculate the Euclidean distance from each primary retrieval document in the primary retrieval document set to the initial clustering cluster centroid, and assign each primary retrieval document to the nearest clustering cluster;

[0016] Step S25: Calculate the average vector of each clustering cluster and use it as the new clustering cluster centroid;

[0017] Step S26: Repeat steps S24 and S25 to iteratively update the clustering clusters until the clustering cluster centroids no longer change to obtain the final clustering cluster set;

[0018] Step S27: Preprocess the core query statement, convert the preprocessed core query statement into a vector representation, and use the cosine similarity method to calculate the first cosine similarity xsd1 between each primary retrieval document in the clustering cluster and the core query statement. The expression is as follows In the formula, A represents the vector representation of the primary retrieval document in the clustering cluster, B represents the vector representation of the core query statement, A·B represents the dot product operation of A and B, ||A|| represents the Euclidean norm of the primary retrieval document vector, and ||B|| represents the Euclidean norm of the core query statement vector;

[0019] Count the number of occurrences \(C1\) where the first cosine similarity between the primary retrieval documents and the core query statement in the clustering cluster is greater than 0 and less than or equal to 1; count the number of occurrences \(C2\) where the first cosine similarity between the primary retrieval documents and the core query statement is less than 0 and greater than or equal to -1, and calculate the similarity ratio \(xsz\). The expression is as follows

[0020] Step S28: Compare the similarity ratio with the preset similarity ratio stage threshold to mark the clustering cluster, identify the redundant risk document set, the retained document set, and the information missing risk document set. The similarity ratio stage threshold includes the first similarity ratio threshold and the second similarity ratio threshold, and the first similarity ratio threshold is less than the second similarity ratio threshold;

[0021] If the similarity ratio is greater than the second similarity ratio threshold, mark the clustering cluster as a redundant risk clustering cluster, and add the primary retrieval documents in the redundant risk clustering cluster to the redundant risk document set;

[0022] If the similarity ratio is less than or equal to the second similarity ratio threshold and greater than or equal to the first similarity ratio threshold, mark the clustering cluster as a normal clustering cluster, and add the primary retrieval documents in the normal clustering cluster to the retained document set;

[0023] If the similarity ratio is less than the first similarity ratio threshold, mark the clustering cluster as an information missing risk clustering cluster, and add the primary retrieval documents in the information missing risk clustering cluster to the information missing risk document set.

[0024] In a preferred embodiment, in step S3, obtain the redundancy detection information and credibility evaluation information of the redundant risk document set, and perform a primary screening on the redundant risk document set to obtain a redundant document screening set, specifically as follows:

[0025] The redundancy detection information includes the second cosine similarity, and the credibility evaluation information includes the credibility evaluation coefficient;

[0026] Use the cosine similarity method to calculate the second cosine similarity between two primary retrieval documents in the redundant risk document set, compare the second cosine similarity with the preset redundancy threshold. If the second cosine similarity is greater than the redundancy threshold, respectively obtain the credibility evaluation coefficients of the two primary retrieval documents, compare the credibility evaluation coefficients of the two primary retrieval documents, and mark the primary retrieval document with the smallest credibility evaluation coefficient as the document to be deleted. If the credibility evaluation coefficients of the two primary retrieval documents are equal, randomly mark one of the primary retrieval documents as the document to be deleted;

[0027] Count the number of times each primary retrieval document in the statistical redundancy risk document set is marked as a document to be deleted, compare the number of times marked as a document to be deleted with a preset marking threshold, and add the primary retrieval documents with a number of times less than the marking threshold to the redundant document screening set.

[0028] In a preferred embodiment, the acquisition logic of the credibility evaluation coefficient is as follows:

[0029] Obtain the self-citation ratio zab of the primary retrieval document, and the expression is as follows In the formula, Z1 represents the number of times the primary retrieval document cites the documents previously published by the author himself, and Z2 represents the total number of documents cited in the primary retrieval document;

[0030] Obtain the number of times byc the primary retrieval document is cited, the number of review rounds psh, and the number of times bcj of being rejected;

[0031] Normalize the self-citation ratio, the number of times cited, the number of review rounds, and the number of times of being rejected, and calculate the credibility evaluation coefficient kxd. The expression is as follows

[0032] In a preferred embodiment, the document retrieval quality information includes a document retrieval quality coefficient, and the document freshness information includes a document freshness coefficient.

[0033] In a preferred embodiment, the acquisition logic of the document retrieval quality coefficient is as follows:

[0034] Preprocess the retrieval documents and the user input query in the final document set, including word segmentation, stop word removal, and stemming;

[0035] Calculate the document retrieval quality coefficient jsz, and the expression is as follows In the formula, N represents the total number of retrieval documents in the final document set, q i represents the i-th word in the user input query, i = {1, 2,..., n}, n is a positive integer, cp(q i , D) represents the word frequency of q i in the retrieval document D, CD represents the document length of the retrieval document D, pjc represents the average length of the retrieval documents in the final document set, and M and L are constant factors.

[0036] In a preferred embodiment, the acquisition logic of the document freshness coefficient is as follows:

[0037] Obtain the release time of the retrieval document and the user query time, and calculate the time difference T1 from the release time of the retrieval document to the user query time;

[0038] Obtain the content update time of each retrieval document, and calculate the time difference T2 between adjacent content updates j , where j represents the sequence number of the content update event trigger;

[0039] Calculate the average value TPJ of the time differences between adjacent content updates. The expression is as follows In the formula, j = {1, 2,..., J};

[0040] Calculate the standard deviation TPB of the time differences between adjacent content updates. The expression is as follows

[0041] Calculate the document freshness coefficient wdj. The expression is as follows

[0042] In a preferred embodiment, normalize the obtained document retrieval quality coefficient and document freshness coefficient, construct a retrieval document screening model based on the normalized document retrieval quality coefficient and document freshness coefficient, and generate a document priority recommendation index wyxt. The formula on which the model is based is as follows wyxt = e (α*jsz+β*wdj) , where α and β respectively represent the preset proportionality coefficients of the document retrieval quality coefficient and the document freshness coefficient, and both α and β are greater than 0;

[0043] Sort the documents in the final document set in descending order according to the document priority recommendation index to obtain a document sorting table, and display the documents to the user in the document sorting order in turn.

[0044] The technical effects and advantages of the present invention are as follows:

[0045] 1. The present invention initializes and parses the query input by the user to generate a core query statement and a primary retrieval document set, more accurately understands the user's query intention, performs clustering analysis on the primary retrieval document set to identify redundant risk document sets, retained document sets, and information missing risk document sets, obtains redundant detection information and credibility evaluation information of the redundant risk document sets, and performs a primary screening on the redundant risk document sets to obtain a redundant document screening set, effectively filtering redundant and less credible retrieval documents, improving the quality and relevance of the final document set, generating supplementary query statements according to the information missing risk document sets, and performing a secondary retrieval in the knowledge base according to the supplementary query statements to obtain an information supplementary document set, more comprehensively covering the user's multi-angle needs and potential fuzzy intentions.

[0046] 2. Through the comprehensive calculation of the document quality coefficient and the document freshness coefficient, the present invention generates a document priority recommendation index, ensuring that in the final document set, documents with high quality and strong timeliness are preferentially displayed to users, effectively improving the user experience, enabling users to obtain high-quality information that meets their needs faster and more accurately during retrieval, and avoiding the problem of incomplete results caused by insufficient information coverage in traditional single-query retrieval methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;

[0048] Figure 1 It is a schematic structural diagram of the method of the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Embodiment: Figure 1 A multi-level retrieval enhancement method for large language models of the present invention is given, including the following steps:

[0051] Step S1, initially parse the query input by the user, generate a core query statement according to the initial parsing result, and perform a primary retrieval in the knowledge base according to the core query statement to obtain a primary retrieval document set;

[0052] Step S2, perform a clustering analysis on the primary retrieval document set and identify a redundant risk document set, a retained document set, and an information missing risk document set;

[0053] Step S3, obtain the redundancy detection information and credibility evaluation information of the redundant risk document set, and perform a primary screening on the redundant risk document set to obtain a redundant document screening set;

[0054] Step S4, generate a supplementary query statement according to the information missing risk document set, and perform a secondary retrieval in the knowledge base according to the supplementary query statement to obtain an information supplementary document set;

[0055] Step S5: Merge the retained document set, the redundant document screening set, and the information supplement document set to obtain the final document set. Obtain the document retrieval quality information and document freshness information of the final document set, generate a document priority recommendation index, sort the documents in the final document set from largest to smallest according to the document priority recommendation index to obtain a document sorting table, and display the documents to the user in the document sorting order;

[0056] In step S1, initially parse the query input by the user, generate a core query statement according to the initial parsing result, and perform a primary retrieval in the knowledge base based on the core query statement to obtain a primary retrieval document set, specifically as follows:

[0057] Perform word segmentation on the query input by the user, extract important keywords and phrases. Common word segmentation tools include, for example, SpaCy, NLTK, etc.

[0058] Perform part-of-speech tagging on the segmented words to identify words of different parts of speech such as nouns, verbs, adjectives, etc., which helps to identify the key components of the input query;

[0059] Input the tagged words into a topic model to extract the query topic in the input query. The query topic includes query topic words and query topic phrases. Common topic models include Latent Dirichlet Allocation, LDA, etc.;

[0060] Use the extracted query topic words and query topic phrases as the core query statement;

[0061] It should be noted that the extracted query topic words and query topic phrases may include multiple groups. Therefore, the core query statement includes one to multiple core query statements;

[0062] Use the generated core query statement to perform a primary retrieval in the knowledge base, and obtain a primary retrieval document set related to the core query statement. Different retrieval models can be selected according to the actual application scenario, such as boolean retrieval, vector space model, BM25, or a retrieval model based on deep learning (such as BERT-based retriever), etc.;

[0063] In step S2, perform cluster analysis on the primary retrieval document set and identify the redundant risk document set, the retained document set, and the information missing risk document set, specifically as follows:

[0064] Step S21: Preprocess the primary retrieval document set, including removing noise and standardizing the text format, and convert each preprocessed primary retrieval document into a vector representation;

[0065] It should be noted that to convert the primary retrieval document into a vector representation, methods such as TF-IDF, word embedding (such as Word2Vec, BERT), etc. can be used to generate the vector representation of the primary retrieval document. Each dimension in the vector can be a corresponding vocabulary in the primary retrieval document;

[0066] Step S22, use the elbow method to determine the number K of initial clustering centroids;

[0067] It should be noted that when using the elbow method to determine the number K of initial clustering centroids, it starts from K = 1 and gradually increases the value of K. For each value of K, the K-means clustering algorithm is executed, and the SSE value after each clustering is recorded. The SSE value refers to the sum of the squares of the distances from each primary retrieval document to its clustering centroid. Taking the K value as the abscissa and the SSE value as the ordinate to draw a curve graph. In the curve graph, find the position where the decline rate of the SSE value slows down significantly as the K value increases. This point usually presents the shape of an "elbow". Calculate the curve slopes at different K values, compare the curve slopes at different K values with the curve elbow slope threshold range, and take the K value corresponding to the curve slope that first appears within the curve elbow slope threshold range as the final number K of initial clustering centroids;

[0068] Step S23, according to the determined number K of initial clustering centroids, randomly select K primary retrieval documents from the primary retrieval document set as the initial clustering cluster centroids;

[0069] Step S24, calculate the Euclidean distance from each primary retrieval document in the primary retrieval document set to the initial clustering cluster centroid, and assign each primary retrieval document to the clustering cluster closest to it;

[0070] Step S25, calculate the average vector of each clustering cluster and use it as the new clustering cluster centroid;

[0071] Step S26, repeat Step S24 and Step S25 to iteratively update the clustering clusters until the clustering cluster centroids no longer change, and obtain the final clustering cluster set;

[0072] The final clustering cluster set contains multiple clustering clusters;

[0073] Step S27, preprocess the core query statement, convert the preprocessed core query statement into a vector representation, and use the cosine similarity method to calculate the first cosine similarity xsd1 between each primary retrieval document in the clustering cluster and the core query statement. The expression is as follows In the formula, A represents the vector representation of the primary retrieval document in the clustering cluster, B represents the vector representation of the core query statement, A·B represents the dot product operation of A and B, ||A|| represents the Euclidean norm of the primary retrieval document vector, and ||B|| represents the Euclidean norm of the core query statement vector;

[0074] The range of the first cosine similarity value obtained by using the cosine similarity method is between -1 and 1, and the specific meanings are as follows:

[0075] 1: It means that the directions of the two vectors are exactly the same, that is, the angle between them is 0 degrees, which means that the two vectors are very similar;

[0076] 0: It means that the two vectors are perpendicular to each other, that is, the angle between them is 90 degrees, which means that the two vectors have no similarity;

[0077] -1: It means that the directions of the two vectors are exactly opposite, that is, the angle between them is 180 degrees, which means that the two vectors are completely dissimilar;

[0078] It should be noted that the preprocessing operation on the core query statement is the same as that on the primary retrieval document set;

[0079] In the clustering cluster, count the number of occurrences C1 where the first cosine similarity between the primary retrieval document and the core query statement is greater than 0 and less than or equal to 1; count the number of occurrences C2 where the first cosine similarity between the primary retrieval document and the core query statement is less than 0 and greater than or equal to -1, and calculate the similarity ratio xsz. The expression is as follows

[0080] The similarity ratio reflects the degree of overall similarity between all primary retrieval documents and the core query statement in the clustering cluster. When the similarity ratio is extremely high, it indicates that there are a large number of primary retrieval documents with high similarity to the core query statement in the clustering cluster. In this case, there may be a high degree of content repetition among the documents within the clustering cluster, and redundant information may be generated; while when the similarity ratio is extremely low, it indicates that there are few primary retrieval documents with high similarity to the core query statement in the clustering cluster, which means that the primary retrieval documents in the clustering cluster may not fully meet the information requirements reflected by the core query statement, and there is a risk of information loss;

[0081] Step S28, compare the similarity ratio with the preset similarity ratio stage threshold, mark the clustering cluster, identify the redundant risk document set, the retained document set, and the information loss risk document set. The similarity ratio stage threshold includes the first similarity ratio threshold and the second similarity ratio threshold, and the first similarity ratio threshold is less than the second similarity ratio threshold;

[0082] If the similarity ratio is greater than the second similarity ratio threshold, mark the clustering cluster as a redundant risk clustering cluster, and add the primary retrieval documents in the redundant risk clustering cluster to the redundant risk document set;

[0083] If the similarity ratio is less than or equal to the second similarity ratio threshold and greater than or equal to the first similarity ratio threshold, mark the cluster as a normal cluster, and add the primary retrieval documents in the normal cluster to the retained document set;

[0084] If the similarity ratio is less than the first similarity ratio threshold, mark the cluster as a cluster at risk of information loss, and add the primary retrieval documents in the cluster at risk of information loss to the document set at risk of information loss;

[0085] In step S3, obtain the redundancy detection information and credibility evaluation information of the redundant risk document set, and perform a primary screening on the redundant risk document set to obtain a redundant document screening set, specifically as follows:

[0086] The redundancy detection information includes the second cosine similarity, and the credibility evaluation information includes the credibility evaluation coefficient;

[0087] Use the cosine similarity method to calculate the second cosine similarity between two primary retrieval documents in the redundant risk document set, compare the second cosine similarity with the preset redundancy threshold. If the second cosine similarity is greater than the redundancy threshold, obtain the credibility evaluation coefficients of the two primary retrieval documents respectively, compare the credibility evaluation coefficients of the two primary retrieval documents, and mark the primary retrieval document with the smallest credibility evaluation coefficient as the document to be deleted. If the credibility evaluation coefficients of the two primary retrieval documents are equal, randomly mark one of the primary retrieval documents as the document to be deleted;

[0088] Count the number of times each primary retrieval document in the redundant risk document set is marked as the document to be deleted, compare the number of times marked as the document to be deleted with the preset marking threshold, and add the primary retrieval documents with a number of times less than the marking threshold to the redundant document screening set;

[0089] It should be noted that the method for calculating the second cosine similarity between two primary retrieval documents in the redundant risk document set using the cosine similarity method is the same as the method for calculating the first cosine similarity, where the two primary retrieval documents are any two primary retrieval documents in the redundant risk document set;

[0090] Traverse the primary retrieval documents in the redundant risk document set, and complete the marking of the primary retrieval documents in the redundant risk document set using the above method. Since all primary retrieval documents in the redundant risk document set need to be traversed, the number of times the same primary retrieval document is marked as the document to be deleted includes zero to multiple times;

[0091] The credibility evaluation coefficient is used to measure the information reliability and authority of the primary retrieval documents in the redundant risk document set, which helps to identify which primary retrieval documents have a high credibility, so as to provide an objective reference basis in the screening process and determine the primary retrieval documents that should be retained or deleted;

[0092] The acquisition logic of the credibility evaluation coefficient is as follows:

[0093] Obtain the self-citation ratio zab of the primary retrieval document, and the expression is as follows In the formula, Z1 represents the number of documents published by the author himself / herself previously cited in the primary retrieval document, and Z2 represents the total number of documents cited in the primary retrieval document;

[0094] Obtain the number of citations byc, the number of review rounds psh, and the number of rejection times bcj of the primary retrieval document;

[0095] Normalize the self-citation ratio, the number of citations, the number of review rounds, and the number of rejection times, and calculate the credibility evaluation coefficient kxd. The expression is as follows

[0096] After the author submits the document to a journal or conference, the editor will send it to peer reviewers for review. The reviewers will give feedback and suggestions, which may include modifications, supplements, or deletions to the content of the document. The author will modify the document according to the reviewers' feedback and then resubmit the modified document to the review institution. This process is called a review round. Multiple review rounds usually indicate that the document has undergone a more rigorous review process. The reviewers have conducted multiple detailed inspections and modifications on the content of the document. The finally approved document may be of high accuracy, integrity, and rigor in terms of content;

[0097] Step S4, generate supplementary query statements according to the information missing risk document set, and perform a secondary retrieval in the knowledge base according to the supplementary query statements to obtain the information supplementary document set;

[0098] Perform word segmentation on the primary retrieval documents in the information missing risk document set, and extract important keywords and phrases.

[0099] Perform part-of-speech tagging on the segmented words to identify words of different parts of speech such as nouns, verbs, and adjectives;

[0100] Input the tagged words into the topic model to extract the document topics of the primary retrieval documents in the information missing document set. The document topics include document topic words and document topic phrases;

[0101] Compare the extracted document themes with the query themes, identify the theme information not covered in the document set at risk of information loss, and determine the theme areas that are not adequately covered. These areas may include specific themes, time, geographical locations, target users, etc.;

[0102] Based on the identified uncovered theme areas, use a multi-query retriever to generate multiple supplementary query statements. These supplementary query statements will expand the original query content from different perspectives to ensure more comprehensive coverage of all aspects of the user's query in the secondary retrieval. For example, if information related to "low budget" is missing under the travel theme, supplementary queries such as "the best travel destinations with low budget" can be generated;

[0103] Use the generated supplementary query statements to perform a secondary retrieval in the knowledge base to obtain an information supplementary document set;

[0104] In step S5, merge the retained document set, the redundant document screening set, and the information supplementary document set to obtain the final document set. Obtain the document retrieval quality information and document freshness information of the final document set, generate a document priority recommendation index, sort the documents in the final document set from largest to smallest according to the document priority recommendation index to obtain a document sorting table, and display the documents to the user in the document sorting order;

[0105] The described document retrieval quality information includes a document retrieval quality coefficient, and the document freshness information includes a document freshness coefficient;

[0106] The document retrieval quality coefficient is used to measure the matching degree between the retrieved documents in the final document set and the user's input query. A higher document retrieval quality coefficient indicates a higher matching degree between the retrieved documents in the final document set and the user's input query, reflecting a higher retrieval quality of the documents and a correspondingly greater probability of being preferentially recommended;

[0107] The acquisition logic of the document retrieval quality coefficient is as follows:

[0108] Preprocess the retrieved documents in the final document set and the user input query, including word segmentation, stop word removal, and stemming;

[0109] Calculate the document retrieval quality coefficient jsz, and the expression is as follows In the formula, N represents the total number of retrieved documents in the final document set, q i represents the i-th word in the user input query, i = {1, 2,..., n}, n is a positive integer, cp(q i , D) represents q iRetrieve the word frequency in document D. The word frequency refers to the number of times a word in the user input query appears in the retrieved document. CD represents the document length of retrieved document D, pjc represents the average length of retrieved documents in the final document set, and M and L are constant factors, which are set according to the actual situation;

[0110] M is a constant factor used to control the influence of word frequency on the document retrieval quality coefficient, which determines the saturation of word frequency in calculating the document retrieval quality coefficient. L is a constant factor used to adjust the influence of the document length of the retrieved document on the word frequency, which controls the normalization effect of the retrieved document length in the calculation;

[0111] The document freshness coefficient is used to measure the timeliness or update degree of retrieved documents in the final document set. The higher the document freshness coefficient, the newer the content of the retrieved document, indicating that its relevance and value at the current time point are greater. In the process of document sorting and recommendation, documents with higher freshness coefficients are more likely to be displayed to users first because they usually contain more timely and practically significant information;

[0112] The acquisition logic of the document freshness coefficient is as follows:

[0113] Obtain the release time of the retrieved document and the user query time, and calculate the time difference T1 from the release time of the retrieved document to the user query time;

[0114] Obtain the content update time of the retrieved document each time, and calculate the time difference T2 between adjacent content updates j , where j represents the sequence number of the content update event trigger;

[0115] Calculate the average value TPJ of the time differences between adjacent content updates. The expression is as follows In the formula, j = {1, 2,..., J};

[0116] Calculate the standard deviation TPB of the time differences between adjacent content updates. The expression is as follows

[0117] Calculate the document freshness coefficient wdj. The expression is as follows

[0118] Normalize the obtained document retrieval quality coefficient and document freshness coefficient, and construct a retrieved document screening model based on the normalized document retrieval quality coefficient and document freshness coefficient to generate a document priority recommendation index wyxt. The formula on which the model is based is as follows wyxt = e (α*jsz+β*wdj) , where α and β respectively represent the preset proportionality coefficients of the document retrieval quality coefficient and the document freshness coefficient, and both α and β are greater than 0;

[0119] It should be noted that α and β are set according to the actual situation. For example, the expert weight assignment method is adopted, that is, experts in the relevant field are invited to determine the preset proportional coefficients of each index through professional opinion surveys and comprehensive evaluations. In addition, various methods such as the analytic hierarchy process and the fuzzy comprehensive evaluation method can also be considered to determine the preset proportional coefficients to ensure the objectivity and scientificity of the preset proportional coefficients, which will not be elaborated here;

[0120] As can be seen from the above calculation expressions, the larger the document retrieval quality coefficient and the larger the document freshness coefficient, the larger the document priority recommendation index, indicating that the probability of the retrieved document being recommended first is greater. On the contrary, the smaller the document retrieval quality coefficient and the smaller the document freshness coefficient, the smaller the document priority recommendation index, indicating that the probability of the retrieved document being recommended first is smaller;

[0121] Sort the documents in the final document set from large to small according to the document priority recommendation index to obtain a document sorting table, and display the documents to the user in the order of the document sorting;

[0122] The present invention initializes and parses the query input by the user to generate a core query statement and a primary retrieved document set, more accurately understands the user's query intention, performs clustering analysis on the primary retrieved document set to identify redundant risk document sets, retained document sets, and information missing risk document sets, obtains redundant detection information and credibility evaluation information of the redundant risk document sets, and performs a primary screening on the redundant risk document sets to obtain a redundant document screening set, effectively filtering redundant and low-credibility retrieved documents, improving the quality and relevance of the final document set, generating a supplementary query statement according to the information missing risk document set, and performing a secondary retrieval in the knowledge base according to the supplementary query statement to obtain an information supplementary document set, more comprehensively covering the user's multi-angle needs and potential fuzzy intentions;

[0123] The present invention generates a document priority recommendation index through the comprehensive calculation of the document quality coefficient and the document freshness coefficient, ensuring that in the final document set, documents with high quality and strong timeliness are preferentially displayed to the user, effectively improving the user experience, enabling the user to obtain high-quality information that meets their needs faster and more accurately during retrieval, and avoiding the problem of incomplete results caused by insufficient information coverage in traditional single-query retrieval methods;

[0124] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0125] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not imply the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0126] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A large language model multi-level retrieval enhancement method, characterized by: The steps include: Step S1, performing an initial analysis on the query input by the user, generating a core query statement based on the initial analysis result, and performing a primary search in the knowledge base based on the core query statement to obtain a primary search document set; Segment the query entered by the user to extract important keywords and phrases; Perform part-of-speech tagging on the words after word segmentation to identify words with different parts of speech such as nouns, verbs, and adjectives; Input the annotated words into the topic model to extract the query topics in the input query; The extracted query subject words and query subject phrases are used as core query statements; Use the generated core query statement to perform a primary search in the knowledge base to obtain a primary search document set related to the core query statement; Step S2, performing cluster analysis on the primary search document set and identifying a redundant risk document set, a retained document set, and an information missing risk document set; Step S3, obtaining redundant detection information and credibility assessment information of the redundant risk document set, performing a primary screening on the redundant risk document set, and obtaining a redundant document screening set; Step S4, generating a supplementary query statement based on the information missing risk document set, and performing a secondary search in the knowledge base based on the supplementary query statement to obtain an information supplementary document set; Step S5, merging the retained document set, the redundant document screening set, and the information supplement document set to obtain a final document set, obtaining document retrieval quality information and document freshness information of the final document set, generating a document priority recommendation index, sorting the documents in the final document set from large to small according to the document priority recommendation index, obtaining a document sorting table, and displaying the documents to the user in order according to the document sorting order; The document retrieval quality information includes a document retrieval quality coefficient, and the document freshness information includes a document freshness coefficient; The logic for obtaining the document retrieval quality coefficient is as follows: Preprocess the retrieved documents and user input queries in the final document set, including word segmentation, stop word removal, and stemming; Calculate the document retrieval quality coefficient jsz, the expression is as follows Where N represents the total number of retrieved documents in the final document set, q i represents the i-th word in the user input query, i = {1, 2, ..., n}, n is a positive integer, cp(q i ,D) represents q i The frequency of the words in the retrieved document D, CD represents the document length of the retrieved document D, pjc represents the average length of the retrieved documents in the final document set, and M and L are constant factors; The M is a constant factor for controlling the effect of word frequency on the document retrieval quality coefficient, and L is a constant factor for adjusting the effect of the document length of the retrieved document on the word frequency; The logic for obtaining the document freshness coefficient is as follows: Obtain the publishing time of the retrieved document and the user query time, and calculate the time difference T1 from the publishing time of the retrieved document to the user query time; Get the content update time of each retrieved document and calculate the time difference T2 between adjacent content updates j , j represents the sequence number of the content update event trigger; Calculate the average TPJ of the time difference between adjacent content updates. The expression is as follows Where j = {1, 2, ..., J}; Calculate the standard deviation TPB of the time difference between adjacent content updates. The expression is as follows Calculate the document freshness coefficient wdj, the expression is as follows 2. The large language model multi-level retrieval enhancement method according to claim 1, characterized in that: In step S2, cluster analysis is performed on the primary search document set to identify redundant risk document sets, retained document sets, and information missing risk document sets, as follows: Step S21, preprocessing the primary search document set, including removing noise, standardizing the text format, and converting each preprocessed primary search document into a vector representation; Step S22, using the elbow method to determine the number K of initial cluster centroids; Step S23, according to the determined number K of initial clustering centroids, randomly selecting K primary retrieval documents from the primary retrieval document set as initial clustering centroids; Step S24, calculating the Euclidean distance between each primary search document in the primary search document set and the centroid of the initial cluster, and assigning each primary search document to the cluster closest to it; Step S25, calculating the average vector of each cluster and taking it as the new cluster centroid; Step S26, repeating steps S24 and S25 to iteratively update the clusters until the centroid of the clusters no longer changes, thereby obtaining a final cluster set; Step S27, preprocess the core query statement, convert the preprocessed core query statement into a vector representation, and use the cosine similarity method to calculate the first cosine similarity xsd1 between each primary search document in the cluster and the core query statement. The expression is as follows: Where A represents the vector representation of the primary retrieval document in the cluster, B represents the vector representation of the core query statement, A·B represents the dot product operation of A and B, ||A|| represents the Euclidean norm of the primary retrieval document vector, and ||B|| represents the Euclidean norm of the core query statement vector; In the cluster, count the number of occurrences C1 where the first cosine similarity between the primary search document and the core query statement is greater than 0 and less than or equal to 1; count the number of occurrences C2 where the first cosine similarity between the primary search document and the core query statement is less than 0 and greater than or equal to -1, and calculate the similarity ratio xsz, which is expressed as follows Step S28, comparing the similarity ratio with a preset similarity ratio stage threshold, marking the clusters, identifying the redundant risk document set, the retained document set, and the information missing risk document set, the similarity ratio stage threshold includes a similarity ratio first threshold and a similarity ratio second threshold, and the similarity ratio first threshold is less than the similarity ratio second threshold; If the similarity ratio is greater than the second similarity ratio threshold, the cluster is marked as a redundant risk cluster, and the primary search document in the redundant risk cluster is added to the redundant risk document set; If the similarity ratio is less than or equal to the second similarity ratio threshold and greater than or equal to the first similarity ratio threshold, the cluster is marked as a normal cluster, and the primary search documents in the normal cluster are added to the reserved document set; If the similarity ratio is less than the first similarity ratio threshold, the cluster is marked as an information missing risk cluster, and the primary search documents in the information missing risk cluster are added to the information missing risk document set.

3. The large language model multi-level retrieval enhancement method according to claim 2, characterized in that: In step S3, redundant detection information and credibility assessment information of the redundant risk document set are obtained, and the redundant risk document set is initially screened to obtain a redundant document screening set, as follows: The redundant detection information includes a second cosine similarity, and the credibility evaluation information includes a credibility evaluation coefficient; The cosine similarity method is used to calculate the second cosine similarity between two primary retrieval documents in the redundant risk document set, and the second cosine similarity is compared with a preset redundancy threshold. If the second cosine similarity is greater than the redundancy threshold, the credibility evaluation coefficients of the two primary retrieval documents are respectively obtained, and the credibility evaluation coefficients of the two primary retrieval documents are compared. The primary retrieval document with the smallest credibility evaluation coefficient is marked as a document to be deleted. If the credibility evaluation coefficients of the two primary retrieval documents are equal, one of the primary retrieval documents is randomly marked as a document to be deleted; The number of times each primary retrieval document in the redundant risk document set is marked as a document to be deleted is counted, the number of times the document is marked as a document to be deleted is compared with a preset marking threshold, and the primary retrieval documents whose number is less than the marking threshold are added to the redundant document screening set.

4. The large language model multi-level retrieval enhancement method according to claim 3, characterized in that: The logic for obtaining the credibility evaluation coefficient is as follows: Get the self-citation ratio of primary search documents zab, the expression is as follows Where Z1 represents the number of times the primary search document cites the author's previously published documents, and Z2 represents the number of all documents cited in the primary search document; Get the number of citations byc, the number of review rounds psh, and the number of rejections bcj of the primary search document; The self-citation ratio, citation count, review rounds, and rejection count are normalized to calculate the credibility evaluation coefficient kxd, which is expressed as follows:

5. The method for enhancing multi-level retrieval using a large language model according to claim 4, characterized in that: The obtained document retrieval quality coefficient and document freshness coefficient are normalized, and a retrieval document screening model is constructed based on the normalized document retrieval quality coefficient and document freshness coefficient to generate a document priority recommendation index wyxt. The model is based on the following formula: wyxt = e (α*jsz+β*wdj) , where α and β represent the preset proportional coefficients of the document retrieval quality coefficient and the document freshness coefficient, respectively, and both α and β are greater than 0; According to the document priority recommendation index, the documents in the final document set are sorted from large to small to obtain a document sorting table, and the documents are displayed to the user in sequence according to the document sorting order.

Citation Information

Patent Citations

  • Method for establishing and searching feature matrix of Web document based on semantics

    CN101251841A

  • Text query method and device, electronic equipment and storage medium

    CN118484517A