Knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval
By adopting the post-blocking adaptive re-search method in knowledge retrieval, the problems of context information loss and insufficient knowledge mastery of professional fields caused by text segmentation in the prior art are solved, and more efficient, accurate and reliable information retrieval is achieved, meeting users' accurate question-and-answer needs.
Patent Information
- Application Number
- CN202510030059.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing technology has unreasonable text segmentation in knowledge retrieval, resulting in loss of context information, insufficient knowledge mastery in professional fields, limited reasoning ability, lack of interpretability and traceability in generated information, prominent privacy and security issues, and traditional search strategies are difficult to accurately capture user intentions, resulting in limited retrieval efficiency and accuracy.
The knowledge retrieval enhancement method based on post-blocking adaptive retrieval is adopted. By setting the confidence threshold of a predetermined large model, the word segmentation below is adaptively retrieval is performed. The knowledge document is divided into text blocks using the post-blocking strategy and the Colbert embedding model, and word vector grouping is stored in the vector database to perform semantic similarity matching and weighting calculations to optimize the search results.
It improves the accuracy and reliability of information retrieval, enhances the relevance and processing efficiency of information, improves the adaptability and robustness of the system, and can better meet users' accurate question-and-answer needs.
Smart Images

Figure CN119961402A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of retrieval enhancement generation, and mainly to a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval. Background Art
[0002] Knowledge retrieval enhancement technology was born in the context of the rapid development of artificial intelligence and natural language processing technology. It aims to meet the challenges of rapid knowledge update, long-tail data processing, answer verification and tracing, and field-specific knowledge needs. This technology combines the advanced results of deep learning, natural language processing and information retrieval, and provides users with more accurate, rich and context-related information generation and retrieval solutions through strategies such as knowledge parameterization, multimodal retrieval, and intelligent guidance and extended retrieval.
[0003] However, the existing technology clusters sentences with similar semantic vectors into text blocks, but often makes unreasonable segmentation when dividing the text, resulting in the loss of contextual information in the text blocks. As a result, the embedded vectors cannot fully represent the semantic information contained in the text blocks, affecting the accuracy of retrieval. In addition, the existing technology lacks in-depth knowledge in professional fields and has limited reasoning ability, making it difficult to meet the needs of accurate question and answering. The generated information lacks explainability and traceability, which reduces user trust. When processing sensitive data, privacy and security issues become prominent. In addition, traditional retrieval strategies may not be able to accurately capture user intentions, and their efficiency and accuracy are limited when faced with massive data and complex queries.
[0004] For example, a Chinese invention patent with publication number "CN117573815A" discloses a "retrieval enhancement generation method based on vector similarity matching optimization", which specifically discloses "determining the feature vector of the text to be processed, obtaining the similarity between the candidate vector and the feature vector of the text to be processed based on the candidate vector in the specified vector database, the feature vector of the text to be processed and the similarity prediction model, determining the reference text associated with the text to be processed based on the similarity, and inputting the prompt text determined by the text to be processed and the reference text into the pre-trained natural language model to obtain the target text". However, this method relies on the similarity between the candidate vector in the vector database and the feature vector of the text to be processed to determine the reference text, but this method does not take into account the problem of insufficient confidence that may occur in the process of generating the target text, resulting in deviations in some key information or expressions in the generated text, affecting the overall quality and practicality of the text; in addition, this method cannot make full use of the semantic information in the text, and after generating the target text, no further feedback and iteration mechanism is provided to optimize the results, resulting in the correlation between the retrieved reference text and the text to be processed is not strong enough. Summary of the invention
[0005] In order to solve the above problems existing in the prior art, the present application provides a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval.
[0006] The technical solution of this application is as follows:
[0007] On the one hand, the present invention proposes a knowledge retrieval enhancement method based on post-block adaptive re-retrieval, the method comprising:
[0008] Acquire user data and set a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; preprocess the knowledge documents, input the preprocessed knowledge documents into a Colbert embedding model using a post-blocking strategy, obtain word vector groups based on sentences, and divide the preprocessed knowledge documents into text blocks based on semantics; store the word vector groups and text blocks in a vector database;
[0009] Input the query statement into the Colbert embedding model to obtain a query vector; match the group of word vectors corresponding to the query vector in the vector database, and use the text block corresponding to the group of the matched word vectors as a recalled document;
[0010] The preset prompt words and the recalled documents are spliced as the input data of the predetermined large model, and the answer result is output; the confidence corresponding to each segmentation when the predetermined large model generates the answer result is calculated, and if the segmentation with the largest confidence is lower than the confidence threshold, the segmentation is marked as a keyword, and adaptive re-retrieval is performed, specifically:
[0011] Perform keyword masking on the keyword to obtain keyword masked word segmentation; perform weighted calculation on the similarity corresponding to the keyword masked word segmentation and the similarity corresponding to the word segmentation without keyword masking using an improved Colbert similarity algorithm to obtain a total similarity, wherein the similarity corresponding to the keyword masked word segmentation is specifically the similarity between the keyword masked word segmentation and the query vector; the similarity corresponding to the word segmentation without keyword masking is specifically the similarity between the word segmentation without keyword masking in the answer result and the query vector;
[0012] In the recalled documents, similarity search is performed on a sentence basis, and different search methods are adopted according to user needs to obtain corresponding re-search results, which are then input into a predetermined large model to regenerate answer results; until the confidence of all word segments in the answer result is higher than the preset confidence threshold, the final answer result is obtained.
[0013] Preferably, the preprocessing of the knowledge document includes document format normalization, document cleaning and word segmentation; the method of inputting the preprocessed knowledge document into the embedding model using the post-blocking strategy to obtain the sentence vector is specifically:
[0014] The preprocessing includes document format normalization, document cleaning and word segmentation; the document cleaning specifically includes removing irrelevant characters in the document, including punctuation marks, numbers and special symbols;
[0015] Identify the text content in the preprocessed knowledge document, use Jieba word segmentation technology to segment the text content, and obtain the word segmentation document T = (S1, S2, S3, ..., S N )=(t1,t2,t3,...,t M ), where T represents the word segmentation document; S i represents the i-th document sentence; N represents the number of sentences; i represents the index value of the i-th document sentence; t j represents the jth participle; M represents the number of participles; j represents the index value of the jth participle;
[0016] The word segmentation document is input into the Colbert embedding model for vectorization to generate a word vector TF=Colbert(t1,t2,t3,...,t M )=(e1,e2,e3,...,e M ), where TF represents document vector; e j represents the jth word vector; M represents the number of word vectors, which are equal to the number of word segments;
[0017] Classify all word vectors belonging to the same sentence in the preprocessed knowledge document to obtain word vector grouping Ed x =(e1,e2,e3,...,e n ), of which Ed x represents the x-th word vector group; x represents the index value of the x-th word vector group; n represents the number of word vector groups.
[0018] Preferably, the method for dividing the preprocessed knowledge document into text blocks based on semantics is specifically:
[0019] Compare the similarity of the document sentences and calculate the document sentence S r With S r+1 The similarity between r Represents the r-th document sentence, and r represents the index value of the r-th document sentence;
[0020] If the similarity is greater than the preset similarity threshold, S r With S r+1 Merge, Update S r+1 =S r ∪S r+1 , the similarity comparison process is repeated until the similarity is less than the preset similarity threshold, then the cycle is stopped and B is updated. b={S r+1}, where B b Represents the bth text block;
[0021] From S r+1 Re-compare the similarity until the preprocessed knowledge document is divided and the text block set B = {B1, B2, ..., B g}, where g represents the number of text blocks.
[0022] Preferably, in the step of calculating the similarity between the query vector and the word vector grouping, the similarity calculation formula is:
[0023]
[0024] In the formula, S q,d represents the similarity between the query vector and the word vector grouping; Eq f represents the f-th query vector; f represents the index value of the f-th query vector; T represents the transpose operation; max represents the maximum value function; |||| represents the modulus symbol;
[0025] The word vector groups whose similarity exceeds a preset threshold are marked, and the text blocks containing the word vector groups are used as recalled documents.
[0026] Preferably, the confidence of the answer result is calculated and expressed as:
[0027] L a =ω·t a +b;
[0028]
[0029] Where, L a represents the original preference logits of the ath word in the answer result; ω represents the preset weight value; b represents the preset bias; p a represents the confidence of the ath word segment in the answer result; A represents the number of word segments in the answer result; a represents the index value of the ath word segment in the answer result;
[0030] Select the word with the highest confidence from all the word segments that generate the answer results. If the highest confidence is less than the confidence threshold, mark the word as a keyword.
[0031] Preferably, keyword masking is performed on the keyword to obtain keyword masked word segmentation, specifically:
[0032] The keyword masking process is specifically to replace the keyword with a mask symbol [mask] in the answer result to indicate the uncertainty of the keyword in the answer result, thereby obtaining the keyword masked word segmentation; input the keyword masked word segmentation into the embedding model to generate a new word segmentation vector; the new word segmentation vector is specifically to replace the marker [CLS] of the query vector with the marker [Q], and replace the marker [CLS] of the word vector grouping in the recalled document with the marker [D], thereby obtaining a new query vector and a recalled document vector, which are expressed as:
[0033] NEq:=Normalize(CNN(BERT("[Q],NEq0,NEq1,...,#,..,NEq l ,..,#")));
[0034] NEd:=Filter(Normalize(CNN(BERT("[D],NEd0,NEd1,...,NEd n "))));
[0035] Among them, NEq: represents the new query vector; NEd: represents the recalled document vector; # represents the mask symbol [mask]; NEq represents the value of the new query vector; NEd represents the value of the recalled document vector; BERT represents the BERT model; CNN represents the convolutional neural network; Normalize represents the normalization function; Filter represents the recalled document vector corresponding to the filtered punctuation marks;
[0036] The improved Colbert similarity algorithm is used to perform weighted calculation on the similarity of keyword masked segmentation and the similarity of segmentation without keyword masking, which can be expressed as:
[0037]
[0038] S=λS k,d +(1-λ)NS q,d ;
[0039] In the formula, S represents the total similarity; S k,d Indicates the similarity of keyword mask segmentation; NS q,d Indicates the similarity of word segments without keyword masking; NEq u represents the value of the u-th new query vector, where u∈[||K||]; NEd c Represents the value of the cth recalled document vector; NEq vrepresents the value of the vth new query vector, where v∈[||NEq\K||]; P represents the number of keywords in the new query vector; H represents the number of word segments in all new query vectors; u represents the index value of the uth keyword; v represents the index value of the vth word segment without keyword masking; c represents the index value of the cth recalled document vector; K represents the set of all keywords; λ represents the preset weight.
[0040] Preferably, similarity retrieval is performed in units of sentences in the recalled documents, and different retrieval methods are adopted according to user needs, wherein:
[0041] If the user's demand is to prioritize accuracy, the word vector grouping with the highest similarity value is selected as the re-search result; if the user's demand is to prioritize completeness, all word vector groups with similarity values greater than the confidence threshold are selected as the re-search result; the re-search result is input into the predetermined large model to regenerate the answer result; until the confidence of all word segments in the answer result is higher than the preset confidence threshold, the final answer result is obtained.
[0042] On the other hand, a knowledge retrieval enhancement generation system based on adaptive re-retrieval is provided, the system comprising a data acquisition module, a recall information generation module, a confidence judgment module and an answer output module, wherein:
[0043] The data acquisition module is used to acquire user data and set a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; preprocess the knowledge documents, input the preprocessed knowledge documents into the Colbert embedding model using a post-blocking strategy, obtain word vector groups based on sentences, and divide the preprocessed knowledge documents into text blocks based on semantics; store the word vector groups and text blocks in a vector database;
[0044] The recall information generation module is used to input the query statement into the Colbert embedding model to obtain a query vector; match the group of word vectors corresponding to the query vector in the vector database, and use the text block corresponding to the group of the matched word vectors as the recall document; and transmit the answer result to the confidence judgment module;
[0045] The confidence judgment module is used to splice the preset prompt word and the recalled document as input data of the predetermined large model, and output the answer result; calculate the confidence corresponding to each segmentation when the predetermined large model generates the answer result, if the segmentation with the largest confidence is lower than the confidence threshold, mark the segmentation as a keyword, and then perform adaptive re-search;
[0046] The re-search module specifically performs keyword masking processing on the keywords to obtain keyword masked word segmentations; uses the improved Colbert similarity algorithm to perform weighted calculation on the similarity corresponding to the keyword masked word segmentations and the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing to obtain the total similarity, wherein the similarity corresponding to the keyword masked word segmentations is specifically the similarity between the keyword masked word segmentations and the query vector; the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing is specifically the similarity between the word segmentations that have not been subjected to keyword masking processing in the answer result and the query vector; performs similarity retrieval in the recalled documents in units of sentences, adopts different retrieval methods according to user needs, obtains corresponding re-search results, and inputs the re-search results into a predetermined large model to regenerate the answer results; until the confidence of all word segmentations in the answer result is higher than the preset confidence threshold, the final answer result is obtained;
[0047] The answer output module is used to output the final answer result.
[0048] On the other hand, the present invention also proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements a knowledge retrieval enhancement method based on post-block adaptive re-retrieval as described in any embodiment of the present invention.
[0049] On the other hand, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a knowledge retrieval enhancement method based on post-block adaptive re-retrieval as described in any embodiment of the present invention.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] 1) The present invention provides a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval, which improves the accuracy of information retrieval and enhances the reliability and relevance of information by setting a confidence threshold of a predetermined large model and adaptively re-retrieves the segmented words below the threshold, continuously iterating and optimizing the retrieval results;
[0052] 2) The present invention provides a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval, which utilizes the post-block strategy and Colbert embedding model to efficiently divide the pre-processed knowledge documents into text blocks, and stores word vector groups in units of sentences, thereby simplifying the data structure, accelerating the matching speed of query vectors in the vector database, and improving the processing efficiency of the entire system;
[0053] 3) The present invention provides a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval. By introducing an adaptive re-retrieval mechanism, keyword masking is performed on word segmentations with insufficient confidence in the answer results, and weighted calculation is performed using an improved Colbert similarity algorithm, thereby enhancing the adaptability and robustness of the system and improving the flexibility of query requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a method flow chart of an embodiment of the present invention;
[0055] Figure 2 4 is a flow chart of the post-blocking strategy of an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0057] The present invention provides the following technical solution: a knowledge retrieval enhancement method and system based on post-block adaptive re-retrieval.
[0058] Example 1
[0059] See Figure 1 This embodiment provides a knowledge retrieval enhancement method based on post-block adaptive re-retrieval, and the specific steps include:
[0060] S1, obtaining user data and setting a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; preprocessing the knowledge documents; the preprocessing includes document format normalization, document cleaning and word segmentation; the document cleaning specifically includes removing irrelevant characters in the document, including punctuation marks, numbers and special symbols;
[0061] S2, see Figure 2 , the preprocessed knowledge documents are input into the Colbert embedding model using the post-blocking strategy to obtain word vector groupings in sentences;
[0062] The preprocessing includes document format normalization, document cleaning and word segmentation; the document cleaning specifically includes removing irrelevant characters in the document, including punctuation marks, numbers and special symbols;
[0063] Identify the text content in the preprocessed knowledge document, use Jieba word segmentation technology to segment the text content, and obtain the word segmentation document T = (S1, S2, S3, ..., S N )=(t1,t2,t3,...,t M ), where T represents the word segmentation document; S i represents the i-th document sentence; N represents the number of sentences; i represents the index value of the i-th document sentence; t j represents the jth participle; M represents the number of participles; j represents the index value of the jth participle;
[0064] The word segmentation document is input into the Colbert embedding model for vectorization to generate a word vector TF=Colbert(t1,t2,t3,...,t M )=(e1,e2,e3,...,e M ), where TF represents document vector; e j represents the jth word vector; M represents the number of word vectors, which are equal to the number of word segments;
[0065] Classify all word vectors belonging to the same sentence in the preprocessed knowledge document to obtain word vector grouping Ed x =(e1,e2,e3,...,e n ), of which Ed x represents the xth word vector group; x represents the index value of the xth word vector group; n represents the number of word vector groups;
[0066] S3, dividing the preprocessed knowledge document into text blocks based on semantics; grouping the word vectors and the text blocks, and storing them in a vector database;
[0067] Compare the similarity of the document sentences and calculate the document sentence S r With S r+1 The similarity between r Represents the r-th document sentence, and r represents the index value of the r-th document sentence;
[0068] If the similarity is greater than the preset similarity threshold, S r With S r+1 Merge, Update S r+1 =S r ∪S r+1 , the similarity comparison process is repeated until the similarity is less than the preset similarity threshold, then the cycle is stopped and B is updated. b ={S r+1}, where B b Represents the bth text block;
[0069] From S r+1Re-compare the similarity until the preprocessed knowledge document is divided and the text block set B = {B1, B2, ..., B g}, where g represents the number of text blocks;
[0070] S4, inputting the query statement into the Colbert embedding model to obtain a query vector; matching the group of the word vector corresponding to the query vector in the vector database, and taking the text block corresponding to the group of the matched word vector as the recalled document;
[0071] In the step of calculating the similarity between the query vector and the word vector grouping, the similarity calculation formula is:
[0072]
[0073] In the formula, S q,d represents the similarity between the query vector and the word vector grouping; Eq f represents the f-th query vector; f represents the index value of the f-th query vector; T represents the transpose operation; max represents the maximum value function; |||| represents the modulus symbol;
[0074] Mark word vector groups whose similarity exceeds a preset threshold, and use the text blocks corresponding to the word vector groups as recall documents;
[0075] S5. The preset prompt word and the recalled document are concatenated as input data of the predetermined large model, and the answer result is output; the confidence corresponding to each segmentation when the predetermined large model generates the answer result is calculated. If the segmentation with the largest confidence is lower than the confidence threshold, the segmentation is marked as a keyword, and an adaptive re-search is performed, specifically:
[0076] S51. Calculate the confidence of the answer result, expressed as:
[0077] L a =ω·t a +b;
[0078]
[0079] Where, L a represents the original preference logits of the ath word in the answer result; ω represents the preset weight value; b represents the preset bias; p a represents the confidence of the ath word segment in the answer result; A represents the number of word segments in the answer result; a represents the index value of the ath word segment in the answer result;
[0080] Select the word with the largest confidence from all the word segments that generate the answer results. If the largest confidence is less than the confidence threshold, mark the word as a keyword.
[0081] S52, performing keyword masking processing on the keyword to obtain keyword masked word segmentation;
[0082] The keyword masking process is specifically to replace the keyword with a mask symbol [mask] in the answer result to indicate the uncertainty of the keyword in the answer result, thereby obtaining the keyword masked word segmentation; input the keyword masked word segmentation into the embedding model to generate a new word segmentation vector; the new word segmentation vector is specifically to replace the marker [CLS] of the query vector with the marker [Q], and replace the marker [CLS] of the word vector grouping in the recalled document with the marker [D], thereby obtaining a new query vector and a recalled document vector, which are expressed as:
[0083] NEq:=Normalize(CNN(BERT("[Q],NEq0,NEq1,...,#,..,NEq l ,..,#")));
[0084] NEd:=Filter(Normalize(CNN(BERT("[D],NEd0,NEd1,...,NEd n "))));
[0085] Among them, NEq: represents the new query vector; NEd: represents the recalled document vector; # represents the mask symbol [mask]; NEq represents the value of the new query vector; NEd represents the value of the recalled document vector; BERT represents the BERT model; CNN represents the convolutional neural network; Normalize represents the normalization function; Filter represents the recalled document vector corresponding to the filtered punctuation marks;
[0086] S53, using the improved Colbert similarity algorithm, weightedly calculate the similarity corresponding to the keyword masked word segmentation and the similarity corresponding to the word segmentation without keyword masking processing, to obtain a total similarity, wherein the similarity corresponding to the keyword masked word segmentation is specifically the similarity between the keyword masked word segmentation and the query vector; the similarity corresponding to the word segmentation without keyword masking processing in the answer result is specifically the similarity between the word segmentation without keyword masking processing in the answer result and the query vector;
[0087] The total similarity is expressed as follows:
[0088]
[0089] S=λS k,d +(1-λ)NS q,d ;
[0090] In the formula, S represents the total similarity; S k,d Indicates the similarity of keyword mask segmentation; NSq,d Indicates the similarity of word segments without keyword masking; NEq u represents the value of the u-th new query vector, where u∈[||K||]; NEd c Represents the value of the cth recalled document vector; NEq v represents the value of the vth new query vector, where v∈[||NEq\K||]; P represents the number of keywords in the new query vector; H represents the number of word segments in all new query vectors; u represents the index value of the uth keyword; v represents the index value of the vth word segment without keyword masking; c represents the index value of the cth recalled document vector; K represents the set of all keywords; λ represents the preset weight;
[0091] S6. Perform similarity search on the recalled documents in units of sentences, and adopt different search methods according to user needs. If the user needs to prioritize accuracy, select the word vector group with the highest similarity value as the re-search result; if the user needs to prioritize completeness, select all word vector groups with similarity values greater than the confidence threshold as the re-search result; input the re-search result into the large model to regenerate the answer result; until the confidence of all word segments in the answer result is higher than the preset confidence threshold, the final answer result is obtained.
[0092] Example 2
[0093] This embodiment provides a knowledge retrieval enhancement system based on post-block adaptive re-retrieval, the system comprising a data acquisition module, a recall information generation module, a confidence judgment module and an answer output module, wherein:
[0094] The data acquisition module is used to acquire user data and set a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; preprocess the knowledge documents, input the preprocessed knowledge documents into the Colbert embedding model using a post-blocking strategy, obtain word vector groups based on sentences, and divide the preprocessed knowledge documents into text blocks based on semantics; store the word vector groups and text blocks in a vector database;
[0095] The recall information generation module is used to input the query statement into the Colbert embedding model to obtain a query vector; match the group of word vectors corresponding to the query vector in the vector database, and use the text block corresponding to the group of the matched word vectors as the recall document; and transmit the answer result to the confidence judgment module;
[0096] The confidence judgment module is used to splice the preset prompt word and the recalled document as input data of the predetermined large model, and output the answer result; calculate the confidence corresponding to each segmentation when the predetermined large model generates the answer result, if the segmentation with the largest confidence is lower than the confidence threshold, mark the segmentation as a keyword, and then perform adaptive re-search;
[0097] The re-search module specifically performs keyword masking processing on the keywords to obtain keyword masked word segmentations; uses the improved Colbert similarity algorithm to perform weighted calculation on the similarity corresponding to the keyword masked word segmentations and the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing to obtain the total similarity, wherein the similarity corresponding to the keyword masked word segmentations is specifically the similarity between the keyword masked word segmentations and the query vector; the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing is specifically the similarity between the word segmentations that have not been subjected to keyword masking processing in the answer result and the query vector; performs similarity retrieval in the recalled documents in units of sentences, adopts different retrieval methods according to user needs, obtains corresponding re-search results, and inputs the re-search results into the large model to regenerate the answer results; until the confidence of all word segmentations in the answer result is higher than the preset confidence threshold, the final answer result is obtained;
[0098] The answer output module is used to output the final answer result.
[0099] Example 3
[0100] This embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, a knowledge retrieval enhancement method based on post-block adaptive re-retrieval as described in any embodiment of the present invention is implemented.
[0101] Example 4
[0102] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the method for enhancing knowledge retrieval based on post-block adaptive re-retrieval as described in any embodiment of the present invention is implemented.
[0103] It is worth noting that the system, electronic device and computer-readable storage medium described in the present invention are based on the same principles as the method described in Example 1, and will not be described in detail here.
[0104] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A knowledge retrieval enhancement method based on post-block adaptive re-retrieval, characterized in that: The method comprises: Acquire user data and set a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; preprocess the knowledge documents, input the preprocessed knowledge documents into a Colbert embedding model using a post-blocking strategy, obtain word vector groups based on sentences, and divide the preprocessed knowledge documents into text blocks based on semantics; store the word vector groups and text blocks in a vector database; Input the query statement into the Colbert embedding model to obtain a query vector; match the group of word vectors corresponding to the query vector in the vector database, and use the text block corresponding to the group of the matched word vectors as a recalled document; The preset prompt words and the recalled documents are spliced as the input data of the predetermined large model, and the answer result is output; the confidence corresponding to each segmentation when the predetermined large model generates the answer result is calculated, and if the segmentation with the largest confidence is lower than the confidence threshold, the segmentation is marked as a keyword, and adaptive re-retrieval is performed, specifically: Perform keyword masking on the keyword to obtain keyword masked word segmentation; perform weighted calculation on the similarity corresponding to the keyword masked word segmentation and the similarity corresponding to the word segmentation without keyword masking using an improved Colbert similarity algorithm to obtain a total similarity, wherein the similarity corresponding to the keyword masked word segmentation is specifically the similarity between the keyword masked word segmentation and the query vector; the similarity corresponding to the word segmentation without keyword masking is specifically the similarity between the word segmentation without keyword masking in the answer result and the query vector; In the recalled documents, similarity search is performed on a sentence basis, and different search methods are adopted according to user needs to obtain corresponding re-search results, which are then input into a predetermined large model to regenerate answer results; until the confidence of all word segments in the answer result is higher than the preset confidence threshold, the final answer result is obtained.
2. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 1 is characterized in that: The preprocessing of the knowledge document includes document format normalization, document cleaning and word segmentation; the method of inputting the preprocessed knowledge document into the embedding model using the post-blocking strategy to obtain the sentence vector is specifically as follows: The preprocessing includes document format normalization, document cleaning and word segmentation; the document cleaning specifically includes removing irrelevant characters in the document, including punctuation marks, numbers and special symbols; Identify the text content in the preprocessed knowledge document, use Jieba word segmentation technology to segment the text content, and obtain the word segmentation document T = (S1, S2, S3, ..., S N )=(t1,t2,t3,...,t M ), where T represents the word segmentation document; S i represents the i-th document sentence; N represents the number of sentences; i represents the index value of the i-th document sentence; t j represents the jth participle; M represents the number of participles; j represents the index value of the jth participle; The word segmentation document is input into the Colbert embedding model for vectorization to generate a word vector TF=Colbert(t1,t2,t3,...,t M )=(e1,e2,e3,...,e M ), where TF represents document vector; e j represents the jth word vector; M represents the number of word vectors, which are equal to the number of word segments; Classify all word vectors belonging to the same sentence in the preprocessed knowledge document to obtain word vector grouping Ed x =(e1,e2,e3,...,e n ), of which Ed x represents the x-th word vector group; x represents the index value of the x-th word vector group; n represents the number of word vector groups.
3. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 1 is characterized in that: The method of dividing the preprocessed knowledge document into text blocks based on semantics is specifically as follows: Compare the similarity of the document sentences and calculate the document sentence S r With S r+1 The similarity between r Represents the r-th document sentence, and r represents the index value of the r-th document sentence; If the similarity is greater than the preset similarity threshold, S r With S r+1 Merge, Update S r+1 =S r ∪S r+1 , the similarity comparison process is repeated until the similarity is less than the preset similarity threshold, then the loop is stopped and B is updated. b ={S r+1 }, where B b Represents the bth text block; From S r+1 Re-compare the similarity until the preprocessed knowledge document is divided and the text block set B = {B1, B2, ..., B g }, where g represents the number of text blocks.
4. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 1 is characterized in that: In the step of calculating the similarity between the query vector and the word vector grouping, the similarity calculation formula is: In the formula, S q,d represents the similarity between the query vector and the word vector grouping; Eq f represents the f-th query vector; f represents the index value of the f-th query vector; T represents the transpose operation; max represents the maximum value function; |||| represents the modulus symbol; The word vector groups whose similarity exceeds a preset threshold are marked, and the text blocks corresponding to the word vector groups are used as recalled documents.
5. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 1 is characterized in that: Calculate the confidence of the answer result, expressed as: L a =ω·t a +b; Where, L a represents the original preference logits of the ath word in the answer result; ω represents the preset weight value; b represents the preset bias; p a represents the confidence of the ath word segment in the answer result; A represents the number of word segments in the answer result; a represents the index value of the ath word segment in the answer result; Select the word with the highest confidence from all the word segments that generate the answer results. If the highest confidence is less than the confidence threshold, mark the word as a keyword.
6. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 5 is characterized in that: The keyword is subjected to keyword masking processing to obtain keyword masked word segmentation, specifically: The keyword masking process is specifically to replace the keyword with a mask symbol [mask] in the answer result to indicate the uncertainty of the keyword in the answer result, thereby obtaining the keyword masked word segmentation; input the keyword masked word segmentation into the embedding model to generate a new word segmentation vector; the new word segmentation vector is specifically to replace the marker [CLS] of the query vector with the marker [Q], and replace the marker [CLS] of the word vector grouping in the recalled document with the marker [D], thereby obtaining a new query vector and a recalled document vector, which are expressed as: NEq:=Normalize(CNN(BERT("[Q],NEq0,NEq1,...,#,..,NEq l ,..,#"))); NEd:=Filter(Normalize(CNN(BERT("[D],NEd0,NEd1,...,NEd n ")))); Among them, NEq: represents the new query vector; NEd: represents the recalled document vector; # represents the mask symbol [mask]; NEq represents the value of the new query vector; NEd represents the value of the recalled document vector; BERT represents the BERT model; CNN represents the convolutional neural network; Normalize represents the normalization function; Filter represents the recalled document vector corresponding to the filtered punctuation marks; The improved Colbert similarity algorithm is used to perform weighted calculation on the similarity of keyword masked segmentation and the similarity of segmentation without keyword masking, which can be expressed as: S=λS k,d +(1-λ)NS q,d ; In the formula, S represents the total similarity; S k,d Indicates the similarity of keyword mask segmentation; NS q,d Indicates the similarity of word segments without keyword masking; NEq u represents the value of the u-th new query vector, where u∈[||K||]; NEd c Represents the value of the cth recalled document vector; NEq v represents the value of the vth new query vector, where v∈[||NEq\K||]; P represents the number of keywords in the new query vector; H represents the number of word segments in all new query vectors; u represents the index value of the uth keyword; v represents the index value of the vth word segment without keyword masking; c represents the index value of the cth recalled document vector; K represents the set of all keywords; λ represents the preset weight.
7. The knowledge retrieval enhancement method based on post-block adaptive re-retrieval according to claim 1 is characterized in that: In the recalled documents, similarity retrieval is performed on a sentence-by-sentence basis. Different retrieval methods are adopted according to user needs, including: If the user's demand is to prioritize accuracy, the word vector grouping with the highest similarity value is selected as the re-search result; if the user's demand is to prioritize completeness, all word vector groups with similarity values greater than the confidence threshold are selected as the re-search result; the re-search result is input into the predetermined large model to regenerate the answer result; until the confidence of all word segments in the answer result is higher than the preset confidence threshold, the final answer result is obtained.
8. A knowledge retrieval enhancement generation system based on adaptive re-retrieval, characterized in that: The system includes a data acquisition module, a recall information generation module, a confidence judgment module and an answer output module, wherein: The data acquisition module is used to acquire user data and set a confidence threshold of a predetermined large model; the user data includes knowledge documents and query statements; the knowledge documents are preprocessed, and the preprocessed knowledge documents are input into the embedding model using a post-blocking strategy to obtain word vector groups based on sentences, and the preprocessed knowledge documents are divided into text blocks based on semantics; the word vector groups based on sentences and the text blocks are stored in a vector database; The recall information generation module is used to input the query statement into the embedding model to obtain a query vector; match the word vector group corresponding to the query vector in the vector database, and use the text block where the word vector group is located as the recall document; splice the preset prompt word and the recall document as input data of the predetermined large model, and output the answer result; transmit the answer result to the confidence judgment module; The confidence judgment module is used to splice the preset prompt word and the recalled document as input data of the predetermined large model, and output the answer result; calculate the confidence corresponding to each segmentation when the predetermined large model generates the answer result, if the segmentation with the largest confidence is lower than the confidence threshold, mark the segmentation as a keyword, and then perform adaptive re-search; The re-search module specifically performs keyword masking processing on the keywords to obtain keyword masked word segmentations; uses the improved Colbert similarity algorithm to perform weighted calculation on the similarity corresponding to the keyword masked word segmentations and the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing to obtain the total similarity, wherein the similarity corresponding to the keyword masked word segmentations is specifically the similarity between the keyword masked word segmentations and the query vector; the similarity corresponding to the word segmentations that have not been subjected to keyword masking processing in the answer result is specifically the similarity between the word segmentations that have not been subjected to keyword masking processing in the answer result and the query vector; performs similarity retrieval in the recalled document in units of sentences, adopts different retrieval methods according to user needs, obtains corresponding re-search results, and inputs the re-search results into a predetermined large model to regenerate the answer result; until the confidence of all word segmentations in the answer result is higher than the preset confidence threshold, the final answer result is obtained; The answer output module is used to output the final answer result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the knowledge retrieval enhancement method based on post-block adaptive re-retrieval as described in claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the knowledge retrieval enhancement method based on post-block adaptive re-retrieval as described in claims 1 to 7 is implemented.
Citation Information
Patent Citations
Retrieval enhancement generation method based on vector similarity matching optimization
CN117573815A
Text Retrieval Method Based on Confidence-Correlation Coefficient Framework Mining Rule Predecessor
CN109299278A
Knowledge graph question answering method based on question graph iterative retrieval
CN113342952A
Text multi-hop reasoning question and answer method using two-stage retrieval
CN116383347A
Processing method and device for adjusting answer text based on adaptive retrieval enhancement mechanism
CN118939766A
Cited By
Retrieval enhancement generation method and device
CN120541206A
Bid invitation information retrieval method and device based on multi-path mixed recall mechanism
CN120763307A
Scientific research information retrieval matching method and device based on large model
CN120929557A