Dense and sparse retrieval fusion method and device based on summary words
The fusion of summary word prediction and dense sparse search through large language models solves the problem of complementary fusion of dense sparse search methods, reduces calculation overhead and improves retrieval accuracy.
Patent Information
- Application Number
- CN202510475350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing dense search and sparse search methods have their own advantages and disadvantages, and no effective technical solutions can be complementary and integrated, resulting in high computing resource overhead and insufficient retrieval accuracy.
A large language model is used to predict the summary word, and the query and document are converted into dense semantic vectors for calculation through dense sparse search fusion method, and the probability distribution vector is used to expand the keyword range for sparse search, and finally the fusion search results are combined.
It reduces the overhead of computing resources, improves the search accuracy, and achieves the complementary fusion effect of dense and sparse searches.
Smart Images

Figure CN120336506A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a dense-sparse retrieval fusion method and device based on summary words. Background Art
[0002] In the field of information retrieval, dense retrieval and sparse retrieval are two common retrieval methods. Dense retrieval is a retrieval method based on semantic features. This type of method converts the query text (referred to as the query) and the retrieval document (referred to as the document) into dense semantic vectors in a high-dimensional feature space respectively, and screens the retrieval results by calculating the vector similarity. Sparse retrieval is a retrieval method based on the inverted index. This type of method matches the query keywords with the documents where the keywords appear, then calculates the relevance score of the documents according to the term matching degree between the documents and the query, and finally determines the documents relevant to the query based on the sorting result of the relevance scores. Compared with the two, the retrieval accuracy of dense retrieval is higher than that of sparse retrieval, and the computational resource overhead of sparse retrieval is lower than that of dense retrieval. These two types of retrieval methods have their own advantages and disadvantages, and there is currently no relatively general technical solution to complement and fuse the two.
[0003] The natural language learning of large language models (LLMs) is very powerful and can fully understand the personalized semantics of each token in the input text and the context-related semantics. The next-word prediction task is a typical pre-training task for LLM models. This task requires the model to predict the next word text of the input text. Before training this task, a prediction network (also called the next-word task prediction head) connected by a linear layer and a Softmax function layer is connected to the output side of the model, and an instruction template that can generate the input text in a configurable manner is customized for the model. The prediction process of this task is as follows: The text configured by the instruction template is used as the input text (regarded as composed of multiple tokens connected together) and sent into the LLM model for forward inference to obtain a text hidden vector H (composed of multiple token hidden vectors h), and the last token hidden vector h of the text hidden vector H end is sent into the next-word task prediction head for forward inference to obtain a probability distribution vector P with the total number of words in the specified vocabulary (the dictionary or vocabulary used by the LLM model) as the vector length (composed of multiple word probabilities p, and the word probability p corresponds one-to-one with the words in the specified vocabulary), and the word corresponding to the maximum probability in the probability distribution vector P is used as the predicted next word text.
[0004] Two insights can be obtained from the next-word prediction task: 1) The feature space of the text hidden vector H is a high-dimensional semantic feature space. Each token hidden vector h can be regarded as a dense vector, and the last token hidden vector h end can be regarded as a dense vector highly relevant to the predicted word; 2) From the probability distribution vector P, the relevance distribution of a vocabulary can be observed.
[0005] Based on the first insight, the following analysis results can be obtained: If the query / document can be condensed into a summary word respectively through the next-word prediction task, and then the dense vectors highly relevant to these two summary words are calculated, it will surely greatly reduce the computational resource overhead of dense retrieval.
[0006] Based on the second insight, the following analysis results can be obtained: If the query can be condensed into a query summary word through the next-word prediction task, then from the probability distribution vector P, a vocabulary distribution related to the query summary word can be observed. Regarding the words with higher probabilities in this vocabulary distribution as pseudo keywords indirectly related to the query, the keyword range can be further expanded. Then, sparse retrieval based on this expanded keyword range can naturally improve the retrieval accuracy.
[0007] Furthermore, if the retrieval results obtained from the above two types of analysis can be fused, it will surely achieve the complementary fusion purpose of improving the retrieval accuracy and reducing the computational overhead.
[0008] The technical problems to be solved by the present invention are exactly: how to perform retrieval based on the above two types of analysis and how to fuse the two types of retrieval results. Summary of the Invention
[0009] The purpose of the present invention is to provide a dense-sparse retrieval fusion method, device, electronic device, and computer-readable storage medium based on summary words in view of the defects of the prior art. The present invention first selects a large language model that has completed the pre-training task as the working model, and uses the next-word task prediction head and instruction template used by the model in its next-word prediction task as the corresponding summary word prediction head and summary word instruction template, and uses the specified vocabulary list of the model as the corresponding working vocabulary list; then, takes the target retrieval library as the first library, and creates a corresponding document summary word dense vector mapping space, i.e., the first space, for the first library using the summary word instruction template, the working model, and the summary word prediction head; then, when receiving the query text (i.e., the first query) input by the user, first uses the summary word instruction template, the working model, and the summary word prediction head to identify the summary word dense vector h q and the probability distribution vector P q of the first query, then performs keyword identification according to the first query, the probability distribution vector P q and the working vocabulary list, and then according to the summary word dense vector hq Densely retrieve the first library against the first space, and sparsely retrieve the first library according to the keyword sequence. Finally, feedback the fusion result of the dense and sparse retrieval results to the user. In the dense retrieval of the present invention, only the dense vectors of the summary words of the query and the document are calculated and compared, thereby reducing the computational resource overhead. In the sparse retrieval, the keyword range is expanded using pseudo keywords, thereby improving the retrieval accuracy. After obtaining the dense and sparse retrieval results, the two are fused to further improve the retrieval accuracy. By the present invention, both the computational overhead can be reduced and the retrieval precision can be improved.
[0010] To achieve the above object, a first aspect of an embodiment of the present invention provides a method for fusing dense and sparse retrieval based on summary words, the method comprising:
[0011] Select a large language model that has completed a pre-training task as the corresponding working model; and use the next-word task prediction head and instruction template used by the working model in its next-word prediction task as the corresponding summary word prediction head and summary word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary; the pre-training task includes the next-word prediction task;
[0012] Use the preset retrieval library as the corresponding first library; and use the summary word instruction template, the working model, and the summary word prediction head to create a corresponding document summary word dense vector mapping space for the first library, denoted as the first space;
[0013] Receive the query text input by the user as the first query;
[0014] Use the summary word instruction template, the working model, and the summary word prediction head to identify the dense vector h q and the probability distribution vector P q of the summary words of the first query; and perform keyword identification according to the first query, the probability distribution vector P q and the working vocabulary to obtain the corresponding first keyword sequence;
[0015] According to the dense vector h q and the first space, densely retrieve the first library to obtain the corresponding first retrieval sequence;
[0016] Sparsely retrieve the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence;
[0017] Fuse the retrieval results according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback it to the current user.
[0018] Preferably, the summary word prediction head is formed by connecting a linear layer and a Softmax function layer;
[0019] The summary word instruction template is a formatted instruction text template; the configurable parameters of the summary word instruction template include text configuration parameters; the summary word instruction template is used to use the text configuration parameters as the corresponding current text and prompt the working model to perform word prediction on the summary words of the current text;
[0020] The working vocabulary includes multiple first words w i , 1 ≤ index i ≤ N W , N W is the total number of words in the working vocabulary; each of the first words w i is an independent word or word text;
[0021] The first library includes multiple first documents D j , 1 ≤ index j ≤ N D , N D is the total number of documents in the first library;
[0022] The spatial dimension of the first space matches the hidden vector feature dimension output by the working model; the first space includes multiple first space points n j , the first space point n j corresponds one-to-one with the first document D j ;
[0023] The first keyword sequence includes multiple first keywords, and each first keyword corresponds to a keyword weight;
[0024] The first retrieval sequence includes multiple first retrieval documents, and each first retrieval document corresponds to a first document score; the score range of the first document score is between 0 and 1;
[0025] The second retrieval sequence includes multiple second retrieval documents, and each second retrieval document corresponds to a second document score; the score range of the second document score is between 0 and 1;
[0026] The third retrieval sequence includes multiple third retrieval documents, and each third retrieval document corresponds to a third document score; the score range of the third document score is between 0 and 1.
[0027] Preferably, using the summary word instruction template, the working model and the summary word prediction head to create a corresponding document summary word dense vector mapping space for the first library, denoted as the first space, specifically includes:
[0028] Step 31: Take each of the first documents D in the first library j as the corresponding current document; set the text configuration parameter of the summary word instruction template to the current document to obtain a document summary word prediction instruction text, denoted as the corresponding first instruction text; input the first instruction text into the working model for forward inference, and denote the text hidden vector output by the current model inference as the hidden vector H j ; and take the last token hidden vector of the hidden vector H j as the corresponding document summary word dense vector, denoted as the dense vector h j,end ;
[0029] Among them, the hidden vector H j includes multiple token hidden vectors h j,k , 1 ≤ index k ≤ N j , N j is the total number of tokens obtained by the working model for tokenizing the first document D j according to its corresponding tokenization rule; the dense vector h j,end is the corresponding token hidden vector
[0030] Step 32: Initialize a multi-dimensional space based on the hidden vector feature dimension output by the working model, denoted as the corresponding first space; and denote the matching point of each dense vector h j,end in the first space as the corresponding first space point n j .
[0031] Preferably, using the summary word instruction template, the working model, and the summary word prediction head to identify the summary word dense vector and probability distribution vector of the first query to obtain the corresponding dense vector h q and probability distribution vector P q , specifically including:
[0032] Set the text configuration parameter of the summary word instruction template to the first query to obtain a query summary word prediction instruction text, denoted as the corresponding second instruction text; input the second instruction text into the working model for forward inference, and denote the text hidden vector output by the current model inference as the hidden vector H q ; and take the last token hidden vector of the hidden vector H q as the corresponding dense vector h q ; input the dense vector h q into the summary word prediction head for forward inference, and take the probability distribution vector output by the current prediction head inference as the corresponding probability distribution vector P q ;
[0033] Among them, the hidden vector H q includes multiple tokenized hidden vectors h q,u , where 1 ≤ index u ≤ N q , and N q is the total number of tokens obtained after the working model tokenizes the first query according to its corresponding tokenization rule; the dense vector h q is the corresponding tokenized hidden vector The probability distribution vector P q consists of N W lexical probabilities p q,i ; the lexical probability p q,i corresponds one-to-one with the first word w of the working vocabulary; the value of the lexical probability p i is between 0 and 1. q,i
[0034] Preferably, the corresponding first keyword sequence is obtained by keyword recognition based on the first query, the probability distribution vector P q and the working vocabulary, and specifically includes:
[0035] Based on a preset keyword recognition rule, keyword recognition is performed on the first query to obtain one or more corresponding keywords A; and a corresponding weight parameter a is set for each keyword A; and each weight parameter a is set to 1;
[0036] The N q lexical probabilities p W of the probability distribution vector P q,i are sorted in descending order of probability value to obtain a corresponding first probability sequence; and the first N P lexical probabilities p q,i in the first probability sequence are extracted to form a corresponding second probability sequence; and each lexical probability p q,i in the second probability sequence is used as a corresponding weight parameter b, and each lexical probability p q,i in the second probability sequence corresponding to the first word w in the working vocabulary i is used as a corresponding pseudo-keyword B; the instruction number N P is a preset positive integer;
[0037] Each keyword A and its corresponding weight parameter a are used as a group of corresponding first keywords and keyword weights; and each pseudo-keyword B and its corresponding weight parameter b are used as a group of corresponding first keywords and keyword weights; and all the obtained first keywords are used to form a corresponding first keyword sequence.
[0038] Preferably, based on the dense vector h q and the first space, perform a dense retrieval on the first library to obtain a corresponding first retrieval sequence, specifically including:
[0039] Denote the matching space point of the dense vector h q in the matching space in the first space as the current space point; and in the first space, retrieve the approximate nearest neighbor space points of the current space point based on a preset approximate nearest neighbor search algorithm to obtain a corresponding first space point set; the approximate nearest neighbor search algorithm includes the LSH algorithm and the VQ algorithm; the first space point set includes multiple first space points n j ;
[0040] Take each of the first space points n in the first space point set j as the corresponding current comparison point one by one; and take the token hidden vector corresponding to the current comparison point as the corresponding current dense vector; and based on the cosine vector similarity algorithm, calculate the similarity between the dense vector h q and the current dense vector to obtain a corresponding current similarity; and when the current similarity is greater than a preset first similarity threshold, take the first document D j corresponding to the current comparison point as a corresponding first retrieval document, and take the current similarity as a corresponding first document score; the first similarity threshold is a preset threshold parameter greater than 0;
[0041] Sort all the obtained first retrieval documents in descending order of the first document score to form the corresponding first retrieval sequence.
[0042] Preferably, based on the first keyword sequence, perform a sparse retrieval on the first library to obtain a corresponding second retrieval sequence, specifically including:
[0043] Count the total number of the first keywords in the first keyword sequence to obtain a corresponding total number N S ; and denote each of the first keywords and its corresponding keyword weight as the corresponding keyword x s , weight y s ; 1 ≤ index s ≤ N S ;
[0044] Take each of the first documents D in the first library j as the corresponding current document one by one; and set an initially zero document cumulative score for the current document; and for all the keywords x sPerform a round of traversal; and during this round of traversal, use the keyword x being currently traversed s and its corresponding weight y s as the corresponding current keyword and current weight; and when the current document contains the current keyword, use the sum of adding the current document cumulative score and the current weight as the latest document cumulative score of the current document;
[0045] Select the maximum and minimum values from the obtained N D document cumulative scores as the corresponding maximum and minimum scores; and perform normalization calculation on each of the document cumulative scores based on the maximum and minimum scores to obtain the corresponding normalized score = (document cumulative score - minimum score) / (maximum score - minimum score); the value of the normalized score ranges from 0 to 1;
[0046] Use each of the normalized scores exceeding a preset first score threshold as a corresponding second document score; and use the first document D corresponding to each of the second document scores j as a corresponding second retrieved document; and sort all the obtained second retrieved documents in descending order of the second document scores to form the corresponding second retrieval sequence.
[0047] Preferably, fusing the retrieval results according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedbacking it to the current user specifically includes:
[0048] Record all the first and second retrieved documents in the first and second retrieval sequences as documents to be screened; and set two score parameters initialized to 0 for each of the documents to be screened, denoted as the corresponding first parameter and second parameter; and identify each of the documents to be screened. If the current document to be screened exists in both the first and second retrieval sequences, reset its corresponding first and second parameters to the corresponding first and second document scores. If the current document to be screened only exists in the first retrieval sequence or the second retrieval sequence, only reset its corresponding first parameter or second parameter to the corresponding first document score or second document score;
[0049] Calculate the mean value of the first and second parameters of each of the documents to be screened to obtain the corresponding first average score; and use each of the first average scores exceeding a preset second score threshold as a corresponding third document score; and use the document to be screened corresponding to each of the third document scores as a corresponding third retrieved document; and sort all the obtained third retrieved documents in descending order of the third document scores to form the corresponding third retrieval sequence and feedback it to the current user.
[0050] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the dense-sparse retrieval fusion method based on summary words described in the first aspect above. The apparatus includes: an LLM model preparation module, a target library preparation module, a query receiving module, an LLM model processing module, a dense retrieval module, a sparse retrieval module, and a retrieval fusion module;
[0051] The LLM model preparation module is used to select a large language model that has completed a pre-training task as the corresponding working model; and use the next-word task prediction head and instruction template used by the working model in its next-word prediction task as the corresponding summary word prediction head and summary word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary; the pre-training task includes the next-word prediction task;
[0052] The target library preparation module is used to use a preset retrieval library as the corresponding first library; and use the summary word instruction template, the working model, and the summary word prediction head to create a corresponding document summary word dense vector mapping space for the first library, denoted as the first space;
[0053] The query receiving module is used to receive the query text input by the user as the first query;
[0054] The LLM model processing module is used to use the summary word instruction template, the working model, and the summary word prediction head to identify the summary word dense vector and probability distribution vector of the first query to obtain the corresponding dense vector h q and probability distribution vector P q ; and perform keyword identification according to the first query, the probability distribution vector P q and the working vocabulary to obtain the corresponding first keyword sequence;
[0055] The dense retrieval module is used to perform dense retrieval on the first library according to the dense vector h q and the first space to obtain the corresponding first retrieval sequence;
[0056] The sparse retrieval module is used to perform sparse retrieval on the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence;
[0057] The retrieval fusion module is used to perform retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback it to the current user.
[0058] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: a memory, a processor, and a transceiver;
[0059] The processor is used to be coupled with the memory, read and execute the instructions in the memory, so as to implement the method steps described in the first aspect above;
[0060] The transceiver is coupled with the processor, and the processor controls the transceiver to send and receive messages.
[0061] A fourth aspect of an embodiment of the present invention provides a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0062] An embodiment of the present invention provides a method, apparatus, electronic device, and computer-readable storage medium for dense and sparse retrieval fusion based on summary words. It can be seen from the above that in an embodiment of the present invention, a large language model that has completed a pre-training task is first selected as a working model, and the next-word task prediction head and instruction template used by the model in its next-word prediction task are used as the corresponding summary-word prediction head and summary-word instruction template, and the specified vocabulary of the model is used as the corresponding working vocabulary; then, the target retrieval library is used as the first library, and a corresponding document summary-word dense vector mapping space, that is, the first space, is created for the first library using the summary-word instruction template, the working model, and the summary-word prediction head; then, when receiving a query text (i.e., the first query) input by the user, the summary-word dense vector h q and the probability distribution vector P q are recognized first, and then keyword recognition is performed according to the first query, the probability distribution vector P q and the working vocabulary, and then the first library is densely retrieved according to the summary-word dense vector h q and the first space, and the first library is sparsely retrieved according to the keyword sequence. Finally, the fusion result of the dense and sparse retrieval results is fed back to the user. In the dense retrieval of the embodiment of the present invention, only the summary-word dense vectors of the query and the document are calculated and compared, thereby reducing the computational resource overhead. In the sparse retrieval, the keyword range is expanded using pseudo-keywords, thereby improving the retrieval accuracy. After obtaining the dense and sparse retrieval results, the two are fused to further improve the retrieval accuracy; through the embodiment of the present invention, both the computational overhead is reduced and the retrieval precision is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 is a schematic diagram of a method for dense and sparse retrieval fusion based on summary words provided in Embodiment 1 of the present invention;
[0064] Figure 2 is a module structure diagram of a device for dense and sparse retrieval fusion based on summary words provided in Embodiment 2 of the present invention;
[0065] Figure 3 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. Detailed implementation manners
[0066] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0067] Embodiment 1 of the present invention provides a method for fusing dense and sparse retrieval based on summary words, as Figure 1 shown in the schematic diagram of the method for fusing dense and sparse retrieval based on summary words provided in Embodiment 1 of the present invention. The method mainly includes the following steps:
[0068] Step 1, select a large language model that has completed a pre-training task as the corresponding working model; and use the next-word task prediction head and instruction template used by the working model in its next-word prediction task as the corresponding summary word prediction head and summary word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary.
[0069] Among them, the pre-training task includes a next-word prediction task.
[0070] Here, the working model in the embodiment of the present invention is any large language model that has completed a pre-training task, such as BERT series models, RoBERTa series models, ALBERT series models, DeBERTa series models, etc. implemented based on the Encoder-Only framework, GPT series models, Llama series models, etc. implemented based on the Decoder-Only framework, T5 series models, BART series models, Pegasus models, Longformer-Encoder-Decoder models, etc. implemented based on the Encoder-Decoder framework. Each large language model is trained based on the next-word prediction task during its pre-training task stage, and a corresponding next-word task prediction head is configured. Moreover, the structures of the next-word task prediction heads of each large language model are approximate and are all connected by a linear layer and a Softmax function layer. Therefore, the summary word prediction head in the embodiment of the present invention is connected by a linear layer and a Softmax function layer.
[0071] The inference process of the working model in the embodiments of the present invention is briefly as follows: In the first step, the input text is tokenized based on a corresponding tokenization rule to obtain a corresponding token sequence; in the second step, the token sequence is embedded and encoded based on a corresponding embedding encoding rule to obtain a corresponding embedding vector; in the third step, if the working model is implemented based on the Encoder-Only framework, then in the third step, an Encoder network extracts high-dimensional features from the embedding vector and outputs a corresponding hidden vector H. If the working model is implemented based on the Decoder-Only framework, then in the third step, a Decoder network extracts high-dimensional features from the embedding vector and outputs a corresponding hidden vector H. If the working model is implemented based on the Encoder-Decoder framework, then in the third step, an Encoder network first performs high-dimensional feature encoding on the embedding vector, and then a Decoder network extracts features based on the feature encoding vector and outputs a corresponding hidden vector H. It should be noted that the hidden vector H is composed of multiple token hidden vectors h, and the token hidden vectors h in the hidden vector H correspond one by one to the tokens in the token sequence.
[0072] The inference process of the next-word task prediction head of the working model in the embodiments of the present invention is briefly as follows: Through a linear layer, the last token hidden vector h of the last token of the hidden vector H output by the working model end is mapped to a vocabulary feature vector, and the Softmax function layer calculates the probability distribution based on the vocabulary feature vector to obtain a corresponding probability distribution vector P.
[0073] The summary word instruction template in the embodiments of the present invention is a formatted instruction text template; the configurable parameters of the summary word instruction template include text configuration parameters; the summary word instruction template is used to use the text configuration parameters as the corresponding current text and prompt the working model to perform word prediction on the summary word of the current text.
[0074] For example, the summary word instruction template can be:
[0075] "Text configuration parameter = ''; Please use a summary word to summarize the content of the text configuration parameter; Summary word = ".
[0076] The working vocabulary in the embodiments of the present invention includes multiple first vocabulary words w i , 1 ≤ index i ≤ N W , N W is the total number of vocabulary words in the working vocabulary; each first vocabulary word w i is an independent word or word text.
[0077] Step 2: Use the preset retrieval library as the corresponding first library; and create a corresponding document summary word dense vector mapping space for the first library using the summary word instruction template, working model, and summary word prediction head, denoted as the first space.
[0078] Specifically, it includes: Step 21: Use the preset retrieval library as the corresponding first library.
[0079] Here, the preset retrieval library is the pre-defined target retrieval library; the first library in the embodiment of the present invention includes multiple first documents D j , where 1 ≤ index j ≤ N D , and N D is the total number of documents in the first library.
[0080] Step 22: Create a corresponding document summary word dense vector mapping space for the first library using the summary word instruction template, working model, and summary word prediction head, denoted as the first space.
[0081] Specifically, it includes: Step 221: Take each first document D in the first library j as the corresponding current document; set the text configuration parameter of the summary word instruction template to the current document to obtain a document summary word prediction instruction text, denoted as the corresponding first instruction text; input the first instruction text into the working model for forward inference, and denote the text hidden vector output by this model inference as the hidden vector H j ; and use the last token hidden vector of the hidden vector H j as the corresponding document summary word dense vector, denoted as the dense vector h j,end .
[0082] Among them, the hidden vector H j includes multiple token hidden vectors h j,k , where 1 ≤ index k ≤ N j , and N j is the total number of tokens obtained by the working model after tokenizing the first document D j according to its corresponding tokenization rule; the dense vector h j,end is the corresponding token hidden vector
[0083] Step 222: Initialize a multi-dimensional space based on the hidden vector feature dimension output by the working model, denoted as the corresponding first space; and denote the matching point of each dense vector h j,end in the first space as the corresponding first space point n j .
[0084] Here, the space dimension of the first space in the embodiment of the present invention matches the hidden vector feature dimension output by the working model; the first space includes multiple first space points n j, the first spatial point n j corresponds to the first document D j one by one.
[0085] Step 3, receive the query text input by the user as the first query.
[0086] Step 4, use the summary word instruction template, the working model, and the summary word prediction head to identify the dense vector and probability distribution vector of the summary words of the first query to obtain the corresponding dense vector h q and probability distribution vector P q ; and according to the first query, the probability distribution vector P q and the working vocabulary to perform keyword identification to obtain the corresponding first keyword sequence;
[0087] Specifically, it includes: Step 41, use the summary word instruction template, the working model, and the summary word prediction head to identify the dense vector and probability distribution vector of the summary words of the first query to obtain the corresponding dense vector h q and probability distribution vector P q ;
[0088] Specifically, set the text configuration parameter of the summary word instruction template to the first query to obtain a query summary word prediction instruction text, denoted as the corresponding second instruction text; and input the second instruction text into the working model for forward inference, and denote the text hidden vector output by this model inference as the hidden vector H q ; and the hidden vector H q The last token hidden vector of is used as the corresponding dense vector h q ; and input the dense vector h q into the summary word prediction head for forward inference, and use the probability distribution vector output by this prediction head inference as the corresponding probability distribution vector P q ;
[0089] Among them, the hidden vector H q includes multiple token hidden vectors h q,u , 1 ≤ index u ≤ N q , N q is the total number of tokens obtained after the working model performs tokenization processing on the first query according to its corresponding tokenization rule; the dense vector h q is the corresponding token hidden vector The probability distribution vector P q is composed of N W lexical probabilities p q,i ; the lexical probability p q,i corresponds one by one to the first vocabulary w of the working vocabulary i ; the value of the lexical probability p q,i is between 0 and 1;
[0090] Step 42, and based on the first query and the probability distribution vector P q and the working vocabulary, perform keyword recognition to obtain a corresponding first keyword sequence;
[0091] Specifically, it includes: Step 421, based on a preset keyword recognition rule, perform keyword recognition on the first query to obtain one or more corresponding keywords A; and set a corresponding weight parameter a for each keyword A; and set each weight parameter a to 1;
[0092] Here, the keyword recognition rule of the embodiment of the present invention is a preset grammar rule, which can be customized according to application requirements; one customization method is to use the subject in the sentence, that is, the first query, as the keyword; another customization method is to use the nouns and / or pronouns in the sentence, that is, the first query, as the keyword;
[0093] Step 422, sort the N q vocabulary probabilities p W of the probability distribution vector P q,i in descending order of probability values to obtain a corresponding first probability sequence; and extract the top N P vocabulary probabilities p q,i from the first probability sequence to form a corresponding second probability sequence; and use each vocabulary probability p q,i in the second probability sequence as a corresponding weight parameter b, and use each vocabulary probability p q,i in the second probability sequence corresponding to the first vocabulary w i in the working vocabulary as a corresponding pseudo-keyword B;
[0094] Here, the instruction number N P of the embodiment of the present invention is a preset positive integer;
[0095] Step 43, take each keyword A and its corresponding weight parameter a as a group of corresponding first keywords and keyword weights; and take each pseudo-keyword B and its corresponding weight parameter b as a group of corresponding first keywords and keyword weights; and form a corresponding first keyword sequence from all the obtained first keywords.
[0096] Here, the first keyword sequence of the embodiment of the present invention includes multiple first keywords, and each first keyword corresponds to a keyword weight.
[0097] Step 5, perform dense retrieval on the first library according to the dense vector h q and the first space to obtain a corresponding first retrieval sequence;
[0098] Among them, the first retrieval sequence includes multiple first retrieval documents, and each first retrieval document corresponds to a first document score; the score range of the first document score is between 0 and 1;
[0099] Specifically, it includes: Step 51, regarding the matching space point of the dense vector h q in the first space as the current space point; and in the first space, retrieving the approximate nearest neighbor space points of the current space point based on a preset approximate nearest neighbor search algorithm to obtain a corresponding first space point set;
[0100] Here, the approximate nearest neighbor search algorithm of the embodiment of the present invention includes a Locality Sensitive Hashing (LSH) algorithm and a Vector Quantization (VQ) algorithm; the first space point set includes multiple first space points n j ;
[0101] Step 52, taking each first space point n in the first space point set j as the corresponding current comparison point one by one; and taking the word segmentation hidden vector corresponding to the current comparison point as the corresponding current dense vector; and based on the cosine vector similarity algorithm, calculating the similarity between the dense vector h q and the current dense vector to obtain the corresponding current similarity; and when the current similarity is greater than a preset first similarity threshold, taking the first document D j corresponding to the current comparison point as a corresponding first retrieval document, and taking the current similarity as a corresponding first document score;
[0102] Here, the first similarity threshold of the embodiment of the present invention is a preset threshold parameter greater than 0;
[0103] Step 53, sorting all the obtained first retrieval documents in descending order of the first document score to form a corresponding first retrieval sequence.
[0104] Step 6, performing a sparse retrieval on the first library according to the first keyword sequence to obtain a corresponding second retrieval sequence;
[0105] Among them, the second retrieval sequence includes multiple second retrieval documents, and each second retrieval document corresponds to a second document score; the score range of the second document score is between 0 and 1;
[0106] Specifically, it includes: Step 61, counting the total number of the first keywords in the first keyword sequence to obtain a corresponding total number N S ; and recording each first keyword and its corresponding keyword weight as the corresponding keyword x s , weight y s ;
[0107] where 1 ≤ index s ≤ N S ;
[0108] Step 62: Take each first document D in the first library j one by one as the corresponding current document; set an initialized document cumulative score of 0 for the current document; and perform a round of traversal on all keywords x s and during this round of traversal, take the currently traversed keyword x s and its corresponding weight y s as the corresponding current keyword and current weight; and when the current document contains the current keyword, take the sum of adding the current document cumulative score and the current weight as the latest document cumulative score of the current document;
[0109] Step 63: Select the maximum and minimum values from the obtained N D document cumulative scores as the corresponding maximum and minimum scores; and perform normalization calculation on each document cumulative score based on the maximum and minimum scores to obtain the corresponding normalized score = (document cumulative score - minimum score) / (maximum score - minimum score);
[0110] where the value of the normalized score ranges from 0 to 1;
[0111] Step 64: Take each normalized score exceeding the preset first score threshold as a corresponding second document score; and take the first document D corresponding to each second document score j as a corresponding second retrieved document; and sort all the obtained second retrieved documents in descending order of the second document score to form the corresponding second retrieval sequence.
[0112] Here, the first score threshold in the embodiment of the present invention is a preset threshold parameter greater than 0.
[0113] Step 7: Perform retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback it to the current user;
[0114] where the third retrieval sequence includes multiple third retrieved documents, and each third retrieved document corresponds to a third document score; the score range of the third document score is between 0 and 1;
[0115] Specifically, it includes: Step 71, marking all the first and second retrieval documents in the first and second retrieval sequences as documents to be screened; setting two score parameters initialized to 0 for each document to be screened, denoted as the corresponding first parameter and second parameter; and identifying each document to be screened. If the current document to be screened exists in both the first and second retrieval sequences, reset its corresponding first and second parameters to the corresponding first and second document scores. If the current document to be screened only exists in the first retrieval sequence or the second retrieval sequence, only reset its corresponding first parameter or second parameter to the corresponding first document score or second document score.
[0116] Step 72, calculating the average value of the first and second parameters of each document to be screened to obtain the corresponding first average score; taking each first average score that exceeds a preset second score threshold as a corresponding third document score; taking the document to be screened corresponding to each third document score as a corresponding third retrieval document; and sorting all the obtained third retrieval documents in descending order of the third document score to form a corresponding third retrieval sequence and feedback it to the current user.
[0117] Figure 2 This is the module structure diagram of a dense-sparse retrieval fusion device based on summary words provided in the second embodiment of the present invention. The device is a terminal device or a server for implementing the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, the device can be a device or a chip system of the foregoing terminal device or server. As Figure 2 shown, the device includes: an LLM model preparation module 201, a target library preparation module 202, a query receiving module 203, an LLM model processing module 204, a dense retrieval module 205, a sparse retrieval module 206, and a retrieval fusion module 207.
[0118] The LLM model preparation module 201 is used to select a large language model that has completed a pre-training task as the corresponding working model; and use the next-word task prediction head and instruction template used by the working model in its next-word prediction task as the corresponding summary word prediction head and summary word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary; the pre-training task includes the next-word prediction task.
[0119] The target library preparation module 202 is used to use the preset retrieval library as the corresponding first library; and create a corresponding document summary word dense vector mapping space for the first library using the summary word instruction template, the working model, and the summary word prediction head, denoted as the first space.
[0120] The query receiving module 203 is used to receive the query text input by the user as the first query.
[0121] The LLM model processing module 204 is used to identify the dense vector h and the probability distribution vector corresponding to the summary word dense vector and the probability distribution vector of the first query by using the summary word instruction template, the working model, and the summary word prediction head. q And the probability distribution vector P q ; And perform keyword identification according to the first query, the probability distribution vector P q And the working vocabulary to obtain the corresponding first keyword sequence.
[0122] The dense retrieval module 205 is used to perform dense retrieval on the first library according to the dense vector h q And the first space to obtain the corresponding first retrieval sequence.
[0123] The sparse retrieval module 206 is used to perform sparse retrieval on the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence.
[0124] The retrieval fusion module 207 is used to perform retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback it to the current user.
[0125] A dense and sparse retrieval fusion device based on summary words provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0126] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the LLM model preparation module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and called and executed by a certain processing element of the above device to perform the functions of the above determined modules. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0127] For example, the above-mentioned modules can be one or more integrated circuits configured to implement the above methods, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).
[0128] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the above computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium can be a magnetic medium (for example, floppy disk, hard disk, magnetic tape), an optical medium (for example, DVD), or a semiconductor medium (for example, solid state disk (SSD)), etc.
[0129] Figure 3 This is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device can be a terminal device or a server that implements the method of the foregoing embodiments, or can be a terminal device or a server that is connected to the foregoing terminal device or server and implements the method of the foregoing embodiments. As Figure 3As shown in the figure, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver actions of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the method of the foregoing embodiments. Preferably, the electronic device according to the embodiment of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0130] In Figure 3 The system bus 305 mentioned may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.
[0131] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0132] It should be noted that the embodiment of the present invention further provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the methods and processing procedures provided in the above embodiments.
[0133] An embodiment of the present invention provides a method, apparatus, electronic device, and computer-readable storage medium for fusing dense and sparse retrieval based on summary words. As can be seen from the above, in the embodiment of the present invention, a large language model that has completed a pre-training task is first selected as the working model, and the next-word task prediction head and instruction template used by the model in its next-word prediction task are used as the corresponding summary-word prediction head and summary-word instruction template, and the specified vocabulary of the model is used as the corresponding working vocabulary; then, the target retrieval library is used as the first library, and a corresponding document summary-word dense vector mapping space, that is, the first space, is created for the first library using the summary-word instruction template, the working model, and the summary-word prediction head; then, when receiving the query text (i.e., the first query) input by the user, the summary-word dense vector h q and the probability distribution vector P q are first recognized, and then keyword recognition is performed according to the first query, the probability distribution vector P q and the working vocabulary, and then the first library is densely retrieved according to the summary-word dense vector h q and the first space, and the first library is sparsely retrieved according to the keyword sequence. Finally, the fusion result of the dense and sparse retrieval results is fed back to the user. In the embodiment of the present invention, only the summary-word dense vectors of the query and the document are calculated and compared during dense retrieval, thereby reducing the computational resource overhead. During sparse retrieval, the keyword range is expanded using pseudo-keywords, thereby improving the retrieval accuracy. After obtaining the dense and sparse retrieval results, the two are fused to further improve the retrieval accuracy; through the embodiment of the present invention, both the computational overhead is reduced and the retrieval precision is improved.
[0134] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0135] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.
Claims
1. A dense and sparse retrieval fusion method based on summary words, characterized in that, The method includes: Select a large language model that has completed a pre-training task as the corresponding working model; and use the next-word task prediction head and instruction template used by the working model in its next-word prediction task as the corresponding summary-word prediction head and summary-word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary; the pre-training task includes the next-word prediction task; Use the preset retrieval library as the corresponding first library; and use the summary-word instruction template, the working model, and the summary-word prediction head to create a corresponding document summary-word dense vector mapping space for the first library, denoted as the first space; Receive the query text input by the user as the first query; Using the summary word instruction template, the working model, and the summary word prediction head to identify the dense vector h and the probability distribution vector of the summary words of the first query, obtaining the corresponding dense vector h q and the probability distribution vector P q ; and identifying the corresponding first keyword sequence according to the first query, the probability distribution vector P q and the working vocabulary According to the dense vector h q and perform a dense retrieval on the first library according to the first space to obtain a corresponding first retrieval sequence; Perform sparse retrieval on the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence; Perform retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback it to the current user.
2. The dense and sparse retrieval fusion method based on summary words according to claim 1, wherein The summary-word prediction head is formed by connecting a linear layer and a Softmax function layer; The summary-word instruction template is a formatted instruction text template; the configurable parameters of the summary-word instruction template include text configuration parameters; the summary-word instruction template is used to use the text configuration parameters as the corresponding current text and prompt the working model to perform word prediction on the summary words of the current text; The working vocabulary includes multiple first vocabulary words w i , where 1 ≤ index i ≤ N W , and N W is the total number of vocabulary words in the working vocabulary; each of the first vocabulary words w i is an independent word or text of a word; The first library includes a plurality of first documents D j , where 1 ≤ index j ≤ N D , N D is the total number of documents in the first library; The spatial dimension of the first space matches the dimension of the latent vector features output by the working model; the first space includes a plurality of first space points n j , the first space point n j corresponds to the first document D j one by one; The first keyword sequence includes multiple first keywords, and each first keyword corresponds to a keyword weight; The first retrieval sequence includes multiple first retrieval documents, and each first retrieval document corresponds to a first document score; the score range of the first document score is between 0 and 1; The second retrieval sequence includes multiple second retrieval documents, and each second retrieval document corresponds to a second document score; the score range of the second document score is between 0 and 1; The third retrieval sequence includes multiple third retrieval documents, and each third retrieval document corresponds to a third document score; the score range of the third document score is between 0 and 1.
3. The method for fusing dense and sparse retrieval based on summary words according to claim 2, wherein The step of using the summary-word instruction template, the working model, and the summary-word prediction head to create a corresponding document summary-word dense vector mapping space for the first library, denoted as the first space, specifically includes: Step 31: Take each of the first documents D in the first library j as the corresponding current document; set the text configuration parameter of the summary word instruction template to the current document to obtain a document summary word prediction instruction text, denoted as the corresponding first instruction text; input the first instruction text into the working model for forward inference, and denote the text hidden vector output by the model inference this time as hidden vector H j ; and take the last token hidden vector of the hidden vector H j as the corresponding document summary word dense vector, denoted as dense vector h j,end ; Among them, the hidden vector H j includes a plurality of tokenized hidden vectors h j,k , where 1 ≤ index k ≤ N j , and N j is the total number of tokens obtained after the working model tokenizes the first document D j according to its corresponding tokenization rule; the dense vector h j,end is the corresponding tokenized hidden vector Step 32, initialize a multi-dimensional space based on the dimension of the hidden vector features output by the working model, denoted as the corresponding first space; and denote each of the dense vectors h j,end The matching points in the first space as the corresponding first space points n j .
4. The method for fusing dense and sparse retrieval based on summary words according to claim 2, wherein Identifying the summary word dense vector and the probability distribution vector of the first query using the summary word instruction template, the working model, and the summary word prediction head to obtain the corresponding dense vector h q and the probability distribution vector P q , specifically including: Set the text configuration parameter of the summary word instruction template to the first query to obtain a query summary word prediction instruction text, denoted as the corresponding second instruction text; and input the second instruction text into the working model for forward inference, and denote the text hidden vector output by the current model inference as the hidden vector H q ; and use the last token hidden vector of the hidden vector H q as the corresponding dense vector h q ; and input the dense vector h q into the summary word prediction head for forward inference, and use the probability distribution vector output by the current prediction head inference as the corresponding probability distribution vector P q ; Among them, the latent vector H q includes multiple tokenized latent vectors h q,u , where 1 ≤ index u ≤ N q , and N q is the total number of tokens obtained after the working model tokenizes the first query according to its corresponding tokenization rule; the dense vector h q is the corresponding tokenized latent vector The probability distribution vector P q consists of N W lexical probabilities p q,i The lexical probability p q,i corresponds one-to-one with the first word w of the working vocabulary i ; the value of the lexical probability p q,i is between 0 and 1.
5. The method for fusing dense and sparse retrieval based on summary words according to claim 4, wherein The first keyword sequence corresponding to the keyword identification is obtained according to the first query, the probability distribution vector P q and the working vocabulary, which specifically includes: Based on the preset keyword recognition rule, perform keyword recognition on the first query to obtain one or more corresponding keywords A; and set a corresponding weight parameter a for each keyword A; and set each weight parameter a to 1; Sort the probability distribution vector P in descending order of probability values q of the N W lexical probabilities p q,i to obtain a corresponding first probability sequence; and extract the first N P lexical probabilities p q,i from the first probability sequence to form a corresponding second probability sequence; and use each of the lexical probabilities p q,i in the second probability sequence as a corresponding weight parameter b, and use each of the lexical probabilities p q,i in the second probability sequence as the corresponding first word w i in the working vocabulary as the corresponding pseudo-keyword B; the number of instructions N P is a preset positive integer; Use each keyword A and its corresponding weight parameter a as a group of corresponding first keywords and keyword weights; and use each pseudo-keyword B and its corresponding weight parameter b as a group of corresponding first keywords and keyword weights; and form the corresponding first keyword sequence from all the obtained first keywords.
6. The method for fusing dense and sparse retrieval based on summary words according to claim 2, wherein The dense retrieval of the first library based on the dense vector h q and the first space to obtain a corresponding first retrieval sequence specifically includes: Take the dense vector h q The matching spatial points in the first space are denoted as the current spatial points; and in the first space, based on a preset approximate nearest neighbor search algorithm, retrieve the approximate nearest neighbor spatial points of the current spatial points to obtain a corresponding first spatial point set; the approximate nearest neighbor search algorithm includes the LSH algorithm and the VQ algorithm; the first spatial point set includes multiple first spatial points n j ; Each of the first spatial points n in the first spatial point set is used as a corresponding current comparison point; and the segmented hidden vector corresponding to the current comparison point is used as the corresponding current dense vector; and based on the cosine vector similarity algorithm, the similarity between the dense vector h j and the current dense vector is calculated to obtain the corresponding current similarity; and when the current similarity is greater than a preset first similarity threshold, the first document D corresponding to the current comparison point q is used as a corresponding first retrieved document, and the current similarity is used as a corresponding first document score; the first similarity threshold is a preset threshold parameter greater than 0; j Sort all the obtained first retrieved documents in descending order according to the first document scores to form the corresponding first retrieval sequence.
7. The method for fusing dense and sparse retrieval based on summary words according to claim 2, wherein The sparse retrieval of the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence specifically includes: Count the total number of the first keywords in the first keyword sequence to obtain the corresponding total number N S ; and record each of the first keywords and its corresponding keyword weight as the corresponding keyword x s , weight y s ; 1 ≤ index s ≤ N S ; Take each of the first documents D in the first library j one by one as the corresponding current document; set an initialized-to-0 document cumulative score for the current document; and perform a round of traversal on all of the keywords x s and, during this round of traversal, take the currently traversed keyword x s and its corresponding weight y s as the corresponding current keyword and current weight; and when the current document contains the current keyword, take the sum of adding the current document cumulative score and the current weight as the latest document cumulative score of the current document; From the obtained N D select the maximum and minimum values from the cumulative scores of the documents as the corresponding maximum and minimum scores; and perform normalization calculation on the cumulative scores of each document based on the maximum and minimum scores to obtain the corresponding normalized score = (cumulative document score - minimum score) / (maximum score - minimum score); the value of the normalized score ranges from 0 to 1; Take each of the normalized scores that exceed a preset first score threshold as a corresponding second document score; and take each first document D corresponding to the second document score j as a corresponding second retrieved document; and sort all the obtained second retrieved documents in descending order of the second document score to form a corresponding second retrieval sequence.
8. The method for fusing dense and sparse retrieval based on summary words according to claim 2, wherein The retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback to the current user specifically includes: All the first and second retrieved documents in the first and second retrieval sequences are regarded as documents to be screened; and two score parameters initialized to 0 are set for each document to be screened, denoted as the corresponding first parameter and second parameter; and each document to be screened is identified. If the current document to be screened exists in both the first and second retrieval sequences, its corresponding first and second parameters are reset to the corresponding first and second document scores. If the current document to be screened only exists in the first retrieval sequence or the second retrieval sequence, only its corresponding first parameter or second parameter is reset to the corresponding first document score or second document score. Calculate the average value of the first and second parameters of each document to be screened to obtain the corresponding first average score; and regard each first average score exceeding the preset second score threshold as a corresponding third document score; and regard the document to be screened corresponding to each third document score as a corresponding third retrieved document; and sort all the obtained third retrieved documents in descending order according to the third document scores to form the corresponding third retrieval sequence and feedback to the current user.
9. An apparatus for performing the method for fusing dense and sparse retrieval based on summary words according to any one of claims 1-8, characterized in that, The device includes: an LLM model preparation module, a target library preparation module, a query receiving module, an LLM model processing module, a dense retrieval module, a sparse retrieval module, and a retrieval fusion module; The LLM model preparation module is used to select a large language model that has completed the pre-training task as the corresponding working model; and use the next word task prediction head and instruction template used by the working model in its next word prediction task as the corresponding summary word prediction head and summary word instruction template; and use the specified vocabulary of the working model as the corresponding working vocabulary; the pre-training task includes the next word prediction task; The target library preparation module is used to use the preset retrieval library as the corresponding first library; and use the summary word instruction template, the working model, and the summary word prediction head to create a corresponding document summary word dense vector mapping space for the first library, denoted as the first space; The query receiving module is used to receive the query text input by the user as the first query; The LLM model processing module is used to identify the dense vector h and the probability distribution vector corresponding to the summary word dense vector of the first query by using the summary word instruction template, the working model, and the summary word prediction head. q And the probability distribution vector P q ; And according to the first query, the probability distribution vector P q And the working vocabulary to perform keyword identification to obtain the corresponding first keyword sequence; The dense retrieval module is used to perform dense retrieval on the first library according to the dense vector h q and the first space to obtain a corresponding first retrieval sequence; The sparse retrieval module is used to perform sparse retrieval on the first library according to the first keyword sequence to obtain the corresponding second retrieval sequence; The retrieval fusion module is used to perform retrieval result fusion according to the first and second retrieval sequences to obtain the corresponding third retrieval sequence and feedback to the current user.
10. An electronic device, characterized in that, Including: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory, so as to implement the method according to any one of claims 1-8; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-8.
Citation Information
Patent Citations
Control method and device for realizing article recommendation based on user behaviors
CN112597389A
Method and device for improving quality of information generated by large model
CN116881398A
Dictionary enhanced self-supervised training for multilingual dense retrieval
CN118193671A
Enhancing diagnosis of disorder through artificial intelligence and mobile health technologies without compromising accuracy
US20140304200A1
Cited By
Large model data grading method and device
CN121681711A