Dense retrieval method and device combining large language model and pseudo query technology
By combining large language model and pseudo-query technology, a pseudo-query-document scoring model is constructed, which solves the search ambiguity problem caused by short query text in dense search, and achieves more accurate search results matching.
Patent Information
- Application Number
- CN202510332989.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-03-20
AI Technical Summary
Existing dense search technology cannot accurately obtain personalized semantic features when the query text is short, resulting in the search results that do not match the search intention and increase the probability of retrieval ambiguity problems.
Combining large language model and pseudo-query technology, a pseudo-query-document scoring model is constructed by selecting pre-trained generative large language model and BERT series models, configuring instruction templates to generate pseudo-feedback text, and performing correlation scores in dense retrieval to improve retrieval accuracy.
It improves the search accuracy, reduces the probability of retrieval ambiguity problems, and enhances the personalized matching of search results.
Smart Images

Figure CN120256573A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a dense retrieval method and device combining a large language model and a pseudo-query technique. Background Art
[0002] In the application fields of big data and artificial intelligence technologies, how to accurately obtain the required information is one of the core problems that information retrieval technologies need to solve. Dense Retrievers is a current mainstream retrieval technology. This technology converts the query object (referred to as query for short) and the information document (referred to as document for short) into dense semantic vectors in a high-dimensional feature space respectively, and filters the retrieval results (also known as feedback) based on the vector similarity between the query and the document. Because dense retrieval can obtain deep semantic features from the high-dimensional feature space, it can achieve the purpose of improving the retrieval accuracy. However, we also found a problem in practical applications: if the query text content is short, even if it can be converted into a dense semantic vector in the high-dimensional feature space, the obtained personalized semantic features are not clear enough, which may result in the situation that the retrieval result does not match the retrieval intention, that is, the so-called retrieval ambiguity problem. For example, if the query text input by the user at this time is "cloud", the related knowledge fields at least include meteorology, computer, and communication. Without knowing the current user's concerned field, it is impossible to avoid the retrieval ambiguity problem even using the dense retrieval technology.
[0003] Pseudo Query is a query expansion technology. Its principle is to pre-synthesize a piece of pseudo-feedback text related to the expected feedback corresponding to the real query, and splice the real query and the pseudo-feedback into a new pseudo-query, and then input the pseudo-query into the retriever for retrieval. In principle, if the semantic correlation between the synthesized pseudo-feedback and the expected real feedback is relatively high, then it is bound to add more and more explicit personalized semantic features to the pseudo-query, and naturally it can also improve the retrieval accuracy and reduce the occurrence probability of the retrieval ambiguity problem. However, if there is no substantial association between the synthesized pseudo-feedback and the expected real feedback, it will not only fail to add effective personalized semantic features, but instead bring more feature interference / noise. In this case, it is very likely to increase the occurrence probability of the retrieval ambiguity problem and reduce the retrieval accuracy. That is to say, if we want to introduce pseudo-query in dense retrieval to improve the retrieval accuracy and reduce the occurrence probability of the retrieval ambiguity problem, we must give a solution to solve the correlation problem between the pseudo-feedback and the expected feedback.
[0004] Large Language Model (LLM) is a task processing system implemented based on deep learning technology. The natural language processing (NLP) ability of the LLM model is very powerful: 1) It can learn and remember a series of personalized features such as the user's retrieval habits and the knowledge fields of interest based on the user's historical retrieval information as context; 2) It can extract formatted query-feedback text pairs from the user's historical retrieval information; 3) It can generate a relatively relevant response text for the new question (i.e., query) specified by the user for the current time based on the question-and-answer task description and question-and-answer samples input by the user. That is to say, based on the known user historical retrieval information, an LLM model can generate a response text related to the user's expected feedback, that is, a pseudo-feedback text.
[0005] Based on the above analysis, we draw the following conclusion: If the LLM model, pseudo-query technology, and dense retrieval technology can be combined, the probability of occurrence of the retrieval ambiguity problem can be reduced more reliably. And how to combine the LLM model, pseudo-query technology, and dense retrieval technology has become the technical problem to be solved in this invention. Summary of the Invention
[0006] The objective of the present invention is to provide a dense retrieval method, apparatus, electronic device, and computer-readable storage medium that combine large language models and pseudo-query technology in view of the deficiencies of the prior art. The present invention pre-selects a generative large language model that has completed pre-training and NLP task fine-tuning as the first large model, selects a pre-trained BERT series model as the first and second encoders, constructs a pseudo-query-document scoring model with the first and second encoders as the core, and uses a preset information document library as the target document library; then randomly extracts a specified number N1 of target documents from each knowledge field of the target document library as seed feedback, sets a corresponding query text for each seed feedback, denoted as the seed query, and forms a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; then configures a large language model instruction template denoted as the first instruction template for generating a corresponding feedback with a given sample set as a reference when the given query is not empty and generating a new query-feedback text pair when the given query is empty, configures a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b, and configures a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context; then constructs a model training dataset for the pseudo-query-document scoring model based on N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model, and trains the pseudo-query-document scoring model based on this dataset; and after the model training is completed, receives the first query text and the first historical retrieval information input from the user side, extracts a first text pair sequence based on the first historical retrieval information, the third instruction template, and the first large model, performs pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the first pseudo-feedback text, forms the first pseudo-query text from the first query text and the first pseudo-feedback text, and performs dense retrieval based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the first retrieval result and feedback it to the current user. The present invention combines the LLM model, pseudo-query technology, and dense retrieval technology, which can not only further improve the retrieval accuracy but also reduce the occurrence probability of retrieval ambiguity problems.
[0007] To achieve the above objective, a first aspect of an embodiment of the present invention provides a dense retrieval method that combines large language models and pseudo-query technology, and the method includes:
[0008] Select a pre-trained and NLP-task fine-tuned generative large language model as the first large model; and select a pre-trained BERT series model as the corresponding first encoder and second encoder, and build a pseudo-query-document scoring model with the first and second encoders as the core; and use a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M; the pseudo-query-document scoring model is used to perform a relevance scoring process on the input pseudo-query Q and document D of the model and output the corresponding score S, and the value range of the score S is between -1 and 1; the target document library includes multiple target documents; each target document corresponds to a knowledge domain; the total number M is a positive integer;
[0009] Randomly select a specified number N1 of target documents from each knowledge domain of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; the specified number N1 is a positive integer; the total number of query-feedback seed samples is N1×M;
[0010] Configure a large language model instruction template, denoted as the first instruction template, for generating a corresponding feedback with a given sample set as a reference when the given query is not empty and generating a new query-feedback text pair when the given query is empty; and configure a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with a given retrieval information as the current context; the specified number N2 is a positive integer; the value range of the text feature similarity and the similarity threshold a is between 0 and 1; the value range of the semantic relevance and the relevance threshold b is between -1 and 1, and when the semantic relevance is positive, the smaller the value, the weaker the relevance, and when it is negative, it indicates a semantic contradiction;
[0011] Construct a model training dataset, denoted as the first dataset, for the pseudo-query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model;
[0012] Perform model training on the pseudo-query-document scoring model based on the first dataset;
[0013] After the model training is completed, receive the first query text and the first historical retrieval information input by the user side; and perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence; and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text; and form the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text; and perform dense retrieval processing based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result; and feedback the first retrieval result to the current user; the first historical retrieval information consists of all the retrieval information of the current user within the most recent specified period.
[0014] Preferably, the generative large language model is any one of the GPT series models, BERT series models, T5 series models, LLaMA series models, Wenxin series models, or Qwen series models.
[0015] Preferably, the first model input end of the pseudo-query-document scoring model is used to receive the pseudo-query Q input by the model, the second model input end is used to receive the document D input by the model, and the model output end is used to output the corresponding score S;
[0016] The pseudo-query-document scoring model includes a first preprocessing module, a second preprocessing module, the first encoder, the second encoder, and a scoring module; the input end of the first preprocessing module is connected to the first model input end, and the output end is connected to the input end of the first encoder; the input end of the second preprocessing module is connected to the second model input end, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the scoring module; the output end of the second encoder is connected to the second input end of the scoring module; the output end of the scoring module is connected to the model output end;
[0017] The first preprocessing module is used to perform stop word removal, special character / symbol, and punctuation removal processing on the pseudo-query Q to obtain the corresponding first preprocessing text; and perform word segmentation processing on the first preprocessing text according to the word segmentation rules of the BERT model to obtain the corresponding first word segmentation sequence; and add a preset classification marker CLS to the head of the first word segmentation sequence to obtain the corresponding second word segmentation sequence; and perform embedding encoding according to the second word segmentation sequence according to the embedding encoding rules of the BERT model to obtain the corresponding embedding encoding vector E QSend to the first encoder; the total number of word segments in the first word segment sequence is denoted as the total number of word segments w1, the total number of word segments in the second word segment sequence is denoted as the total number of word segments w2, and w2 = 1 + w1; the embedded encoding vector E Q consists of w2 sub-encoding vectors e Q,i where 1 ≤ vector index i ≤ w2;
[0018] The second preprocessing module is used to perform sentence splitting on the document D to obtain corresponding multiple first sentences A j , and count the total number of the first sentences A j to obtain the corresponding total number of sentences N3, where 1 ≤ sentence index j ≤ N3; and perform stop word, special character / symbol, and punctuation removal processing on each of the first sentences A j to obtain the corresponding second preprocessing text B j ; and perform word segmentation processing on each of the second preprocessing texts B j according to the word segmentation rules of the BERT model to obtain the corresponding third word segment sequence C j ; and add the classification marker CLS to the head of the first third word segment sequence C j=1 ; and add the preset separator marker SEP to the head of the second to N3-th third word segment sequences C j ; and perform sequential splicing on the N3 third word segment sequences C j with the markers added to obtain the corresponding fourth word segment sequence; and perform embedded encoding on the fourth word segment sequence according to the embedded encoding rules of the BERT model to obtain the corresponding embedded encoding vector E D Send to the second encoder; the total number of word segments in the fourth word segment sequence is denoted as the total number of word segments w3; the embedded encoding vector E D consists of w3 sub-encoding vectors e D,k where 1 ≤ vector index k ≤ w3;
[0019] The first encoder is used to perform high-dimensional semantic feature encoding on the input embedded encoding vector E Q to obtain the corresponding encoding tensor H Q Send to the scoring module; the encoding tensor H Q consists of w2 sub-encoding vectors h Q,i ; the first sub-encoding vector h Q,i=1 corresponds to the classification marker CLS of the second word segment sequence;
[0020] The second encoder is used to perform high-dimensional semantic feature encoding on the input embedded encoding vector E D to obtain the corresponding encoding tensor H D Send to the scoring module; the encoding tensor H DComposed of w3 sub-encoded vectors h D,k ; The feature dimension of the sub-encoded vector h D,k is consistent with the feature dimension of the sub-encoded vector h Q,i ; The first sub-encoded vector h D,k=1 corresponds to the classification token CLS of the fourth tokenized sequence;
[0021] The scoring module is used to extract the first sub-encoded vector h Q of the encoded tensor H Q,i=1 as the corresponding first vector, and extract the first sub-encoded vector h D of the encoded tensor H D,k=1 as the corresponding second vector; and calculate the similarity of the first and second vectors based on the cosine similarity algorithm to obtain the corresponding first similarity; and output the first similarity as the corresponding score S.
[0022] Preferably, the configurable parameters of the first instruction template include the given sample set and the given query; the given sample set is composed of multiple first samples; the first sample is composed of a first sample query and a first sample feedback;
[0023] The first instruction template is a formatted text, composed of multiple text lines; the line text format of the first line is: "Task:" + the first task description text, and the first task description text is used to prompt the first large model to generate a corresponding feedback for the given query with the given sample set as a reference; the line text of the second to the penultimate line corresponds one by one to the first samples in the given sample set, and the line text format of each line is: "Sample x:" + "Query:" + the first sample query + the first delimiter + "Feedback:" + the first sample feedback + the second delimiter, where x is the sample number starting from 1 and incrementing by 1 successively, and the + in the line text format is a character sequence concatenation symbol; the first delimiter is default to the semicolon ";", and the second delimiter is default to the period "."; the line text format of the last line is: "Query:" + the given query;
[0024] When the given query of the first instruction template is empty, the first large model will output a corresponding query-feedback text pair in the sample format; when the given query of the first instruction template is not empty, the first large model will only output the corresponding feedback text;
[0025] The configurable parameters of the second instruction template include the given positive feedback;
[0026] The second instruction template is a formatted text consisting of 4 text lines; the line text format of the first line is: "Task: " + the second task description text, and the second task description text is used to prompt the first large model to generate N2 negative feedbacks with similar text features but weakly related or semantically contradictory to the given positive feedback; the line text format of the second line is: "Requirement 1: " + the first requirement text, and the first requirement text is used to require that the text feature similarity between each negative feedback generated by the first large model and the given positive feedback should be higher than the similarity threshold a, and the value range of the text feature similarity is between 0 and 1; the line text format of the third line is: "Requirement 2: " + the second requirement text, and the second requirement text is used to require that the semantic relevance between each negative feedback generated by the first large model and the given positive feedback should be lower than the relevance threshold b, and the value range of the semantic relevance is between -1 and 1, and it is pointed out that the semantic relevance is positive when weakly related and negative when there is semantic contradiction; the line text format of the fourth line is: "Positive feedback: " + the given positive feedback;
[0027] The configurable parameters of the third instruction template include the given retrieval information;
[0028] The third instruction template is a formatted text consisting of 2 paragraphs; the line text format of the first line is: "Task: " + the third task description text, and the third task description text is used to prompt the first large model to extract query-feedback text pairs with the given retrieval information as the context; the line text format of the second line is: "Retrieval information: " + the given retrieval information;
[0029] The first data set includes multiple first data records; the first data record includes a first training query, a first positive feedback, and a first negative feedback sequence; the first negative feedback sequence consists of N2 first negative feedbacks; the first positive feedback is similar in text features to each of the first negative feedbacks but weakly related or semantically contradictory.
[0030] Preferably, a model training data set is constructed for the pseudo query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model, denoted as the first data set, specifically including:
[0031] Step 51, set a first counter initialized to 0;
[0032] Step 52, obtain a corresponding first instruction by setting the given sample set of the first instruction template as the corresponding N1×M query-feedback seed samples and setting the given query as empty; input the first instruction into the first large model to perform query-feedback text pair generation processing to obtain a corresponding first query-feedback text pair; use the query text of the first query-feedback text pair as a corresponding first training query, and use the feedback text of the first query-feedback text pair as a corresponding first positive feedback;
[0033] Wherein, the first query-feedback text pair consists of a query text and a feedback text;
[0034] Step 53, obtain a corresponding second instruction by setting the given positive feedback of the second instruction template as the current first positive feedback; input the second instruction into the first large model to perform negative feedback generation processing to obtain N2 negative feedbacks that are similar in text features to the current first positive feedback but weakly semantically related or semantically contradictory; use each negative feedback obtained this time as a corresponding first negative feedback, and form a corresponding first negative feedback sequence from the N2 first negative feedbacks obtained this time;
[0035] Step 54, form a corresponding first data record from the current first training query and its corresponding first positive feedback and the first negative feedback sequence;
[0036] Step 55, increment the first counter by 1; and identify whether the incremented first counter exceeds a preset record total threshold; if not, return to Step 52; if it has exceeded, form a corresponding first data set from all the first data records obtained;
[0037] Preferably, the model training of the pseudo query-document scoring model based on the first data set specifically includes:
[0038] Step 61, divide the first data set into two sub-data sets according to a preset first segmentation ratio and denote them as the corresponding first training set and first evaluation set;
[0039] Wherein, both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio;
[0040] Step 62, calculate a minimum loss l based on a preset temperature parameter τ and the specified quantity N2 min ;
[0041] Wherein, the minimum loss lmin The calculation formula is as follows:
[0042]
[0043] The minimum loss l min In the calculation formula, + is the plus sign;
[0044] Step 63: Extract the first first data record of the first training set as the corresponding current training record;
[0045] Step 64: Use the first training query and the first positive feedback of the current training record as the corresponding pseudo-query Q and document D to input into the pseudo-query-document scoring model for relevance scoring processing, and use the score S output by the model in this processing as the corresponding positive sample score s p ; And use the first training query and each first negative feedback of the first negative feedback sequence of the current training record as the corresponding pseudo-query Q and document D to input into the pseudo-query-document scoring model for relevance scoring processing, and use the score S output by the model in this processing as a corresponding negative sample score s n,g , 1 ≤ negative sample index g ≤ N2;
[0046] Step 65: Set the maximum positive sample score to 1; and use the positive sample score s obtained this time p and N2 negative sample scores s n,g Substitute into the preset loss function L; and based on the preset first model optimizer, in the direction of making the positive sample score s p reach the maximum positive sample score and the loss function L reach the minimum loss l min Perform one round of fine-tuning on the model parameters of the first and second encoders of the pseudo-query-document scoring model;
[0047] Among them, the loss function L is:
[0048]
[0049] In the loss function L, + is the plus sign;
[0050] The first model optimizer includes at least the Adam optimizer and the SGD optimizer;
[0051] Step 66: Identify whether the current training record is the last first data record of the first training set; if so, go to Step 67; if not, use the next first data record of the first training set as the new current training record and return to Step 64;
[0052] Step 67, take the minimum loss l min as the corresponding first label loss; traverse all the first data records in the first evaluation set once; during this traversal, take the currently traversed first data record as the corresponding current evaluation record; take the first training query and the first positive feedback of the current evaluation record as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model this time as the corresponding positive sample score s p ; take the first training query of the current evaluation record and each first negative feedback in the first negative feedback sequence as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model this time as a corresponding negative sample score s n,g ; take the positive sample score s obtained this time p and N2 negative sample scores s n,g and substitute them into the loss function L for calculation, and take the calculation result as a corresponding first prediction loss; combine the current first prediction loss and the first label loss to form a corresponding first prediction-label pair; at the end of this traversal, take all the obtained first prediction-label pairs and substitute them into a preset model evaluation function for calculation to obtain a corresponding first evaluation value;
[0053] The model evaluation function is implemented based on the MAE function, the MSE function or the RMSE function;
[0054] Step 68, identify whether the first evaluation value meets a preset first evaluation value range; if not, return to Step 63 to continue training; if so, stop training and confirm that the model training is completed.
[0055] Preferably, the extraction process of the query-feedback text pair based on the first historical retrieval information, the third instruction template and the first large model to obtain the corresponding first text pair sequence specifically includes:
[0056] Obtain a corresponding third instruction by setting the given retrieval information of the third instruction template as the current first historical retrieval information; input the third instruction into the first large model to extract query-feedback text pairs to obtain one or more corresponding query-feedback text pairs; form a corresponding first text pair from the query text and the feedback text of each extracted query-feedback text pair; and form the corresponding first text pair sequence from all the obtained first text pairs.
[0057] Preferably, the process of generating pseudo-feedback for the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text specifically includes:
[0058] Obtain a corresponding fourth instruction by setting the given sample set of the first instruction template as the current first text pair sequence and setting the given query as the current first query text; and input the fourth instruction into the first large model for feedback text generation processing, and use the feedback text generated by the model in this processing as the corresponding first pseudo-feedback text.
[0059] Preferably, the process of forming the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text specifically includes:
[0060] Take the first query text as the corresponding first string, take the first pseudo-feedback text as the corresponding second string, and use the third string formed by sequentially concatenating the first string, the preset separator token SEP, and the second string as the corresponding first pseudo-query text.
[0061] Preferably, the process of performing dense retrieval on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result specifically includes:
[0062] Take each target document in the target document library as the corresponding current document, and input the first pseudo-query text and the current document as the corresponding pseudo-query Q and document D into the pseudo-query-document scoring model for relevance scoring processing, and use the score S output by the model in this processing as a corresponding first document score; and take the largest first document score as the corresponding maximum score; and extract one or more target documents in the target document library corresponding to the maximum score to form the corresponding first retrieval result.
[0063] In the second aspect of the embodiments of the present invention, there is provided an apparatus for implementing the dense retrieval method combining a large language model and a pseudo-query technique described in the first aspect above. The apparatus includes: a model preparation module, a seed sample preparation module, an instruction template preparation module, a data set preparation module, a scoring model training module, and a dense retrieval module;
[0064] The model preparation module is used to select a pre-trained and NLP task fine-tuned generative large language model as the first large model; and select a pre-trained BERT series model as the corresponding first encoder and second encoder, and build a pseudo query-document scoring model with the first and second encoders as the core; and use a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M; the pseudo query-document scoring model is used to perform relevance scoring on the pseudo query Q and the document D input to the model and output the corresponding score S, and the value range of the score S is between -1 and 1; the target document library includes multiple target documents; each target document corresponds to a knowledge domain; the total number M is a positive integer;
[0065] The seed sample preparation module is used to randomly select a specified number N1 of target documents from each knowledge domain of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; the specified number N1 is a positive integer; the total number of query-feedback seed samples is N1×M;
[0066] The instruction template preparation module is used to configure a large language model instruction template, denoted as the first instruction template, which is used to generate a corresponding feedback with a given sample set as a reference when the given query is not empty, and generate a new query-feedback text pair when the given query is empty; and configure a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback, and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context; the specified number N2 is a positive integer; the value range of the text feature similarity and the similarity threshold a is between 0 and 1; the value range of the semantic relevance and the relevance threshold b is between -1 and 1, and when the semantic relevance is positive, the smaller the value, the weaker the relevance, and when it is negative, it indicates a semantic contradiction;
[0067] The dataset preparation module is used to construct a model training dataset, denoted as the first dataset, for the pseudo query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model;
[0068] The scoring model training module is used to train the pseudo query-document scoring model based on the first dataset;
[0069] The dense retrieval module is used to receive the first query text and the first historical retrieval information input by the user side after the model training is completed; and extract and process the query-feedback text pairs based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence; and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text; and form the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text; and perform dense retrieval processing based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result; and feedback the first retrieval result to the current user; the first historical retrieval information consists of all the retrieval information of the current user within the most recent specified period.
[0070] In a third aspect of the embodiments of the present invention, an electronic device is provided, including: a memory, a processor, and a transceiver;
[0071] The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method steps described in the first aspect above;
[0072] The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
[0073] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the instructions of the method described in the first aspect above.
[0074] An embodiment of the present invention provides a dense retrieval method, device, electronic device, and computer-readable storage medium that combines a large language model and a pseudo-query technique. As can be seen from the above content, in the embodiment of the present invention, a pre-trained and NLP task fine-tuned generative large language model is pre-selected as the first large model, and a pre-trained BERT series model is selected as the first and second encoders. A pseudo-query-document scoring model is constructed with the first and second encoders as the core, and a preset information document library is used as the target document library. Then, a specified number N1 of target documents are randomly selected from each knowledge field of the target document library as seed feedback, and a corresponding query text, denoted as a seed query, is set for each seed feedback. A corresponding query-feedback seed sample is formed by each seed query and the corresponding seed feedback. Then, a large language model instruction template, denoted as the first instruction template, is configured to generate a corresponding feedback with a given sample set as a reference when the given query is not empty, and generate a new query-feedback text pair when the given query is empty. A second instruction template is configured to generate a specified number N2 of negative feedbacks for a given positive feedback, and require that the text feature similarity between the positive and negative feedbacks is higher than a preset similarity threshold a but the semantic relevance is lower than a preset relevance threshold b. A third instruction template is configured to extract query-feedback text pairs with a given retrieval information as the current context. Then, a model training data set is constructed for the pseudo-query-document scoring model according to N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model, and the pseudo-query-document scoring model is trained based on this data set. After the model training is completed, the first query text and the first historical retrieval information input from the user side are received, and a first text pair sequence is obtained by extracting query-feedback text pairs based on the first historical retrieval information, the third instruction template, and the first large model. A first pseudo-feedback text is obtained by performing pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model. A first pseudo-query text is formed by the first query text and the first pseudo-feedback text, and a first retrieval result is obtained by performing dense retrieval based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library and fed back to the current user. The embodiment of the present invention combines the LLM model, the pseudo-query technique, and the dense retrieval technique, which not only improves the retrieval accuracy but also reduces the occurrence probability of retrieval ambiguity problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 FIG. is a schematic diagram of a dense retrieval method that combines a large language model and a pseudo-query technique provided by Embodiment 1 of the present invention;
[0076] Figure 2 FIG. is a module structure diagram of the pseudo-query-document scoring model provided by Embodiment 1 of the present invention;
[0077] Figure 3 Schematic diagram of the template formats of the first instruction template, the second instruction template, and the third instruction template provided in the first embodiment of the present invention;
[0078] Figure 4 Module structure diagram of a dense retrieval device combining a large language model and a pseudo-query technique provided in the second embodiment of the present invention;
[0079] Figure 5 Schematic diagram of the structure of an electronic device provided in the third embodiment of the present invention. Detailed implementation manners
[0080] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0081] The first embodiment of the present invention provides a dense retrieval method combining a large language model and a pseudo-query technique, as Figure 1 shown in the schematic diagram of a dense retrieval method combining a large language model and a pseudo-query technique provided in the first embodiment of the present invention. The method mainly includes the following steps:
[0082] Step 1, select a generative large language model that has completed pre-training and NLP task fine-tuning as the first large model; and select a BERT series model that has completed pre-training as the corresponding first encoder and second encoder, and construct a pseudo-query-document scoring model with the first and second encoders as the core; and use a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M.
[0083] Here, the generative large language model in the embodiment of the present invention is any one of the GPT series models, BERT series models, T5 series models, LLaMA series models, Wenxin series models, or Qwen series models. The target document library in the embodiment of the present invention includes multiple target documents; each target document corresponds to a knowledge domain; the total number M of knowledge domains in the target document library is a positive integer.
[0084] The pseudo-query-document scoring model in the embodiment of the present invention is used to perform relevance scoring processing on the pseudo-query Q and the document D input to the model and output the corresponding score S. The value range of the score S is between -1 and 1. As Figure 2As shown in the module structure diagram of the pseudo-query-document scoring model provided in the first embodiment of the present invention, the first model input end of the pseudo-query-document scoring model is used to receive the pseudo-query Q input by the model, the second model input end is used to receive the document D input by the model, and the model output end is used to output the corresponding score S. The model components of the pseudo-query-document scoring model include: a first preprocessing module, a second preprocessing module, a first encoder, a second encoder, and a scoring module.
[0085] As Figure 2 shown, the connection relationships of the components of the pseudo-query-document scoring model are as follows: the input end of the first preprocessing module is connected to the first model input end, and the output end is connected to the input end of the first encoder; the input end of the second preprocessing module is connected to the second model input end, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the scoring module; the output end of the second encoder is connected to the second input end of the scoring module; the output end of the scoring module is connected to the model output end.
[0086] The functions of the components of the pseudo-query-document scoring model are as follows.
[0087] 1) First preprocessing module:
[0088] The first preprocessing module of the embodiment of the present invention is used to remove stop words, special characters / symbols, and punctuation marks from the pseudo-query Q to obtain the corresponding first preprocessed text; and perform word segmentation on the first preprocessed text according to the word segmentation rules of the BERT model to obtain the corresponding first word segmentation sequence; and add a preset classification marker CLS to the head of the first word segmentation sequence to obtain the corresponding second word segmentation sequence; and perform embedding encoding on the second word segmentation sequence according to the embedding encoding rules of the BERT model to obtain the corresponding embedding encoding vector E Q and send it to the first encoder.
[0089] Here, the total number of word segments of the first word segmentation sequence is denoted as the total number of word segments w1, the total number of word segments of the second word segmentation sequence is denoted as the total number of word segments w2, and w2 = 1 + w1; the embedding encoding vector E Q consists of w2 sub-encoding vectors e Q,i where 1 ≤ vector index i ≤ w2.
[0090] 2) Second preprocessing module:
[0091] The second preprocessing module of the embodiment of the present invention is used to perform sentence splitting on the document D to obtain the corresponding multiple first sentences A j , and count the total number of the first sentences A j to obtain the corresponding total number of sentences N3, where 1 ≤ sentence index j ≤ N3; and for each first sentence A jPerform stop word, special character / symbol, and punctuation removal processing to obtain the corresponding second preprocessed text B j ; and tokenize each second preprocessed text B j according to the tokenization rules of the BERT model to obtain the corresponding third token sequence C j ; and add the classification marker CLS to the head of the first third token sequence C j=1 ; and add the preset separator marker SEP to the heads of the second to N3rd third token sequences C j ; and perform sequential concatenation on the N3 third token sequences C with markers added to obtain the corresponding fourth token sequence; and perform embedding encoding on the fourth token sequence according to the embedding encoding rules of the BERT model to obtain the corresponding embedding encoding vector E j Send it to the second encoder. D Here, the total number of tokens in the fourth token sequence is denoted as the total number of tokens w3; the embedding encoding vector E
[0092] is composed of w3 sub-encoding vectors e D where 1 ≤ vector index k ≤ w3. D,k
[0093] 3) First encoder:
[0094] The first encoder of the embodiment of the present invention is used to perform high-dimensional semantic feature encoding on the input embedding encoding vector E Q to obtain the corresponding encoding tensor H Q Send it to the scoring module.
[0095] Here, since the first encoder is implemented based on the BERT model, the high-dimensional semantic feature encoding implementation process of the first encoder is similar to the encoding implementation process of the BERT model and will not be further elaborated here. It should be noted that the encoding tensor H Q is composed of w2 sub-encoding vectors h Q,i ; among them, the first sub-encoding vector h Q,i=1 corresponds to the classification marker CLS of the second token sequence.
[0096] 4) Second encoder:
[0097] The second encoder of the embodiment of the present invention is used to perform high-dimensional semantic feature encoding on the input embedding encoding vector E D to obtain the corresponding encoding tensor H D Send it to the scoring module.
[0098] Here, since the second encoder is implemented based on the BERT model, the high-dimensional semantic feature encoding implementation process of the second encoder is similar to the encoding implementation process of the BERT model and will not be further elaborated here. It should be noted that the encoding tensor HD Composed of w3 sub-encoded vectors h D,k ; among them, the first sub-encoded vector h D,k=1 corresponds to the classification token CLS of the fourth token sequence. And, the feature dimension of the sub-encoded vector h D,k is the same as the feature dimension of the sub-encoded vector h Q,i .
[0099] 5) Scoring module:
[0100] The scoring module of the embodiment of the present invention is used to extract the first sub-encoded vector h Q of the encoded tensor H Q,i=1 as the corresponding first vector, and extract the first sub-encoded vector h D of the encoded tensor H D,k=1 as the corresponding second vector; and calculate the similarity of the first and second vectors based on the cosine similarity algorithm to obtain the corresponding first similarity; and output the first similarity as the corresponding score S.
[0101] Here, because the calculation result of the cosine similarity algorithm ranges from -1 to 1, the value of the score S also ranges from -1 to 1.
[0102] Step 2, randomly extract a specified number N1 of target documents from each knowledge field of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback.
[0103] Among them, the specified number N1 is a positive integer; the total number of query-feedback seed samples is N1×M.
[0104] Here, in the current step 2, when setting a corresponding query text for each seed feedback, it can be set manually, or can be obtained by the first large model performing reverse question / query generation processing with each seed feedback as the answer / retrieval result.
[0105] Step 3, configure a large language model instruction template denoted as the first instruction template for generating a corresponding feedback when the given query is not empty and generating a new query-feedback text pair when the given query is empty with the given sample set as a reference; and configure a second instruction template for generating a specified number N2 of negative feedbacks for the given positive feedback and requiring the text feature similarity between the positive and negative feedbacks to be higher than the preset similarity threshold a but the semantic relevance to be lower than the preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context.
[0106] Here, the specified quantity N2 is a positive integer; the value ranges of the text feature similarity and the similarity threshold a are between 0 and 1; the value ranges of the semantic relevance and the relevance threshold b are between -1 and 1. When the semantic relevance is a positive number, the smaller the value, the weaker the correlation. When the semantic relevance is a negative number, it indicates a semantic contradiction.
[0107] The configurable parameters of the first instruction template in the embodiments of the present invention include a given sample set and a given query; among them, the given sample set consists of multiple first samples; the first sample consists of a first sample query and a first sample feedback. The first instruction template is a formatted text, and the specific template format is as Figure 3 shown in the schematic diagram of the template formats of the first instruction template, the second instruction template, and the third instruction template provided in the first embodiment of the present invention, which consists of multiple text lines, where:
[0108] 1) The line text format of the first line is: "Task:" + the first task description text, and the first task description text is used to prompt the first large model to generate a corresponding feedback for the given query with the given sample set as a reference;
[0109] 2) The line text of the second to the second-to-last line corresponds one by one to the first samples in the given sample set, and the line text format of each line is: "Sample x:" + "Query:" + the first sample query + the first separator + "Feedback:" + the first sample feedback + the second separator;
[0110] Here, x is the sample number starting from 1 and incrementing by 1 successively, and the + in the line text format is a string concatenation symbol; the first separator is default to be the semicolon ";", and the second separator is default to be the period ".";
[0111] 3) The line text format of the last line is: "Query:" + the given query.
[0112] It should be noted that when the given query of the first instruction template in the embodiments of the present invention is empty, the first large model will output a corresponding query-feedback text pair in the format of the query-feedback text pair of sample x; when the given query of the first instruction template is not empty, the first large model only outputs the corresponding feedback text.
[0113] The configurable parameters of the second instruction template in the embodiments of the present invention include a given positive feedback. The second instruction template is a formatted text, and the specific template format is as Figure 3 shown, consisting of 4 text lines, where:
[0114] 1) The line text format of the first line is: "Task:" + the second task description text;
[0115] Here, the second task description text is used to prompt the first large model to generate N2 text features similar but semantically weakly relevant or semantically contradictory negative feedbacks for the given positive feedback;
[0116] 2) The line text format of the second line is: "Requirement 1:" + the first requirement text;
[0117] Here, the first requirement text is used to require that the text feature similarity between each negative feedback generated by the first large model and the given positive feedback should be higher than the similarity threshold a, and the value range of the text feature similarity is between 0 and 1;
[0118] 3) The line text format of the third line is: "Requirement 2:" + the second requirement text;
[0119] Here, the second requirement text is used to require that the semantic relevance between each negative feedback generated by the first large model and the given positive feedback should be lower than the relevance threshold b, and the value range of the semantic relevance is between -1 and 1, and it is pointed out that the semantic relevance is positive when there is weak relevance and negative when there is semantic contradiction;
[0120] 4) The line text format of the fourth line is: "Positive feedback:" + the given positive feedback.
[0121] The configurable parameters of the third instruction template in the embodiment of the present invention include the given retrieval information. The third instruction template is a formatted text, and the specific template format is as Figure 3 shown, consisting of 2 paragraphs of text, where:
[0122] 1) The line text format of the first line is: "Task:" + the third task description text;
[0123] Here, the third task description text is used to prompt the first large model to extract query-feedback text pairs with the given retrieval information as the context;
[0124] 2) The line text format of the second line is: "Retrieval information:" + the given retrieval information.
[0125] Step 4, construct a model training dataset denoted as the first dataset for the pseudo query-document scoring model according to N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model;
[0126] Among them, the first dataset includes multiple first data records; the first data record includes the first training query, the first positive feedback, and the first negative feedback sequence; the first negative feedback sequence consists of N2 first negative feedbacks; the first positive feedback is textually similar but semantically weakly relevant or semantically contradictory to each first negative feedback;
[0127] Specifically, it includes: Step 41, set a first counter initialized to 0;
[0128] Step 42, obtain a corresponding first instruction by setting the given sample set of the first instruction template as the corresponding N1×M query-feedback seed samples and setting the given query as empty; input the first instruction into the first large model for query-feedback text pair generation processing to obtain the corresponding first query-feedback text pair; use the query text of the first query-feedback text pair as a corresponding first training query and the feedback text of the first query-feedback text pair as a corresponding first positive feedback;
[0129] Here, the first query-feedback text pair consists of a query text and a feedback text;
[0130] Step 43, obtain a corresponding second instruction by setting the given positive feedback of the second instruction template as the current first positive feedback; input the second instruction into the first large model for negative feedback generation processing to obtain N2 negative feedbacks that are similar in text features but weakly related or semantically contradictory to the current first positive feedback text; use each of the negative feedbacks obtained this time as a corresponding first negative feedback, and form a corresponding first negative feedback sequence from the N2 first negative feedbacks obtained this time;
[0131] Step 44, form a corresponding first data record from the current first training query and its corresponding first positive feedback and first negative feedback sequence;
[0132] Step 45, increment the first counter by 1; and identify whether the incremented first counter exceeds a preset record total threshold; if it does not exceed, return to Step 42; if it has exceeded, form a corresponding first data set from all the first data records obtained;
[0133] Here, the record total threshold is a preset threshold parameter.
[0134] Step 5, train the pseudo query-document scoring model based on the first data set;
[0135] Specifically, it includes: Step 51, divide the first data set into two sub-data sets denoted as the corresponding first training set and first evaluation set based on a preset first splitting ratio;
[0136] Here, the first splitting ratio is a preset ratio parameter, such as 8:2; both the first training set and the first evaluation set consist of multiple first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first splitting ratio;
[0137] Step 52, calculate a minimum loss l based on a preset temperature parameter τ and a specified number N2 min ;
[0138] Here, the temperature parameter τ is a preset positive value parameter;
[0139] The minimum loss l of the embodiment of the present invention min has the following calculation formula:
[0140]
[0141] In the calculation formula of the minimum loss l min the + that appears is a plus sign;
[0142] Step 53: Extract the first first data record of the first training set as the corresponding current training record;
[0143] Step 54: Use the first training query and the first positive feedback of the current training record as the corresponding pseudo-query Q and document D to input into the pseudo-query-document scoring model for relevance scoring processing, and use the score S output by the model in this processing as the corresponding positive sample score s p ; and use the first training query of the current training record and each first negative feedback in the first negative feedback sequence as the corresponding pseudo-query Q and document D to input into the pseudo-query-document scoring model for relevance scoring processing, and use the score S output by the model in this processing as a corresponding negative sample score s n,g ;
[0144] Here, 1 ≤ negative sample index g ≤ N2;
[0145] Step 55: Set the maximum value of the positive sample score to 1; and substitute the positive sample score s p obtained this time and the N2 negative sample scores s n,g into the preset loss function L; and based on the preset first model optimizer, adjust the model parameters of the first and second encoders of the pseudo-query-document scoring model in one round in the direction of making the positive sample score s p reach the maximum value of the positive sample score and the loss function L reach the minimum loss l min ;
[0146] Here, the loss function L of the embodiment of the present invention is:
[0147]
[0148] The + that appears in the loss function L is a plus sign;
[0149] The first model optimizer of the embodiment of the present invention at least includes an Adam optimizer and an SGD optimizer;
[0150] Step 56: Identify whether the current training record is the last first data record of the first training set; if so, go to Step 57; if not, use the next first data record of the first training set as the new current training record and return to Step 54;
[0151] Step 57, take the minimum loss l min as the corresponding first label loss; and traverse all the first data records in the first evaluation set once; and during this traversal, take the currently traversed first data record as the corresponding current evaluation record; and take the first training query and the first positive feedback of the current evaluation record as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model this time as the corresponding positive sample score s p ; and take the first training query of the current evaluation record and each first negative feedback in the first negative feedback sequence as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model this time as a corresponding negative sample score s n,g ; and take the positive sample score s obtained this time p and N2 negative sample scores s n,g and substitute them into the loss function L for calculation, and take the calculation result as a corresponding first prediction loss; and form a corresponding first prediction-label pair with the current first prediction loss and the first label loss; and at the end of this traversal, substitute all the obtained first prediction-label pairs into the preset model evaluation function for calculation to obtain the corresponding first evaluation value;
[0152] Here, the model evaluation function of the embodiment of the present invention is implemented based on the MAE function, the MSE function or the RMSE function;
[0153] Step 58, identify whether the first evaluation value meets the preset first evaluation value range; if not, return to step 53 to continue training; if so, stop training and confirm that the model training is completed.
[0154] Here, the first evaluation value range of the embodiment of the present invention is a preset numerical range.
[0155] Step 6, after the model training is completed, receive the first query text and the first historical retrieval information input by the user side; and perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template and the first large model to obtain the corresponding first text pair sequence; and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template and the first large model to obtain the corresponding first pseudo-feedback text; and form the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text; and perform dense retrieval processing based on the first pseudo-query text, the pseudo-query-document scoring model and the target document library to obtain the corresponding first retrieval result; and feedback the first retrieval result to the current user;
[0156] Specifically, it includes: Step 61, after the model training is completed, receive the first query text and the first historical retrieval information input from the user side;
[0157] Here, the first query text in the embodiment of the present invention is the current query text of the current user; the first historical retrieval information consists of all the retrieval information of the current user within the most recent specified period;
[0158] Step 62, and perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence;
[0159] Specifically, it includes: obtaining a corresponding third instruction by setting the given retrieval information of the third instruction template as the current first historical retrieval information; inputting the third instruction into the first large model to perform query-feedback text pair extraction to obtain the corresponding one or more query-feedback text pairs; forming a corresponding first text pair from the query text and the feedback text of each extracted query-feedback text pair; and forming the corresponding first text pair sequence from all the obtained first text pairs;
[0160] Step 63, and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text;
[0161] Specifically, it includes: obtaining a corresponding fourth instruction by setting the given sample set of the first instruction template as the current first text pair sequence and setting the given query as the current first query text; inputting the fourth instruction into the first large model to perform feedback text generation processing and using the feedback text generated by the model in this processing as the corresponding first pseudo-feedback text;
[0162] Step 64, and form the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text;
[0163] Specifically, it includes: taking the first query text as the corresponding first string, taking the first pseudo-feedback text as the corresponding second string, and taking the third string formed by sequentially splicing the first string, the preset separator tag SEP, and the second string as the corresponding first pseudo-query text;
[0164] Here, the first pseudo-query text = the first query text + 'SEP' + the first pseudo-feedback text, where + is the string concatenation operator;
[0165] Step 65, and perform dense retrieval processing based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result;
[0166] Specifically, it includes: taking each target document in the target document library as the corresponding current document, and taking the first pseudo-query text and the current document as the corresponding pseudo-query Q and document D and inputting them into the pseudo-query-document scoring model for relevance scoring processing, and taking the score S output by the model in this processing as a corresponding first document score; and taking the maximum first document score as the corresponding maximum score; and extracting one or more target documents corresponding to the maximum score in the target document library to form a corresponding first retrieval result;
[0167] Step 66, and feedback the first retrieval result to the current user.
[0168] Figure 4 It is a module structure diagram of a dense retrieval device combining a large language model and pseudo-query technology provided in the second embodiment of the present invention. This device is a terminal device or a server for implementing the foregoing method embodiment, or can also be a device that enables the foregoing terminal device or server to implement the foregoing method embodiment. For example, this device can be a device or a chip system of the foregoing terminal device or server. As Figure 4 shown, this device includes: a model preparation module 201, a seed sample preparation module 202, an instruction template preparation module 203, a data set preparation module 204, a scoring model training module 205, and a dense retrieval module 206.
[0169] The model preparation module 201 is used to select a pre-trained and NLP task fine-tuned generative large language model as the first large model; and select a pre-trained BERT series model as the corresponding first encoder and second encoder, and build a pseudo-query-document scoring model with the first and second encoders as the core; and take a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M; the pseudo-query-document scoring model is used to perform relevance scoring processing on the input pseudo-query Q and document D of the model and output the corresponding score S, and the value range of the score S is between -1 and 1; the target document library includes multiple target documents; each target document corresponds to a knowledge domain; the total number M is a positive integer.
[0170] The seed sample preparation module 202 is used to randomly extract a specified number N1 of target documents from each knowledge domain of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; the specified number N1 is a positive integer; the total number of query-feedback seed samples is N1×M.
[0171] The instruction template preparation module 203 is used to configure a large language model instruction template for generating a corresponding feedback when a given query is not empty with a given sample set as a reference and generating a new query-feedback text pair when the given query is empty, denoted as the first instruction template; and configure a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback, and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context; the specified number N2 is a positive integer; the value ranges of the text feature similarity and the similarity threshold a are between 0 and 1; the value ranges of the semantic relevance and the relevance threshold b are between -1 and 1, and when the semantic relevance is a positive number, the smaller the value, the weaker the correlation, and when it is a negative number, it indicates a semantic contradiction.
[0172] The dataset preparation module 204 is used to construct a model training dataset for the pseudo query-document scoring model, denoted as the first dataset, according to N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model.
[0173] The scoring model training module 205 is used to perform model training on the pseudo query-document scoring model based on the first dataset.
[0174] The dense retrieval module 206 is used to, after the model training is completed, receive the first query text and the first historical retrieval information input by the user side; and perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence; and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text; and form the corresponding first pseudo query text from the first query text and the first pseudo-feedback text; and perform dense retrieval processing based on the first pseudo query text, the pseudo query-document scoring model, and the target document library to obtain the corresponding first retrieval result; and feedback the first retrieval result to the current user; the first historical retrieval information consists of all the retrieval information of the current user within the most recent specified period.
[0175] The dense retrieval device combining a large language model and a pseudo query technology provided by an embodiment of the present invention can execute the method steps in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here.
[0176] It should be noted that it should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the model preparation module can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above determined module can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0177] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Signal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a System-on-a-chip (SOC).
[0178] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the foregoing method embodiments are generated in whole or in part. The above computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the above computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or wireless (such as infrared, wireless, Bluetooth, microwave, etc.). The above computer-readable storage medium may be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The above available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
[0179] Figure 5 FIG. 4 is a schematic structural diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device may be a terminal device or a server for implementing the method of the foregoing embodiments, or may be a terminal device or a server for implementing the method of the foregoing embodiments connected to the foregoing terminal device or server. As Figure 5 shown, the electronic device may include: a processor 301 (such as a CPU), a memory 302, and a transceiver 303; the transceiver 303 is coupled to the processor 301, and the processor 301 controls the transceiver operations of the transceiver 303. Various instructions may be stored in the memory 302 for completing various processing functions and implementing the processing steps described in the foregoing method embodiments. Preferably, the electronic device according to the embodiments of the present invention further includes: a power supply 304, a system bus 305, and a communication port 306. The system bus 305 is used to implement communication connections between components. The above communication port 306 is used for the electronic device to connect and communicate with other peripherals.
[0180] In Figure 5The system bus 305 mentioned above can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The system bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used to implement communication between the database access device and other devices (such as clients, read-write libraries, and read-only libraries). The memory may include Random Access Memory (RAM), and may also include non-volatile memory, such as at least one disk memory.
[0181] The above-mentioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0182] It should be noted that the embodiments of the present invention also provide a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is made to execute the methods and processing procedures provided in the above embodiments.
[0183] An embodiment of the present invention provides a dense retrieval method, apparatus, electronic device, and computer-readable storage medium that combine a large language model and pseudo-query technology. As can be seen from the above, the embodiment of the present invention pre-selects a generative large language model that has completed pre-training and NLP task fine-tuning as the first large model, selects a BERT series model that has completed pre-training as the first and second encoders, constructs a pseudo-query-document scoring model with the first and second encoders as the core, and uses the preset information document library as the target document library; then randomly extracts a specified number N1 of target documents from each knowledge field of the target document library as seed feedback, sets a corresponding query text for each seed feedback, denoted as the seed query, and forms a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; then configures a large language model instruction template for generating a corresponding feedback with a given sample set as a reference when the given query is not empty, and generating a new query-feedback text pair when the given query is empty, denoted as the first instruction template, configures a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback, and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b, configures a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context; then constructs a model training data set for the pseudo-query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model, and trains the pseudo-query-document scoring model based on this data set; and after the model training is completed, receives the first query text and the first historical retrieval information input by the user side, extracts a first text pair sequence based on the first historical retrieval information, the third instruction template, and the first large model, performs pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain a first pseudo-feedback text, forms a first pseudo-query text from the first query text and the first pseudo-feedback text, and performs dense retrieval based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain a first retrieval result and feedback it to the current user. The embodiment of the present invention combines the LLM model, pseudo-query technology, and dense retrieval technology, which not only improves the retrieval accuracy but also reduces the occurrence probability of retrieval ambiguity problems.
[0184] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be implemented in hardware, software modules executed by a processor, or a combination of both. The software modules may be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0185] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A dense retrieval method combining large language models and pseudo-query technology, characterized in that, The method includes: Select a pre-trained and NLP task fine-tuned generative large language model as the first large model; and select a pre-trained BERT series model as the corresponding first encoder and second encoder, and build a pseudo-query-document scoring model with the first and second encoders as the core; and use a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M; the pseudo-query-document scoring model is used to perform relevance scoring processing on the pseudo-query Q and the document D input to the model and output the corresponding score S, and the value range of the score S is between -1 and 1; the target document library includes multiple target documents; each target document corresponds to a knowledge domain; the total number M is a positive integer; Randomly select a specified number N1 of target documents from each knowledge domain of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; the specified number N1 is a positive integer; the total number of query-feedback seed samples is N1×M; Configure a large language model instruction template denoted as the first instruction template for generating a corresponding feedback with a given sample set as a reference when the given query is not empty and generating a new query-feedback text pair when the given query is empty; and configure a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with a given retrieval information as the current context; the specified number N2 is a positive integer; the value range of the text feature similarity and the similarity threshold a is between 0 and 1; the value range of the semantic relevance and the relevance threshold b is between -1 and 1, and the smaller the value of the semantic relevance when it is positive, the weaker the correlation, and when it is negative, it indicates a semantic contradiction; Construct a model training data set denoted as the first data set for the pseudo-query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model; Perform model training on the pseudo-query-document scoring model based on the first data set; After the model training is completed, receive the first query text and the first historical retrieval information input by the user side; and perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence; and perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text; and form the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text; and perform dense retrieval processing based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result; and feedback the first retrieval result to the current user; the first historical retrieval information consists of all the retrieval information of the current user within the most recent specified period.
2. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 1, wherein the generative large language model is any one of the GPT series models, BERT series models, T5 series models, LLaMA series models, Wenxin series models, or Qwen series models.
3. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 1, wherein the first model input end of the pseudo-query-document scoring model is used to receive the pseudo-query Q input by the model, the second model input end is used to receive the document D input by the model, and the model output end is used to output the corresponding score S; the pseudo-query-document scoring model includes a first preprocessing module, a second preprocessing module, the first encoder, the second encoder, and a scoring module; the input end of the first preprocessing module is connected to the first model input end, and the output end is connected to the input end of the first encoder; the input end of the second preprocessing module is connected to the second model input end, and the output end is connected to the input end of the second encoder; the output end of the first encoder is connected to the first input end of the scoring module; the output end of the second encoder is connected to the second input end of the scoring module; the output end of the scoring module is connected to the model output end; the first preprocessing module is used to perform stop word, special character / symbol, and punctuation removal processing on the pseudo-query Q to obtain the corresponding first preprocessed text; Segment the first preprocessed text according to the tokenization rules of the BERT model to obtain the corresponding first token sequence; add a preset classification marker CLS to the head of the first token sequence to obtain the corresponding second token sequence; and perform embedding encoding according to the second token sequence according to the embedding encoding rules of the BERT model to obtain the corresponding embedding encoding vector E Q Send it to the first encoder; the total number of tokens in the first token sequence is denoted as the total number of tokens w1, and the total number of tokens in the second token sequence is denoted as the total number of tokens w2, where w2 = 1 + w1; the embedding encoding vector E Q Consists of w2 sub-encoding vectors e Q,i Where 1 ≤ vector index i ≤ w2; The second preprocessing module is used to perform sentence splitting on the document D to obtain a plurality of corresponding first sentences A j , and count the total number of the first sentences A j to obtain a corresponding total number of sentences N3, where 1 ≤ sentence index j ≤ N3; and perform stop word, special character / symbol, and punctuation removal processing on each of the first sentences A j to obtain a corresponding second preprocessed text B j ; and perform word segmentation processing on each of the second preprocessed texts B j according to the word segmentation rules of the BERT model to obtain a corresponding third word segmentation sequence C j ; and add the classification marker CLS to the head of the first third word segmentation sequence C j=1 ; And add a preset separator marker SEP to the head of the second to N3rd third token sequences C j ; And for the N3 third token sequences C with the completion markers added j perform sequential concatenation to obtain the corresponding fourth token sequence; And perform embedding encoding according to the fourth token sequence according to the embedding encoding rules of the BERT model to obtain the corresponding embedding encoding vector E D and send it to the second encoder; The total number of tokens in the fourth token sequence is denoted as the total number of tokens w3; The embedding encoding vector E D consists of w3 sub-encoding vectors e D,k where 1 ≤ vector index k ≤ w3; The first encoder is used to perform high-dimensional semantic feature encoding on the input embedded encoding vector E Q to obtain a corresponding encoding tensor H Q and send it to the scoring module; the encoding tensor H Q is composed of w2 sub-encoding vectors h Q,i The first sub-encoding vector h Q,i=1 corresponds to the classification token CLS of the second tokenized sequence; The second encoder is used to perform high-dimensional semantic feature encoding on the input embedded encoding vector E D to obtain a corresponding encoding tensor H D and send it to the scoring module; the encoding tensor H D is composed of w3 sub-encoding vectors h D,k ; the feature dimension of the sub-encoding vector h D,k is the same as that of the sub-encoding vector h Q,i ; the first sub-encoding vector h D,k=1 corresponds to the classification token CLS of the fourth tokenized sequence The scoring module is used to take the first sub-encoding vector h Q of the encoding tensor H Q,i=1 and extract it as the corresponding first vector, and take the first sub-encoding vector h D of the encoding tensor H D,k=1 and extract it as the corresponding second vector; and calculate the similarity of the first and second vectors based on the cosine similarity algorithm to obtain the corresponding first similarity; and output the first similarity as the corresponding score S.
4. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 1, wherein the configurable parameters of the first instruction template include the given sample set and the given query; the given sample set consists of multiple first samples; the first sample consists of a first sample query and a first sample feedback. The first instruction template is a formatted text composed of multiple text lines; the line text format of the first line is: "Task: " + the first task description text, and the first task description text is used to prompt the first large model to generate a corresponding feedback for the given query with reference to the given sample set; the line text of the second to the penultimate line corresponds one by one to the first samples in the given sample set, and the line text format of each line is: "Sample x: " + "Query: " + the first sample query + the first delimiter + "Feedback: " + the first sample feedback + the second delimiter, where x is the sample number starting from 1 and incrementing by 1 successively, and the + in the line text format is a character sequence concatenation symbol; the first delimiter is default to be the semicolon ";", and the second delimiter is default to be the period "."; the line text format of the last line is: "Query: " + the given query; When the given query of the first instruction template is empty, the first large model will output a corresponding query-feedback text pair in the sample format; when the given query of the first instruction template is not empty, the first large model only outputs the corresponding feedback text; The configurable parameters of the second instruction template include the given positive feedback; The second instruction template is a formatted text composed of 4 text lines; the line text format of the first line is: "Task: " + the second task description text, and the second task description text is used to prompt the first large model to generate N2 negative feedbacks with similar text features but weakly semantically related or semantically contradictory to the given positive feedback; the line text format of the second line is: "Requirement 1: " + the first requirement text, and the first requirement text is used to require that the text feature similarity between each negative feedback generated by the first large model and the given positive feedback should be higher than the similarity threshold a, and the value range of the text feature similarity is between 0 and 1; the line text format of the third line is: "Requirement 2: " + the second requirement text, and the second requirement text is used to require that the semantic relevance between each negative feedback generated by the first large model and the given positive feedback should be lower than the relevance threshold b, and the value range of the semantic relevance is between -1 and 1, and it is pointed out that the semantic relevance is a positive number when weakly related and a negative number when there is semantic contradiction; the line text format of the fourth line is: "Positive feedback: " + the given positive feedback; The configurable parameters of the third instruction template include the given retrieval information; The third instruction template is a formatted text composed of 2 paragraphs; the line text format of the first line is: "Task: " + the third task description text, and the third task description text is used to prompt the first large model to extract query-feedback text pairs with the given retrieval information as the context; the line text format of the second line is: "Retrieval information: " + the given retrieval information; The first data set includes a plurality of first data records; the first data records include first training queries, first positive feedback, and first negative feedback sequences; the first negative feedback sequences are composed of N2 first negative feedbacks; the first positive feedback is textually similar to each of the first negative feedbacks but weakly semantically related or semantically contradictory.
5. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 4, wherein Construct a model training data set for the pseudo query-document scoring model according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model, denoted as the first data set, specifically including: Step 51, set a first counter initialized to 0; Step 52, obtain a corresponding first instruction by setting the given sample set of the first instruction template as the corresponding N1×M query-feedback seed samples and setting the given query to empty; input the first instruction into the first large model for query-feedback text pair generation processing to obtain a corresponding first query-feedback text pair; use the query text of the first query-feedback text pair as a corresponding first training query, and use the feedback text of the first query-feedback text pair as a corresponding first positive feedback; Wherein, the first query-feedback text pair consists of a query text and a feedback text; Step 53, obtain a corresponding second instruction by setting the given positive feedback of the second instruction template as the current first positive feedback; input the second instruction into the first large model for negative feedback generation processing to obtain N2 negative feedbacks that are textually similar to the current first positive feedback but weakly semantically related or semantically contradictory; use each of the negative feedbacks obtained this time as a corresponding first negative feedback, and form a corresponding first negative feedback sequence from the N2 first negative feedbacks obtained this time; Step 54, form a corresponding first data record from the current first training query and its corresponding first positive feedback and first negative feedback sequence; Step 55, increment the first counter by 1; and identify whether the incremented first counter exceeds a preset record total threshold; if not, return to Step 52; if so, form the corresponding first data set from all the first data records obtained.
6. The dense retrieval method combining large language models and pseudo-query technology according to claim 4, wherein Perform model training on the pseudo query-document scoring model based on the first data set, specifically including: Step 61, divide the first data set into two sub-data sets according to a preset first segmentation ratio, denoted as the corresponding first training set and first evaluation set; Wherein, both the first training set and the first evaluation set are composed of a plurality of first data records; the ratio of the total number of records in the first training set to the total number of records in the first evaluation set satisfies the first segmentation ratio; Step 62, calculate a minimum loss l based on a preset temperature parameter τ and the specified quantity N2 min ; Among them, the minimum loss l min is calculated by the following formula: The minimum loss l min in the calculation formula is the plus sign; Step 63, extract the first first data record in the first training set as the corresponding current training record; Step 64: Input the first training query and the first positive feedback of the current training record as the corresponding pseudo-query Q and document D into the pseudo-query-document scoring model for relevance scoring, and use the score S output by the model in this processing as the corresponding positive sample score s p ; and input the first training query and each of the first negative feedbacks in the first negative feedback sequence of the current training record as the corresponding pseudo-query Q and document D into the pseudo-query-document scoring model for relevance scoring, and use the score S output by the model in this processing as a corresponding negative sample score s n,g , 1 ≤ negative sample index g ≤ N2; Step 65, set the maximum positive sample score to 1; and use the positive sample score s obtained this time p and the N2 negative sample scores s n,g to be input into a preset loss function L; and based on a preset first model optimizer, in the direction of making the positive sample score s p reach the maximum positive sample score and the loss function L reach the minimum loss l min perform one round of fine-tuning on the model parameters of the first and second encoders of the pseudo query-document scoring model; Wherein, the loss function L is: In the loss function L, + is the plus sign; The first model optimizer at least includes an Adam optimizer and an SGD optimizer; Step 66: Identify whether the current training record is the last first data record in the first training set; if so, go to step 67; if not, take the next first data record in the first training set as the new current training record and return to step 64; Step 67, take the minimum loss l min as the corresponding first label loss; traverse all the first data records of the first evaluation set once; during this traversal, take the currently traversed first data record as the corresponding current evaluation record; take the first training query and the first positive feedback of the current evaluation record as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model in this processing as the corresponding positive sample score s p ; take the first training query of the current evaluation record and each first negative feedback of the first negative feedback sequence as the corresponding pseudo-query Q and document D and input them into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model in this processing as a corresponding negative sample score s n,g ; take the positive sample score s obtained this time p and N2 negative sample scores s n,g and substitute them into the loss function L for calculation, and take the calculation result as a corresponding first prediction loss; form a corresponding first prediction-label pair with the current first prediction loss and the first label loss; at the end of this traversal, substitute all the obtained first prediction-label pairs into a preset model evaluation function for calculation to obtain the corresponding first evaluation value; The model evaluation function is implemented based on an MAE function, an MSE function, or an RMSE function; Step 68: Identify whether the first evaluation value meets a preset first evaluation value range; if not, return to step 63 to continue training; if so, stop training and confirm that the model training is completed.
7. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 4, wherein The extraction process of the query-feedback text pair based on the first historical retrieval information, the third instruction template, and the first large model to obtain the corresponding first text pair sequence specifically includes: Obtaining a corresponding third instruction by setting the given retrieval information of the third instruction template as the current first historical retrieval information; inputting the third instruction into the first large model to extract one or more query-feedback text pairs; forming a corresponding first text pair from the query text and the feedback text of each extracted query-feedback text pair; and forming the corresponding first text pair sequence from all the obtained first text pairs.
8. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 4, wherein The process of generating a pseudo-feedback based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain the corresponding first pseudo-feedback text specifically includes: Obtaining a corresponding fourth instruction by setting the given sample set of the first instruction template as the current first text pair sequence and setting the given query as the current first query text; inputting the fourth instruction into the first large model to generate a feedback text and taking the feedback text generated by the model this time as the corresponding first pseudo-feedback text.
9. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 1, wherein The composition of the corresponding first pseudo-query text from the first query text and the first pseudo-feedback text specifically includes: Taking the first query text as the corresponding first string, taking the first pseudo-feedback text as the corresponding second string, and taking the third string formed by sequentially concatenating the first string, the preset separator SEP, and the second string as the corresponding first pseudo-query text.
10. The dense retrieval method combining a large language model and a pseudo-query technique according to claim 1, wherein The process of performing dense retrieval based on the first pseudo-query text, the pseudo-query-document scoring model, and the target document library to obtain the corresponding first retrieval result specifically includes: Take each of the target documents in the target document library as the corresponding current document, and use the first pseudo-query text and the current document as the corresponding pseudo-query Q and document D to input into the pseudo-query-document scoring model for relevance scoring processing, and take the score S output by the model in this processing as a corresponding first document score; and take the maximum of the first document scores as the corresponding maximum score; and extract one or more target documents in the target document library corresponding to the maximum score to form the corresponding first retrieval result.
11. An apparatus for performing the dense retrieval method of combining a large language model and a pseudo-query technique according to any one of claims 1-10, characterized in that, The device includes: a model preparation module, a seed sample preparation module, an instruction template preparation module, a dataset preparation module, a scoring model training module, and a dense retrieval module; The model preparation module is used to select a pre-trained and NLP-task fine-tuned generative large language model as the first large model; and select a pre-trained BERT series model as the corresponding first encoder and second encoder, and build a pseudo-query-document scoring model with the first and second encoders as the core; and use a preset information document library as the corresponding target document library, and count the total number of knowledge domains in the target document library to obtain the corresponding total number M; the pseudo-query-document scoring model is used to perform relevance scoring processing on the input pseudo-query Q and document D of the model and output the corresponding score S, and the value range of the score S is between -1 and 1; the target document library includes multiple target documents; each target document corresponds to a knowledge domain; the total number M is a positive integer; The seed sample preparation module is used to randomly select a specified number N1 of target documents from each knowledge domain of the target document library as the corresponding N1 seed feedbacks; and set a corresponding query text for each seed feedback, denoted as the seed query; and form a corresponding query-feedback seed sample from each seed query and the corresponding seed feedback; the specified number N1 is a positive integer; the total number of the query-feedback seed samples is N1×M; The instruction template preparation module is used to configure a large language model instruction template, denoted as the first instruction template, for generating a corresponding feedback with a given sample set as a reference when the given query is not empty, and generating a new query-feedback text pair when the given query is empty; and configure a second instruction template for generating a specified number N2 of negative feedbacks for a given positive feedback, and requiring the text feature similarity between the positive and negative feedbacks to be higher than a preset similarity threshold a but the semantic relevance to be lower than a preset relevance threshold b; and configure a third instruction template for extracting query-feedback text pairs with the given retrieval information as the current context; the specified number N2 is a positive integer; the value range of the text feature similarity and the similarity threshold a is between 0 and 1; the value range of the semantic relevance and the relevance threshold b is between -1 and 1, and when the semantic relevance is positive, the smaller the value, the weaker the relevance, and when it is negative, it indicates a semantic contradiction; The dataset preparation module is used to construct a model training dataset for the pseudo query-document scoring model, denoted as the first dataset, according to the N1×M query-feedback seed samples, the first instruction template, the second instruction template, and the first large model; The scoring model training module is used to train the pseudo query-document scoring model based on the first dataset; The dense retrieval module is configured to, after the model training is completed, receive a first query text and first historical retrieval information input by the user side; perform query-feedback text pair extraction processing based on the first historical retrieval information, the third instruction template, and the first large model to obtain a corresponding first text pair sequence; perform pseudo-feedback generation processing based on the first text pair sequence, the first query text, the first instruction template, and the first large model to obtain a corresponding first pseudo-feedback text; form a corresponding first pseudo query text from the first query text and the first pseudo-feedback text; perform dense retrieval processing based on the first pseudo query text, the pseudo query-document scoring model, and the target document library to obtain a corresponding first retrieval result; and feedback the first retrieval result to the current user; the first historical retrieval information consists of all retrieval information of the current user within a recently specified period.
12. An electronic device, characterized in that, Comprising: A memory, a processor, and a transceiver; The processor is used to be coupled with the memory, read and execute the instructions in the memory to implement the method according to any one of claims 1-10; The transceiver is coupled with the processor, and the processor controls the transceiver to perform message sending and receiving.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1-10.
Citation Information
Patent Citations
Document retrieval method based on dense pseudo query vector representation
CN112732864A
Lexical item level query expansion method based on collaborative enhancement
CN118467708A
Dense retrieval method based on query expansion and delayed interaction
CN118550988A
Electronic archive retrieval method and system based on large language model
CN118643148A
Construction method of expert system based on RAG technology
CN119227792A