Dialogue-based Generation Retrieval Method and System with Large Model Document Index Perception

By adopting the generation and search method of large-model document index perception in dialogue information retrieval, using the cross attention layer and two-stage training strategy, the information retrieval problem under complex context and noise is solved, and more efficient and accurate information retrieval effect is achieved.

CN119903161BActive Publication Date: 2025-06-17SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510386510.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-17
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Existing dialogue-based information retrieval methods have limitations in dealing with complex contexts and noises, especially when user queries contain pronoun references and intentions are masked by unrelated contexts, making it difficult to achieve accurate and efficient information retrieval.

Method used

A dialogue-based search method based on large-model document index perception is adopted. Through the trained large language model, a cross-attention layer and a two-stage training strategy is combined with a document identifier and bundle search and decode it, and a sorting list of information that is most relevant to user queries is output.

Benefits of technology

It significantly improves the performance of the dialogue search system in complex dialogue scenarios, enhances resistance to noise and irrelevant contexts, and can more accurately understand user intentions and retrieve highly relevant information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119903161B_ABST
    Figure CN119903161B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of information retrieval, and provides a conversational generation retrieval method and system based on large model document index perception. The conversational generation retrieval method based on large model document index perception includes: according to the user query input, using a trained large language model to output an information ranking list related to the user query input; based on the corpus, using a cross-attention layer to extract key information of the context to obtain a proposition and obtain a document identifier; in the first-stage training, introducing a generation loss to generate information related to the current query and generate a document identifier; in the second-stage training, introducing a comprehensive loss to optimize the ranking list of the retrieved document identifiers; through a beam search decoding strategy, outputting a paragraph ranking list corresponding to the document identifier. The present invention realizes more effective context understanding and denoising through innovative document identifier design and training strategies, improving the accuracy of conversational retrieval and the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval, and in particular, to a conversational generation retrieval method and system based on large model document index perception. Background Art

[0002] The statements in this part only provide background technical information related to the present invention and do not necessarily constitute prior art.

[0003] With the continuous development of artificial intelligence and natural language processing technologies, conversational information retrieval systems have gradually become an important part of the modern information retrieval field. Conversational information retrieval aims to more accurately understand the user's needs through natural language interaction with the user and provide relevant and accurate information. Such a system can not only provide a more natural and intuitive search experience, but also dynamically adjust the retrieval results according to the user's context information and intentions, thereby improving user satisfaction and retrieval efficiency. In traditional information retrieval systems, users usually obtain information by entering an independent query. However, in conversational information retrieval, the user's queries are often context-related, which means that the model needs to extract key information from the conversation history to understand the user's current information needs. This poses unique challenges to traditional ad-hoc retrieval methods, prompting the continuous development and innovation of conversational retrieval technologies to better capture the user's search intent and improve the accuracy of retrieval results. Two main challenges faced by conversational retrieval are: First, the user's query may contain pronoun references. For example, in a conversation, the user may use "it" or "the latter place" to refer to "Dunkirk" and "Wouerck" mentioned previously. Second, the intent of the current query may be masked by irrelevant context. For example, in the last question about "Wouerck", the retrieval model may be misled by the historical context related to "Dunkirk" and wrongly rank documents about "Dunkirk" higher than those about "Wouerck".

[0004] To address the above challenges, two main methods have been developed: Contextual Query Rewrite (CQR) and Crash Data Retrieval (CDR). The method of contextual query rewrite aims to train a rewrite model to transform the current query into an independent query based on the dialogue context. However, due to the independence between the rewrite process and the retrieval process, these methods face difficulties in end-to-end training and are restricted by the query rewrite training data. The method of conversational dense retrieval trains a dual-encoder model to embed all contexts and documents into a high-dimensional space to calculate the relevance score. Although some studies have designed complex training strategies to improve the context representation, due to the limitations of dense retrieval, the method of conversational dense retrieval cannot achieve the optimal denoising effect. Specifically, dense retrieval uses fixed-length encoding to represent the context, resulting in the embedding bottleneck problem. In addition, the supervision signal can only propagate through these fixed-length encodings, restricting the learning ability of the model. These inherent limitations are further exacerbated in conversational retrieval scenarios with complex and noisy contexts.

[0005] To overcome these limitations, generative retrieval is an emerging retrieval paradigm that uses a sequence-to-sequence framework to generate document identifiers (docids) for retrieval. Compared with dense retrieval, the main advantages of generative retrieval lie in two aspects. First, generative retrieval enables true end-to-end training, allowing the supervision signal to directly propagate to the model parameters. This eliminates the embedding bottleneck problem and enables the model to more effectively learn the context denoising ability. Second, generative retrieval GR adopts an encoder-decoder architecture. In this framework, the cross-attention layer promotes the interaction between query tokens and document tokens, helping the decoder better understand the input query, which is beneficial for context denoising. Although existing studies have demonstrated the advantages of generative retrieval in understanding complex user queries, previous methods only rely on document titles and do not consider the specific content of the documents, resulting in poor performance on widely used conversational retrieval datasets. In summary, existing conversational retrieval methods have limitations in dealing with complex contexts and noise. Summary of the Invention

[0006] To solve the technical problems existing in the above background art, the present invention provides a conversational generative retrieval method and system based on large model document index awareness. The present invention aims to improve the accuracy and robustness of conversational search in the presence of context noise based on large model document index awareness through a generative retrieval framework, so as to better meet the information needs of users in the dynamic interaction process.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The first aspect of the present invention provides a conversational generation retrieval method based on large model document index awareness.

[0009] A conversational generation retrieval method based on large model document index awareness, comprising:

[0010] According to the user query input, using a trained large language model, output a sorted list of information related to the user query input;

[0011] Among them, the process of training the large language model includes: based on a corpus, using a cross-attention layer, extracting key information of the context, obtaining a proposition, obtaining a document identifier; in the first stage of training, introducing a generation loss to generate information related to the current query and generating a document identifier; in the second stage of training, introducing a comprehensive loss to optimize the ranking list of retrieved document identifiers; through a beam search decoding strategy, output a paragraph ranking list corresponding to the document identifier.

[0012] Further, the document identifier is a combination of the document title followed by the proposition.

[0013] Further, during the decoding process of the large language model, the document title is used as a prefix to guide the encoding process.

[0014] Further, the generation loss uses the maximum likelihood estimation loss.

[0015] Further, the comprehensive loss is represented by the following formula:

[0016]

[0017]

[0018]

[0019] Among them, represents the comprehensive loss, represents the generation loss, represents the ranking loss, represents any document identifier, represents the th hard negative sample document identifier of the current training data instance, is the margin hyperparameter, is the number of hard negative sample document identifiers, is the weight hyperparameter for balancing the two losses.

[0020] Further, through the beam search decoding strategy, a paragraph ranking list corresponding to the document identifier is output; the method includes: constraining the decoding space through the FM-index to ensure the validity of the generated document identifier, ranking the paragraphs according to the generation probability, and providing the user with information most relevant to the user's query input.

[0021] The second aspect of the present invention provides a conversational generation retrieval system based on large model document index awareness.

[0022] A conversational generation retrieval system based on large model document index awareness includes:

[0023] A query module, which is configured to: according to the user's query input, adopt a trained large language model and output a sorted list of information related to the user's query input;

[0024] Among them, the process of training the large language model includes: based on the corpus, using the cross-attention layer to extract key information of the context, obtaining propositions, and obtaining document identifiers; in the first stage of training, introducing a generation loss to generate information related to the current query and generating document identifiers; in the second stage of training, introducing a comprehensive loss to optimize the ranking list of the retrieved document identifiers; through the beam search decoding strategy, outputting a paragraph ranking list corresponding to the document identifier.

[0025] The third aspect of the present invention provides a computer-readable storage medium.

[0026] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the conversational generation retrieval method based on large model document index awareness described in the first aspect above.

[0027] The fourth aspect of the present invention provides a computer device.

[0028] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the conversational generation retrieval method based on large model document index awareness described in the first aspect above.

[0029] The fifth aspect of the present invention provides a computer program product or a computer program.

[0030] The present invention provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the conversational generation retrieval method based on large model document index awareness as described in the first aspect above.

[0031] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0032] The present invention proposes a conversational generation retrieval method and system based on large model document index awareness. According to a user query input, a trained large language model is used to output a sorted list of information related to the user query input. Among them, the process of training the large language model includes: based on a corpus, using a cross-attention layer to extract key information of the context, obtaining propositions, and obtaining document identifiers; in the first stage of training, a generation loss is introduced to generate information related to the current query and generate document identifiers; in the second stage of training, a comprehensive loss is introduced to optimize the ranking list of retrieved document identifiers; through a beam search decoding strategy, a ranking list of paragraphs corresponding to the document identifiers is output. By generating document identifiers, the present invention not only reduces the training burden, but also can capture fine-grained information in paragraphs, providing a new perspective for the application of generative retrieval in conversational search, and helping to more accurately locate and retrieve relevant information; through a two-stage training method, the capabilities of the large language model in context denoising and paragraph ranking have been comprehensively improved, and the most relevant information is provided to users under the constraint search decoding strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention.

[0034] Figure 1 It is a working flowchart of the conversational generation retrieval framework CGR4CD shown in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The present invention will be further described below in conjunction with the drawings and embodiments.

[0036] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0037] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0038] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and systems according to various embodiments of the present disclosure. It should be noted that each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the flowchart and / or block diagram, and the combinations of blocks in the flowchart and / or block diagram, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0039] Embodiment 1

[0040] This embodiment provides a conversational generation and retrieval method based on large model document index awareness. In this embodiment, taking the application of this method to a server as an example, it can be understood that this method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, web servers, cloud communications, middleware services, domain name services, security services CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. In this embodiment, the method includes the following steps:

[0041] According to the user query input, use the trained large language model to output a sorted list of information related to the user query input;

[0042] Among them, the process of large language model training includes: based on a corpus, using a cross-attention layer to extract key information from the context, obtaining propositions, and obtaining document identifiers; in the first-stage training, introducing a generation loss to generate information related to the current query and generating document identifiers; in the second-stage training, introducing a comprehensive loss to optimize the ranking list of retrieved document identifiers; through a beam search decoding strategy, outputting the paragraph ranking list corresponding to the document identifiers.

[0043] The present invention designs the document identifier as a combination of a title and a proposition. Then, prompts the large language model (LLM) to summarize paragraphs into multiple propositions to obtain document identifiers. During training, a two-stage training strategy is adopted to gradually train the model. For inference, beam search is used to generate a ranking list, and the FM-index is applied to constrain the decoding space to ensure that the generated document identifiers are valid, as Figure 1 shown:

[0044] The following is a detailed description of this embodiment:

[0045] In some embodiments, the method of extracting key information from the context, obtaining propositions, and obtaining document identifiers based on a corpus using a cross-attention layer includes:

[0046] The existing designs of document identifiers in generative retrieval mainly face two problems: difficulty in scaling to large-scale corpora and inability to capture fine-grained information in paragraphs. To solve these problems, the present invention proposes using a title plus a proposition as the document identifier. A proposition is defined as an independent atomic expression that conveys the key information in a paragraph. The inventors found that directly using propositions as document identifiers in generative retrieval would lead to unstable performance. This is because beam search can only explore a limited number of next tokens at each step, and in the initial decoding step, the true paragraph may be pruned incorrectly. To solve this problem, a title is used as a prefix to guide the decoding process, achieving retrieval from coarse-grained (title) to fine-grained (proposition). This greatly reduces the risk of incorrectly pruning the true paragraph.

[0047] To obtain propositions from paragraphs, a prompt is designed to guide the large language model to summarize paragraphs into propositions, ensuring that each proposition is context-independent. For the training data, the large language model is further instructed to select several propositions that can be used to answer the current question from the true paragraph as the training target.

[0048] Among them, in practical applications, the prompt in the above content can be: "Your task is to summarize the content of the document into multiple clear and concise propositions, ensuring that these propositions are interpretable without context."

[0049] In some embodiments, during the first-stage training, a generation loss is introduced to generate information related to the current query and generate document identifiers. The method includes:

[0050] The present invention adopts a two-stage training strategy to train the model in a progressive manner. In the first stage, like other generative retrieval methods, a standard maximum likelihood estimation (MLE) loss is used to train an autoregressive generative model. The input is the dialogue context, and the target is the document identifier. The goal of this stage is to train the model to identify information related to the current query and generate the corresponding document identifier. For each training instance, the maximum likelihood estimation loss can be expressed as:

[0051]

[0052] where, represents the maximum likelihood estimation loss (generation loss), is the number of true document identifiers, is the th length of the document identifier, represents the th document identifier, then represents the th th token in the th document identifier, represents the generation probability of the generative retrieval model parameters represents the dialogue context.

[0053] In some embodiments, during the second-stage training, a comprehensive loss is introduced to optimize the ranking list of retrieved document identifiers. The method includes:

[0054] If the model trained only with the maximum likelihood estimation loss tends to generate short and low-information propositions, such as "Dunkirk is a movie". To address this bias, a ranking loss is introduced in the second-stage training. Specifically, for each query of the training data instance, there is a corresponding positive example document identifier , and to construct the training data for the second stage, a hard negative sample needs to be found for each training data instance. The method is to sample document identifiers with the same title from all the document identifier sets in the corpus as this hard negative sample. In each training step, the generation probability of the document identifier is calculated, and the model is trained to assign a higher probability to the positive example document identifier. In addition, to make the training process more stable, the generation loss of the first stage is reintroduced. The loss in the second stage can be expressed as:

[0055]

[0056]

[0057]

[0058] Among them, represents any document identifier, represents the -th hard negative sample document identifier of the current training data instance, represents the ranking loss, is the margin hyperparameter, is the number of hard negative sample document identifiers, represents the comprehensive loss, is the weight hyperparameter for balancing the two losses.

[0059] In some embodiments, the beam search decoding strategy is used to output a list of ranked paragraphs corresponding to the document identifiers; the method includes:

[0060] In the inference stage, a constrained beam search decoding algorithm is adopted to generate the ranked list. To ensure that the generated document identifiers are valid, a data structure is needed to store all propositions and constrain the decoding space. This is achieved using the FM-index. The FM-index supports returning a list of candidate tokens that satisfy the input prefix from any position. Therefore, the title and propositions of the paragraph are concatenated into the format of "document title@proposition 1|document title@proposition 2...", and the decoding is forced to start from the title. Finally, the paragraphs are ranked according to the probabilities of the generated document identifiers. The score of the -th paragraph can be calculated as:

[0061]

[0062] Among them, represents the score of the -th paragraph, represents the set of all document identifiers of the -th paragraph.

[0063] The present invention proposes a conversational generative retrieval method based on large model document index perception, aiming to overcome the limitations of the prior art. Through innovative model architectures and training strategies, the performance of the dialogue search system in complex dialogue scenarios is significantly improved. First, a novel conversational generative retrieval framework (CGR4CD) is proposed: this framework utilizes a sequence-to-sequence generative retrieval model, and effectively captures key information in the context through the cross-attention layer during the decoding process, significantly enhancing the model's resistance to noise and irrelevant context, thereby more accurately understanding the user's intent and retrieving highly relevant information. At the same time, leveraging the powerful text understanding and perception ability of the large model, a fine-grained document index is constructed; a proposition-based document identifier (docids) is designed: compared with traditional document identifiers, the document identifier proposed in the present invention not only reduces the training burden, but also can capture fine-grained information in paragraphs, providing a new perspective for the application of generative retrieval in dialogue search and helping to more accurately locate and retrieve relevant information; a two-stage training strategy is adopted: this strategy first trains the model through a generative loss, enabling it to identify information relevant to the current query from the context and generate corresponding document identifiers; then a ranking loss is introduced to further enhance the model's ability to distinguish similar paragraphs and optimize the ranking of retrieval results. This gradually enhanced training method comprehensively improves the model's capabilities in context denoising and paragraph ranking; a constrained search decoding strategy is adopted in the inference stage: the decoding space is constrained by the FM-index to ensure the validity of the generated document identifiers, and the paragraphs are ranked according to the generation probability to provide the most relevant information to the user.

[0064] Embodiment 2

[0065] This embodiment provides a conversational generative retrieval system based on large model document index perception.

[0066] A conversational generative retrieval system based on large model document index perception, comprising:

[0067] A query module, which is configured to: according to the user query input, adopt a trained large language model and output a sorted list of information related to the user query input;

[0068] Wherein, the process of training the large language model includes: based on a corpus, adopting a cross-attention layer to extract key information of the context, obtaining propositions, and obtaining document identifiers; in the first stage of training, a generative loss is introduced to generate information related to the current query and generate document identifiers; in the second stage of training, a comprehensive loss is introduced to optimize the ranking list of the retrieved document identifiers; through a beam search decoding strategy, a paragraph ranking list corresponding to the document identifiers is output.

[0069] In some embodiments, the document identifier is a combination of the document title followed by a proposition.

[0070] In some embodiments, during the decoding process of the large language model, the document title is used as a prefix to guide the encoding process.

[0071] In some embodiments, the generation loss uses the maximum likelihood estimation loss.

[0072] In some embodiments, the comprehensive loss is represented by the following formula:

[0073]

[0074]

[0075]

[0076] where, represents the comprehensive loss, represents the generation loss, represents the ranking loss, represents any document identifier, represents the th hard negative sample document identifier of the current training data instance, is the margin hyperparameter, is the number of hard negative sample document identifiers, is the weight hyperparameter for balancing the two losses.

[0077] In some embodiments, through the beam search decoding strategy, a list of paragraph rankings corresponding to the document identifier is output; the method includes: constraining the decoding space through the FM-index to ensure the validity of the generated document identifier, and ranking the paragraphs according to the generation probability to provide the user with the information most relevant to the user's query input.

[0078] Through the innovative document identifier design and training strategy of the present invention, more effective context understanding and denoising are achieved, thereby improving the accuracy of conversational retrieval and the user experience.

[0079] Embodiment III

[0080] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the conversational generation retrieval method based on large model document index perception described in Embodiment I above are implemented.

[0081] Embodiment IV

[0082] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the conversational generation and retrieval method based on large model document index perception described in Embodiment 1 above.

[0083] Embodiment Five

[0084] This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the conversational generation and retrieval method based on large model document index perception described in Embodiment 1 above.

[0085] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) that contain computer-usable program code.

[0086] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, as well as the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0087] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0088] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks. Figure 1 One process or a plurality of processes and / or blocks Figure 1 Steps for realizing the functions specified in one block or a plurality of blocks.

[0089] Those of ordinary skill in the art can understand that all or part of the processes of the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), or the like.

[0090] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A conversational generative retrieval method based on large-model document index perception, characterized in that: include: Based on the user query input, the trained large language model is used to output a ranked list of information related to the user query input; The training process of the large language model includes: based on the corpus, using the cross attention layer to extract the key information of the context, obtain the proposition, and obtain the document identifier; in the first stage of training, the generation loss is introduced to generate information related to the current query and generate the document identifier; in the second stage of training, the comprehensive loss is introduced to optimize the ranking list of the retrieved document identifiers; through the beam search decoding strategy, the paragraph ranking list corresponding to the document identifier is output; The comprehensive loss is expressed by the following formula: in, represents the comprehensive loss, represents the generation loss, represents the sorting loss, represents an arbitrary document identifier, Represents the current training data instance hard negative document identifiers, is the margin hyperparameter, is the number of hard negative document identifiers, is the weight hyperparameter used to balance the two losses, Represents the conversation context.

2. The conversational generation and retrieval method based on large model document index perception according to claim 1 is characterized in that: The document identifier is a combination of the document title followed by a proposition.

3. The conversational generation and retrieval method based on large model document index perception according to claim 1 is characterized in that: In the large language model decoding process, the document title is used as a prefix to guide the encoding process.

4. The conversational generation and retrieval method based on large model document index perception according to claim 1 is characterized in that: The generation loss adopts the maximum likelihood estimation loss.

5. The conversational generation and retrieval method based on large model document index perception according to claim 1 is characterized in that: The beam search decoding strategy is used to output a paragraph ranking list corresponding to the document identifier; the method includes: constraining the decoding space through FM-index to ensure that the generated document identifier is valid, and ranking the paragraphs according to the generation probability to provide the user with the most relevant information to the user query input.

6. A conversational generative retrieval system based on large-model document index perception, characterized in that: include: A query module is configured to: based on a user query input, use a trained large language model to output a ranked list of information related to the user query input; The training process of the large language model includes: based on the corpus, using the cross attention layer to extract the key information of the context, obtain the proposition, and obtain the document identifier; in the first stage of training, the generation loss is introduced to generate information related to the current query and generate the document identifier; in the second stage of training, the comprehensive loss is introduced to optimize the ranking list of the retrieved document identifiers; through the beam search decoding strategy, the paragraph ranking list corresponding to the document identifier is output; The comprehensive loss is expressed by the following formula: in, represents the comprehensive loss, represents the generation loss, represents the sorting loss, represents an arbitrary document identifier, Represents the current training data instance hard negative document identifiers, is the margin hyperparameter, is the number of hard negative document identifiers, is the weight hyperparameter used to balance the two losses, Represents the conversation context.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, it implements the steps in the conversational generation and retrieval method based on large model document index perception as described in any one of claims 1-5.

8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the steps in the conversational generation and retrieval method based on large model document index perception as described in any one of claims 1-5.

9. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the conversational generation and retrieval method based on large model document index perception as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Universal semantic retrieval based on unified object indices and representations

    CN118193670A

  • Generative retrieval model training method, information retrieval method and device

    CN119599080A