Pseudo document generation method and system based on multiple models

The method enhances search performance by using multiple language models to decompose, filter, and reassemble pseudodocuments, addressing the instability and semantic inconsistency issues in traditional query expansion methods, resulting in more accurate and complete search results.

CN120316201AActive Publication Date: 2025-07-15SHIJIAZHUANG TIEDAO UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510427311.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-15
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing sparse and intensive search methods perform poorly in strong semantic interpretation scenarios and are difficult to deal with wording changes and synonyms. Traditional pseudo-document generation methods lead to incomplete or inaccurate search results when prompting poor design or lack of domain-specific knowledge.

Method used

The multi-model pseudo-document generation method is adopted to generate pseudo-document by inputting query information into multiple language models, decomposing it into multiple information strips, filtering relevant information strips using the correlation evaluator, and reorganizing it into target pseudo-document, combining sparse and dense search methods to optimize the connection relationship of query information.

Benefits of technology

It improves the accuracy and search effect of pseudo-document, ensures the correlation between the information strip and the query information, and improves the accuracy and efficiency of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316201A_ABST
    Figure CN120316201A_ABST
Patent Text Reader

Abstract

The invention provides a pseudo document generation method and system based on multiple models, and belongs to the technical field of data generation. The method comprises the steps that query information is input into multiple language models to generate multiple pseudo documents; decomposing each pseudo document into a plurality of information bars based on the length of each pseudo document, wherein the number of the information bars is greater than or equal to the number of the pseudo documents; screening the plurality of information bars based on a correlation evaluator to obtain a plurality of related information bars; and recombining the plurality of related information strips into a target pseudo document. According to the multi-model-based pseudo document generation method and system provided by the invention, the pseudo document generation effect can be improved, so that the retrieval performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of data generation, and more particularly, relates to a multi-model-based pseudo-document generation method and system. Background Art

[0002] In the field of information retrieval, there are generally two main methods: sparse retrieval and dense retrieval. Sparse retrieval performs poorly in scenarios that require strong semantic interpretation and has difficulty dealing with changes in wording and synonym problems. Dense retrieval has significantly improved performance in semantic complex queries, but it is still challenged by lexical inconsistency and subtle context changes, resulting in incomplete or inaccurate retrieval results.

[0003] Traditional query expansion usually relies on shallow information, such as word frequency and co-occurrence relationships, which limits its ability to capture deep semantic relationships behind queries. In addition, traditional methods introduce irrelevant terms, thus reducing the accuracy of retrieval. To solve the above problems, current research uses pseudo-documents generated by large language models for query expansion, but there are still problems with unstable pseudo-document quality, especially when the prompt design is poor or the model lacks domain-specific knowledge. At this time, the generated pseudo-documents will deviate from the actual semantic requirements of the query, resulting in a lack of semantic consistency between its content and the original query, and even introducing irrelevant or misleading information, leading to poor retrieval performance.

[0004] Therefore, an accurate and reliable pseudo-document generation method is needed to improve retrieval performance. Summary of the Invention

[0005] The purpose of the present disclosure is to provide a multi-model-based pseudo-document generation method and system to improve the generation effect of pseudo-documents, thereby improving retrieval performance.

[0006] In the first aspect of the embodiments of the present disclosure, a multi-model-based pseudo-document generation method is provided, including: Inputting query information into multiple language models to generate multiple pseudo-documents; Decomposing each pseudo-document into multiple information items based on the length of each pseudo-document, and the number of information items is greater than or equal to the number of pseudo-documents; Filtering multiple information items based on a relevance evaluator to obtain multiple relevant information items; Recombining multiple relevant information items into a target pseudo-document.

[0007] In the second aspect of the embodiments of the present disclosure, a multi-model-based pseudo-document generation system is provided, including: A pseudo-document generation module for inputting query information into multiple language models to generate multiple pseudo-documents; A pseudo-document decomposition module, configured to decompose each pseudo-document into a plurality of information items based on the length of each pseudo-document, where the number of information items is greater than or equal to the number of pseudo-documents; A screening module, configured to screen a plurality of information items based on a relevance evaluator to obtain a plurality of relevant information items; A pseudo-document recombination module, configured to recombine a plurality of relevant information items into a target pseudo-document.

[0008] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the above-mentioned multi-model-based pseudo-document generation method are implemented.

[0009] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned multi-model-based pseudo-document generation method are implemented.

[0010] The beneficial effects of the multi-model-based pseudo-document generation method and system provided by the embodiments of the present disclosure are as follows: In the present disclosure, by decomposing a pseudo-document into a plurality of information items and calculating the relevance between each information item and the query information to screen out information items with strong relevance, the accuracy of the information items is improved, thereby improving the accuracy of the pseudo-document composed of them, and further improving the retrieval effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 It is a schematic flowchart of a multi-model-based pseudo-document generation method provided by an embodiment of the present disclosure; Figure 2 It is a schematic flowchart of a second multi-model-based pseudo-document generation method provided by an embodiment of the present disclosure; Figure 3 It is a structural block diagram of a multi-model-based pseudo-document generation system provided by an embodiment of the present disclosure; Figure 4 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0013] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.

[0014] The following explains the terms that appear in the embodiments of the present disclosure: Query expansion: The process of expanding and optimizing a query by adding relevant words, phrases, or semantic information based on the original query entered by the user. The purpose is to solve problems such as the possible ambiguity and single expression of the user's initial query, enabling the retrieval system to more accurately understand the user's intention and return more comprehensive and relevant results. For example, when the user enters "artificial intelligence", through query expansion, the system may supplement relevant words such as "machine learning" and "deep learning" to make the retrieval more accurate.

[0015] Sparse retrieval: A way of information retrieval that represents text as sparse vectors. Regarding documents and queries as sets of terms, an index is constructed based on discrete information such as the frequency of term occurrence and position in the document. During retrieval, the relevance between the document and the query is evaluated by calculating the matching degree and related weights between the query terms and the terms in the document index, so as to determine the retrieval results.

[0016] Dense retrieval: Using deep learning technology to transform text into low-dimensional dense vector representations. In the vector space, text vectors with similar semantics are close in distance. By calculating the similarity (such as cosine similarity) between the query vector and the document vector, the semantic relevance between texts is measured, and then the retrieval results are determined, emphasizing the in-depth understanding of text semantics and similarity measurement in the vector space.

[0017] To make the purpose, technical solutions, and advantages of the present disclosure clearer, the following will be described through specific embodiments in conjunction with the accompanying drawings.

[0018] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a multi-model-based pseudo-document generation method provided by an embodiment of the present disclosure. The method includes: S101: Input query information into multiple language models to generate multiple pseudo-documents.

[0019] In this embodiment, the query information refers to the information input for retrieving relevant content. The form of the query information is preferably in text form. The language model is an algorithm model based on deep learning, which has been trained with a large amount of text data and has the ability to understand and generate natural language, such as GPT, Qwen, Llama, GLM, or DeepSeek, etc. The pseudo-document is not an original document that actually exists, but a text generated by the language model according to the query information. That is, the previously generated text can simulate a real document, which contains content related to the query information and is used to assist in information retrieval or other natural language processing tasks. The number of pseudo-documents output by a language model can be one or more, that is, there is no clear size relationship between the number of language models and the number of pseudo-documents.

[0020] For example, if the query information is "The latest breakthroughs in electric vehicle battery technology", and it is input into multiple large language models, GPT, based on its training data and algorithms, generates a text containing content such as "In recent years, significant breakthroughs have been made in the energy density of electric vehicle battery technology. The research and development of solid-state batteries have accelerated, which is expected to greatly improve the driving range." This text is a pseudo-document. DeepSeek generates a similar pseudo-document with content such as "Regarding the charging speed of batteries, new fast-charging technologies are emerging continuously, and some technologies can achieve 80% charge in 15 minutes." The pseudo-document provides more relevant text data for subsequent information retrieval and processing, assisting in improving the retrieval and analysis effects.

[0021] S102: Decompose each pseudo-document into multiple information items based on the length of each pseudo-document, where the number of information items is greater than or equal to the number of pseudo-documents.

[0022] In this embodiment, the length of the pseudo-document refers to metrics for measuring the scale of the text, such as the number of characters, words, or sentences contained in the pseudo-document, and is used to determine whether the pseudo-document needs to be decomposed and how to decompose it. An information item is a smaller text unit split from the pseudo-document, and each information item contains relatively independent information.

[0023] The specific decomposition method can be: Decompose each pseudo-document into multiple information items based on the length of each pseudo-document, including: In response to the length of the pseudo-document being less than or equal to the first length threshold, use the pseudo-document as an information item; In response to the length of the pseudo-document being greater than the first length threshold, decompose the pseudo-document into information items based on heuristic rules; where , is the length of the pseudo-document, is the first length threshold, represents rounding up.

[0024] In this embodiment, if the length of the pseudo-document is less than or equal to the first length threshold, it is regarded as an independent information item. When the length of the pseudo-document is greater than the first length threshold, it can be decomposed into multiple information items by heuristic rules.

[0025] Heuristic rules are a method based on experience or specific strategies. For example, they can be split according to granularity, sentence boundaries, paragraph structure, topic relevance, etc.

[0026] S103: Screen multiple information items based on a relevance evaluator to obtain multiple relevant information items In this embodiment, the relevance evaluator is a tool or algorithm for evaluating the degree of association between information items and query information. Considering that the pseudo-document itself is generated data information, after splitting it, the data in the information items may have a low relevance to the initial query information. Therefore, the present disclosure screens the information items after splitting and uses the data that meets the screening conditions as relevant information items.

[0027] Specifically, screening multiple information items based on a relevance evaluator to obtain multiple relevant information items includes: In response to the relevance value of the information item being greater than the first relevance value, the information item is used as a relevant information item.

[0028] The calculation formula for the relevance value of the relevance evaluator is: , where represents the relevance value of the information item and the query information , represents the first relevance value, represents the second relevance value, is the first correlation coefficient, representing the weight of the first relevance value, is the second correlation coefficient, representing the weight of the second relevance value; represents the number of information items; , where is the dense embedding vector of the query information , is the dense embedding vector of the information item . and can be determined based on the experimental effect during the experiment or set according to experience.

[0029] In this embodiment, , where represents the inverse document frequency of the th word in the query information , represents the number of terms in the query information Q, Indicates information bar Length, Indicates query information The words, express In the information bar The frequency in represents the average length of all information bars, are the default parameters of the algorithm. , To include The number of information bars, Indicates the number of all information bars. It refers to the inverse document frequency.

[0030] S104: Reorganize the multiple related information pieces into a target pseudo document.

[0031] In this embodiment, the related information obtained in S103 can be represented as a set ,in, represents the first relevant value. The reorganized target pseudo document can be expressed as ,in, Represents a reorganized set , Indicates the first information item after reorganization. Indicates the second information item after reorganization. After the reorganization Information bars, that is, related information bars. The reorganization process can be random or according to the logical relationship between the information bars. For example, information bars that describe events in chronological order will be arranged in chronological order; if it is information bars that describe different aspects of a certain topic, the information bars describing different aspects such as causes, current situation, and impact can be sorted according to the internal logic of the topic.

[0032] From the above, it can be concluded that the present disclosure improves the accuracy of the information bars by decomposing the pseudo-document into multiple information bars and calculating the relevance of each information bar with the query information to screen out information bars with strong relevance, thereby improving the accuracy of the pseudo-documents composed of the information bars and further improving the retrieval effect.

[0033] Considering that different language models use different data domains and focuses during training, some are good at scientific and technological knowledge, while others have a better understanding of humanities and social sciences. Therefore, in the process of generating pseudo documents based on multilingual models, the query information should be matched with the areas of expertise of each model. Specifically, in one embodiment of the present disclosure, the query information is input into multiple language models to generate multiple pseudo documents, including: Determine the selection probability of each language model based on the first formula; Input the query information into multiple language models based on the selection probability of each language model to generate multiple pseudo-documents; The first formula is: , where is the query information, is the query information in the case of the language model selection probability, is the sensitivity parameter of the probability of selecting each language model, is the language model corresponding data domain keyword converted into a dense vector, is the similarity between the query information and the language model data domain, .

[0034] In this embodiment, the probability of selecting each language model can be calculated through the first formula in the case of a given query information , that is, the selection probability. In the formula, uses the vector dot product and norm to calculate the similarity between the query information and the dense vector converted from the data domain keywords of the language model , reflecting the degree of fit between the two in semantics and domain. The hyperparameter is used to adjust the sensitivity of the model selection. If is larger, the model selection will be more sensitive to the similarity difference and tend to select a model with a high similarity to the query information; if is smaller, the model selection is relatively loose. The first formula obtains the probability of each model being selected by performing an exponential operation on the similarity and normalizing, ensuring that the sum of the probabilities is 1. It can be concluded from the above that the present disclosure determines the probability of selecting each language model by calculating the similarity between the proficient fields of each model and the query information, which helps to ensure that the generated pseudo-documents are closer to the actual needs of the query information, improve the relevance and quality of the pseudo-documents, and thus improve the retrieval effect.

[0035] In an embodiment of the present disclosure, the pseudo-document is decomposed into

[0036] information items based on heuristic rules, including: The pseudo-document is decomposed into N information items with a granularity of based on heuristic rules; ; , where is the basic granularity threshold, is a regulatory factor, is a query complexity metric, , is the number of distinct words in the query information, is the total number of words in the query information, is the number of named entities in the query, is the weight coefficient.

[0037] In this embodiment, considering that if the query uses a large number of different words (high lexical richness) or contains multiple entities, the partitioning granularity should be finer (i.e., becomes smaller). If the query is relatively simple (such as a single keyword), the threshold can be appropriately increased (making the information strip longer). can be determined based on prior knowledge, represents the number of distinct words in the query information, reflecting the lexical richness of the query. The more distinct words, the more extensive the concepts involved in the query; is the total number of words in the query information; is the number of named entities in the query, representing the number of specific objects or concepts in the query; is the weight coefficient, used to adjust the importance of the number of named entities in the complexity calculation.

[0038] If the query uses a large number of different words (high lexical richness) or contains multiple entities, it means that the query complexity is high. At this time has a larger value. According to the calculation formula of, will become smaller, that is, the partitioning granularity is finer, and the pseudo-document is decomposed into more shorter information strips to more accurately match and process the information; if the query is relatively simple, such as only a single keyword, has a smaller value, will become larger, the information strip is longer, reducing the number of decompositions and improving the processing efficiency.

[0039] It can be concluded from the above that the present disclosure helps to capture key information in the query, reduce information omission and misunderstanding, and thus improve the accuracy and efficiency of information retrieval by dynamically adjusting the partitioning granularity of the information strip based on the query complexity.

[0040] In one embodiment of the present disclosure, the multi-model based pseudo-document generation method further includes: Determining a target connection relationship based on the retrieval method; the target connection relationship is the connection relationship between the target pseudo-document and the query information; Connecting the target pseudo-document and the query information based on the target connection relationship to obtain target query information; Performing a retrieval based on the target query information to obtain a retrieval result.

[0041] Determine the target connection relationship based on the retrieval method, including: In response to the retrieval method being sparse retrieval, determine the first connection relationship as the target connection relationship; In response to the retrieval method being dense retrieval, determine the second connection relationship as the target connection relationship; The number of repetitions of the query information in the first connection relationship and the second connection relationship is different.

[0042] In this embodiment, considering that sparse retrieval is based on the discrete representation of text, an index is mainly constructed for retrieval through information such as the occurrence frequency and position of terms. Since it relies more on the exact matching of words, the weight of the query information is relatively important. In actual applications, the query information is usually short and the pseudo-document is long. In order to prevent the query information from being ignored during the matching process with the pseudo-document during retrieval, the weight of the query information can be increased. For example, the query information can be repeated multiple times and connected to the pseudo-document, so that the retrieval system can pay more attention to the matching situation of the query terms in the pseudo-document and improve the retrieval accuracy.

[0043] Dense retrieval is based on deep learning to map text into a low-dimensional dense vector space representation, and the similarity between vectors is calculated to measure the relevance between the document and the query. It focuses on the overall semantic relationship of the text and has less dependence on the exact matching of individual words. Therefore, simply connecting the original query and the pseudo-document allows the model to effectively calculate the semantic similarity between them in the vector space, without the need to repeat the query information like sparse retrieval to enhance the weight.

[0044] Specifically, since the query information in sparse retrieval is usually much shorter than the target pseudo-document in order to balance the relative weights between the query information and the target pseudo-document before the query, the two are connected according to the first connection relationship. The first connection relationship can be expressed as , where represents the string concatenation operation. The extended query is the target query information and can be used as the new query information for BM25 retrieval. is a positive integer, preferably 5.

[0045] Dense retrieval: The connection method is similar to sparse retrieval, which is a simple connection of the original query q and the pseudo-document d, separated by [SEP]. Specifically, the second connection relationship can be . SEP is the abbreviation of Separator, which means separator and is used to distinguish different text segments.

[0046] As can be seen from the above, for sparse retrieval, the present disclosure effectively improves the weight of query information in the retrieval process by increasing the repetition times of query information to connect with the target pseudo-document, which helps to ensure that the query information will not be ignored when matching with the pseudo-document and improves the accuracy of retrieval.

[0047] Figure 2 This is a schematic flowchart of the second multi-model-based pseudo-document generation method provided by an embodiment of the present disclosure. Refer to Figure 2 , input the query information into multiple models to obtain multiple pseudo-documents, and the target pseudo-document can be obtained by decomposing, screening and reorganizing them , then input the query information and the target pseudo-document for connection to obtain the target query information. Figure 2 The information items obtained after screening in Figure 2 are relevant information items, and their essence is information items, so information items are selected for expression in

[0048] Corresponding to the multi-model-based pseudo-document generation method in the above embodiment, Figure 3 This is a structural block diagram of a multi-model-based pseudo-document generation system provided by an embodiment of the present disclosure. For the sake of illustration, only the parts related to the embodiments of the present disclosure are shown. Refer to Figure 3 , the multi-model-based pseudo-document generation system 20 includes: a pseudo-document generation module 21, a pseudo-document decomposition module 22, a screening module 23 and a pseudo-document recombination module 24.

[0049] Among them, the pseudo-document generation module 21 is used to input the query information into multiple language models to generate multiple pseudo-documents; The pseudo-document decomposition module 22 is used to decompose each pseudo-document into multiple information items based on the length of each pseudo-document, and the number of information items is greater than or equal to the number of pseudo-documents; The screening module 23 is used to screen multiple information items based on a relevance evaluator to obtain multiple relevant information items; The pseudo-document recombination module 24 is used to recombine multiple relevant information items into a target pseudo-document.

[0050] In an embodiment of the present disclosure, the pseudo-document generation module 21 is specifically used to determine the selection probability of each language model based on the first formula; Input the query information into multiple language models to generate multiple pseudo-documents based on the selection probability of each language model; The first formula is: , where is the query information, is the language model in the case of the query information Q The selection probability, is the sensitivity parameter of the probability of selecting each language model, for the language model is the dense vector converted from the keyword of the corresponding data field of the language model, is the similarity between the query information and the data field of the language model of. .

[0051] In an embodiment of the present disclosure, the pseudo-document decomposition module 22 is specifically configured to, in response to the length of the pseudo-document being less than or equal to the first length threshold, use the pseudo-document as an information item; In response to the length of the pseudo-document being greater than the first length threshold, decompose the pseudo-document into information items based on heuristic rules; where , is the length of the pseudo-document, is the first length threshold, represents rounding up.

[0052] In an embodiment of the present disclosure, the pseudo-document decomposition module 22 is specifically further configured to decompose the pseudo-document into N information items with a granularity of based on heuristic rules; where , where is the basic granularity threshold, is the adjustment factor, is the query complexity index, , is the number of non-repeating words in the query information, is the total number of words in the query information, is the number of named entities in the query, is the weight coefficient.

[0053] In an embodiment of the present disclosure, the screening module 23 is specifically configured to, in response to the correlation value of the information item being greater than the first correlation value, use the information item as a relevant information item.

[0054] The calculation formula for the correlation value of the correlation evaluator is: , where represents the correlation between the information item and the query information , represents the first correlation value, represents the second correlation value, is the first correlation coefficient, representing the weight of the first correlation value, is the second correlation coefficient, representing the weight of the second correlation value; Indicates the number of information bars; , where is the query information of the dense embedding vector, is the information bar of the dense embedding vector.

[0055] In an embodiment of the present disclosure, the multi-model based pseudo-document generation system 20 further includes: an information connection module, configured to determine a target connection relationship based on a retrieval method; the target connection relationship is the connection relationship between the target pseudo-document and the query information; Connect the target pseudo-document and the query information based on the target connection relationship to obtain the target query information; Retrieve based on the target query information to obtain a retrieval result.

[0056] In an embodiment of the present disclosure, the information connection module is specifically configured to determine the first connection relationship as the target connection relationship in response to the retrieval method being sparse retrieval; Determine the second connection relationship as the target connection relationship in response to the retrieval method being dense retrieval; The number of repetitions of the query information in the first connection relationship and the second connection relationship is different.

[0057] See Figure 4 , Figure 4 is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 4 shown, the electronic device 300 in this embodiment may include: one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The above-mentioned processors 301, input devices 302, output devices 303, and memories 304 communicate with each other through a communication bus 305. The memory 304 is used to store computer programs, and the computer programs include program instructions. The processor 301 is configured to execute the program instructions stored in the memory 304. Among them, the processor 301 is configured to call the program instructions to execute the functions of each module / unit in the above-mentioned system embodiments, such as Figure 3 the functions of the modules 21 to 24 shown.

[0058] It should be understood that in the embodiments of the present disclosure, the so-called processor 301 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0059] The input device 302 may include a touchpad, a fingerprint acquisition sensor (for acquiring the fingerprint information and the direction information of the fingerprint of the user), a microphone, etc., and the output device 303 may include a display (such as an LCD), a speaker, etc.

[0060] The memory 304 may include a read-only memory and a random access memory, and provide instructions and data to the processor 301. A part of the memory 304 may also include a non-volatile random access memory. For example, the memory 304 may also store information about the device type.

[0061] In specific implementation, the processor 301, the input device 302, and the output device 303 described in the embodiments of the present disclosure may execute the implementation manners described in the first embodiment and the second embodiment of the multi-model-based pseudo-document generation method provided by the embodiments of the present disclosure, and may also execute the implementation manner of the electronic device described in the embodiments of the present disclosure, which will not be elaborated herein.

[0062] In another embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, all or part of the processes in the method of the above embodiment are implemented. It can also be completed by instructing relevant hardware through the computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0063] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the electronic device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0064] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present disclosure.

[0065] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described electronic devices and units can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.

[0066] In several embodiments provided by this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces or units, and can also be in the form of electrical, mechanical or other connections.

[0067] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present disclosure.

[0068] In addition, each functional unit in various embodiments of the present disclosure can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0069] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A method for generating pseudo-documents based on multiple models, characterized in that Including: Inputting query information into multiple language models to generate multiple pseudo-documents; Based on the lengths of the respective pseudo-documents, decomposing the respective pseudo-documents into multiple information items, the number of information items being greater than or equal to the number of pseudo-documents; Screening the multiple information items based on a relevance evaluator to obtain multiple relevant information items; Recombining the multiple relevant information items into a target pseudo-document.

2. The method for generating pseudo documents based on multiple models according to claim 1, wherein The inputting query information into multiple language models to generate multiple pseudo-documents includes: Determining the selection probability of each language model based on a first formula; Inputting query information into multiple language models to generate multiple pseudo-documents based on the selection probability of each language model; The first formula is as follows: , where is the query information, is the selection probability of the language model in the case of query information Q, is the sensitivity parameter of the probability of selecting each language model, is the language model corresponding dense vector converted from the data domain keyword, is the data domain similarity between the query information and the language model , .

3. The method for generating pseudo-documents based on multiple models according to claim 1, wherein, The decomposing the respective pseudo-documents into multiple information items based on the lengths of the respective pseudo-documents includes: In response to the length of the pseudo-document being less than or equal to a first length threshold, using the pseudo-document as an information item; In response to the length of the pseudo-document being greater than the first length threshold, decompose the pseudo-document into information items based on heuristic rules; Among them, , is the length of the pseudo-document, is the first length threshold, represents rounding up.

4. The method for generating pseudo-documents based on multiple models according to claim 3, wherein Decompose the pseudo-document into information items based on heuristic rules, including: Decompose the pseudo-document into N information items with a granularity of according to heuristic rules; Among them, , where, is the basic granularity threshold, is the adjustment factor, is the query complexity index, , is the number of non-repeating words in the query information, is the total number of words in the query information, is the number of named entities in the query, is the weight coefficient.

5. The method for generating pseudo-documents based on multiple models according to claim 1, wherein The screening the multiple information items based on a relevance evaluator to obtain multiple relevant information items includes: In response to the relevance value of the information item being greater than a first relevance value, using the information item as a relevant information item. The calculation formula for the relevance value of the relevance evaluator is: , where represents an information bar and the query information correlation, represents the first correlation value, represents the second correlation value, is the first correlation coefficient, representing the weight of the first correlation value, is the second correlation coefficient, representing the weight of the second correlation value; represents the number of information bars; , where is the dense embedding vector of the query information , is the dense embedding vector of the information bar .

6. The method for generating pseudo documents based on multiple models according to claim 1, wherein Also including: Determining a target connection relationship based on a retrieval method; The target connection relationship is the connection relationship between the target pseudo-document and the query information; Connecting the target pseudo-document and the query information based on the target connection relationship to obtain target query information; Performing a retrieval based on the target query information to obtain a retrieval result.

7. The method for generating pseudo-documents based on multiple models according to claim 6, wherein, The determining a target connection relationship based on a retrieval method includes: In response to the retrieval method being a sparse retrieval, determining a first connection relationship as the target connection relationship; In response to the retrieval method being a dense retrieval, determining a second connection relationship as the target connection relationship; The number of repetitions of the query information in the first connection relationship and the second connection relationship is different.

8. A multi-model-based pseudo-document generation system, characterized in that, Including: A pseudo-document generation module for inputting query information into multiple language models to generate multiple pseudo-documents; A pseudo-document decomposition module for decomposing the respective pseudo-documents into multiple information items based on the lengths of the respective pseudo-documents, the number of information items being greater than or equal to the number of pseudo-documents; A screening module for screening the multiple information items based on a relevance evaluator to obtain multiple relevant information items; A pseudo-document recombination module for recombining the multiple relevant information items into a target pseudo-document.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Information retrieval method and system and medium

    CN112163065A

  • Chain-based retrieval enhancement generation method and device and readable storage medium

    CN118939776A

  • Special vehicle operation and maintenance knowledge retrieval method based on large model

    CN119066144A

  • Power document intelligent question and answer method and system based on large language model

    CN119577082A

  • Machine learning architecture for contextual data retrieval

    US12254029B1