Document retrieval method and device, electronic equipment and storage medium
By obtaining and fusion of text information and picture content, and using visual language models for document search, the problem of insufficient retrieval efficiency and accuracy of mixed graphic documents in the prior art is solved, and efficient and accurate meeting of diversified retrieval needs is achieved.
Patent Information
- Application Number
- CN202411881844.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is difficult to meet the different search needs of users in document retrieval, especially when processing mixed graphic documents, the search efficiency and accuracy are insufficient.
By obtaining the text information and picture content in the search request, the visual language model extracts the picture content information and fuses it with the text information to generate query content for search. At the same time, a variety of search methods are used to search different databases, and combined with the reciprocal sorting and fusion method with weights, the final search results are determined.
It realizes efficient retrieval of mixed graphic documents, meets users' diverse search needs, and improves search efficiency and accuracy.
Smart Images

Figure CN119961491A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to natural language processing, computer vision, deep learning and large models, and specifically to document retrieval methods, devices, electronic devices and storage media. Background Art
[0002] In practical applications, users often have the need to retrieve documents. Traditional retrieval methods mainly include text retrieval through natural language and image retrieval based on the text title of the image. Summary of the invention
[0003] The present disclosure provides a document retrieval method, device, electronic device and storage medium.
[0004] A document retrieval method, comprising:
[0005] Obtaining a search request, wherein the search request includes second text information and / or a second image, and determining query content according to the search request;
[0006] Searching a target database according to the query content, wherein the target database stores predetermined document information of each original document as a search object, wherein the original document includes first text information and / or a first picture;
[0007] A target document as a target search result is determined from the retrieved original documents, and the target document is returned.
[0008] A document retrieval device comprises: a request processing module, a document retrieval module and a result generation module;
[0009] The request processing module is used to obtain a search request, wherein the search request includes the second text information and / or the second image, and determine the query content according to the search request;
[0010] The document retrieval module is used to search a target database according to the query content, wherein the target database stores predetermined document information of each original document as a search object, wherein the original document includes first text information and / or a first image;
[0011] The result generation module is used to determine a target document as a target search result from the retrieved original documents, and return the target document.
[0012] An electronic device, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described above.
[0016] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0017] A computer program product comprises a computer program / instruction, wherein the computer program / instruction implements the method described above when executed by a processor.
[0018] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0020] Figure 1 A flowchart of an embodiment of the document retrieval method disclosed in the present invention;
[0021] Figure 2 This is a schematic diagram of the implementation process of the pretreatment process described in the present disclosure;
[0022] Figure 3 A schematic diagram of the overall implementation process of the document retrieval method disclosed in the present invention;
[0023] Figure 4 It is a schematic diagram of the composition structure of the first embodiment 400 of the document retrieval device described in the present disclosure;
[0024] Figure 5 It is a schematic diagram of the composition structure of the second embodiment 500 of the document retrieval device described in the present disclosure;
[0025] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0026] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] In addition, it should be understood that the term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0028] Figure 1 Flow chart of an embodiment of the document retrieval method described in the present disclosure. Figure 1 As shown, the following specific implementation methods are included.
[0029] In step 101, a search request is obtained, the search request includes second text information and / or a second image, and query content is determined according to the search request.
[0030] In step 102, a target database is searched according to the query content. The target database stores predetermined document information of each original document as a search object. The original document includes first text information and / or a first picture.
[0031] In step 103, a target document as a target search result is determined from the retrieved original documents, and the target document is returned.
[0032] The search request may be a search request input by a user. It can be seen that the scheme described in the above method embodiment can support users to perform various forms of search, such as mixed image and text search. Accordingly, the original document can be a mixed image and text document, which is not limited to the traditional method of searching text by natural language and searching pictures according to the text title of the picture, thereby meeting the different search needs of users and improving the search efficiency.
[0033] The target database may store predetermined document information of each original document. In some embodiments of the present disclosure, the predetermined document information may include: first document information and second document information. The target database may include: a first database and a second database. Accordingly, before obtaining a retrieval request, a preprocessing process may be performed, that is, for each original document, the first document information and the second document information of the original document may be obtained respectively, and the first document information may be stored in the first database, and the second document information may be stored in the second database.
[0034] The preprocessing process can be completed offline in advance. Different document information is stored in different databases. Accordingly, different databases can be searched separately later, that is, the original document can be searched based on different document information, thereby improving the comprehensiveness and accuracy of the search.
[0035] In some embodiments of the present disclosure, for any original document, a method for obtaining the first document information of the original document may include: in response to determining that the original document only includes first text information, determining the first text information as the first document information of the original document; in response to determining that the original document only includes a first picture or determining that the original document includes both the first text information and the first picture, extracting first picture content information (i.e., the main content description information of the picture) from the first picture, and replacing the first picture in the original document with the corresponding first picture content information, thereby obtaining the first document information of the original document.
[0036] Assuming that an original document includes both the first text information and the first image, the original document can be parsed to obtain the first text information and the first image therein. The first text information may include multiple paragraphs of text. In addition, the number of first images may be one or more. In order to improve the subsequent processing efficiency and the accuracy of the processing results, the first text information may be cleaned, such as cleaning abnormal characters, continuous line breaks and spaces therein. Afterwards, the first image content information can be extracted from the first image. For example, assuming that the original document includes two first images, namely, image 1 and image 2, the image content information 1 of image 1 and the image content information 2 of image 2 can be obtained respectively, and then the image content information 1 can be used to replace the image 1 in the original document, and the image content information 2 can be used to replace the image 2 in the original document, thereby obtaining a full-text document in plain text form, that is, the first document information of the original document.
[0037] By adopting the above processing method, the first document information in the form of plain text can be quickly and accurately obtained through operations such as image content information extraction and image replacement, thereby laying a good foundation for subsequent processing.
[0038] In some embodiments of the present disclosure, when extracting the first picture content information from the first picture, a vision-language model may be used to extract the first picture content information from the first picture.
[0039] The visual language big model is a big model that can process text and images / videos at the same time. It can use deep learning technology to convert text and image / video data into high-dimensional vector representations, thereby realizing cross-modal information interaction and fusion.
[0040] The visual language model can use a predetermined prompt word (prompt) to extract the first picture content information. The prompt word can be manually preset, and by reasonably setting the prompt word, the accuracy of the extracted first picture content information can be improved.
[0041] The visual language large model can be a pre-trained large model. Since the content of image understanding is included in the pre-training process, there is no need for secondary training and it can be used directly, thereby reducing the implementation cost.
[0042] In some embodiments of the present disclosure, for any original document, target key information corresponding to the first document information of the original document may be obtained, and the target key information may be determined as the second document information of the original document.
[0043] Specifically, in some embodiments of the present disclosure, in response to determining that the first document information only includes first picture content information, first key information (key information can be understood as summary information) can be extracted from the first picture content information, and the first key information can be determined as target key information; in response to determining that the first document information only includes first text information, second key information can be extracted from the first text information, and the second key information can be determined as target key information; in response to determining that the first document information includes both first text information and first picture content information, the extracted first key information and second key information can be determined as target key information.
[0044] Assuming that an original document includes both a first text information and a first picture, in order to generate the second document information of the original document, the first key information can be extracted from the first picture content information in the first document information of the original document. For example, the first key information can be extracted using a large model, and the large model can be a visual language large model or other large models. In addition, the second key information can also be extracted from the first text information in the first document information of the original document, and then the extracted first key information and the second key information can be determined as the target key information of the original document, that is, determined as the second document information of the original document.
[0045] It can be seen that for any original document, when generating the corresponding second document information, the first document information corresponding to the original document can be used as a basis, thereby improving the generation efficiency of the second document information.
[0046] In some embodiments of the present disclosure, a method for extracting second key information from first text information may include: slicing the first text information to obtain text fragments, extracting fragment key information from each text fragment, splicing the fragment key information in order of the appearance position of the corresponding text fragment in the first text information from first to last to obtain the second key information, or inputting the fragment key information into an information fusion model to obtain the output second key information.
[0047] After slicing the first text information, a plurality of text fragments can usually be obtained. For each text fragment, the key information of the fragment can be extracted using the large model respectively, and then the key information of each fragment can be spliced to obtain the second key information. Alternatively, the key information of each fragment can be input into the information fusion model (large model) to obtain the second key information output by the information fusion model after integrating the key information of each fragment. The specific method to be adopted can be determined according to actual needs, which is very flexible and convenient. Preferably, the latter method can be adopted to improve the quality of the generated second key information.
[0048] After obtaining the corresponding first document information and second document information for each original document, the first document information can be stored in the first database and the second document information can be stored in the second database. Depending on the information stored, the first database can also be called a full-text database and the second database can be called a key information database.
[0049] In some embodiments of the present disclosure, storing the first document information into the first database may include: performing word segmentation processing on the first document information, storing each word segmentation result in the first database, and obtaining a first vector representation corresponding to the first document information, and storing the first vector representation in the first database; storing the second document information into the second database may include: performing word segmentation processing on the second document information, storing each word segmentation result in the second database, and obtaining a second vector representation corresponding to the second document information, and storing the second vector representation in the second database.
[0050] That is, for each original document, each database can use two storage forms to store the corresponding document information, namely, the word segmentation storage form and the vector storage form. Among them, for any original document, the word segmentation storage form means that the document information of the original document can be processed by word segmentation to obtain each word segmentation result, and each word segmentation result can be stored, and the vector storage form means that the document information of the original document can be converted into a corresponding vector representation and then stored.
[0051] Since the above two storage forms are used at the same time, each database can support different retrieval methods, such as word segmentation fuzzy retrieval method, similarity distance retrieval method based on vector representation, etc.
[0052] Combined with the above introduction, Figure 2 Schematic diagram of the implementation process of the pre-processing process described in the present disclosure. Figure 2 As shown, assuming that the original document includes both the first text information and the first image, each original document can be first parsed to obtain the first text information and the first image therein, and then the first text information can be cleaned, and the first image content information can be extracted from the first image, and then the first image can be replaced with the corresponding first image content information in the original document to obtain the first document information of the original document, and the first document information can be stored in the first database. In addition, the first text information after data cleaning can be sliced to obtain various text fragments, and the fragment key information can be extracted from each text fragment respectively, and then the key information of each fragment can be fused to obtain the second key information, and the first key information can be extracted from the first image content information, and then the first key information and the second key information can be used to generate the second document information of the original document, and further, the second document information can be stored in the second database.
[0053] After that, an online search can be performed based on the processing results of the preprocessing process. For example, a search request input by a user can be obtained, the search request can include the second text information and / or the second image, and the query content can be determined according to the search request, and then the target database can be searched according to the query content, and a target document as a target search result can be determined from the original documents retrieved, and then the target document can be returned.
[0054] Among them, the specific sources of the second text information and the second picture are not limited. For example, the second text information can be text information directly input by the user, or it can be text information obtained by recognizing the voice after obtaining the voice input by the user. For another example, the second picture can be a picture downloaded from a website, a picture stored this time, or a picture captured from a video, etc.
[0055] In some embodiments of the present disclosure, in response to determining that the search request only includes the second text information, the second text information may be determined as the query content; in response to determining that the search request only includes the second picture, the second picture content information may be extracted from the second picture, and the second picture content information may be determined as the query content; in response to determining that the search request includes both the second text information and the second picture, the second text information and the extracted second picture content information may be spliced in a predetermined order, and the splicing result may be determined as the query content, or the second text information and the second picture content information may be input into an information fusion model to obtain the output query content. The predetermined order may refer to the order in which the second text information comes first and the second picture content information comes later, or vice versa.
[0056] The second text information may be first cleaned, such as removing abnormal characters, continuous line breaks, and spaces.
[0057] In addition, in some embodiments of the present disclosure, the visual language model can be used to extract the second picture content information from the second picture, wherein the visual language model can use the same prompt words when extracting the first picture content information and the second picture content information.
[0058] Assuming that the search request includes both the second text information and the second image, the second text information and the second image content information can be fused as the query content. For example, the second text information and the second image content information can be input into an information fusion model to obtain the output query content.
[0059] Through the above processing, the query content can include various information input by the user, that is, graphic and text information, thereby improving the accuracy of subsequent search results.
[0060] After the query content is determined, the target database can be searched. In some embodiments of the present disclosure, the first database can be searched using M search methods according to the query content to obtain initial search results corresponding to each search method, where M is a positive integer greater than 1, and the second database can be searched using the M search methods according to the query content to obtain initial search results corresponding to each search method.
[0061] The specific value of M can be determined according to actual needs. Assuming that the value of M is 3, for the convenience of description, the three search methods are respectively called search method 1, search method 2 and search method 3. Then, the first database can be searched by using search method 1, search method 2 and search method 3 respectively, and the second database can be searched by using search method 1, search method 2 and search method 3 respectively.
[0062] In order to save search time, various search methods can be executed in parallel, that is, a parallel hybrid search method can be adopted in the scheme described in the present disclosure, thereby improving the search efficiency, etc.
[0063] Further, a target document as a target search result may be determined from the retrieved original documents. In some embodiments of the present disclosure, a candidate document may be determined from the original documents in each initial search result, and a comprehensive score of the candidate document may be determined, and then the candidate documents may be sorted in descending order of the comprehensive score, and the candidate documents in the top N after sorting may be determined as the target document, where N is a positive integer and is less than or equal to the number of candidate documents.
[0064] Assume that after searching the first database using retrieval method 1, initial retrieval result 11 is obtained, after searching the first database using retrieval method 2, initial retrieval result 12 is obtained, after searching the first database using retrieval method 3, initial retrieval result 13 is obtained, and assume that after searching the second database using retrieval method 1, initial retrieval result 21 is obtained, after searching the second database using retrieval method 2, initial retrieval result 22 is obtained, and after searching the second database using retrieval method 3, initial retrieval result 23 is obtained. Assume that initial retrieval result 11, initial retrieval result 12, initial retrieval result 13, initial retrieval result 21, initial retrieval result 22 and initial retrieval result 23 include 50 original documents in total. Then, candidate documents can be first determined from these 50 original documents. Assume that the number of candidate documents is 30, and the comprehensive score of each candidate document can be determined respectively. Then, these 30 candidate documents can be sorted in descending order according to the comprehensive score, and the candidate documents in the top 10 (assuming that the value of N is 10) after sorting are determined as target documents.
[0065] It can be seen that the above processing method determines the target document by performing two-level screening, that is, first screening candidate documents from the original documents in each initial search result, and then screening the target document from the candidate documents, thereby improving the accuracy of the determined target document.
[0066] In some embodiments of the present disclosure, the comprehensive scores of the original documents in each initial search result can be obtained respectively, and the original documents in each initial search result can be used to form a first document set, and then the first document set can be determined as a target document set, and the following first processing can be performed: any original document in the target document set is determined as a document to be processed, and the original documents in the target document set other than the document to be processed are determined as selected documents, in response to determining that there is no predetermined document in the selected documents, the predetermined document is a selected document that is the same document as the document to be processed, the document to be processed is determined as a candidate document, the comprehensive score of the document to be processed is determined as the comprehensive score of the candidate document, and the document to be processed is deleted from the first document set; in response to determining that there is a predetermined document in the selected documents, the document to be processed is determined as a candidate document, the sum of the comprehensive score of the document to be processed and the comprehensive score of the predetermined document is obtained, the sum is determined as the comprehensive score of the candidate document, and the document to be processed and the predetermined document are deleted from the first document set; in response to determining that the updated first document set is not empty, the updated first document set is determined as the target document set, and the first processing is repeated.
[0067] For example, assuming that initial search result 11 includes original document 1, original document 2, original document 3, original document 4, original document 5, and original document 6, initial search result 12 includes original document 1, original document 2, original document 3, original document 4, original document 5, and original document 6, initial search result 13 includes original document 2, original document 3, original document 4, original document 5, original document 6, and original document 7, initial search result 21 includes original document 3, original document 4, original document 5, original document 6, original document 7, and original document 8, initial search result 22 includes original document 3, original document 4, original document 5, original document 7, original document 8, and original document 9, and initial search result 23 includes original document 1, original document 2, original document 3, original document 5, original document 6, and original document 7, then each initial search result includes 36 original documents in total, and the comprehensive scores of the 36 original documents can be obtained respectively, with the original documents Taking original document 1 as an example, original document 1 can be determined as a candidate document. Since original document 1 appears 3 times in the 6 initial search results, the sum of the comprehensive scores of the 3 original documents 1 can be used as the comprehensive score of the candidate document. Taking original document 3 as an example, original document 3 can be determined as another candidate document. Since original document 3 appears 6 times in the 6 initial search results, the sum of the comprehensive scores of the 6 original documents 3 can be used as the comprehensive score of the candidate document. Taking original document 9 as an example, original document 9 can be determined as another candidate document. Since original document 9 appears once in the 6 initial search results, the comprehensive score of original document 9 can be directly used as the comprehensive score of the candidate document. According to the above method, a total of 9 candidate documents can be obtained, namely original document 1, original document 2, original document 3, original document 4, original document 5, original document 6, original document 7, original document 8 and original document 9, and the comprehensive score of each candidate document can be obtained respectively.
[0068] In some embodiments of the present disclosure, a method for respectively obtaining a comprehensive score of an original document in each initial search result may include: for any original document in each initial search result, respectively performing the following processing: respectively obtaining the retrieval score of the original document corresponding to 2M retrieval operations, the 2M retrieval operations including: retrieval operations performed on a first database using M retrieval methods and retrieval operations performed on a second database using M retrieval methods, and determining the sum of the 2M retrieval scores as the comprehensive score of the original document.
[0069] For example, assuming that the value of M is 3, for a certain original document, a retrieval score can be obtained for each retrieval operation, that is, a total of 6 retrieval scores can be obtained, and then the 6 retrieval scores can be added together, and the sum is used as the comprehensive score of the original document.
[0070] In some embodiments of the present disclosure, for any retrieval operation, a method for obtaining a retrieval score of an original document corresponding to the retrieval operation may include: in response to determining that the original document is included in the initial retrieval results corresponding to the retrieval operation, and determining that the ranking position of the original document is less than a predetermined threshold, obtaining the sum of the ranking position and K, where K is a positive integer, and obtaining the ratio of the weight corresponding to the retrieval operation to the sum, determining the ratio as the retrieval score, the ranking position being the ranking position of the original document in the initial retrieval results corresponding to the retrieval operation, and in response to determining that the original document is not included in the initial retrieval results corresponding to the retrieval operation, or determining that the ranking position of the original document is greater than or equal to the predetermined threshold, determining 0 as the retrieval score.
[0071] That is:
[0072]
[0073] Where d represents any original document d, which may be included in at least one initial search result, D represents the set of original documents in each initial search result, i.e., the initial first document set, R represents the set of 2M search operations, r represents any search operation, and rank r,d Indicates the ranking position value, where if it is determined that the initial search result corresponding to the search operation r includes the original document d, and it is determined that the ranking position of the original document d in the initial search result corresponding to the search operation r is less than a predetermined threshold, the ranking position of the original document d in the initial search result corresponding to the search operation r can be used as the rank r,d For example, assuming that the initial search results corresponding to the search operation r include 20 original documents, and the original document d ranks 5th, which is less than the threshold 15, then 5 can be used as the rank r,d , w r Represents the weight corresponding to the search operation r. Accordingly, the rank can be obtained. r,d The sum of K and w can be obtained r With rank r,d +K, and use the ratio as the retrieval score of the original document d corresponding to the retrieval operation r. If it is determined that the initial retrieval result corresponding to the retrieval operation r does not include the original document d, or if it is determined that the initial retrieval result corresponding to the retrieval operation r includes the original document d, but the ranking position of the original document d in the initial retrieval result corresponding to the retrieval operation r is greater than or equal to a predetermined threshold, then the rank r,d It is regarded as infinite, and accordingly, 0 can be determined as the retrieval score of the original document d corresponding to the retrieval operation r. After obtaining the retrieval scores of the original document d corresponding to each retrieval operation in R respectively, each retrieval score can be added, and the sum can be used as the comprehensive score of the original document d.
[0074] The specific values of the above-mentioned predetermined threshold, K, and the weights corresponding to each search operation can be determined according to actual needs.
[0075] It can be seen that a weighted reciprocal rank fusion method is proposed in the scheme described in the present disclosure. The weighted reciprocal rank fusion is based on the traditional reciprocal rank fusion (RRF, Reciprocal Rank Fusion) and has made improvements such as weight setting, wherein the weight represents the contribution value of each sorting method to the final calculation result. Accordingly, the contribution value can be controlled by adjusting the weight, which is very flexible and convenient, and improves controllability. In addition, the scheme described in the present disclosure combines a variety of different sorting results (corresponding to different initial retrieval results) to finally determine the required target document, thereby improving the accuracy of the determined target document.
[0076] Combined with the above introduction, Figure 3 FIG. 1 is a schematic diagram of the overall implementation process of the document retrieval method described in the present disclosure. Figure 3 As shown, the search request input by the user can be first obtained. Assuming that the search request includes both the second text information and the second image, the second text information and the second image can be parsed from the search request respectively, and then the second text information can be cleaned, and the second image content information can be extracted from the second image, and then the second text information and the second image content information after data cleaning can be fused to obtain the query content, and then according to the query content, M kinds of search methods (corresponding to the search method 1 to the search method M shown in the figure) can be used to search the first database respectively to obtain the initial search results corresponding to each search method, and the M kinds of search methods can be used to search the second database respectively to obtain the initial search results corresponding to each search method, and further, based on the original documents in each initial search result, the target document as the target search result can be finally determined through the weighted inverse sorting fusion method described in the present disclosure, and the target document can be returned to the user.
[0077] The M types of search methods may include a word segmentation fuzzy search method and a similarity distance search method based on vector representation, etc. Word segmentation fuzzy search is a fuzzy search algorithm after Chinese word segmentation, which is different from "precise search". When the query content is partially similar to the content to be searched, it can still be retrieved. The first database and the second database described in the present disclosure both store the word segmentation results of the document information. Accordingly, when the first database and the second database are searched using the word segmentation fuzzy search method, the query content also needs to be processed by word segmentation. The similarity distance search method based on vector representation refers to representing the query content and the content to be searched with vectors respectively, and then selecting several results with higher similarity or closer distance as the search results by calculating the similarity or distance of the two vectors. The first database and the second database described in the present disclosure both store the vector representation of the document information. Accordingly, when the first database and the second database are searched using the similarity distance search method based on vector representation, the query content also needs to be converted into the corresponding vector representation.
[0078] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the described order of actions, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0079] The above is an introduction to the method embodiment. The following is a further explanation of the scheme disclosed in the present invention through an apparatus embodiment.
[0080] Figure 4 FIG. 4 is a schematic diagram of the composition structure of the first embodiment 400 of the document retrieval device described in the present disclosure. Figure 4 As shown, it includes: a request processing module 401, a document retrieval module 402 and a result generation module 403.
[0081] The request processing module 401 is used to obtain a search request, where the search request includes the second text information and / or the second image, and determine the query content according to the search request.
[0082] The document retrieval module 402 is used to search the target database according to the query content. The target database stores predetermined document information of each original document as the retrieval object. The original document includes the first text information and / or the first picture.
[0083] The result generation module 403 is used to determine a target document as a target search result from the retrieved original documents and return the target document.
[0084] The scheme described in the above-mentioned device embodiment can support users to perform various forms of retrieval, such as mixed image and text retrieval. Accordingly, the original document can be a mixed image and text document, which is not limited to the traditional method of searching text through natural language and searching pictures according to the text title of the picture, thereby meeting the different retrieval needs of users and improving retrieval efficiency.
[0085] Figure 5 FIG. 5 is a schematic diagram of the structure of the second embodiment 500 of the document retrieval device described in the present disclosure. Figure 5 As shown, it includes: a request processing module 401, a document retrieval module 402, a result generation module 403 and a pre-processing module 404.
[0086] In some embodiments of the present disclosure, the predetermined document information may include: first document information and second document information, and the target database may include: a first database and a second database. Accordingly, before obtaining a retrieval request, the preprocessing module 404 may also perform a preprocessing process, that is, for each original document, the first document information and the second document information of the original document may be obtained respectively, and the first document information may be stored in the first database, and the second document information may be stored in the second database.
[0087] In some embodiments of the present disclosure, for any original document, the way in which the preprocessing module 404 obtains the first document information of the original document may include: in response to determining that the original document only includes first text information, determining the first text information as the first document information of the original document; in response to determining that the original document only includes a first picture or determining that the original document includes both the first text information and the first picture, extracting first picture content information from the first picture, and replacing the first picture in the original document with the corresponding first picture content information, thereby obtaining the first document information of the original document.
[0088] In some embodiments of the present disclosure, when extracting the first picture content information from the first picture, the pre-processing module 404 may use a large visual language model to extract the first picture content information from the first picture.
[0089] In some embodiments of the present disclosure, for any original document, the preprocessing module 404 may also obtain target key information corresponding to the first document information of the original document, and may determine the target key information as the second document information of the original document.
[0090] Specifically, in some embodiments of the present disclosure, in response to determining that the first document information only includes the first picture content information, the preprocessing module 404 may extract the first key information from the first picture content information, and may determine the first key information as the target key information; in response to determining that the first document information only includes the first text information, the preprocessing module 404 may extract the second key information from the first text information, and may determine the second key information as the target key information; in response to determining that the first document information includes both the first text information and the first picture content information, the extracted first key information and the second key information may be determined as the target key information.
[0091] In some embodiments of the present disclosure, when the preprocessing module 404 extracts the second key information from the first text information, the first text information can be sliced to obtain text fragments, and fragment key information can be extracted from each text fragment respectively, and the fragment key information can be spliced in the order of the appearance position of the corresponding text fragment in the first text information from first to last to obtain the second key information, or the fragment key information can be input into the information fusion model to obtain the output second key information.
[0092] After obtaining the corresponding first document information and second document information for each original document, the preprocessing module 404 may store the first document information into the first database and the second document information into the second database.
[0093] In some embodiments of the present disclosure, the preprocessing module 404 can perform word segmentation on the first document information, store each word segmentation result in the first database, and obtain a first vector representation corresponding to the first document information, store the first vector representation in the first database, and can perform word segmentation on the second document information, store each word segmentation result in the second database, and can obtain a second vector representation corresponding to the second document information, and store the second vector representation in the second database.
[0094] After that, online retrieval can be performed based on the processing results of the preprocessing process. For example, the request processing module 401 can obtain the retrieval request input by the user, the retrieval request includes the second text information and / or the second image, and can determine the query content according to the retrieval request, and then the document retrieval module 402 can search the target database according to the query content, and then the result generation module 403 can determine the target document as the target retrieval result from the retrieved original documents, and can return the target document.
[0095] In some embodiments of the present disclosure, in response to determining that the retrieval request includes only the second text information, the request processing module 401 may determine the second text information as the query content; in response to determining that the retrieval request includes only the second picture, the second picture content information may be extracted from the second picture, and the second picture content information may be determined as the query content; in response to determining that the retrieval request includes both the second text information and the second picture, the second text information and the extracted second picture content information may be spliced in a predetermined order, and the splicing result may be determined as the query content, or the second text information and the second picture content information may be input into an information fusion model to obtain output query content.
[0096] In addition, in some embodiments of the present disclosure, the request processing module 401 may utilize the visual language model to extract the second picture content information from the second picture, wherein the visual language model may utilize the same prompt words when extracting the first picture content information and the second picture content information.
[0097] After the query content is determined, the target database can be searched. In some embodiments of the present disclosure, the document retrieval module 402 can use M retrieval methods to search the first database according to the query content, and obtain initial retrieval results corresponding to each retrieval method, where M is a positive integer greater than 1, and can use the M retrieval methods to search the second database according to the query content, and obtain initial retrieval results corresponding to each retrieval method.
[0098] Further, the result generation module 403 may determine the target document as the target search result from the retrieved original documents. In some embodiments of the present disclosure, the result generation module 403 may determine the candidate documents from the original documents in each initial search result, and may determine the comprehensive scores of the candidate documents, and then may sort the candidate documents in descending order of the comprehensive scores, and determine the candidate documents that are in the top N after sorting as the target documents, where N is a positive integer, and N is less than or equal to the number of candidate documents.
[0099] In some embodiments of the present disclosure, the result generation module 403 can respectively obtain the comprehensive scores of the original documents in each initial retrieval result, and can use the original documents in each initial retrieval result to form a first document set, and then determine the first document set as a target document set, and perform the following first processing: determine any original document in the target document set as a document to be processed, and determine the original documents in the target document set other than the document to be processed as selected documents, in response to determining that there is no predetermined document in the selected documents, the predetermined document is a selected document that is the same document as the document to be processed, determine the document to be processed as a candidate document, determine the comprehensive score of the document to be processed as the comprehensive score of the candidate document, and delete the document to be processed from the first document set; in response to determining that there is a predetermined document in the selected documents, determine the document to be processed as a candidate document, obtain the sum of the comprehensive score of the document to be processed and the comprehensive score of the predetermined document, determine the sum as the comprehensive score of the candidate document, and delete the document to be processed and the predetermined document from the first document set; in response to determining that the updated first document set is not empty, determine the updated first document set as the target document set, and repeat the first processing.
[0100] In some embodiments of the present disclosure, the method in which the result generation module 403 obtains the comprehensive score of the original document in each initial search result may include: for any original document in each initial search result, performing the following processing respectively: obtaining the retrieval score of the original document corresponding to 2M retrieval operations respectively, the 2M retrieval operations including: retrieval operations performed on the first database using M retrieval methods and retrieval operations performed on the second database using M retrieval methods, and determining the sum of the 2M retrieval scores as the comprehensive score of the original document.
[0101] In addition, in some embodiments of the present disclosure, for any retrieval operation, the way in which the result generation module 403 obtains the retrieval score of a certain original document corresponding to the retrieval operation may include: in response to determining that the original document is included in the initial retrieval result corresponding to the retrieval operation, and determining that the ranking position of the original document is less than a predetermined threshold, obtaining the sum of the ranking position and K, where K is a positive integer, and obtaining the ratio of the weight corresponding to the retrieval operation to the sum, determining the ratio as the retrieval score, the ranking position being the ranking position of the original document in the initial retrieval result corresponding to the retrieval operation, and in response to determining that the original document is not included in the initial retrieval result corresponding to the retrieval operation, or determining that the ranking position of the original document is greater than or equal to the predetermined threshold, determining 0 as the retrieval score.
[0102] Figure 4 and Figure 5 The specific working process of the device embodiment shown can refer to the relevant description in the aforementioned method embodiment, which will not be repeated here.
[0103] The scheme disclosed in the present invention can be applied to the field of artificial intelligence, especially to the fields of natural language processing, computer vision, deep learning and large models. Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It has both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0104] In addition, the original documents and search requests in the embodiments of the present disclosure are not for a specific user and do not reflect the personal information of a specific user. Moreover, the execution subject of the method described in the present disclosure can obtain the search request after obtaining the user's authorization. In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0105] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0106] Figure 6 A schematic block diagram of an electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0107] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 to a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0108] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the electronic device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0109] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the methods described in the present disclosure may be executed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).
[0110] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0111] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0112] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Electronically Programmable Read-Only Memory), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM, Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0113] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0114] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication (e.g., a communication network) in any form or medium. Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0115] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0116] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0117] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A document retrieval method, comprising: Obtaining a search request, wherein the search request includes second text information and / or a second image, and determining query content according to the search request; Searching a target database according to the query content, wherein the target database stores predetermined document information of each original document as a search object, wherein the original document includes first text information and / or a first picture; A target document as a target search result is determined from the retrieved original documents, and the target document is returned.
2. The method according to claim 1, wherein: The predetermined document information includes: first document information and second document information; The target database includes: a first database and a second database; The method further includes: before obtaining the retrieval request, respectively obtaining the first document information and the second document information of each original document, storing the first document information in the first database, and storing the second document information in the second database.
3. The method according to claim 2, wherein: The acquiring the first document information of the original document includes: In response to determining that the original document includes only the first text information, determining the first text information as the first document information; In response to determining that the original document only includes the first picture or determining that the original document includes both the first text information and the first picture, first picture content information is extracted from the first picture, and the first picture is replaced with the corresponding first picture content information in the original document to obtain the first document information.
4. The method according to claim 3, wherein: Acquiring the second document information of the original document includes: Target key information corresponding to the first document information is acquired, and the target key information is determined as the second document information.
5. The method according to claim 4, wherein: The step of obtaining target key information corresponding to the first document information includes: In response to determining that the first document information only includes the first picture content information, extracting first key information from the first picture content information, and determining the first key information as the target key information; In response to determining that the first document information only includes the first text information, extracting second key information from the first text information, and determining the second key information as the target key information; In response to determining that the first document information includes both the first text information and the first image content information, the extracted first key information and the second key information are determined as the target key information.
6. The method according to claim 5, wherein: The extracting the second key information from the first text information comprises: Slicing the first text information to obtain text segments; Extract key information from each text segment respectively; The key information of each fragment is spliced in the order of the appearance position of the corresponding text fragment in the first text information from first to last to obtain the second key information, or the key information of each fragment is input into the information fusion model to obtain the output second key information.
7. The method according to claim 4, wherein: The storing the first document information into the first database includes: performing word segmentation processing on the first document information, storing each word segmentation result into the first database, and obtaining a first vector representation corresponding to the first document information, and storing the first vector representation into the first database; Storing the second document information in the second database includes: performing word segmentation processing on the second document information, storing each word segmentation result in the second database, obtaining a second vector representation corresponding to the second document information, and storing the second vector representation in the second database.
8. The method according to claim 3, wherein: Determining the query content according to the search request includes: In response to determining that the search request only includes the second text information, determining the second text information as the query content; In response to determining that the search request only includes the second image, extracting second image content information from the second image, and determining the second image content information as the query content; In response to determining that the retrieval request includes both the second text information and the second image, the second text information and the extracted second image content information are spliced in a predetermined order, and the splicing result is determined as the query content, or the second text information and the second image content information are input into an information fusion model to obtain the output query content.
9. The method according to claim 8, wherein: The extracting the first picture content information from the first picture includes: extracting the first picture content information from the first picture using a visual language large model; The extracting the second picture content information from the second picture comprises: extracting the second picture content information from the second picture using the visual language large model; Wherein, when extracting the first picture content information and the second picture content information, the visual language large model uses the same prompt words.
10. The method according to any one of claims 1 to 9, wherein: The target database includes: a first database and a second database; The searching the target database according to the query content includes: According to the query content, the first database is searched using M search methods to obtain initial search results corresponding to each search method, where M is a positive integer greater than 1; And, according to the query content, the second database is searched respectively using the M search methods to obtain initial search results corresponding to each search method.
11. The method according to claim 10, wherein: Determining a target document as a target search result from the retrieved original document includes: Determine candidate documents from the original documents in each initial search result, and determine a comprehensive score of the candidate documents; The candidate documents are sorted in descending order of the comprehensive scores, and the candidate documents in the top N after sorting are determined as the target documents, where N is a positive integer and is less than or equal to the number of the candidate documents.
12. The method according to claim 11, wherein: The step of determining candidate documents from the original documents in each initial search result and determining a comprehensive score of the candidate documents includes: Obtaining comprehensive scores of original documents in each initial search result respectively, and using the original documents in each initial search result to form a first document set; The first document set is determined as a target document set, and the following first processing is performed: any original document in the target document set is determined as a document to be processed, and original documents in the target document set other than the document to be processed are determined as selected documents; in response to determining that a predetermined document does not exist in the selected documents, the predetermined document is a selected document that is the same as the document to be processed, the document to be processed is determined as a candidate document, the comprehensive score of the document to be processed is determined as the comprehensive score of the candidate document, and the document to be processed is deleted from the first document set; in response to determining that the predetermined document exists in the selected documents, the document to be processed is determined as a candidate document, the sum of the comprehensive score of the document to be processed and the comprehensive score of the predetermined document is obtained, the sum is determined as the comprehensive score of the candidate document, and the document to be processed and the predetermined document are deleted from the first document set; in response to determining that the updated first document set is not empty, the updated first document set is determined as the target document set, and the first processing is repeatedly performed.
13. The method according to claim 12, wherein: The step of obtaining the comprehensive scores of the original documents in each initial search result comprises: For any original document in each initial search result, the following processing is performed respectively: Respectively obtaining search scores of the original document corresponding to 2M search operations, the 2M search operations including: a search operation on the first database using M search methods and a search operation on the second database using the M search methods; The sum of the 2M search scores is determined as the comprehensive score of the original document.
14. The method according to claim 13, wherein: For any retrieval operation, obtaining the retrieval score of the original document corresponding to the retrieval operation includes: In response to determining that the original document is included in the initial search result corresponding to the search operation, and determining that the ranking position of the original document is less than a predetermined threshold, obtaining a sum of the ranking position and K, where K is a positive integer, and obtaining a ratio of a weight corresponding to the search operation to the sum, determining the ratio as the search score, and the ranking position is the ranking position of the original document in the initial search result corresponding to the search operation; In response to determining that the original document is not included in the initial search results corresponding to the search operation, or determining that the ranking position of the original document is greater than or equal to the predetermined threshold, 0 is determined as the search score.
15. A document retrieval device, comprising: Request processing module, document retrieval module and result generation module; The request processing module is used to obtain a search request, wherein the search request includes the second text information and / or the second image, and determine the query content according to the search request; The document retrieval module is used to search a target database according to the query content, wherein the target database stores predetermined document information of each original document as a search object, wherein the original document includes first text information and / or a first image; The result generation module is used to determine a target document as a target search result from the retrieved original documents, and return the target document.
16. The device according to claim 15, wherein: The predetermined document information includes: first document information and second document information; The target database includes: a first database and a second database; The device also includes: a preprocessing module, which is used to obtain the first document information and the second document information of each original document, respectively, and store the first document information in the first database and store the second document information in the second database.
17. The device according to claim 16, wherein: In response to determining that the original document only includes the first text information, the preprocessing module determines the first text information as the first document information; in response to determining that the original document only includes the first picture or determining that the original document includes both the first text information and the first picture, the preprocessing module extracts first picture content information from the first picture, and replaces the first picture in the original document with the corresponding first picture content information to obtain the first document information.
18. The device according to claim 17, wherein: The preprocessing module obtains target key information corresponding to the first document information, and determines the target key information as the second document information.
19. The device according to claim 18, wherein: In response to determining that the first document information only includes the first image content information, the preprocessing module extracts first key information from the first image content information and determines the first key information as the target key information; in response to determining that the first document information only includes the first text information, the preprocessing module extracts second key information from the first text information and determines the second key information as the target key information; in response to determining that the first document information includes both the first text information and the first image content information, the extracted first key information and the second key information are determined as the target key information.
20. The device according to claim 19, wherein The preprocessing module slices the first text information to obtain text segments, extracts segment key information from each text segment, and splices the segment key information in the order of the appearance positions of the corresponding text segments in the first text information from first to last to obtain the second key information, or inputs the segment key information into an information fusion model to obtain the output second key information.
21. The device according to claim 18, wherein The preprocessing module performs word segmentation on the first document information, stores each word segmentation result in the first database, obtains a first vector representation corresponding to the first document information, stores the first vector representation in the first database, and performs word segmentation on the second document information, stores each word segmentation result in the second database, obtains a second vector representation corresponding to the second document information, and stores the second vector representation in the second database.
22. The device according to claim 17, wherein: In response to determining that the retrieval request only includes the second text information, the request processing module determines the second text information as the query content; in response to determining that the retrieval request only includes the second image, the request processing module extracts second image content information from the second image and determines the second image content information as the query content; in response to determining that the retrieval request includes both the second text information and the second image, the request processing module splices the second text information and the extracted second image content information in a predetermined order and determines the splicing result as the query content, or inputs the second text information and the second image content information into an information fusion model to obtain the output query content.
23. The device according to claim 22, wherein: The preprocessing module extracts the first picture content information from the first picture using a visual language large model; The request processing module extracts the second picture content information from the second picture using the visual language large model; Wherein, when extracting the first picture content information and the second picture content information, the visual language large model uses the same prompt words.
24. The device according to any one of claims 15 to 23, wherein: The target database includes: a first database and a second database; The document retrieval module uses M retrieval methods to search the first database according to the query content, and obtains initial retrieval results corresponding to each retrieval method, where M is a positive integer greater than 1; and, based on the query content, uses the M retrieval methods to search the second database, and obtains initial retrieval results corresponding to each retrieval method.
25. The device according to claim 24, wherein: The result generation module determines candidate documents from the original documents in each initial search result, determines the comprehensive score of the candidate documents, sorts the candidate documents in descending order of the comprehensive score, and determines the candidate documents that are in the top N after sorting as the target documents, where N is a positive integer and is less than or equal to the number of the candidate documents.
26. The device according to claim 25, wherein The result generation module obtains the comprehensive scores of the original documents in each initial search result respectively, and uses the original documents in each initial search result to form a first document set, determines the first document set as a target document set, and performs the following first processing: determines any original document in the target document set as a document to be processed, and determines the original documents in the target document set other than the document to be processed as selected documents, in response to determining that there is no predetermined document in the selected documents, the predetermined document is a selected document that is the same as the document to be processed, determines the document to be processed as a candidate document, determines the comprehensive score of the document to be processed as the comprehensive score of the candidate document, and deletes the document to be processed from the first document set; In response to determining that the predetermined document exists in the documents to be selected, determining the document to be processed as a candidate document, obtaining a sum of a comprehensive score of the document to be processed and a comprehensive score of the predetermined document, determining the sum as the comprehensive score of the candidate document, and deleting the document to be processed and the predetermined document from the first document set; In response to determining that the updated first document set is not empty, the updated first document set is determined as the target document set, and the first process is repeatedly performed.
27. The device according to claim 26, wherein: The result generation module obtains, for any original document in each initial search result, the retrieval score of the original document corresponding to 2M retrieval operations, wherein the 2M retrieval operations include: retrieval operations on the first database using M retrieval methods and retrieval operations on the second database using the M retrieval methods, and the sum of the 2M retrieval scores is determined as the comprehensive score of the original document.
28. The device according to claim 27, wherein For any retrieval operation, the result generation module obtains the retrieval score of the original document corresponding to the retrieval operation, including: in response to determining that the original document is included in the initial retrieval result corresponding to the retrieval operation, and determining that the ranking position of the original document is less than a predetermined threshold, obtaining the sum of the ranking position and K, K is a positive integer, and obtaining the ratio of the weight corresponding to the retrieval operation to the sum, determining the ratio as the retrieval score, the ranking position is the ranking position of the original document in the initial retrieval result corresponding to the retrieval operation, and in response to determining that the original document is not included in the initial retrieval result corresponding to the retrieval operation, or determining that the ranking position of the original document is greater than or equal to the predetermined threshold, determining 0 as the retrieval score.
29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.
30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 14.
31. A computer program product, comprising a computer program / instructions, which implement the method according to any one of claims 1 to 14 when executed by a processor.