Scientific and technological intelligence information retrieval result reordering method and equipment
By vectorizing the query problem and searching in multiple vector libraries, the comprehensive factor is calculated in combination with word frequency-inverse document rate factor, date factor and searched frequency factor, and the searched frequency factor are reordered, and the searched search results are solved, which is the problem of inaccurate search, inaccurate correlation sorting and priority of the latest information sorting, and more accurate and relevant search results are achieved.
Patent Information
- Application Number
- CN202510080640.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-13
AI Technical Summary
There are problems in multi-vector library searches that inaccurate retrieval, inaccurate relevance sorting and prioritize the latest information sorting.
By vectorizing the text of the query problem, searching in each vector library of the big model based on the query vector to obtain the query result set of each vector library. Then, each query result set is preprocessed, the word frequency-inverse document rate factor, date factor and searched frequency factor of each to-determined query result are calculated, the comprehensive factor is determined based on the vectorized correlation coefficient, and the to-determined query result is arranged in descending order, and the preset number of to-determined query results are finally determined as the target query result.
It solves the problems of inaccurate search, inaccurate relevance sorting, and prioritization of latest information sorting in multi-vector library search, improves the accuracy and relevance of search results, and ensures priority of latest information.
Smart Images

Figure CN119988585A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information retrieval technology, and in particular to a method and device for reordering scientific and technological intelligence information retrieval results. Background Art
[0002] The content of scientific and technological intelligence is required to be both timely and accurate. At present, there are two main problems in the application of big model generation technology to the field of scientific and technological intelligence: First, big model knowledge relies on pre-trained corpora: because the pre-training of big models is time-consuming, computationally intensive and power-consuming, and the maintenance cost is high, the existing big model pre-training is generally spaced half a year or longer; as a result, the knowledge acquired through training can only be up to the time of the most recent pre-training, so the timeliness of scientific and technological intelligence content cannot be guaranteed. Second, the retrieval is based on limited information: the big model generates text based on user questions, and the information contained in the user's questions is limited. In order to maintain the accuracy of the retrieval results, it is necessary to set relevant information for the big model to assist in the generation of retrieval results. Based on the above two reasons, the big model service of scientific and technological intelligence generally adopts the big model + information enhancement mode to provide question-and-answer services.
[0003] Information retrieval enhancement takes semantic search and speed into consideration. Generally, intelligence-related data is converted into vectors through a vectorization model, and stored and retrieved in a vectorized manner. Because scientific and technological intelligence involves a wide range of fields and the correlation between fields is small, a multi-vector library is generally used. At present, there are the following problems with multi-vector library retrieval: 1) Vectorized retrieval takes efficiency into consideration. Generally, the inner product between the problem vector and the text block vector is used as a relevance measure. The relevance does not involve the date of the text block information. 2) Multi-vector library retrieval generally searches in each vector library, which leads to the relevance of the retrieval results in one library. The relevance can be measured with the retrieval problem, but the relevance of the retrieval results across libraries cannot be measured with the retrieval problem. 3) Vectorized retrieval determines the retrieval result by calculating the distance between two vectors; but sometimes, even if the query and text block vectors are close in space, their meanings may not match; for example, "Shandong University of Science and Technology" and "Beijing University of Science and Technology" are close in space, but they are essentially two completely different results. Summary of the invention
[0004] The purpose of this application is to provide a method and device for re-ranking scientific and technological intelligence information retrieval results to solve the problems of inaccurate retrieval, inaccurate relevance ranking, and the need to prioritize the latest information in multi-vector library retrieval.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a method for re-ranking scientific and technological intelligence information retrieval results, comprising:
[0007] Vectorize the query text to obtain the query vector;
[0008] Based on the query vector, a search is performed in each vector library of the large model to obtain a query result set corresponding to each vector library; any query result in any query result set includes a text block identifier, text block text content, a corresponding document, a corresponding document generation date, a vectorization correlation coefficient, and a number of times the corresponding document has been retrieved;
[0009] Preprocessing each of the query result sets respectively to obtain multiple optimized result sets; the optimized result sets include multiple pending query results;
[0010] Determine the term frequency-inverse document rate factor for each pending query result;
[0011] Determine the date factor for each pending query result;
[0012] Determine the retrieval frequency factor of the document to which each pending query result belongs;
[0013] Determine the comprehensive factor corresponding to the pending query result according to the vectorized correlation coefficient, the word frequency-reverse document rate factor, the date factor and the retrieval frequency factor of the corresponding document corresponding to the same pending query result;
[0014] Arrange all pending query results in descending order according to the comprehensive factor;
[0015] A preset number of pending query results are determined to be target query results of the query question.
[0016] Optionally, the large model is used for scientific and technological information retrieval;
[0017] The large model includes a plurality of vector libraries;
[0018] Different vector libraries correspond to different fields;
[0019] Any of the vector libraries are built using multiple documents in the corresponding field;
[0020] Any of the vector libraries includes a summary vector library and a plurality of text block vector libraries;
[0021] The summary vectors in the summary vector library correspond one-to-one to the text block vector library;
[0022] The summary vector and the corresponding text block vector library are constructed using the same document;
[0023] The summary vector is obtained by vectorizing the summary text of the document;
[0024] The text block vector library is constructed by dividing the main text of the document into multiple text blocks and performing vectorization processing on each text block to obtain multiple text block vectors.
[0025] Optionally, preprocess each of the query result sets to obtain multiple optimized result sets, including:
[0026] Determine the proportion of hyperlinks for each query result;
[0027] Delete the query results in each query result set whose hyperlink ratio is greater than the hyperlink ratio threshold;
[0028] Determine any query result set as the current query result set;
[0029] Determine any query result in the current query result set as the first query result;
[0030] Determine all query results except the first query result in the current query result set as query results to be matched;
[0031] Determine any to-be-matched query result as a second query result;
[0032] Determine a coincidence determination result between the first query result and the second query result;
[0033] Update the second query result, and return to step "determine the coincidence determination result of the first query result and the second query result" until all query results to be matched are traversed;
[0034] Update the first query result, and return to the step of "determining all query results other than the first query result in the current query result set as query results to be matched" until all query results in the current query result set are traversed to obtain the current optimized result set;
[0035] The current query result set is updated, and the process returns to the step of "determining any query result in the current query result set as the first query result" until all query result sets are traversed to obtain multiple optimized result sets.
[0036] Optionally, determining a result of a coincidence between the first query result and the second query result includes:
[0037] Determine whether the document to which the first query result belongs is the same as the document to which the second query result belongs, and obtain a first determination result;
[0038] If the first judgment result is no, determining that the first query result and the second query result do not overlap;
[0039] If the first judgment result is yes, determining the edit distance similarity between the text block corresponding to the first query result and the text block corresponding to the second query result;
[0040] Determine whether the edit distance similarity is greater than the edit distance similarity threshold, and obtain a second determination result;
[0041] If the second judgment result is no, determining that the first query result and the second query result do not overlap;
[0042] If the second judgment result is yes, then determining that the first query result and the second query result overlap, and determining whether the vectorized correlation coefficient corresponding to the first query result and the vectorized correlation coefficient corresponding to the second query result are equal, to obtain a third judgment result;
[0043] If the third judgment result is yes, the first query result is deleted and the second query result is used as the updated first query result; or the second query result is deleted and the first query result is used as the updated first query result;
[0044] If the third judgment result is no, then it is determined whether the vectorized correlation coefficient corresponding to the first query result is greater than the vectorized correlation coefficient corresponding to the second query result, and a fourth judgment result is obtained;
[0045] If the fourth judgment result is yes, deleting the second query result and using the first query result as the updated first query result;
[0046] If the fourth judgment result is no, the first query result is deleted and the second query result is used as the updated first query result.
[0047] Optionally, the hyperlink ratio is:
[0048]
[0049] Among them, hyper_weight is the proportion of hyperlinks;
[0050] The edit distance similarity is:
[0051] similarity=1-distance / max(len(text1),len(text2));
[0052] Wherein, similarity is the edit distance similarity, distance is the edit distance between the text block text1 corresponding to the first query result and the text block text2 corresponding to the second query result, len(text1) is the length of the text block text1, and len(text2) is the length of the text block text2.
[0053] Optionally, the term frequency-inverse document rate factor is:
[0054]
[0055] in,
[0056] In the formula, r_TF-IDF is the term frequency-inverse document rate factor, n is the number of words in the text block after word segmentation, TF-IDF i is the term frequency-inverse document rate factor of the i-th word, freq i is the number of times the i-th word appears in this text block, df(i) is the number of text blocks where the i-th word appears in rule result1, and N is the number of text blocks in rule result1.
[0057] Optionally, the date factor is:
[0058]
[0059] Among them, r date is the date factor.
[0060] Optionally, the retrieved frequency factor is:
[0061]
[0062] in, is the search frequency factor.
[0063] Optionally, the comprehensive factor is:
[0064]
[0065] Among them, x+y+z+w=100%;
[0066] Among them, r_total is the comprehensive factor, r_vector is the vectorized correlation coefficient, x is the vectorized correlation coefficient weight, y is the word frequency-inverse document rate factor weight, z is the date factor weight, and w is the retrieval frequency factor weight.
[0067] In a second aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for reordering scientific and technological intelligence information retrieval results.
[0068] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0069] The present application provides a method and device for reordering the results of scientific and technological intelligence information retrieval, which vectorizes the text of a query question to obtain a query vector; searches each vector library of a large model based on the query vector to obtain a query result set corresponding to each vector library; any query result in any query result set includes a text block identifier, text block text content, a corresponding document, a date of generation of the corresponding document, a vectorized correlation coefficient, and the number of times the corresponding document is retrieved; preprocesses each query result set to obtain multiple optimized result sets; the optimized result set includes multiple pending query results; determines the word frequency-reverse document rate factor, date factor, and the frequency factor of the retrieved document of each pending query result, and determines the comprehensive factor of each pending query result in combination with the vectorized correlation; arranges all pending query results in descending order according to the comprehensive factor; determines a preset number of pending query results as the target query results of the query question; the present application completes the reordering of scientific and technological intelligence information retrieval results by calculating the comprehensive factor of the query result, thereby solving the problems of inaccurate retrieval, inaccurate relevance sorting, and the need to prioritize the latest information sorting in multi-vector library retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0071] Figure 1 This is a flow chart of a method for reordering scientific and technological intelligence information retrieval results in one embodiment of the present application. DETAILED DESCRIPTION
[0072] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0073] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0074] In an exemplary embodiment, Figure 1 As shown, a method for re-ranking scientific and technological intelligence information retrieval results is provided, comprising:
[0075] Step 101: vectorize the text of the query question to obtain a query vector.
[0076] Step 102: Based on the query vector, search each vector library of the large model to obtain the query result set corresponding to each vector library. Any query result in any query result set includes a text block identifier, text block text content, the document to which it belongs, the date when the document to which it belongs was generated, the vectorization correlation coefficient, and the number of times the document to which it belongs was retrieved. The large model is used for scientific and technological information retrieval. The large model includes multiple vector libraries. Different vector libraries correspond to different fields. Any vector library is built using multiple documents in the corresponding field. Any vector library includes an abstract vector library and multiple text block vector libraries. The abstract vector in the abstract vector library corresponds to the text block vector library one by one. The abstract vector and the corresponding text block vector library are built using the same document. The abstract vector is obtained by vectorizing the abstract text of the document. The text block vector library is built by dividing the main text of the document into multiple text blocks and vectorizing each text block to obtain multiple text block vectors.
[0077] In the construction of the vector library, a document is divided into several text blocks according to the specified length, and the text blocks are vectorized and stored, so the retrieval results are also text blocks. Perform multi-library vectorized retrieval: For the query question (text query), the query is vectorized to form a query vector (query_vector) through the same vectorization model as the vector library construction, and then searched in the 18 vector libraries of the science and technology intelligence model. Each library mainly consists of two files, one is composed of abstracts and the other is composed of document blocks. Then the search is carried out in two steps. First, the relevant documents are filtered out through the abstract, and then only searched in this relevant group. According to the pre-set vector library retrieval result return number TOPK and knowledge base matching relevance threshold SCORE_THRESHOLD (knowledge base matching relevance threshold, the value range is between 0-1, the smaller the SCORE, the higher the relevance, and 1 is equivalent to no screening), search is performed in each library. Eighteen query result sets are returned through the query resultlt0. Each piece of data includes the text block identifier (id), the text content of the text block (text), the document to which it belongs (documentid), the date the document was generated (d_date), the number of times the document was retrieved, and the correlation coefficient (id, text, documentid, d_date, d_frequency, r_vector).
[0078] Step 103: pre-process each query result set respectively to obtain multiple optimized result sets. The optimized result sets include multiple pending query results.
[0079] Step 103 includes:
[0080] Step 103-1: Determine the hyperlink ratio of each query result. The hyperlink ratio is: Among them, hyper_weight is the proportion of hyperlinks.
[0081] Step 103 - 2: Delete the query results in each query result set whose hyperlink ratio is greater than the hyperlink ratio threshold.
[0082] Step 103 - 3: Determine any query result set as the current query result set.
[0083] Step 103 - 4 : Determine any query result in the current query result set as the first query result.
[0084] Step 103 - 5 : Determine that all query results other than the first query result in the current query result set are query results to be matched.
[0085] Step 103 - 6 : Determine that any query result to be matched is the second query result.
[0086] Step 103 - 7 : Determine the coincidence judgment result of the first query result and the second query result.
[0087] Step 103-8: Update the second query result, and return to step 103-7 until all query results to be matched are traversed.
[0088] Step 103-9: Update the first query result, and return to step 103-5 until all query results in the current query result set are traversed to obtain the current optimized result set.
[0089] Step 103-10: Update the current query result set, and return to step 103-4 until all query result sets are traversed to obtain multiple optimized result sets.
[0090] Step 103-7 includes:
[0091] Step 103-7-1: Determine whether the document to which the first query result belongs is the same as the document to which the second query result belongs, and obtain a first determination result.
[0092] Step 103-7-2: If the first judgment result is no, determine that the first query result and the second query result do not overlap.
[0093] Step 103-7-3: If the first judgment result is yes, determine the edit distance similarity between the text block corresponding to the first query result and the text block corresponding to the second query result. The edit distance similarity is:
[0094] similarity=1-distance / max(len(text1), len(text2)).
[0095] Wherein, similarity is the edit distance similarity, distance is the edit distance between the text block text1 corresponding to the first query result and the text block text2 corresponding to the second query result, len(text1) is the length of the text block text1, and len(text2) is the length of the text block text2.
[0096] Step 103-7-4: Determine whether the edit distance similarity is greater than the edit distance similarity threshold, and obtain a second determination result.
[0097] Step 103-7-5: If the second judgment result is no, it is determined that the first query result and the second query result do not overlap.
[0098] Step 103-7-6: If the second judgment result is yes, determine whether the first query result and the second query result overlap, and determine whether the vectorized correlation coefficient corresponding to the first query result is equal to the vectorized correlation coefficient corresponding to the second query result, to obtain a third judgment result.
[0099] Step 103-7-7: If the third judgment result is yes, the first query result is deleted and the second query result is used as the updated first query result, or the second query result is deleted and the first query result is used as the updated first query result.
[0100] Step 103-7-8: If the third judgment result is no, determine whether the vectorized correlation coefficient corresponding to the first query result is greater than the vectorized correlation coefficient corresponding to the second query result, and obtain a fourth judgment result.
[0101] Step 103-7-9: If the fourth judgment result is yes, the second query result is deleted and the first query result is used as the updated first query result.
[0102] Step 103-7-10: If the fourth judgment result is no, the first query result is deleted and the second query result is used as the updated first query result.
[0103] De-duplicate the query results: Because a document may belong to more than one field, the documents in the eighteen vector libraries may be repeated. Therefore, first deduplicate the query result text blocks. Compare the query result text blocks in pairs. First, compare the documentid of the text blocks. If they are the same, it means that they are from the given document. Then compare the edit distance similarity of the two text blocks (text1, text2). If the edit distance similarity is > 0.9, they are considered to have a high degree of duplication. According to the r_vector of the two text blocks, the smaller one is kept and the larger one is removed from the result data set. If the r_vector is the same, one of them is selected to be kept. Through this step, the de-duplicated result data set result1 is generated.
[0104] According to the characteristics of the query result data, remove the blocks with little value to the question-answering task. According to the analysis of multiple query results, short texts and text blocks containing hyperlinks have smaller correlation coefficients in vectorized retrieval than other text blocks. However, if the text is too segmented and the hyperlinks are too long, they are not very meaningful as reference information for large model question-answering. Therefore, set the text block text length threshold LEN_THRESHOLD and the hyper_weigh threshold HYPER_THRESHOLD of the proportion of hyperlinks in the text block. If the text length of the text block is less than LEN_THRESHOLD, it is removed from the result data set result1. If the hyper_weigh calculated for the text block is greater than HYTX_THRESHOLD, the text block is removed from the result data set. Through this step, the result data set (optimized result set) is represented as result2.
[0105] Step 104: Determine the term frequency-reverse document rate factor of each pending query result. The term frequency-reverse document rate factor is:
[0106]
[0107] in,
[0108] In the formula, r_TF-IDF is the term frequency-inverse document rate factor, n is the number of words in the text block after word segmentation, TF-IDF i is the term frequency-inverse document rate factor of the i-th word, freq i is the number of times the i-th word appears in this text block, df(i) is the number of text blocks where the i-th word appears in rule result1, and N is the number of text blocks in rule result1.
[0109] TF (Term Frequency) is the frequency of the search term appearing in a result text block, which is the number of search terms divided by the total number of words in the result text block. The higher the frequency, the higher the score. DF (Document Frequency) is the frequency of the search term appearing in all text blocks of the search result 2. IDF (Inverse Document Frequency) is the number of text blocks divided by the number of text blocks in which the search term appears. The higher the frequency, the lower the score.
[0110] Step 105: Determine the date factor of each pending query result. The date factor is:
[0111]
[0112] Among them, r date is the date factor. The date factor is used to modify the relevance of the search results to ensure that documents with recent dates are ranked at the top.
[0113] Step 106: Determine the retrieval frequency factor of the document to which each pending query result belongs. If the data block where the document is located is retrieved and eventually used as the relevant information of the large model, it is counted as the document being retrieved once, and the total number of times the document has been retrieved is recorded in the information table of the vector database. The retrieval frequency factor is:
[0114]
[0115] in, is the search frequency factor.
[0116] Step 107: Determine the comprehensive factor corresponding to the pending query result according to the vectorized correlation coefficient, word frequency-reverse document rate factor, date factor and the retrieved frequency factor of the corresponding document corresponding to the same pending query result. The comprehensive factor is:
[0117]
[0118] Among them, x+y+z+w=100%.
[0119] Among them, r_total is the comprehensive factor, r_vector is the vectorized correlation coefficient, r_TF-IDF is the term frequency-inverse document rate factor, x is the vectorized correlation coefficient weight, y is the term frequency-inverse document rate factor weight, z is the date factor weight, and w is the retrieved frequency factor weight.
[0120] Step 108: Arrange all pending query results in descending order according to the comprehensive factors.
[0121] Step 109: Determine the first preset number of pending query results as the target query results of the query question. Sort the search results in result2 according to the relevance score, and take the top TOPK as the return set of this search. Use the relevant information and the user question query to give to the big model for this question and answer.
[0122] In an exemplary embodiment, a computer device is provided, which may be a server or a terminal. The computer device includes a processor, a memory, an input / output interface (I / O for short) and a communication interface. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for reordering scientific and technological intelligence information retrieval results is implemented.
[0123] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0124] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0125] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0126] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0127] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0128] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0129] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for re-ranking scientific and technological intelligence information retrieval results, characterized in that: include: Vectorize the query text to obtain the query vector; Based on the query vector, a search is performed in each vector library of the large model to obtain a query result set corresponding to each vector library; any query result in any query result set includes a text block identifier, text block text content, a corresponding document, a corresponding document generation date, a vectorization correlation coefficient, and a number of times the corresponding document has been retrieved; Preprocessing each of the query result sets respectively to obtain multiple optimized result sets; the optimized result sets include multiple pending query results; Determine the term frequency-inverse document rate factor for each pending query result; Determine the date factor for each pending query result; Determine the retrieval frequency factor of the document to which each pending query result belongs; Determine the comprehensive factor corresponding to the pending query result according to the vectorized correlation coefficient, the word frequency-reverse document rate factor, the date factor and the retrieval frequency factor of the corresponding document corresponding to the same pending query result; Arrange all pending query results in descending order according to the comprehensive factor; A preset number of pending query results are determined to be target query results of the query question.
2. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 1, characterized in that: The large model is used for scientific and technological information retrieval; The large model includes a plurality of vector libraries; Different vector libraries correspond to different fields; Any of the vector libraries are built using multiple documents in the corresponding field; Any of the vector libraries includes a summary vector library and a plurality of text block vector libraries; The summary vectors in the summary vector library correspond one-to-one to the text block vector library; The summary vector and the corresponding text block vector library are constructed using the same document; The summary vector is obtained by vectorizing the summary text of the document; The text block vector library is constructed by dividing the main text of the document into multiple text blocks and performing vectorization processing on each text block to obtain multiple text block vectors.
3. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 1, characterized in that: Preprocess each of the query result sets respectively to obtain multiple optimized result sets, including: Determine the proportion of hyperlinks for each query result; Delete the query results in each query result set whose hyperlink ratio is greater than the hyperlink ratio threshold; Determine any query result set as the current query result set; Determine any query result in the current query result set as the first query result; Determine all query results except the first query result in the current query result set as query results to be matched; Determine any to-be-matched query result as the second query result; Determine a coincidence determination result between the first query result and the second query result; Update the second query result, and return to step "determine the coincidence determination result of the first query result and the second query result" until all query results to be matched are traversed; Update the first query result, and return to step "determine all query results other than the first query result in the current query result set as query results to be matched" until all query results in the current query result set are traversed to obtain the current optimized result set; Update the current query result set, and return to step "determine any query result in the current query result set as the first query result" until all query result sets are traversed to obtain multiple optimized result sets.
4. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 3, characterized in that: Determining a coincidence determination result of the first query result and the second query result includes: Determine whether the document to which the first query result belongs is the same as the document to which the second query result belongs, and obtain a first determination result; If the first judgment result is no, determining that the first query result and the second query result do not overlap; If the first judgment result is yes, determining the edit distance similarity between the text block corresponding to the first query result and the text block corresponding to the second query result; Determine whether the edit distance similarity is greater than the edit distance similarity threshold, and obtain a second determination result; If the second judgment result is no, determining that the first query result and the second query result do not overlap; If the second judgment result is yes, then determining that the first query result and the second query result overlap, and determining whether the vectorized correlation coefficient corresponding to the first query result and the vectorized correlation coefficient corresponding to the second query result are equal, to obtain a third judgment result; If the third judgment result is yes, the first query result is deleted and the second query result is used as the updated first query result; or the second query result is deleted and the first query result is used as the updated first query result; If the third judgment result is no, then it is determined whether the vectorized correlation coefficient corresponding to the first query result is greater than the vectorized correlation coefficient corresponding to the second query result, and a fourth judgment result is obtained; If the fourth judgment result is yes, deleting the second query result and using the first query result as the updated first query result; If the fourth judgment result is no, the first query result is deleted and the second query result is used as the updated first query result.
5. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 4, characterized in that: The proportion of hyperlinks is: Among them, hyper_weight is the proportion of hyperlinks; The edit distance similarity is: similarity=1-distance / max(len(text1),len(text2)); Wherein, similarity is the edit distance similarity, distance is the edit distance between the text block text1 corresponding to the first query result and the text block text2 corresponding to the second query result, len(text1) is the length of the text block text1, and len(text2) is the length of the text block text2.
6. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 1, characterized in that: The term frequency-inverse document rate factor is: in, In the formula, r_TF-IDF is the term frequency-inverse document rate factor, n is the number of words in the text block after word segmentation, TF-IDF i is the term frequency-inverse document rate factor of the i-th word, freq i is the number of times the i-th word appears in this text block, df(i) is the number of text blocks where the i-th word appears in rule result1, and N is the number of text blocks in rule result1.
7. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 6, characterized in that: The date factors are: Among them, r date is the date factor.
8. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 7, characterized in that: The retrieved frequency factor is: in, is the search frequency factor.
9. The method for re-ranking scientific and technological intelligence information retrieval results according to claim 8, characterized in that: The comprehensive factors are: Among them, x+y+z+w=100%; Among them, r_total is the comprehensive factor, r_vector is the vectorized correlation coefficient, x is the vectorized correlation coefficient weight, y is the word frequency-inverse document rate factor weight, z is the date factor weight, and w is the retrieval frequency factor weight.
10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for reordering scientific and technological intelligence information retrieval results as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Search ranking method, apparatus, computer device, and storage medium
CN109033386A
Reordering method and device for improving retrieval performance of AI large language model
CN117725183A
Military question and answer method and system based on large language model
CN118839008A
Method for discovering network hot topics in science and technology field
CN118897894A
Retrieval enhancement generation-based retrieval method, product, equipment and medium
CN119003795A