Information retrieval methods, devices, readable storage media and electronic devices

By calculating the comprehensive index values ​​of multiple retrieval systems, filtering and integrating the initial results of the target retrieval system, the problem of users needing to make multiple adjustments to obtain satisfactory search results in existing technologies is solved, thus achieving efficient and accurate information retrieval.

CN117290576BActive Publication Date: 2026-04-07CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing information retrieval processes, there is a discrepancy between the user's query request and the system's response. Users need to make multiple adjustments to obtain satisfactory results, resulting in a large workload and low accuracy.

Method used

By receiving user query requests, obtaining initial search results from multiple search systems, calculating the comprehensive index value of each search system, selecting the target search system, and merging the initial results of the target search system, the comprehensive index value is used to characterize search performance and differences, thereby improving search accuracy.

Benefits of technology

It reduced the workload of retrieval, improved the efficiency and accuracy of retrieval, reduced repeated retrievals, and enhanced fusion performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117290576B_ABST
    Figure CN117290576B_ABST
Patent Text Reader

Abstract

This application discloses an information retrieval method, apparatus, readable storage medium, and electronic device, relating to the field of cloud computing technology. The method includes: receiving a query request sent by a user; acquiring initial search results obtained by each of multiple search systems in response to the query request; calculating a comprehensive index value for each search system, and filtering a target search system based on the comprehensive index value among the multiple search systems; and fusing the initial search results corresponding to the target search system to obtain the target search result corresponding to the query request. This disclosure can acquire the initial search results of each search system based on the query request, calculate a comprehensive index value to determine the target search system, and fuse the target search results, thereby improving fusion performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular to an information retrieval method, apparatus, readable storage medium, and electronic device. Background Technology

[0002] With the development of big data and the continuous enrichment of network resources, the total amount of information on the Internet has expanded rapidly. Information retrieval has become the main way for users to query and obtain information, enabling users to quickly and accurately obtain the desired information from the massive amount of information on the Internet with the help of search engines.

[0003] In existing information retrieval processes, there is a certain discrepancy between user query requests and system response results. Users often need to make multiple queries and adjustments to obtain satisfactory results, resulting in a large workload and low accuracy, failing to guarantee the depth of retrieval technology. Summary of the Invention

[0004] In view of this, this application provides an information retrieval method, apparatus, readable storage medium, and electronic device, which can filter target retrieval systems that match the query request from multiple retrieval systems by calculating comprehensive index values, and use the target retrieval system to determine the target retrieval results corresponding to the query request, thereby reducing the workload of retrieval and ensuring the accuracy of retrieval.

[0005] According to the first aspect of this application, an information retrieval method is provided, comprising:

[0006] Receive query requests sent by users;

[0007] Obtain the initial search results obtained by each of the multiple search systems in response to the query request;

[0008] Calculate the comprehensive index value for each retrieval system, and filter the target retrieval system among the multiple retrieval systems based on the comprehensive index value, wherein the comprehensive index value is used to characterize the retrieval performance of each retrieval system and the difference between each retrieval system and other retrieval systems;

[0009] By integrating the initial search results corresponding to the target retrieval system, the target retrieval result corresponding to the query request is obtained.

[0010] According to a second aspect of this application, an information retrieval device is provided, comprising:

[0011] The receiving module is used to receive query requests sent by users;

[0012] The acquisition module is used to acquire the initial search results obtained by each of the multiple search systems in response to the query request.

[0013] The calculation module is used to calculate the comprehensive index value of each retrieval system and filter the target retrieval system among the multiple retrieval systems based on the comprehensive index value, wherein the comprehensive index value is used to characterize the retrieval performance of each retrieval system and the difference between each retrieval system and other retrieval systems;

[0014] The fusion module is used to fuse the initial search results corresponding to the target retrieval system to obtain the target retrieval result corresponding to the query request.

[0015] According to a third aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the information retrieval method described in the first aspect.

[0016] According to a fourth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the information retrieval method described in the first aspect.

[0017] According to a fifth aspect of this application, a chip is provided, including one or more interface circuits and one or more processors; the interface circuits are configured to receive signals from the memory of an electronic device and send signals to the processors, the signals including computer instructions stored in the memory, which, when executed by the processors, cause the electronic device to perform the methods described in the first aspect of this disclosure.

[0018] By employing the above technical solutions, this application provides an information retrieval method, apparatus, readable storage medium, and electronic device. Compared with current information retrieval methods, this application can receive query requests sent by users; obtain initial search results from each of multiple retrieval systems in response to the query request; calculate a comprehensive index value for each retrieval system; and filter target retrieval systems among the multiple retrieval systems based on the comprehensive index value; and fuse the initial search results corresponding to the target retrieval systems to obtain the target search results corresponding to the query request. The technical solution in this disclosure utilizes multiple retrieval systems simultaneously, which can reduce the workload of retrieval and improve retrieval efficiency. Furthermore, by calculating a comprehensive index value and filtering target retrieval systems that match the query request from multiple retrieval systems, and using the target retrieval systems to determine the target search results corresponding to the query request, the accuracy of the retrieval can be guaranteed.

[0019] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objectives, features and advantages more obvious and understandable, specific embodiments of this application are described below. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating an information retrieval method provided in an embodiment of this disclosure;

[0023] Figure 2 A flowchart illustrating an information retrieval method according to another embodiment of this disclosure;

[0024] Figure 3 A graph showing the relationship between the number of retrieval systems and the fusion performance of information retrieval, provided in an embodiment of this disclosure;

[0025] Figure 4 A flowchart of an information retrieval system selection algorithm provided in this embodiment of the disclosure;

[0026] Figure 5 This is a schematic diagram of the structure of an information retrieval device provided in an embodiment of the present disclosure;

[0027] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure;

[0028] Figure 7 This is a schematic diagram of the chip structure provided in an embodiment of this disclosure. Detailed Implementation

[0029] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details to aid understanding, which should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0030] The information retrieval method, apparatus, readable storage medium, and electronic device according to embodiments of the present disclosure are described below with reference to the accompanying drawings.

[0031] In related technologies, there is a certain discrepancy between the user's query request and the query result returned by the system. Users often need to make multiple queries and adjustments to obtain satisfactory search results. The performance of the retrieval system is poor and cannot guarantee the depth of the retrieval technology.

[0032] To address the aforementioned technical problems, this disclosure provides an information retrieval method, apparatus, readable storage medium, and electronic device, which can reduce the workload of retrieval and improve the accuracy of retrieval.

[0033] like Figure 1 As shown, embodiments of this disclosure provide an information retrieval method, including:

[0034] Step 101: Receive the query request sent by the user.

[0035] In specific application scenarios, cloud computing platforms can receive query requests sent by users. Query requests can include, but are not limited to, query-based queries. Users can change the query-based queries and use search tools to perform multiple queries to find the information they need from the information set.

[0036] The execution entity of this embodiment can be a cloud computing platform, capable of obtaining initial search results from each search system based on query requests, calculating comprehensive index values ​​to determine the target search system, and fusing the target search results to improve fusion performance. Furthermore, based on the true relevance score and the predicted relevance score, search performance and average variance can be calculated, and search systems with lower search performance and higher average variance can be selected for fusion, reducing the complexity of the fusion system and improving fusion efficiency.

[0037] Step 102: Obtain the initial search results obtained by each of the multiple search systems in response to the query request.

[0038] In this embodiment of the disclosure, the cloud computing platform can distribute user-sent query requests to multiple retrieval systems. Each retrieval system can reorder all documents in the document dataset based on the relevance of each document in the document dataset to the query request, thereby obtaining the initial retrieval results for each retrieval system in response to the query request. The retrieval system processes, organizes, and stores information in a certain way, and then accurately retrieves relevant information according to the user's query request. The initial retrieval results can be a list of documents retrieved by the retrieval system in response to the query request.

[0039] Step 103: Calculate the comprehensive index value for each retrieval system, and filter the target retrieval system among multiple retrieval systems based on the comprehensive index value. The comprehensive index value is used to characterize the retrieval performance of each retrieval system and the differences between each retrieval system and other retrieval systems.

[0040] In specific application scenarios, Euclidean distance can be used to calculate the retrieval performance of each retrieval system and the differences between them. Based on the retrieval performance of each system and its differences from other retrieval systems, a comprehensive index value is calculated for each system. The retrieval systems are then ranked according to their comprehensive index values. A filtering algorithm is used to select a predetermined number of target retrieval systems from among the multiple systems. This helps to extract search results that better match the query request and improves the accuracy of the search. Specifically, the comprehensive index value characterizes the retrieval performance of each system and its differences from other retrieval systems. Retrieval performance characterizes the accuracy of the search results for each system, while the differences characterize the variations in search results between different retrieval systems.

[0041] Step 104: Merge the initial search results corresponding to the target retrieval system to obtain the target retrieval results corresponding to the query request.

[0042] One possible approach is to use data fusion algorithms to assign weights to the participating retrieval systems, improving the accuracy of search results. These fusion algorithms include, but are not limited to, clustering algorithms, random forests, genetic algorithms, and particle swarm optimization. Another possible approach is to calculate a comprehensive index value for each retrieval system based on Euclidean distance, representing the search results of each system as a point in a multi-dimensional space. This transforms the fusion problem into a geometric problem. Based on the comprehensive index value, a target retrieval system is selected from multiple retrieval systems. The initial search results of the target retrieval systems are then fused to form the target search result, which is then fed back to the user's query interface, reducing time complexity.

[0043] In summary, according to the information retrieval method disclosed herein, compared with current information retrieval methods, this application can receive query requests sent by users; obtain initial search results obtained by each of multiple retrieval systems in response to the query request; calculate a comprehensive index value for each retrieval system; and filter target retrieval systems among multiple retrieval systems based on the comprehensive index value; and fuse the initial search results corresponding to the target retrieval systems to obtain the target retrieval results corresponding to the query request. The technical solution in this disclosure utilizes multiple retrieval systems simultaneously, which can reduce the workload of retrieval and improve retrieval efficiency. Furthermore, by calculating a comprehensive index value and filtering target retrieval systems that match the query request from multiple retrieval systems, and using the target retrieval systems to determine the target retrieval results corresponding to the query request, the accuracy of the retrieval can be guaranteed.

[0044] Based on the above architecture, this public application example also provides an information retrieval method, which includes:

[0045] Step 201: Receive the query request sent by the user.

[0046] For the specific implementation process of the embodiments disclosed herein, please refer to the relevant description in step 101 of the embodiment, which will not be repeated here.

[0047] Step 202: Obtain the initial search results obtained by each of the multiple search systems in response to the query request.

[0048] In this embodiment of the disclosure, the information retrieval system may include multiple retrieval systems for searching a document dataset in response to a query request and obtaining initial retrieval results. For example, the information retrieval system IR may include m retrieval systems, which can be represented as IR = (ir1, ir2, ..., ir...). m ),ir j (1≤j≤m) represents the j-th retrieval system, and the document dataset D can contain p documents, which can be represented as D=(d1, d2, ..., dm). p ), d k (1≤k≤p) represents the kth document.

[0049] Step 203: Obtain the sample query request, which carries the true relevance score of each document in the document dataset.

[0050] In specific application scenarios, sample query requests and actual search results matching the sample query requests can be obtained, and the true relevance score of each document can be calculated. This helps to compare whether the initial search results of various search systems match the user's query request. The true relevance score can be used to grade the relevance between the actual search results and the query request. For example, a sample query request O can contain the true relevance scores of p documents, which can be represented as O = (o 1 o 2 , ···, o p ), o k (1≤k≤p) represents the k-th document d. k The true relevance score.

[0051] Step 204: Determine the predicted relevance score for each document in the document dataset when each retrieval system executes a sample query request.

[0052] In specific application scenarios, each retrieval system can retrieve initial retrieval results after retrieving a sample query request, and calculate the predicted relevance score for each document data in the initial retrieval results. The predicted relevance score can be used to assess the relevance between the initial retrieval results and the query request. For example, the predicted relevance score of the j-th retrieval system for the sample query request can be expressed as: For the j-th retrieval system regarding the sample query request for the k-th document d k The predictive relevance score can be set. The range is between 0 and 1, which normalizes the predicted relevance score of each retrieval system, facilitating comparison of the retrieval performance of each system. Specifically, a multi-dimensional hypercube can be defined, distributing the predicted relevance score of each retrieval system for the document dataset within the hypercube. This is abstracted as a point within the hypercube, making it easier to calculate the performance value of the retrieval system. Here, the hypercube X(q, D) can be the retrieval result space formed by the sample query request q and the document dataset D, and the dimension can be the number of documents p in the document dataset D.

[0053] Step 205: Calculate the retrieval performance value of each retrieval system based on the true relevance score and the predicted relevance score.

[0054] As one possible implementation, the first Euclidean distance between the predicted relevance score of each retrieval system and the true relevance score in the sample query request is calculated. This first Euclidean distance can be determined as the retrieval performance value of the corresponding retrieval system. Here, the first Euclidean distance refers to the Euclidean distance between the predicted relevance score of each retrieval system and the true relevance score in the sample query request within a multidimensional space. The formula for calculating the first Euclidean distance can be expressed as:

[0055]

[0056] In the formula, dist(S,O) represents the retrieval performance value of the retrieval system S relative to the sample query request O, and s k (1≤k≤p) represents the predicted relevance score of the retrieval system S for the k-th document, o k (1≤k≤p) represents the true relevance score of the sample query request O with respect to the k-th document.

[0057] Accordingly, the implementation steps may include: calculating a first Euclidean distance between the true relevance score and the predicted relevance score; and determining the first Euclidean distance as the retrieval performance value of the corresponding retrieval system.

[0058] Step 206: Based on the predicted relevance score, calculate the average difference value between each retrieval system and other retrieval systems.

[0059] As one possible implementation, based on the predicted relevance score, the second Euclidean distance between each retrieval system and other retrieval systems can be calculated one by one, obtaining multiple second Euclidean distances between the corresponding retrieval system and other retrieval systems. This can be used to analyze the differences between the retrieval results of the same query request between two retrieval systems. Then, the average of the multiple second Euclidean distances between each retrieval system and other retrieval systems is calculated, and the average of these multiple second Euclidean distances is determined as the average difference value between the corresponding retrieval system and other retrieval systems. This can be used to analyze the average difference between the corresponding retrieval system and other retrieval systems. Here, the second Euclidean distance refers to the Euclidean distance between the predicted relevance score of each retrieval system and the predicted relevance scores of other retrieval systems in a multidimensional space. The formula for calculating the second Euclidean distance can be expressed as:

[0060]

[0061] In the formula, dist(S1, S2) represents the difference value between retrieval system S1 and retrieval system S2. This represents the predicted relevance score of retrieval system S1 for the k-th document. This represents the predicted relevance score of retrieval system S2 for the k-th document.

[0062] Accordingly, the implementation steps may include: calculating multiple second Euclidean distances between each retrieval system and other retrieval systems regarding the predicted relevance score; calculating the average of the multiple second Euclidean distances, and determining the average of the multiple second Euclidean distances as the average difference value between the corresponding retrieval system and other retrieval systems.

[0063] Step 207: Calculate the comprehensive index value for each retrieval system based on the retrieval performance value and the average difference value.

[0064] The retrieval performance value and average dissimilarity value of each retrieval system can be substituted into the formula for calculating the comprehensive index value to calculate the Euclidean distance between the centroid C of each retrieval system and the sample query request O, thus obtaining the comprehensive index value dist(C,O) of each retrieval system. The formula for calculating the comprehensive index value can be expressed as:

[0065]

[0066] The above formula is derived as follows:

[0067]

[0068] Let dist(S) i O) 2 =p 2 ,dist(S i S j ) 2=d 2 d = θ * p, thus obtaining

[0069]

[0070] Let index i =dist(C,O) 2 p i =p 2 d i =θ 2 *p 2 ,get

[0071] index i =p i -d i ×(m-1) / 2m, (1≤i,m≤n)

[0072] In the formula, index i Let C be the Euclidean distance between the centroid C of the i-th retrieval system and the sample query request O, which is the comprehensive index value of the i-th retrieval system; m be the number of retrieval systems randomly selected from multiple retrieval systems; n be the total number of retrieval systems; and p be the total number of retrieval systems. i (1≤i≤n) represents the retrieval performance value of the i-th retrieval system, d i (1≤i≤n) represents the average difference between the i-th retrieval system and other retrieval systems, where the formula for calculating the centroid C of each retrieval system can be expressed as:

[0073]

[0074] c k (1≤k≤p) represents the j-th retrieval system S. j The center point.

[0075] Based on the above formula, the selectable retrieval performance value dist(S) i The difference between the O and other retrieval systems should be minimized, and the difference between the O and other retrieval systems should be minimized. i ,S j The goal is to maximize the size of the retrieval system and minimize the overall index value dist(C,0) of the corresponding retrieval system. This results in search results that better match the query request and are unique to other retrieval systems, thus improving fusion performance. Accordingly, the implementation steps may include: substituting the retrieval performance value and the average dissimilarity value into the overall index value calculation formula to obtain the overall index value for each retrieval system; the formula characteristics of the index value calculation formula are described as follows:

[0076] index i =p i -d i ×(m-1) / 2m, (1≤i,m≤n)

[0077] In the formula, index i Let p be the comprehensive index value of the i-th retrieval system. i Let d be the retrieval performance value of the i-th retrieval system. i Let be the average difference value of the i-th retrieval system, n be the number of retrieval systems, and m be the number of retrieval systems randomly selected from the multiple retrieval systems.

[0078] Step 208: Select the target retrieval system from multiple retrieval systems based on the comprehensive index value.

[0079] To facilitate understanding of the disclosed technical solutions, this section combines... Figure 3 The scheme in this disclosure is described in full. Figure 3 This paper presents the trends of fusion performance and the average dissimilarity ratio coefficient θ with respect to the number of retrieval systems under four different values. It can be observed that fusion performance decreases as the number of retrieval systems increases, and as the number of retrieval systems approaches infinity, the fusion performance tends to... For example, when θ is 0.25, 0.5, 0.75, and 1, the values ​​of dist(C,O) are 0.984p, 0.935p, 0.848p, and 0.707p, respectively. When the number of retrieval systems involved in the fusion reaches 30 or more, the improvement in fusion performance is not significant.

[0080] As one possible implementation, during each query, m retrieval systems can be randomly selected from n retrieval systems. A comprehensive index value is calculated for each of the m retrieval systems. The m retrieval systems are then ranked based on their comprehensive index values, and the m systems with the smallest comprehensive index values ​​are identified as the target retrieval systems, thereby obtaining the retrieval systems with better fusion performance. Here, m is any value less than or equal to n. Accordingly, the implementation steps may include: randomly selecting m retrieval systems from n retrieval systems and obtaining the comprehensive index value calculated for each of the m retrieval systems, where m is any value less than or equal to n; and identifying the m retrieval systems with the smallest corresponding comprehensive index values ​​as the target retrieval systems.

[0081] As one possible implementation, such as Figure 4As shown, during each query, the retrieval system with the largest comprehensive index value can be repeatedly eliminated from multiple retrieval systems, and the number of systems n among the multiple retrieval systems can be updated until a preset number of retrieval systems remain. The remaining preset number of retrieval systems are then determined as the target retrieval systems. Specifically, m retrieval systems can be randomly selected from n retrieval systems, and the comprehensive index values ​​of the m retrieval systems can be sorted. The retrieval system with the largest comprehensive index value is then eliminated, thereby removing retrieval systems with poor fusion performance. Here, m is any value less than or equal to n. Correspondingly, the steps of the embodiment may include: repeatedly executing the following retrieval system elimination process and updating the number of systems n among the multiple retrieval systems until a preset number of retrieval systems remain, and determining the remaining preset number of retrieval systems as the target retrieval systems; wherein, the retrieval system elimination process includes: randomly selecting m retrieval systems from n retrieval systems and obtaining the comprehensive index value calculated by each retrieval system in the m retrieval systems, where m is any value less than or equal to n; eliminating the retrieval system with the largest corresponding comprehensive index value among the multiple retrieval systems.

[0082] Step 209: Merge the initial search results corresponding to the target retrieval system to obtain the target retrieval results corresponding to the query request.

[0083] In specific application scenarios, the initial search results corresponding to the target retrieval system can be obtained, and the document data in each initial search result can be sorted according to relevance to generate the target search results. The target search results can then be fed back to the user's query interface, reducing the complexity of the fusion system and improving fusion efficiency.

[0084] In summary, the information retrieval method provided in this disclosure, compared with current information retrieval methods, can obtain the initial retrieval results of each retrieval system based on the query request, calculate the comprehensive index value to determine the target retrieval system, merge the target retrieval results, improve fusion performance, and avoid repeated retrievals, thus reducing the workload of retrieval. Furthermore, based on the true relevance score and the predicted relevance score, retrieval performance and average variance can be calculated, and retrieval systems with lower retrieval performance and higher average variance can be selected for fusion, thereby ensuring the accuracy of the retrieval.

[0085] Based on the above Figure 1 , Figure 2 The specific implementation of the method shown in this embodiment provides an information retrieval device, such as... Figure 5 The device includes: a receiving module 31, an acquisition module 32, a calculation module 33, and a fusion module 34;

[0086] The receiving module 31 can be used to receive query requests sent by users;

[0087] The acquisition module 32 can be used to acquire the initial search results obtained by each of the multiple search systems in response to the query request.

[0088] The calculation module 33 can be used to calculate the comprehensive index value of each retrieval system and filter the target retrieval system among multiple retrieval systems based on the comprehensive index value. The comprehensive index value is used to characterize the retrieval performance of each retrieval system and the differences between each retrieval system and other retrieval systems.

[0089] The fusion module 34 can be used to fuse the initial search results corresponding to the target retrieval system to obtain the target retrieval results corresponding to the query request.

[0090] In specific application scenarios, the calculation module 33 can be used to obtain sample query requests, which carry the true relevance score of each document in the document dataset; determine the predicted relevance score output by each retrieval system for each document in the document dataset when executing the sample query request; calculate the retrieval performance value of each retrieval system based on the true relevance score and the predicted relevance score; calculate the average difference value between each retrieval system and other retrieval systems based on the predicted relevance score; and calculate the comprehensive index value of each retrieval system based on the retrieval performance value and the average difference value.

[0091] In specific application scenarios, the calculation module 33 can be used to calculate the first Euclidean distance between the true relevance score and the predicted relevance score; and the first Euclidean distance is determined as the retrieval performance value of the corresponding retrieval system.

[0092] In specific application scenarios, the calculation module 33 can be used to calculate multiple second Euclidean distances between each retrieval system and other retrieval systems regarding the predicted relevance score; calculate the average of the multiple second Euclidean distances, and determine the average of the multiple second Euclidean distances as the average difference value between the corresponding retrieval system and other retrieval systems.

[0093] In specific application scenarios, the calculation module 33 can be used to substitute the retrieval performance value and the average difference value into the comprehensive index value calculation formula to obtain the comprehensive index value of each retrieval system; the formula characteristics of the index value calculation formula are described as follows:

[0094] index i =p i -d i ×(m-1) / 2m, (1≤i,m≤n)

[0095] In the formula, index i Let p be the comprehensive index value of the i-th retrieval system. i Let d be the retrieval performance value of the i-th retrieval system. iLet be the average difference value of the i-th retrieval system, n be the number of retrieval systems, and m be the number of retrieval systems randomly selected from the multiple retrieval systems.

[0096] In specific application scenarios, the calculation module 33 can be used to randomly select m retrieval systems from n retrieval systems and obtain the comprehensive index value calculated by each of the m retrieval systems, where m is any value less than or equal to n; the m retrieval systems with the smallest corresponding comprehensive index value are determined as the target retrieval systems.

[0097] In specific application scenarios, the calculation module 33 can be used to repeatedly execute the following retrieval system elimination process and update the number n of multiple retrieval systems until a preset number of retrieval systems remain, and determine the remaining preset number of retrieval systems as the target retrieval systems; wherein, the retrieval system elimination process includes: randomly selecting m retrieval systems from n retrieval systems, and obtaining the comprehensive index value calculated by each retrieval system in the m retrieval systems, where m is any value less than or equal to n; and eliminating the retrieval system with the largest corresponding comprehensive index value among the multiple retrieval systems.

[0098] Since the apparatus provided in this embodiment corresponds to the methods provided in the above embodiments, the implementation of the methods is also applicable to the apparatus provided in this embodiment, and will not be described in detail in this embodiment.

[0099] The methods and apparatus provided in the embodiments of this application have been described above. To implement the functions of the methods provided in the embodiments of this application, the electronic device may include a hardware structure and software modules, and may implement the above functions in the form of a hardware structure, software modules, or a hardware structure plus software modules. One of the above functions may be executed in the form of a hardware structure, software modules, or a hardware structure plus software modules.

[0100] Figure 6 This is a block diagram illustrating an electronic device 600 for implementing the above-described information retrieval method, according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, computer, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.

[0101] Reference Figure 6 The electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.

[0102] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.

[0103] Memory 604 is configured to store various types of data to support the operation of electronic device 600. Examples of this data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0104] Power supply component 606 provides power to various components of electronic device 600. Power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.

[0105] Multimedia component 608 includes a screen that provides an output interface between electronic device 600 and user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When electronic device 600 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0106] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.

[0107] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0108] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 may detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0109] Communication component 616 is configured to facilitate wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio), or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0110] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0111] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0112] Embodiments of this disclosure also propose a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the information retrieval method described in the above embodiments of this disclosure.

[0113] Embodiments of this disclosure also provide a computer program product, including a computer program that is executed by a processor using the information retrieval method described in the above embodiments of this disclosure.

[0114] Embodiments of this disclosure also propose a chip, which can be found in [reference]. Figure 7 The diagram shows the structure of the chip. Figure 7 The chip shown includes a processor 701 and an interface circuit 702. The number of processors 701 and the number of interface circuits 702 can be one or more.

[0115] Optionally, the chip also includes a memory 703 for storing necessary computer programs and data; an interface circuit 702 for receiving signals from the memory 703 and sending signals to the processor 701, the signals including computer instructions stored in the memory 703, which, when executed by the processor 701, cause the electronic device to perform the information retrieval method described in the above embodiments of this disclosure.

[0116] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0117] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with an embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0118] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (control method), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic device, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which programs can be printed, because programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0120] It should be understood that various parts of the embodiments of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0121] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0122] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc.

[0123] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. An information retrieval method, characterized in that, include: Receive query requests sent by users; Obtain the initial search results obtained by each of the multiple search systems in response to the query request; Calculate a comprehensive index value for each retrieval system, and filter target retrieval systems among multiple retrieval systems based on the comprehensive index value. The comprehensive index value characterizes the retrieval performance of each retrieval system and its differences from other retrieval systems. This includes: obtaining a sample query request, which carries the true relevance score for each document in the document dataset; determining the predicted relevance score output by each retrieval system for each document in the document dataset when executing the sample query request; calculating the retrieval performance value of each retrieval system based on the true relevance score and the predicted relevance score; calculating the average difference value between each retrieval system and other retrieval systems based on the predicted relevance score; and calculating the comprehensive index value of each retrieval system based on the retrieval performance value and the average difference value. By integrating the initial search results corresponding to the target retrieval system, the target retrieval result corresponding to the query request is obtained.

2. The method according to claim 1, characterized in that, The calculation of the retrieval performance value of each retrieval system based on the true relevance score and the predicted relevance score includes: Calculate the first Euclidean distance between the true relevance score and the predicted relevance score; The first Euclidean distance is determined as the retrieval performance value of the corresponding retrieval system.

3. The method according to claim 1, characterized in that, The step of calculating the average difference value between each retrieval system and other retrieval systems based on the predicted relevance score includes: Calculate multiple second Euclidean distances between each retrieval system and other retrieval systems with respect to the predicted relevance score; Calculate the average of the plurality of second Euclidean distances, and determine the average of the plurality of second Euclidean distances as the average difference value between the corresponding retrieval system and other retrieval systems.

4. The method according to claim 1, characterized in that, The step of calculating the comprehensive index value of each retrieval system based on the retrieval performance value and the average difference value includes: Substituting the retrieval performance value and the average difference value into the comprehensive index value calculation formula, the comprehensive index value of each retrieval system is obtained; The formula characteristics of the index value calculation formula are described as follows: In the formula, Let i be the comprehensive index value of the i-th retrieval system. Let be the retrieval performance value of the i-th retrieval system. Let be the average difference value of the i-th retrieval system. The number of systems in multiple retrieval systems. The number of retrieval systems randomly selected from the multiple retrieval systems.

5. The method according to claim 4, characterized in that, The process of filtering target search systems among multiple search systems based on the comprehensive index value includes: exist Randomly selected from a search system A retrieval system, and obtain the The comprehensive index value calculated by each retrieval system in the retrieval system, wherein... Less than or equal to Any value; The one with the smallest corresponding comprehensive index value The retrieval system is identified as the target retrieval system.

6. The method according to claim 4, characterized in that, The process of filtering target search systems among multiple search systems based on the comprehensive index value includes: Repeat the removal process for the search systems described below, and update the number of systems for the plurality of search systems. Until a preset number of retrieval systems remain, the remaining preset number of retrieval systems will be identified as the target retrieval system; The elimination process of the retrieval system includes: exist Randomly selected from a search system A retrieval system, and obtain the The comprehensive index value calculated by each retrieval system in the retrieval system, wherein... Less than or equal to Any value; The retrieval system with the highest comprehensive index value among the multiple retrieval systems is removed.

7. An information retrieval device, characterized in that, include: The receiving module is used to receive query requests sent by users; The acquisition module is used to acquire the initial search results obtained by each of the multiple search systems in response to the query request. A calculation module is used to calculate a comprehensive index value for each retrieval system and to filter target retrieval systems among multiple retrieval systems based on the comprehensive index value. The comprehensive index value characterizes the retrieval performance of each retrieval system and the difference between each retrieval system and other retrieval systems. The module includes: obtaining a sample query request, which carries the true relevance score for each document in the document dataset; determining the predicted relevance score output by each retrieval system for each document in the document dataset when executing the sample query request; calculating the retrieval performance value of each retrieval system based on the true relevance score and the predicted relevance score; calculating the average difference value between each retrieval system and other retrieval systems based on the predicted relevance score; and calculating the comprehensive index value of each retrieval system based on the retrieval performance value and the average difference value. The fusion module is used to fuse the initial search results corresponding to the target retrieval system to obtain the target retrieval result corresponding to the query request.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 6.

9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 6.

10. A chip comprising one or more interface circuits and one or more processors; the interface circuits being configured to receive signals from a memory of an electronic device and to send the signals to the processors, the signals including computer instructions stored in the memory, wherein when the processor executes the computer instructions, the electronic device performs the method of any one of claims 1 to 6.

Citation Information

Patent Citations

  • Information retrieval data fusion method based on retrieval result diversification

    CN103838874A

  • Data integration method supporting diversification of information retrieving results

    CN104408089A