Data retrieval method and device
By generating and storing detailed structured information during the data retrieval process, the problem of delay in obtaining detailed information in the prior art is solved, the search efficiency and applicability are improved, and it is suitable for complex business scenarios.
Patent Information
- Application Number
- CN202311531812.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-15
AI Technical Summary
When obtaining detailed information of search results, existing data retrieval technology leads to a long delay and cannot be applied to complex business scenarios.
By generating and pre-stored detailed structured information during the search process, including scoring description information and hit information, the number of subsequent queries is reduced and the search efficiency is improved.
It reduces the delay in obtaining detailed information, improves retrieval efficiency and system performance, is suitable for business scenarios with huge data volume, and increases its applicability to complex businesses.
Smart Images

Figure CN120045521A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data retrieval, and in particular, to a method and apparatus for data retrieval. Background Art
[0002] Data Retrieval is an important branch in the fields of computer science and information retrieval, which involves finding information related to search conditions from large data sets. Data retrieval can be applied to various types of data, including text, images, audio, video, etc.
[0003] In the existing data retrieval field, data can be retrieved using any conditions, and the conditions can be combined in various ways and multiple fields can be combined to obtain the corresponding retrieved data. For the retrieved search results, detailed information, such as explanatory information, can be obtained by performing an interpretive query operation.
[0004] However, in the prior art, since obtaining detailed information such as explanatory information requires performing an interpretive query operation on each hit document in the search results separately, it may lead to a long delay, and only the explanatory information of the search results in simple services can be obtained through the interpretive query, which may not be applicable to some complex business scenarios. Therefore, how to solve the problem of long delay caused by obtaining the explanatory information of search results during data retrieval and increase the applicability to various business scenarios is an urgent problem to be solved. Summary of the Invention
[0005] Embodiments of this application provide a method and apparatus for data retrieval to solve the problem of long delay caused by obtaining detailed information of search results during data retrieval and improve the retrieval efficiency.
[0006] In a first aspect, embodiments of this application provide a method for data retrieval, which may include:
[0007] Receiving a search condition and generating a first structured object, where the first structured object includes a query statement included in the search condition; creating a scoring device based on the first structured object, where the scoring device includes the first structured object, M document identifiers matching the query statement, and a scoring component, and M is a positive integer greater than 0; obtaining M first documents matching the M document identifiers, and generating scoring description information corresponding to each of the M first documents through the scoring component; generating M second structured objects based on the scoring description information corresponding to each first document and the first structured object, where one second structured object corresponds to one first document.
[0008] In the prior art, when a user conducts a content search, the search engine usually queries and returns preliminary search results according to the search conditions (for example, preliminary query results including one or more hit documents). When the user needs to further view one or more of the hit documents in the search results, the search engine needs to re-query and extract detailed information (such as the reason for the hit, hit details, detailed content, explanatory information, etc.) for the one or more hit documents. In this way, query waiting latency will be generated, resulting in a poor user experience and also causing a waste of system resources. To address this technical problem, in the embodiments of the present application, by setting the search engine to, after receiving the user's search conditions for the first time, not only obtain preliminary search results during the search process, but also further obtain more in-depth detailed information for the search results and pre-generate corresponding structured information. As a result, when the user needs to conduct an in-depth query on a certain hit document in the search results, the search engine can immediately feedback and present the pre-generated structured information to the user, greatly saving the latency caused by the need for the search engine in the prior art to query multiple times to feedback the results, saving system resources, and improving system performance. Specifically, in the embodiments of the present application, during the search process, the search engine creates a scoring calculator for the search conditions. The scoring calculator may include preliminary structured information (i.e., the first structured object including the query statement), the identifier of the hit document, and a scoring component. Further, the search engine obtains the corresponding candidate documents according to the identifier of each hit document, generates the scoring description information for each hit document through the scoring component, and further generates in-depth structured information (i.e., the second structured object) including the above first structured object and the scoring description information for each hit document, so as to accurately and efficiently obtain relevant information when the user needs to query later. In summary, in the embodiments of the present application, the required detailed information can be obtained and saved while searching for hit documents for the first time. When the user conducts a further query, there is no need to perform repeated explanatory query operations. Therefore, the latency of obtaining detailed information during the search process is greatly reduced, and the performance of the search is improved. It can be applied to services with a larger amount of data, increasing the applicability to more business scenarios.
[0009] In a possible implementation, the search engine includes an inverted index, and the inverted index includes a mapping relationship between a document identifier and a corresponding document; the obtaining of the M first documents matching the M document identifiers includes: searching for the M first documents corresponding to the M document identifiers from the inverted index. In the embodiments of the present application, by applying the inverted index structure, the search engine can retrieve the M candidate documents matching the M document identifiers faster, thereby improving the retrieval efficiency. Specifically, through the mapping relationship between the document identifier (i.e., ID) and the candidate document (i.e., the first document), the inverted index can quickly locate and retrieve the M candidate documents matching the search condition; when dealing with a large amount of data, this method significantly reduces the search time, enabling users to obtain relevant search results in a short time. In the embodiments of the present application, finding the corresponding document based on the inverted index structure is particularly important in business scenarios involving complex queries and large-scale document libraries, which can ensure that the search engine can still maintain high accuracy and response speed under high load.
[0010] In a possible implementation, the scoring description information corresponding to each of the first documents includes hit information and scoring information; the generating of the scoring description information corresponding to each of the M first documents by the scoring component includes: for the scoring calculator, screening out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q; generating the hit information and scoring information corresponding to each of the M first documents by the scoring component, and the scoring information is the score of the matching degree between the corresponding first document and the search condition. In the embodiments of the present application, through the detailed description of the scoring description information in the search engine and the classification of the structure of each candidate document in the scoring calculator, different processing can be performed on documents with different structures when obtaining detailed information subsequently, improving the relevance and accuracy of the search results. Specifically, in the embodiments of the present application, the scoring calculator can identify nested and non-nested document structures (i.e., when dealing with documents containing multi-level data, such as documents in nested JSON or XML formats, the scoring calculator can distinguish and identify whether the document is a nested structure), which enables the search engine to process complex queries more accurately; in the prior art, different processing is not performed on documents of different structure types. In addition, the scoring component in the scoring calculator can generate the scoring description information (i.e., hit information and scoring information) corresponding to each candidate document, and the scoring information is the score based on the matching degree between the corresponding candidate document and the search condition, thereby improving the accuracy of the search results.
[0011] In a possible implementation, the hit information in the J first documents includes one or more of the hit fields, hit words, and hit positions that match the search condition in the corresponding first document. In the embodiments of the present application, after determining that the J candidate documents are non-nested structures, the scoring component can determine the possible information included in the hit information based on the matching situation between the search condition and the non-nested candidate documents, such as one or more of the information such as hit fields, hit words, and hit positions. It can indicate the specific position where the query term appears based on the hit information, and can also be used to calculate the scoring information together with other information in the matching situation (for example, document frequency (DF) of the term, term frequency (TF), document length, query term weight, etc.).
[0012] In a possible implementation, the first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document in each of the Q first documents, the set of document identifier sequences of the one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document. In the prior art, when a user needs to further view the hit reasons of one or more hit documents in the search results, for hit documents with complex structures, only shallow hit reasons (such as the hit reasons of the main document in a nested document) can be obtained, and deeper hit reasons (such as the hit reasons of the sub-documents nested in the main document) cannot be further obtained. In this way, obtaining detailed information such as explanation information can only be applied to simple services and has no practical business use significance. To address this technical problem, in the embodiments of the present application, the scoring component in the scorer can obtain the explanation information of the main document and the sub-documents in the candidate document with a nested structure, rather than only being able to obtain the explanation information of the main document, improving the applicability to more business scenarios. Specifically, if the Q candidate documents are nested documents, based on the documents with a nested structure, by modifying the scorer (such as the scorer in the lucene search engine), the scorer can obtain the hit information of the internal documents in each candidate document with a nested structure after detecting that the candidate document is a document with a nested structure (for example, the ID of the main document, the set of ID sequences of one or more sub-documents, the ID of the sub-document hit in the main document, and one or more of the hit positions of the hit sub-documents can be obtained), enabling the embodiments of the present application to be applicable to business scenarios with complex structures (such as documents with a nested structure, which can include videos, note-taking APPs, rich texts, etc.). In summary, in the embodiments of the present application, when detecting a complex service with a nested document structure, by enabling the scorer to obtain more and more detailed information (such as the hit information of deep-level sub-documents), the embodiments of the present application can be applicable to more complex business scenarios, increasing the applicability to various business scenarios.
[0013] In a possible implementation, the generation of the first structured object includes: obtaining K sub-query statements included in the query statement, where K is a positive integer greater than 0; generating K third structured objects based on the K sub-query statements, where each of the third structured objects includes a corresponding sub-query statement, and the first structured object includes the K third structured objects. In the embodiments of the present application, by obtaining multiple sub-query statements included in the query statement and generating corresponding structured objects for each sub-query statement, the search engine can more precisely understand and process complex queries. The independent structuring of each sub-query statement enables the search engine to process each sub-query statement separately (in addition, in the query statement, data can be filtered according to a combination of multiple conditions, for example, using relational words such as logical operators "OR" and "AND" to combine multiple sub-query statements to form a combined query). When querying sub-documents under nested objects through sub-query statements, the parent documents hit due to the sub-document search conditions can be returned, improving the relevance and accuracy of search results; at the same time, decomposing the complex query into a combination of multiple sub-query statements enables the search engine to adopt different strategies and algorithms for different parts of the query conditions, enhancing the adaptability of the search engine to complex information requirements and the ability to process complex queries, which can help improve the efficiency of information retrieval, especially in the case of processing queries containing logical operators (such as AND / OR / NOT) or nested queries, making the search results returned by data retrieval more accurate.
[0014] In a possible implementation, K corresponding sub-scorers are created based on the K third structured objects, and each of the sub-scorers includes the corresponding third structured object, L document identifiers matching the corresponding third structured object, and the scoring component; the scorer is created based on the K corresponding sub-scorers, where L is a positive integer greater than 0. In the embodiments of the present application, corresponding sub-scorers are created based on each sub-query statement in the query statement, and then different sub-query statements are processed separately based on different sub-scorers, improving the flexibility of complex query processing. Specifically, in the embodiments of the present application, corresponding sub-scorers are created for each sub-query statement, and the above sub-scorers exist in the corresponding main scorer (i.e., the scorer corresponding to the first structured object generated based on the query statement). Each sub-scorer operates on a specific query fragment (for example, it can score and sort each one separately), enabling the search engine to better interpret and optimize the query results, while improving the ability to process nested queries and multi-condition queries.
[0015] In a possible implementation, for each scoring device, the corresponding first document is sorted first based on the scoring description information; based on the result of the first sorting, N first documents whose scoring description information meets the preset conditions are selected from the first documents corresponding to each scoring device, where N is a positive integer less than or equal to M. In the embodiments of the present application, after obtaining detailed information such as explanatory information, the candidate documents are sorted according to the magnitudes of the scoring information of each scoring device and its sub-scoring devices, and then the hit documents that meet the preset conditions are selected based on the sorting result, which can effectively distinguish, preferentially display, and screen out the data with the highest relevance to the query condition among a large number of candidate documents. Specifically, by first considering the score level during the sorting process, the candidate documents (i.e., the first documents) corresponding to each scoring device and its sub-scoring devices are sorted (for example, the candidate documents are sorted in sequence according to the relevance level), and then the documents that meet specific criteria are selected according to the preset conditions (for example, there are 100,000 candidate documents in the search results, and the business party only needs the top ten documents with the highest relevance, then the top ten documents are taken as the hit documents according to the sorting result), thereby ensuring the relevance and accuracy of the search results, making the finally presented result the information that best meets the search conditions, and thus improving the user's search experience.
[0016] In a possible implementation, a preset storage component is included in the search engine; for each of the first documents in the scoring device, the corresponding second structured object is saved to the storage component; the method further includes: based on the document identifiers of the N first documents, the corresponding N second structured objects are retrieved from the storage component; based on the N second structured objects, N corresponding structured information is generated. In the embodiments of the present application, during the preliminary retrieval process, by pre-storing the detailed information (such as explanation information) of all candidate documents, only the required explanation information is retrieved when outputting the search results. Compared with the prior art where multiple queries and retrievals are required for the returned search results, the latency caused by the explanation query is avoided; and in the prior art, the storage method of the explanation information is to splice the explanation information into a string, which cannot be directly used, while in the embodiments of the present application, the explanation information is constructed into structured information, which is more convenient for downstream services to directly use. Specifically, based on the explanation information of the candidate documents in each scoring device, the corresponding explanation information object can be generated and saved to the explanation information object collector. During the stage of outputting the search results, based on the ID of the hit document in the search results, the corresponding explanation information object is retrieved from the above-mentioned explanation information object collector to generate the corresponding structured information of the explanation, without retrieving all the explanation information of the candidate documents, thus avoiding waste of memory resources; while in the prior art, the explanation information usually needs to be obtained after multiple explanation queries, and the explanation information is spliced into a string form, and the hit information it contains is not rich enough, making it difficult to extract structured information for application in subsequent business logics, and this explanation is mainly used for development and debugging; in addition, all the explanation reasons are spliced in the string, which will also lead to a high memory occupancy rate and no practical business application value. In summary, in the embodiments of the present application, by obtaining all the explanation information at once during the search process and retrieving it when needed, the latency caused by the explanation query is avoided, and the explanation information is constructed into structured information, which can be directly used by downstream services, reducing memory, and being more convenient to use, increasing the practical business use significance.
[0017] In a possible implementation, for each scoring device, the corresponding first document is sorted first based on the scoring description information; based on the result of the first sorting, N first documents whose scoring description information meets the preset conditions are selected from the first documents corresponding to each scoring device, where N is a positive integer less than or equal to M. In the embodiments of the present application, when the detailed information (such as the explanatory information) is more comprehensive structured information, the explanatory information can be better utilized during the secondary sorting for in-depth screening, so as to screen out the documents whose matching degree with the query condition hits, thereby improving the accuracy of the search results. Specifically, through the explanatory information in the structured information, the search engine can understand in detail how the score of each hit document is calculated, including the matching degree of each word in the query with the document, word frequency, location information, etc.; based on the above information, the search engine can more accurately understand the relevance between the document and the query, and then achieve a more refined sorting.
[0018] In a possible implementation, based on the similarity scores between each of the N structured information and the corresponding first document, the N corresponding first documents are sorted second; alternatively, based on the ratio scores of the hit words in each of the N structured information in the corresponding hit fields, the N corresponding first documents are sorted second. In the embodiments of the present application, since the detailed information (such as the explanatory information) is structured information, it can be directly utilized by downstream services, and more explanatory information can be obtained more conveniently for more detailed secondary sorting of the hit documents, thereby performing a more in-depth analysis and optimization of the retrieval results. Specifically, compared with the string-concatenated explanatory information in the prior art, the structured explanatory information in the embodiments of the present application can be directly utilized more conveniently, provides transparency on how the records match the search conditions, and reveals how the scores of each condition affect the final ranking. The score can be calculated based on the similarity between the explanatory information and the corresponding candidate document (for example, in the field of semantic retrieval, the similarity between the query term and the document can be calculated using the vector space model), and the search results are sorted second based on the similarity score result; alternatively, the weight ratio of the hit words in the hit fields can also be used to find the hit documents with a high ratio, further refining the sorting logic to screen out more relevant search results and improve the adaptability of the search engine in various business scenarios.
[0019] Second, the embodiments of the present application provide a data retrieval device, which may include:
[0020] A first parsing unit, configured to receive a search condition and generate a first structured object, where the first structured object includes a query statement included in the search condition;
[0021] The first creation unit creates a scoring calculator based on the first structured object. The scoring calculator includes the first structured object and M document identifiers that match the first structured object, where M is a positive integer greater than 0.
[0022] The information acquisition unit acquires M first documents that match the M document identifiers, and generates scoring description information corresponding to each of the M first documents through the scoring component.
[0023] The first generation unit generates M second structured objects based on the scoring description information corresponding to each of the first documents and the first structured object, where one second structured object corresponds to one first document.
[0024] In the prior art, when a user conducts a content search, the search engine usually queries and returns preliminary search results according to the search conditions (for example, preliminary query results including one or more hit documents). When the user needs to further view one or more of the hit documents in the search results, the search engine needs to re-query and extract detailed information (such as the reason for the hit, hit details, detailed content, explanatory information, etc.) for the one or more hit documents. In this way, query waiting latency will occur, resulting in a poor user experience and also causing waste of system resources. To address this technical problem, in the embodiments of the present application, by setting the search engine to, after receiving the user's search conditions for the first time, not only obtain preliminary search results during the search process, but also further obtain more detailed and in-depth detailed information (such as explanatory information) for the search results, and pre-generate corresponding structured information. Thus, when the user needs to conduct an in-depth query on a certain hit document in the search results, the search engine can immediately feedback and present the pre-generated structured information to the user, greatly saving the latency caused by the need for the search engine in the prior art to query multiple times to feedback the results, saving system resources, and improving system performance. Specifically, in the embodiments of the present application, during the search process, the search engine creates a scoring calculator for the search conditions. The scoring calculator may include preliminary structured information (i.e., a first structured object including a query statement), the identifier of the hit document, and a scoring component. Further, the search engine obtains corresponding candidate documents according to the identifier of each hit document, generates scoring description information for each hit document through the scoring component, and further generates in-depth structured information (i.e., a second structured object) including the above first structured object and the scoring description information of each hit document, so as to accurately and efficiently obtain relevant information when the user needs to query later. In summary, in the embodiments of the present application, the required explanatory information can be obtained and saved when the hit document is first searched, and when the user conducts a further query, there is no need to perform repeated explanatory query operations. Therefore, the latency of obtaining explanatory information during the search process is greatly reduced, and the performance of the search is improved. It can be applied to services with a larger amount of data, increasing the applicability to more business scenarios.
[0025] In a possible implementation manner, the search engine includes an inverted index, and the inverted index includes a mapping relationship between a document identifier and the corresponding document;
[0026] The information acquisition unit is specifically configured to:
[0027] Search for M first documents corresponding to the M document identifiers from the inverted index.
[0028] In a possible implementation manner, the scoring description information corresponding to each of the first documents includes hit information and scoring information;
[0029] The information acquisition unit is specifically configured to:
[0030] For each scoring device, screen out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q;
[0031] Generate the hit information and scoring information corresponding to each first document in the M first documents through the scoring component, and the scoring information is the score of the matching degree between the corresponding first document and the search condition.
[0032] In a possible implementation manner, the hit information in the J first documents includes one or more of a hit field, a hit word, and a hit position that match the search condition in the corresponding first document.
[0033] In a possible implementation manner, the first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document of each of the Q first documents, the set of document identifier sequences of the one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document.
[0034] In a possible implementation manner, the first generation unit is specifically configured to:
[0035] Obtain K sub-query statements included in the query statement, where K is a positive integer greater than 0;
[0036] Based on the K sub-query statements, generate K third structured objects, where each third structured object includes a corresponding sub-query statement, and the first structured object includes the K third structured objects.
[0037] In a possible implementation manner, the first creation unit is specifically configured to:
[0038] Create K corresponding sub-scoring devices based on the K third structured objects, and each sub-scoring device includes the corresponding third structured object, L document identifiers that match the corresponding third structured object, and the scoring component;
[0039] Create the scoring device based on the K corresponding sub-scoring devices, where L is a positive integer greater than 0.
[0040] In a possible implementation manner, the device further includes:
[0041] A first sorting unit, for each score calculator, performs a first sorting on the corresponding first document based on the score description information;
[0042] A screening unit, based on the result of the first sorting, screens out N first documents whose score description information meets a preset condition from the first documents corresponding to each score calculator, where N is a positive integer less than or equal to M.
[0043] In a possible implementation, the search engine includes a preset storage component; for each of the first documents in the score calculator, the corresponding second structured object is saved to the storage component; the apparatus further includes:
[0044] An extraction unit, based on the document identifiers of the N first documents, retrieves the corresponding N second structured objects from the storage component;
[0045] An information generation unit, based on the N second structured objects, generates N corresponding structured information.
[0046] In a possible implementation, the apparatus further includes:
[0047] A second sorting unit, based on the N corresponding structured information, performs a second sorting on the N first documents, and outputs the sorting result to the search result.
[0048] In a possible implementation, the second sorting unit is specifically configured to:
[0049] Based on the similarity score between each of the N structured information and the corresponding first document, perform a second sorting on the N corresponding first documents;
[0050] Or, based on the occupancy ratio score of the hit words in each of the N structured information in the corresponding hit fields, perform a second sorting on the N corresponding first documents.
[0051] In a third aspect, an embodiment of the present application provides a computer storage medium, characterized in that the computer storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of claims 1-11 above is implemented.
[0052] In a fourth aspect, an embodiment of the present application provides a computer program, characterized in that the computer program includes instructions, and when the computer program is executed by a computer, the computer is caused to execute the method described in any one of claims 1-11. Description of the Drawings
[0053] To more clearly illustrate the technical solutions in the embodiments of the present application or the background art, the following will describe the drawings required to be used in the embodiments of the present application or the background art.
[0054] Figure 1A It is a schematic diagram of a system architecture provided by an embodiment of the present application.
[0055] Figure 1B It is another schematic diagram of a system architecture provided by an embodiment of the present application.
[0056] Figure 2 It is a schematic diagram of a search engine module provided by an embodiment of the present application.
[0057] Figure 3A It is a schematic flowchart of a search method provided by an embodiment of the present application.
[0058] Figure 3B It is a schematic flowchart of another search method provided by an embodiment of the present application.
[0059] Figure 4A It is a schematic diagram of a search step provided by an embodiment of the present application.
[0060] Figure 4B It is a schematic diagram of a score calculator structure provided by an embodiment of the present application.
[0061] Figure 4C It is another schematic diagram of a score calculator structure provided by an embodiment of the present application.
[0062] Figure 4D It is a schematic diagram of a nested structure document provided by an embodiment of the present application.
[0063] Figure 4E It is a schematic flowchart of a structure in a search engine provided by an embodiment of the present application.
[0064] Figure 4F It is a schematic diagram of an application scenario provided by an embodiment of the present application.
[0065] Figure 4G It is another schematic diagram of an application scenario provided by an embodiment of the present application.
[0066] Figure 4H It is still another schematic diagram of an application scenario provided by an embodiment of the present application.
[0067] Figure 5 It is a schematic diagram of the structure of another search device provided by an embodiment of the present application.
[0068] Figure 6A It is a hardware structure block diagram of an electronic device provided by an embodiment of the present application.
[0069] Figure 6B It is the software architecture of the electronic device provided by the embodiments of the present application. Specific embodiments
[0070] Next, the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.
[0071] The terms "first", "second", "third", "fourth", etc. in the specification, claims and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0072] Referring to "embodiment" herein means that a particular feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0073] The terms "component", "module", "system", etc. used in this specification are used to represent computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, an application running on a computing device and the computing device can both be components. One or more components can reside in a process and / or execution thread, and the components can be located on one computer and / or distributed between two or more computers. In addition, these components can execute from various computer-readable media on which various data structures are stored. For example, a component can communicate through local and / or remote processes according to a signal having one or more data packets (e.g., data from two components interacting with another component in a local system, a distributed system, and / or a network, e.g., the Internet interacting with other systems through a signal).
[0074] First, some terms in the present application are explained to facilitate understanding by those skilled in the art.
[0075] (1) Data Retrieval refers to the process of retrieving relevant or specific data from a large data collection through specific query conditions. It is an important part of the fields of computer science and information technology and is applied in search engines, database management, information retrieval systems, etc. The data retrieval process usually includes steps such as query parsing, index lookup, result ranking, and result return, aiming to quickly and accurately find the data required by users and meet the users' information needs.
[0076] (2) Lucene (Apache Lucene) is an open-source full-text search engine library provided in the form of a Java library, which can provide full-text search functions for application programs. It has powerful indexing and search functions, can handle a large amount of data, and at the same time provides a variety of scoring models and extension mechanisms. The core functions of Lucene include index construction, query parsing, search execution, and result sorting, etc., and are widely used in fields such as search engines, databases, and information retrieval systems. Through Lucene, developers can build high-performance and highly scalable search solutions to meet various complex search needs.
[0077] (3) Candidate Document, in the field of information retrieval, refers to a set of documents that are initially screened out from a large number of documents according to specific query conditions and may be relevant to the query. Candidate documents are the documents that are considered likely to meet the user's query after the index lookup is completed but before the final scoring and sorting. The size and quality of the candidate document set directly affect the efficiency and accuracy of the retrieval system. Usually, the retrieval system will use techniques such as inverted indexes to quickly find a set of candidate documents, and then use more refined scoring and sorting algorithms to find the final retrieval results from the candidate documents.
[0078] (4) Hit Document refers to the documents that meet the conditions found from the dataset according to the user's query conditions during the data retrieval process. Hit documents are usually sorted according to the scores of a certain scoring model and then returned to the user. It reflects the relevance between the document and the user's query and is the direct output of the core task of the retrieval system. In retrieval systems such as Lucene or Elasticsearch, hit documents usually contain document identifiers, scores, and other possible metadata information to help users understand and evaluate the quality and relevance of the retrieval results.
[0079] (5) Indexed Document refers to a document that has been indexed for quick retrieval in an information retrieval system. During the indexing process, the system parses the document content, extracts keywords, and establishes a mapping relationship from keywords to documents, usually stored in an inverted index. In this way, when a user submits a query, the system can quickly find the documents that match the query conditions by looking up the index, improving the retrieval efficiency and accuracy.
[0080] First, analyze and propose the specific technical problems to be solved by this application. In the prior art, the technologies related to data retrieval include the following solutions:
[0081] In the existing data retrieval technologies, such as Lucene or Elasticsearch, an Explain function is provided to help users understand the source of the scores of search results. Specifically, Lucene provides the IndexSearcher.explain() method, which can generate explanation information for a given query and document ID; while in Elasticsearch, by setting the explain parameter to true in the query request, explanation information can be generated for each returned document. These explanation information include the scores of the documents, the sources of the scores, and the detailed information on how the scores are calculated. Through these explanation information, developers can understand why certain documents are matched and can debug and optimize their queries.
[0082] Solution: Based on the method of search explanation in data retrieval, it can specifically include the following steps 1 and 2.
[0083] Step 1: Conduct a separate explanation query based on each returned record.
[0084] The search engine retrieves the search conditions and obtains multiple returned records. If explanations need to be returned, a separate explanation query needs to be executed for each returned document. Exemplarily, assume that 100 records are returned in the search results, then 101 queries need to be executed (1 result query + 100 explanation queries); in Lucene, the explanation query can be implemented through the IndexSearcher.explain(Query query, int doc) method. For each record, this method will be called once. Among the input parameters, query is the original query object, and doc is the ID of the document; after executing this method, an Explanation object will be returned, which contains the detailed explanation of the document score, including the calculation method of the score, the scoring model used, and the contributions of various scorers.
[0085] Step 2: The returned explanations are returned in a complete character description manner.
[0086] In the prior art, it is possible to perform an interpretive query on a hit document with a simple structure (for example, the hit document is not a nested document including multiple nested objects), obtain the interpretive information therefrom, and splice the interpretive information into a complete string and return it to the result display interface.
[0087] The above solution is currently mainly applicable to the scenario of performing an interpretive query based on the search results to obtain interpretive information after the search is completed, but there are also the following disadvantages:
[0088] Disadvantage 1:
[0089] The above method of interpretive query is usually applicable to some data retrieval services with a relatively small amount of data. However, for data retrieval services with a relatively large amount of data, since an interpretive query needs to be executed for each returned record, when the result set obtained after retrieval is large, it will cause a large number of additional queries, increasing the burden and response time of the system, resulting in the latency not meeting the requirements of real-time services; in addition, if the retrieval service is in a high-concurrency environment, each query requires an additional interpretive query, which will greatly affect the throughput of the system; finally, although the implementation method of the interpretive query can provide detailed interpretive information to help understand the query results, it has relatively large defects in terms of efficiency and scalability and cannot support actual business requirements, and is generally only used in the debugging and testing stages of search.
[0090] Disadvantage 2:
[0091] First of all, the information returned by the interpretive query will be spliced into a string, and the reason for the hit therein is not detailed enough (for example, the hit information of the sub-documents in the nested document is not included). Such interpretive information may not be rich enough, for example, it cannot provide sufficient debugging information for the debugger; secondly, since the string is presented as unstructured text, it is extremely difficult to extract structured information therefrom. Specifically, for complex nested objects, when retrieving nested sub-objects, it is also very difficult to know which sub-object in the nested object the current search is due to, and which attribute in the interpretive information caused the hit; the unstructured interpretive information returned by the interpretive query makes it difficult to extract structured information therefrom and is almost inapplicable to complex business logics, lacking flexibility and applicability.
[0092] Generally speaking, although the interpretive function is relatively useful, due to their latency and unstructured output format, it is very difficult to use them in real-time services, and they are currently usually used in the development and debugging stages. Moreover, the format and richness of the interpretive information may also not meet the needs of users to deeply understand and optimize queries.
[0093] To address the problem that the current interpretation query technology fails to meet the actual business requirements, achieve the goal of quickly obtaining detailed information and being applicable to complex real-time business tasks, and thus improve the speed and applicability of interpretation queries in data retrieval, considering the drawbacks of existing technologies, the technical problems actually to be solved in this application include one or more of the following three aspects:
[0094] 1. Reduce the latency of obtaining detailed information (such as interpretation information) (Drawback 1). Based on the current wide application of the interpretation query scheme in data retrieval, it partially solves the problem of obtaining interpretation information. However, since only multiple return records can be obtained after a result query, and obtaining interpretation information requires a separate interpretation query for each return record, when the data volume is large, multiple interpretation queries are often required, resulting in a long latency for obtaining interpretation information and affecting the data retrieval business. Therefore, an efficient method for obtaining interpretation information needs to be proposed to reduce the number of interpretation queries and the latency when obtaining interpretation information.
[0095] 2. Improve the practicality of detailed information (such as interpretation information) (Drawback 2). Based on concatenating strings to return interpretation information, it is currently mainly applicable to scenarios where professionals conduct testing and debugging. In existing technologies, the interpretation information obtained based on return records is usually concatenated into a string, which is unstructured information. It is difficult to clearly extract the required information from it for application in subsequent business logics. Downstream businesses cannot directly retrieve values for use, and it occupies too much memory, making it of no practical use in business scenarios where ordinary users want to obtain detailed information (such as the reason for a hit, details of a hit, detailed content, interpretation information, etc.).
[0096] 3. Improve the applicability of detailed information (such as interpretation information) (Drawback 3). Based on the interpretation information obtained through multiple interpretation queries, it is currently mainly applicable to scenarios of simple businesses. In existing technologies, the interpretation information obtained based on return records is usually concatenated into a string, which is unstructured information and can only be used to store the interpretation information of simple businesses and cannot be applied to complex business requirements (for example, the retrieved candidate document is a nested document, and this candidate document does not include the nested document); for example, when a user searches for "cat", it is necessary to retrieve a nested document including the word "cat", and at the same time, it is necessary to know that it is the first sub-object in the nested document that is hit, and it is hit because the tag_name in the first sub-object specifically includes "cat", and at the same time, the score of this hit condition can also be known. Therefore, a structure applicable to complex businesses is needed, which can be used to display the interpretation information of complex businesses and improve the applicability of interpretation information in data retrieval.
[0097] In summary, the existing data retrieval solutions cannot meet the actual business requirements for obtaining explanatory information. Therefore, the data retrieval method provided in this application is used to solve the above technical problems.
[0098] Based on the above-mentioned technical problems raised, and for the convenience of understanding the embodiments of this application, first, one of the system architectures on which the embodiments of this application are based will be described below. Please refer to Figure 1A , Figure 1A which is a schematic diagram of a system architecture provided by an embodiment of this application. The structure in this application may include Figure 1A the search engine 01, the search library 02 ( Figure 1A the search library 02 in
[0099] may include the web search database 22), and the terminal device 03.
[0100] Search library 02, which is used to store the data to be retrieved. The search library 02 can be a network search database 22, referring to database resources accessed through the network, which can be one or more of a distributed database, a cloud database, or any other type of database, and mainly provides data storage and retrieval functions through the network. Such a database usually consists of multiple database servers and supports remote data access and management. In the embodiments of the present application, no specific limitation is made; after receiving the Query object obtained by parsing the search conditions by the parser, the search library 02 will quickly locate and retrieve the matching documents in the inverted index library according to the query conditions, and provide accurate search results for users; at the same time, the search library 02 also supports operations such as updating, deleting, and adding index data to ensure the timeliness and accuracy of the index data.
[0101] Terminal device 03 can be a smart phone (mobile phone), a tablet computer (Tablet), a personal computer (Personal Computer, PC), a laptop computer (Laptop), a network terminal (NetworkTerminal), a multifunction printer (Multifunction Printer, MFP), a smart TV (Smart TV), a game console (Game Console), an in-vehicle infotainment system (In-Vehicle Infotainment, IVI), a smart watch (SmartWatch), a smart bracelet (Smart Bracelet), and other intelligent devices with search functions, etc. The search engines in these devices can be Lucene or Elasticsearch, which can provide users with convenient, fast, and accurate search services and meet the information retrieval functions of terminal devices under different scenarios and requirements, etc.
[0102] For different business scenarios, the search engine 01 is also different. If the business scenario of the search engine 01 is to retrieve local data, please refer to Figure 1B , where Figure 1B is another schematic diagram of the system architecture provided by the embodiments of the present application. The structure in the present application can include Figure 1B the terminal device 03-A in Figure 1A (please refer to the description of the embodiment of the terminal device 03 in Figure 1A(The embodiments in the search engine 01 are described and will not be elaborated here), a local search database 21. Among them, the user operates the local data 04 in the terminal device 03-A (for example, it can be an APP that supports local search services). After inputting the search condition, the search condition is parsed by the search engine 01-A, and the retrieval result is obtained from the local search database 21. Then, the search engine 01-A performs a preset operation (for example, it can be scoring the candidate documents, screening out the hit documents based on the scoring result, storing the detailed information of the hit documents (such as the explanation information), and constructing the explanation information of the hit documents into structured information during the result extraction stage). The search engine 01-A returns the retrieval result to the result display interface in the local data 04.
[0103] Among them, the local data 04 can be various types of software or applications, such as document editing software (DocumentEditor), spreadsheet software (Spreadsheet Software), image editing software (Image EditingSoftware), video player software (Video Player), music player software (Music Player), email client (Email Client), contacts (Contacts), calendar application (Calendar App), maps and navigation software (Mapsand Navigation Software), database management system (Database Management System, DBMS), file manager (File Manager), and other software that can generate, store, or manage data. Through search engines such as Lucene or Elasticsearch, users can conveniently and quickly search and retrieve the required local data in these software or applications to meet different information retrieval needs, etc.
[0104] It can be understood that Figure 1A - Figure 1B the system architecture in is only one or more exemplary implementation manners in the embodiments of the present application. The structure in the embodiments of the present application includes but is not limited to the above system architecture.
[0105] Based on the above system architecture, the embodiments of the present application provide a search engine 01 applied to the above system architecture. Please refer to Figure 2 , Figure 2 which is a schematic diagram of a search engine module provided by the embodiments of the present application. The search engine 01-A may include a parsing module 1001, an index reading module 1002, a scoring module 1003, a storage module 1004, and a result display module 1005. Among them:
[0106] Parsing Module 1001: Responsible for parsing the search conditions into a Query object that Lucene can understand. The Query object can include multiple nested sub-query statement objects. Specifically, during the parsing process, the Parsing Module 1001 may perform lexical analysis, syntactic analysis, and semantic analysis to ensure understanding of the query intent of the search conditions. By parsing the search conditions, the Parsing Module 1001 can identify important keywords, phrases, and other elements, and then convert these elements into a structured query object (such as a nested query or a combined query, where each sub-query in the combined query can be related by words such as or, and). This query object is the basis for Lucene to perform the search and defines the parameters and conditions of the search. The Parsing Module 1001 can handle various complex queries, including boolean queries, phrase queries, wildcard queries, etc., providing accurate guidance for the subsequent retrieval process.
[0107] Index Reading Module 1002: Responsible for reading the document data that matches the query from the inverted index library. This includes the inverted index. After receiving a query object, the Index Reading Module 1002 will look for documents that match the query conditions in the inverted index (such as documents that include the characters of the query conditions) and mark them as candidate documents. The inverted index (InvertedIndex) is a data structure designed for search engines, used to store and retrieve keywords in a large number of documents and quickly find documents that contain specific keywords. In the inverted index library, each unique keyword is associated with a list of documents, and this list contains the identifiers of all documents that contain the keyword. Through the inverted index library, the search engine can find documents that match the query conditions in a short time. The Index Reading Module 1002 will quickly locate relevant documents in the inverted index library according to the conditions in the query object and read the data of these documents. It can process a large amount of data and return results in a short time. In addition, the Index Reading Module 1002 will also work in coordination with the Scoring Module 1003 to ensure that the returned documents are highly relevant to the query conditions.
[0108] Scoring Module 1003: Responsible for calculating the matching degree between each retrieved document and the query, and assigning a score to each document. The scoring module may include a scorer, and the scorer may further include a query statement object, a scoring component, and a document identifier (such as the ID of the candidate document stored in the inverted index obtained from the index reading module 1002). Among them, the scoring component adopts a specific scoring model, such as the TF-IDF model or the BM25 model, to calculate the score based on the content of the document and the query conditions. Through the score, the scoring module 1003 can measure the relevance between the document and the query. The higher the score, the higher the matching degree between the document and the query. In the embodiments of the present application, the Scorer is modified. The Scorer will save the hit fields, hit words, source query, score, and the position of the hit object of each candidate document in a corresponding explainquery object in detail. In addition, when it is detected that a certain candidate document is a nested document, the Scorer will also save the position of the hit object of the candidate document and the position of the nested sub-object in the above query object. Then, the query object corresponding to each candidate document is stored in a preset collector for subsequent retrieval and construction of structured information (such as explanation information).
[0109] Storage Module 1004: This module is a core component of the search engine, including a preset storage component, which is responsible for managing and storing all second structured objects. These second structured objects are structured objects generated for the detailed information (such as explanation information) of each candidate document. The above storage component is responsible for maintaining the storage logic of the second structured objects within this module. Before the scoring module 1003 screens out the documents that meet the preset conditions, each second structured object is stored in the preset storage component in the storage module 1004 and retrieved after the collection of the hit documents is completed.
[0110] Result Display Module 1005: Responsible for retrieving the paging results in the fetch phase and constructing the retrieved detailed information (such as explanation information) into structured information. After the scoring and screening of the candidate documents are completed, the result display module can respond quickly and retrieve and provide the corresponding second structured object from the preset storage component in the storage module 1004 according to the ID of the retrieved document. Specifically, based on the ID of the retrieved document, in the above preset storage component, find the query object corresponding to the ID of the retrieved document and construct the query object into a structured information. Compared with the prior art in which the explanation information is spliced into a string, the structured information can be displayed more intuitively for the user to view. And the structured information can be directly used by downstream services, reducing the memory occupancy rate compared with the string. In addition, the structured information can explain documents including complex nested objects, and based on this feature, it can be applied to more complex business scenarios.
[0111] It is understandable that Figure 2 the structure of the search engine in [] is only an exemplary implementation manner in the embodiments of the present application. The structure of the search engine in the embodiments of the present application includes but is not limited to the above structure.
[0112] Based on Figure 1A - Figure 1B the provided system architecture, and Figure 2 the modules of the provided search engine, combined with the explanatory search method provided in the present application, the technical problems proposed in the present application are specifically analyzed and solved.
[0113] See Figure 3A , Figure 3A which is a schematic flowchart of a search method provided by an embodiment of the present application. This method can be applied to the system architecture described in the above Figure 1A - Figure 1B wherein the search engine 01 can be used to support and execute the method flow steps S300 - S303 shown in Figure 3A . The following will be described from the side of the search engine 01. This method may include the following steps S300 - S303. Figure 3A
[0114] Step S300: Receive a search condition and generate a first structured object.
[0115] Specifically, when the search engine receives the user's search condition, it will first parse and process the search condition and convert it into a query statement object that the search engine can understand and execute (for example, the first structured object converted by the lucene search engine from the query statement included in the search condition); the query statement object may include a query statement, as well as some possible additional conditions and parameters, such as filtering conditions, sorting conditions, etc. The query statement object is the core of the search process, defining the scope and manner of the search. Exemplarily, through the generated query statement object, the Lucene search engine can accurately locate and retrieve documents (such as candidate documents) in the index library that match the search condition, thereby quickly and accurately retrieving the search results.
[0116] In a possible implementation, K sub-query statements included in the query statement are obtained, where K is a positive integer greater than 0; based on the above K sub-query statements, K third structured objects are generated. Each third structured object includes a corresponding sub-query statement, and the first structured object includes the above K third structured objects. Specifically, if the search condition includes K sub-query statements, then the query1 object (i.e., the query statement object) converted based on the search condition may include K query2 objects (i.e., sub-query statement objects), where the above K query2 objects are sub-objects stored in the nested object query1 object. Further, the independent structuring of each sub-query statement enables the search engine to process each sub-query statement separately (in addition, in the query statement, data can be filtered by combining multiple conditions for querying, for example, using relational words such as logical operators "OR" and "AND" to combine multiple sub-query statements to form a combined query), so as to perform nested queries, query sub-documents under the nested object, and return the parent documents hit due to the sub-document search conditions, improving the relevance and accuracy of search results; at the same time, decomposing complex queries into combinations of multiple sub-query statements enables the search engine to adopt different strategies and algorithms for different parts of the query conditions, enhancing the adaptability of the search engine to complex information requirements and the ability to process complex queries, which can help improve the efficiency of information retrieval, especially in the case of processing queries containing logical operators (such as AND / OR / NOT) or nested queries, making the search results returned by data retrieval more accurate.
[0117] Step S301: Create a scoring calculator based on the first structured object.
[0118] Specifically, create a scoring calculator based on the first structured object. The scoring calculator may include one or more of the first structured object, M document identifiers matching the query statement, and a scoring component, where M is a positive integer greater than 0. After the search engine parses the query statement in the search condition to generate a query statement object (i.e., the first structured object), a scoring calculator (such as the scorer in the lucene search engine) is created based on the query statement object. The scoring calculator may include M document identifiers (i.e., the IDs of candidate documents) matching the query statement object and a scoring component (such as the scoring component in the scorer). Exemplarily, the corresponding M candidate documents can be found based on the above M document identifiers, and then the above M candidate documents are scored and filtered based on the scoring calculator.
[0119] In a possible implementation, K corresponding sub-scorers are created based on K third structured objects. Each sub-scorer includes a corresponding third structured object, L document identifiers that match the corresponding third structured object, and a scoring component. A scorer is created based on the above K corresponding sub-scorers, where L is a positive integer greater than 0. Specifically, in the embodiments of the present application, a corresponding sub-scorer can be created based on each sub-query statement in the query statement. These sub-scorers also include corresponding sub-query statements, scoring components, and document identifiers. The scorer generated based on the query statement includes all the sub-scorers generated based on the sub-query statements. Exemplarily, please refer to Figure 4A Step 3 in a schematic diagram of a search step provided, and Figure 4B A schematic diagram of a scorer structure provided. If the query1 object parsed from the search condition conversion is a nested object (i.e., the query1 object is a nested object structure, including N query2 objects, where the above N query2 objects are sub-objects generated based on the sub-query statements in the query statement), first create a corresponding parentscorer based on the query1 object. The parentscorer can include the query1 object, N childscorers (i.e., sub-scorers) generated based on the query2 objects, the ID information of each candidate document in the parentscorer, and a scoring component. Further, a corresponding sub-scorer can be created for each sub-query statement. The above sub-scorers exist in the corresponding main scorer (i.e., the scorer corresponding to the first structured object generated based on the query statement). Each sub-scorer operates on a specific query segment (for example, it can perform individual scoring and sorting), enabling the search engine to better interpret and optimize the query results, while improving the ability to handle nested queries and multi-condition queries.
[0120] Step S302: Obtain M first documents that match the M document identifiers, and generate scoring description information corresponding to each of the M first documents through the scoring component.
[0121] Specifically, the search engine can find the corresponding M candidate documents based on the M document identifiers (for example, it can find the candidate documents corresponding to each ID from the inverted index), and then obtain the scoring description information corresponding to each candidate document based on the scoring component (for example, the scoring description information can include hit information and scoring information. The hit information can be one or more of the hit words, hit fields, and hit positions that the scoring component obtains based on each candidate document and matches the search condition. The scoring information can be the score of the scoring component based on the matching degree of each candidate document and the query statement object).
[0122] In a possible implementation, the search engine includes an inverted index, and the inverted index includes the mapping relationship between the document identifier and the corresponding document; obtaining M first documents matching M document identifiers includes: searching for the M first documents corresponding to the M document identifiers in the inverted index. Specifically, in the search engine, data retrieval can be implemented through a query object query1 generated by a query statement. A key step in this process is to use the identifier (i.e., ID) of the candidate document to retrieve in the inverted index. The inverted index is a data structure in Lucene. In the inverted index, each document has a unique ID as its identification code in the document collection; Exemplarily, when the scorer determines the candidate documents matching the query conditions in Lucene, it can use the ID of each document to retrieve the specific candidate documents from the inverted index. In addition, through these IDs, the scorer can query the index and retrieve the document information related to each keyword in the query statement, including their storage locations, the frequencies and distributions of the keywords in the documents, etc.; this information can be used as a scoring component in the scorer to calculate the relevance score of the document, and then determine its ranking in the search results.
[0123] In a possible implementation, the scoring description information corresponding to each first document includes hit information and scoring information; the scoring component generates the scoring description information corresponding to each of the M first documents, including: for the scorer, screening out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q; the scoring component generates the hit information and scoring information corresponding to each of the M first documents, and the scoring information is the score of the matching degree between the corresponding first document and the search condition. Specifically, in the embodiments of the present application, the scorer can be used to identify nested and non-nested document structures (i.e., when processing documents containing multi-level data, such as documents in nested JSON or XML formats, the scorer can distinguish and identify whether the document is a nested structure), which enables the search engine to process complex queries more accurately; in addition, the scoring component in the scorer can generate the scoring description information (i.e., hit information and scoring information) corresponding to each candidate document, and the scoring information is the score based on the matching degree between the corresponding candidate document and the search condition, so as to improve the accuracy of search results; exemplarily, the score of each candidate document can be that the scoring component in the scorer obtains the hit information of the corresponding candidate document based on the inverted index, and scores each candidate document based on the hit information. For example, it can be calculated by TF-IDF or BM25 to calculate the relevance score of a document for a query; further, first, the scoring component in the Scorer creates weights for multiple terms included in each search condition in the query statement. Based on the weights of the terms, the frequency in the document, the document length, and other relevance signals, the scoring component scores each candidate document one by one until all relevant documents are scored.
[0124] In a possible implementation, the hit information in the J first documents includes one or more of the hit fields, hit words, and hit positions that match the search condition in the corresponding first document. Specifically, for each candidate document that matches the query statement, the hit information of each candidate document may include, but is not limited to: one or more of the hit fields (e.g., the field in the candidate document where the query term appears), the hit words (e.g., the words or phrases that actually match the query statement), and the hit positions (e.g., the specific positions of the hit words parsed from the query statement in the candidate document or in a specific field). By collecting the hit information of each candidate document in the embodiments of the present application, it is possible to indicate the specific position where the query term appears in the document based on this hit information, and it can also be used to calculate the scoring information together with other information in the matching situation (e.g., the document frequency (DF) of the term, the term frequency (TF), the document length, the query term weight, etc.). Moreover, this hit information can be used as a type of explanatory information to indicate the specific hit situation of the query statement in the corresponding document.
[0125] In a possible implementation, the first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document of each of the Q first documents, the set of document identifier sequences of one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document. Specifically, first, it is confirmed whether the candidate documents include candidate documents of the nested type (i.e., the second documents). In a search engine, the documents may have a complex nested structure to support richer information expression. Exemplarily, a report may contain multiple chapters, and each chapter may contain multiple paragraphs; or, a video may include multiple paragraphs or highlights, and each paragraph may include multiple frames. Such nested structure documents can support richer content. In the embodiments of the present application, the nested structure document may be a nested structure including a main document and multiple sub-documents. If the Q candidate documents are nested structure documents, then the hit information that can be obtained by the modified scorer is such that each scorer (e.g., a scorer generated for the query statement or a sub-scorer generated for a sub-query statement) can obtain the ID (i.e., the document identifier) of the main document in each candidate document, the set of ID sequences of one or more sub-documents, the ID of the sub-document hit in the main document, and the hit position of the hit sub-document as the hit information; further, please refer to Figure 4A Step 3 in the schematic diagram of a search step shown in Figure 4C and another schematic diagram of the scorer structure shown in Figure 4CIn the scorer shown, the main document IDs hit by frames.ocr_text are 10 and 20. The sub-documents with ID 10 are 7, 8, and 9, and among them, 8 is the sub-document ID hit by frames.ocr_text. By parsing this information, it can be known that this condition hits the second one among the frames nested objects in the main document. In the embodiments of the present application, by transforming the scorer, the explanation information of complex nested documents can be obtained, enabling the embodiments of the present application to be applied to more application scenarios. It can be understood that the nested document can be a nested structure composed of multiple layers of text, or a nested document composed of each specific frame included in the video, or a nested structure of rich text type. The embodiments of the present application do not make specific limitations on this, as long as the document can be reflected as a nested structure.
[0126] Step S303: Generate M second structured objects based on the scoring description information and the first structured object corresponding to each first document.
[0127] Specifically, please refer to Figure 4A Step 4 in. Before collecting the candidate documents in each scorer (that is, before sorting and filtering the scoring information corresponding to each candidate document), the scoring description information of each candidate document and the corresponding query statement object (such as the first structured object generated by the query statement, or the third structured object generated by the sub-query statement) are used as detailed information (such as explanation information) to generate the corresponding explanation information object (that is, the second structured object). Exemplarily, each explanation information object can correspond to a candidate document, and each explanation information object can include the query object (that is, the first structured object or the third structured object) converted from the query statement corresponding to the scorer, the scoring information and the hit information of the corresponding candidate document. In the embodiments of the present application, by storing the explanation information of each candidate document as an explanation object during the data retrieval process, when the explanation information needs to be obtained subsequently, it is not necessary to perform multiple explanation queries again, and the explanation information can be obtained only by finding the explainquery object corresponding to the target candidate document.
[0128] In a possible implementation manner, a storage method for the detailed information (such as explanation information) object of complex nested documents is implemented based on the Lucene search engine: the nested sub-objects can be stored as independent index documents, and the construction method of its fields is in the form of "parent field"."sub field"; in this way, the sub-objects also have independent retrieval capabilities; please refer to Figure 4DA schematic diagram of a nested structure document. When retrieving the parent document through the sub-document, fields such as frames.tag_name can be retrieved to first retrieve the sub-document, and then find the first parent document backward (i.e., the parent document of this record). In the fetcher, based on the position of the parent document and the number of sub-documents recorded in the frames field, the nested objects in the frames field are restored sequentially forward.
[0129] Optionally, when the above method steps S300 - S303 obtain the detailed information (such as explanatory information) of each candidate document during the data retrieval process, and generate corresponding structured objects based on the explanatory information of each candidate document, the following can be referred to subsequently Figure 4E , before collecting the hit documents during the search process, by injecting an explanation collector, various hit conditions of each document are collected, and an explanation object is constructed during the final fetch stage; compared with the prior art, explanations can be obtained simultaneously with just one search, as Figure 3B shown, the following steps S400 - step S403 can also be included:
[0130] Step S400: Save the second structured objects of the candidate documents in each scorer to a preset storage component.
[0131] Specifically, refer to Figure 4A step 4 in. During the data retrieval process, the detailed information (such as explanatory information) of each candidate document is obtained, and after generating the corresponding structured object, each explanatory information object (i.e., the second structured object) can be saved to explainquerycolletcor (i.e., the preset storage component). After collecting all the candidate documents, the hit documents that meet the preset conditions can be screened out from the candidate documents based on the scores. By pre-storing the explanatory information of all the candidate documents before screening out the hit documents, the required explanatory information can be retrieved from explainquerycolletcor when outputting the search results, without having to return to each hit document in the search results for explanation query, thus avoiding waste of resources and reducing latency.
[0132] Step S401: Based on the result of the first sorting, screen out N first documents from the first documents corresponding to each scorer whose scoring information meets the preset conditions.
[0133] Specifically, for each scoring device, the corresponding first document is sorted first based on the scoring description information; based on the result of the first sorting, N first documents whose scoring description information meets the preset conditions are selected from the first documents corresponding to each scoring device, where N is a positive integer less than or equal to M. Exemplarily, after the scoring component calculates the scoring information of the candidate documents in each scoring device (such as the scorer in Lucene), the corresponding candidate documents are sorted based on the level of the scoring information, so that the candidate documents with higher relevance are ranked in the front; then, based on the preset conditions, the candidate documents in each scoring device are screened, and N first documents whose scoring information meets the preset conditions are selected, where N is a positive integer less than or equal to M. Through this process, documents with higher scores, better matching the search conditions, and meeting the requirements of the preset conditions are selected from the original candidate document set as the hit documents; Exemplarily, if the business party only needs the top ten documents with the highest accuracy as the output result, all candidate documents can be sorted by the scoring information, and then by setting paging display, only the first ten are displayed, so as to meet the business requirements. In the embodiment of the present application, by comparing with the preset scoring conditions, the search engine can effectively narrow the search scope, improve the search efficiency and the accuracy of the results, so as to meet more business requirements.
[0134] Step S402: Based on the document identifiers of the N first documents, the corresponding N second structured objects are retrieved from the storage component to generate N corresponding structured information.
[0135] Specifically, after sorting the candidate documents based on the scoring information, the candidate document with the highest relevance to the corresponding query statement is displayed in the front, and then the N candidate documents with the highest relevance are selected according to the preset conditions as the hit documents. Exemplarily, refer to Figure 4A Step 6 in. After screening out the hit documents that meet the business requirements through scoring, in the fetch stage (i.e., the stage of fetching the paging results), the corresponding hit documents can be found from the explainquerycollector through the IDs of the hit documents. Exemplarily, refer to Figure 4AIn step 6, during the fetch phase (i.e., the phase of retrieving the paged results), since the document IDs stored in the explain query collector and their corresponding explain queries form a dictionary structure (such as a hash set), where the key is the ID and the value is the explain query, in the fetch phase, it is not necessary to iterate through each explain query. Instead, the corresponding explain query can be directly found through the document ID, and then the explanatory information can be extracted from each explain query (for example, it can include a query object constructed based on the corresponding search conditions, hit fields, hit words, source query, score, location of the hit document or location of nested hit documents). The extracted explanatory information is organized in a specific format (such as XML, JSON, or a database table). If there are hierarchical or linked relationships between the target hit documents (for example, the hit document is a nested structure), these relationships are correctly expressed in the structured information to better describe the explanatory information, and downstream services can thus utilize this information to perform more advanced logic.
[0136] In a possible implementation manner, based on the structured information constructed from the detailed information (such as explanatory information) of each hit document described above, on the search result page, the user can be prompted with the detailed information of the record (such as the reason for the hit, hit details, detailed content, explanatory information, etc.). Exemplarily, when the searched picture is hit because of text, a text recognition corner mark can be displayed; when the searched picture is hit because of a label, a label corner mark can be displayed; when the picture is hit because of a location, a location corner mark can be displayed. Taking contacts as an example, there is often not enough space to display the index information of contacts in the search result display. Therefore, through the explanation, corresponding data can be displayed according to different hit reasons during the display. The explanation includes the hit fields and hit words. Through this information, the user can obtain the position information of the hit words in the original text, and based on this information, the hit words can be highlighted during the search display to prompt the user. For the case where the input word is not a direct hit in the original text, the explanation provides the position information of the hit word in the original text. For example, in scenarios such as hitting by the first letter, pinyin, synonyms, etc., compared with the traditional string inclusion explanatory information that cannot accurately display the hit reason, the embodiments of the present application can achieve high - light prompting.
[0137] In a possible implementation, through the interpretation of nested objects, more advanced display effects can be applied. Specifically, for video and highlight retrieval: a video can contain multiple nested sub - paragraphs, and a highlight includes multiple sub - photos and sub - video clips; the content in each nested sub - paragraph of the video can be searched; for example, tags, semantic vectors, captions, etc. can be used to return the video; at the same time, through the interpretation, it will return which attribute of which field is hit; further, assume there is a video. When a user searches for the word "protection" in a certain piece of text in a certain frame of the video based on local data, if the keyword exists in a certain clip of a certain video, the video including the keyword can be returned. At the same time, through the interpretation, it can be known that it is the text in the second paragraph of the video that is hit; in this way, the retrieved hit position can be directly displayed as the cover, and playback can continue from this position; compared with existing video retrieval, it can dynamically display the cover of the hit position and the hit content, and implement the jump function.
[0138] For the retrieval of rich - text structures: the rich - text structure can be saved as a complex object, which contains paragraphs, tables, pictures, etc.; when searching for a search term, the rich - text document can be returned, and detailed information (interpretation information, such as which table, which paragraph, and which picture in the rich - text is hit) can be returned. Through this information, it is possible to directly jump to the corresponding position; for note applications, it can be analogous to the above - mentioned rich - text structure. Notes contain multiple elements: titles, text, handwriting, attachments, tables, recordings, pictures, etc., which can be constructed as a complex nested object; the search can be performed on all elements of the note. When a hit occurs, through the interpretation, it can be known which element and which one is hit. In this way, when presenting the result display interface to the user finally, it can be directly positioned to the specified position and a prompt can be given.
[0139] Step S403: Perform a second sorting on the N first documents based on the N corresponding structured information.
[0140] Generate N corresponding structured information based on the N second structured objects.
[0141] Specifically, in a possible implementation, secondary filtering and sorting can be performed through detailed information (such as explanatory information); the explanatory information may include which search condition the record comes from and the score of that condition. Exemplarily, in the field of semantic retrieval, similarity scores between search terms and retrieved documents are often calculated using vector calculations; after recall, the similarity score for a specified semantic retrieval condition can be obtained through explanation, and based on this score, secondary sorting and relevance filtering can be performed. Or secondary sorting can be performed on the retrieved documents based on the weight ratio score; for example, when the user searches for "ni", it can be learned through explanation that two title fields are hit, namely title1 and title2. Secondary sorting can be performed by comparing the length ratios of the hit word "ni" in title1 and title2, and the document with a shorter title can be ranked ahead.
[0142] Based on Figure 3A - Figure 3B the steps of the embodiments described in Figure 4E , when the search engine receives the search condition input by the user, it parses the search condition and determines whether the query statement included in the search condition is a nested query. A nested query is used to query sub-documents under a nested object, and the parent documents hit due to the sub-document search condition are returned; a non-nested query is used to query non-nested objects or the parent document part of a nested object; next, the cases of nested queries and non-nested queries will be introduced in detail, where
[0143] the case of nested queries may include the following steps S511 - step S517:
[0144] Step S511: The search engine parses the query statement included in the search condition. The query statement may include one or more sub-query statements, and the independent structuring of each sub-query statement enables the search engine to process each sub-query statement separately. In addition, in the query statement, data can be filtered according to multiple condition combinations. For example, logical operators such as "OR" and "AND" are used to combine multiple sub-query statements to form a combined query.
[0145] Step S512: Based on the query statement, a query statement object query1 that the search engine can understand and execute is generated (for example, the structured object converted from the query statement included in the search condition by the lucene search engine), and corresponding sub-query statement objects are generated for each sub-query statement. The query statement object query1 may include multiple sub-query statement objects query2. The main document matched by the query statement object query1 can be the first parent document found forward from the sub-documents hit by multiple sub-query statement objects query2, which is the parent document of this search record, to obtain a candidate document list.
[0146] Step S513: Create a corresponding sub-scorer Childids scorer for each sub-query statement. The existing corresponding main scorer parent scorer (i.e., the scorer based on the query statement object query1) includes the Childidsscorer sub-scorer. Each scorer in the main scorer parent scorer and the sub-scorer Childids scorer can include the query statement object of each candidate document and the document identifier (i.e., document ID) of each candidate document. Based on each document identifier, the corresponding candidate document is obtained, and the scoring component generates corresponding scoring description information for each candidate document (for example, it can include the hit field, hit word, score information, and the location of the hit object (nested sub-object)).
[0147] Step S514: Use the scoring description information of each candidate document and the corresponding query statement object query (such as the first structured object generated by the query statement or the third structured object generated by the sub-query statement) as the explanation information to generate a corresponding explanation information object explainquery.
[0148] Step S515: Before collecting the candidate documents in each scorer (for example, it can be the main scorer parent scorer or the sub-scorer Childids scorer) (before sorting and filtering the scoring information corresponding to each candidate document), store each explanation information object explainquery in the preset storage component Explain query collecor.
[0149] Step S516: After sorting the candidate documents based on the scoring information, the candidate document with the highest relevance to the corresponding query statement is displayed first, and then the top N candidate documents with the highest relevance are filtered out according to the preset conditions as the hit documents.
[0150] Step S517: In the stage of outputting the search results, based on the ID of the hit document in the search results, take out the corresponding explanation information object explainquery from the above-mentioned preset storage component Explain query collecor to generate the corresponding structured information of the explanation. Compared with the existing storage method of explanation information, which concatenates the explanation information into a string and cannot be directly retrieved, in the embodiment of the present application, the explanation information is constructed as structured information, which is more convenient for downstream services to directly retrieve.
[0151] The non-nested query situation can include the following steps S521 - S527:
[0152] Step S521: The search engine parses the query statement included in the search condition.
[0153] Step S522: Generate a query statement object query that can be understood and executed by the search engine based on the query statement (for example, the first structured object converted by the lucene search engine from the query statement included in the search condition). The query statement object query can only query the documents of non-nested objects or the parent document part of nested objects to obtain a list of candidate documents.
[0154] Step S523: Generate a corresponding scorer based on the query statement object query. The scorer can include the query statement object query of each candidate document and the document identifier (i.e., document ID) of each candidate document. Based on each document identifier, the corresponding candidate document is obtained. The scoring component generates corresponding scoring description information based on each candidate document (for example, it can include the hit field, hit word, score information, and the location of the hit object (nested sub-object)).
[0155] Step S524: Use the scoring description information of each candidate document and the corresponding query statement object query (for example, the first structured object generated by the query statement or the third structured object generated by the sub-query statement) as the explanation information to generate a corresponding explanation information object explainquery.
[0156] Step S525: Before collecting the candidate documents in each scorer (that is, before sorting and filtering the scoring information corresponding to each candidate document), store each explanation information object explainquery in the preset storage component Explain query collecor.
[0157] Step S526: After sorting the candidate documents based on the scoring information, the candidate document with the highest relevance to the corresponding query statement is displayed first, and then the top N candidate documents with the highest relevance are filtered out according to the preset conditions as the hit documents.
[0158] Step S527: In the stage of outputting the search results, based on the ID of the hit document in the search results, the corresponding explanation information object explainquery is retrieved from the above-mentioned preset storage component Explain query collecor to generate the corresponding structured information of the explanation. Compared with the existing technology, the storage method of the explanation information is to splice the explanation information into a string and cannot be directly retrieved. In the embodiment of the present application, the explanation information is constructed as structured information, which is more convenient for downstream services to directly retrieve.
[0159] For ease of understanding Figure 4EFor the steps of the embodiments described, the following are exemplary scenarios to which the data retrieval method in this application can be applied, which may include the following three scenarios:
[0160] Scenario 1: On the search result page, prompt the reason for the record hit:
[0161] Since the detailed information (such as explanatory information) obtained in the embodiments of this application is structured rather than concatenated in string form, it is possible to better obtain the required explanatory information. After retrieving based on the search conditions, it is possible to more accurately perform special marking on the detailed information (such as the reason for the hit, the details of the hit, the detailed content, the explanatory information, etc.). For example, when the search condition entered by the user is the two words "Hello", and finally a picture including this text is obtained in the search result, the retrieved picture (i.e., a certain hit document in the search result) is because the text (i.e., the first structured object or the third structured object generated based on the search condition) hits; when the search condition entered by the user is the two words "Happy", if the retrieved picture (i.e., the hit document in the search result) is because of the label hit, the label corner mark can be displayed because the search engine detects that the picture is hit because of the label related to "Happy". This corner mark can be a small icon or a visual hint indicating that the picture contains the text searched by the user. This way enhances the intuitiveness of the user interface, enabling the user to easily identify which pictures are retrieved because of the label information; similarly, if a picture is retrieved because the location information matches the search condition, for example, the user searches for a specific geographical location as the search condition, then a "location corner mark" will be displayed next to the picture in the search result. This corner mark usually takes the form of a map pin or an icon of the location, indicating that the picture is related to a specific location. The location corner mark helps the user identify the pictures related to the search condition due to geographical information, further providing a direct visualization association indicating that there is an actual connection between the picture content and a certain location.
[0162] Further, in the address book application, taking the search with the contact name as the search condition as an example, there is often not enough space to display the index information of the contacts in the display of search results. Therefore, through the explanatory information, corresponding data can be displayed based on different hit reasons during display. For example, when the user searches for the name "Zhang San" as the search condition, the search engine may return contact records containing this name (i.e., the hit documents in the search results). Next to each record, the system can display a small part of the explanatory information (i.e., the structured information generated based on the second structured object), such as "Name field match: Zhang San", which means that the hit word "Zhang San" exactly appears in the "Name" field of the contact record. Through such explanatory information, the user not only knows why this record is retrieved but also can understand the specific position of the hit word "Zhang San" in the original contact data. In addition, the search engine can further emphasize this match through the highlighting function. In the display of search results, the name "Zhang San" can be highlighted to distinguish it from other texts to prompt the user.
[0163] Scenario 2: Secondary filtering and sorting of hit documents based on explanatory information:
[0164] When the explanatory information is more comprehensive structured information, during secondary sorting, the explanatory information can be better utilized based on this structure, and more explanatory information can be obtained for in-depth screening to screen out the hit documents with a matching degree to the query condition. The explanatory information (i.e., the structured information of the explanatory information generated based on the second structured object) in the embodiments of the present application may include which search condition the record or the nested document in the record comes from (i.e., which query statement or sub-query statement the hit document of a certain nested structure in the search results is matched based on), and the scoring information of this condition. In the field of semantic retrieval, the similarity score between the search term and the hit document can be obtained through vector calculation; after recall, through explanation, the similarity score of the specified semantic retrieval condition can be obtained. Based on this score, secondary sorting and relevance filtering can be performed, thereby improving the accuracy of data retrieval. For example, when the search condition included in the obtained explanatory information is "renewable energy", the system will convert this search condition into a semantic vector. Then, the system will calculate the similarity score between this vector and the document vectors stored in the database, which have also been converted into vector form. This score reflects the proximity degree between the search term and the document at the semantic level, and a high score means a high correlation; once all the documents are given a similarity score relative to the search term, the system will perform a "recall" operation, that is, select a group of documents with the highest scores to form a preliminary search result set. After the recall stage, the search results can be further sorted and filtered for relevance.
[0165] The search results can also be sorted and filtered a second time based on the weight of the length ratio of the hit words included in the explanation information in the hit fields. For example, if the title fields of two documents, document A and document B, both contain "ni" (i.e., the search condition), but the title of document A is "Nike Shoes" (i.e., the hit field in document A), and the title of document B is "Nikon Camera Lens" (i.e., the hit field in document B), in this case, since "ni" accounts for a larger proportion in the title of document A, document A will be given a higher sorting priority. The above sorting method enhances the accuracy of the search results and ensures that when faced with numerous documents containing short query words, the content with the highest relevance can be found more quickly.
[0166] Scenario 3 is applicable to obtaining explanation information for complex operations:
[0167] Compared with the prior art, the embodiments of the present application can obtain the explanation information in the documents of the nested structure type, so it can be applicable to obtaining the explanation information for complex operations. For nested documents, the most fundamental hit reason can be specifically located. For example, when retrieving videos, highlights, etc., a video can include multiple nested sub-paragraphs, a highlight can include multiple sub-photos and sub-video clips. The embodiments of the present application can search the content in each nested sub-paragraph in the video: such as tags, semantic vectors, subtitles, etc. to return the video; at the same time, the explanation will return which attribute of which field is hit specifically; as Figure 4F shown, when searching for a search term (i.e., the first structured object generated by parsing the search condition), the rich text document can be returned, and at the same time, it can be returned which table, which paragraph, and which picture in the rich text are hit (i.e., the hit position in the explained structured information). Through this information, it is possible to directly jump to the corresponding position.
[0168] As Figure 4G shown, for a note application, it can be analogized to the above rich text structure. A note contains multiple elements: title, text, handwriting, attachments, tables, recordings, pictures, etc., which can be constructed as a complex nested object (i.e., the hit document of the nested structure). The search can search all elements of this note. When hitting, through the explanation information (i.e., the explained structured information), it can be known which element and which one in it is hit. In this way, when displaying, it is possible to directly locate to the specified position (i.e., the hit position in the hit information) and give a prompt.
[0169] As Figure 4HAs shown, when we search for the "protection red line" (i.e., the first structured object constructed based on the search conditions), we can return this video (i.e., the main document in the hit document with a nested structure), and through the explanatory information, we can know that it is the text of the second paragraph of this video (i.e., the sub-document in the hit document with a nested structure); in this way, we can directly display the cover as shown above, and click to continue playing from this position; compared with the existing video retrieval, we can dynamically display the cover of the hit position and the hit content.
[0170] It can be understood that the above three application scenarios are just several exemplary implementation manners in the embodiments of the present application, and the application scenarios in the embodiments of the present application include but are not limited to the above application scenarios.
[0171] The method of the embodiments of the present application is elaborated in detail above, and the related devices of the embodiments of the present application are provided below.
[0172] Please refer to Figure 5 , Figure 5 FIG. is a schematic structural diagram of a search engine device provided by an embodiment of the present application. The search engine device 01-C may include a first parsing unit 601, a first creating unit 602, an information obtaining unit 603, a first generating unit 604, a first sorting unit 605, a screening unit 606, an extraction unit 607, an information generating unit 608, and a second sorting unit 609. The detailed descriptions of each unit are as follows.
[0173] The first parsing unit 601 is configured to receive a search condition and generate a first structured object, where the first structured object includes a query statement included in the search condition;
[0174] The first creating unit 602 creates a scoring calculator based on the first structured object. The scoring calculator includes the first structured object and M document identifiers that match the first structured object, where M is a positive integer greater than 0;
[0175] In a possible implementation manner, the first creating unit 602 is specifically configured to:
[0176] Create K corresponding sub-scoring calculators based on the K third structured objects. Each sub-scoring calculator includes a corresponding third structured object, L document identifiers that match the corresponding third structured object, and the scoring component;
[0177] Create the scoring calculator based on the K corresponding sub-scoring calculators, where L is a positive integer greater than 0.
[0178] The information obtaining unit 603 obtains M first documents that match the M document identifiers, and generates scoring description information corresponding to each of the M first documents through the scoring component;
[0179] In a possible implementation, the search engine includes an inverted index, and the inverted index includes a mapping relationship between a document identifier and a corresponding document;
[0180] The information acquisition unit 603 is specifically configured to:
[0181] Search for M first documents corresponding to the M document identifiers from the inverted index.
[0182] In a possible implementation, the scoring description information corresponding to each of the first documents includes hit information and scoring information;
[0183] The information acquisition unit 603 is specifically configured to:
[0184] For each scorer, screen out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q;
[0185] Generate the hit information and scoring information corresponding to each of the M first documents through the scoring component, and the scoring information is the score of the matching degree between the corresponding first document and the search condition.
[0186] In a possible implementation, the hit information in the J first documents includes one or more of a hit field, a hit word, and a hit position that match the search condition in the corresponding first document.
[0187] In a possible implementation, the first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document of each of the Q first documents, the set of document identifier sequences of the one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document.
[0188] The first generation unit 604 generates M second structured objects based on the scoring description information corresponding to each of the first documents and the first structured object, where one second structured object corresponds to one first document;
[0189] In a possible implementation, the first generation unit 604 is specifically configured to:
[0190] Obtain K sub-query statements included in the query statement, where K is a positive integer greater than 0;
[0191] Based on the K sub-query statements, K third structured objects are generated, where each of the third structured objects includes a corresponding sub-query statement, and the first structured object includes the K third structured objects.
[0192] The first sorting unit 605 performs a first sorting on the corresponding first document for each scoring device based on the scoring description information.
[0193] The screening unit 606 screens out N first documents that meet the preset conditions from the first documents corresponding to each scoring device based on the result of the first sorting, where N is a positive integer less than or equal to M.
[0194] The extraction unit 607 retrieves the corresponding N second structured objects from the storage component based on the document identifiers of the N first documents.
[0195] The information generation unit 608 generates N corresponding structured information based on the N second structured objects.
[0196] The second sorting unit 609 performs a second sorting on the N first documents based on the N corresponding structured information and outputs the sorting result to the search result.
[0197] In a possible implementation manner, the second sorting unit is specifically configured to:
[0198] Perform a second sorting on the N corresponding first documents based on the similarity score between each of the N structured information and the corresponding first document.
[0199] Or, perform a second sorting on the N corresponding first documents based on the occupancy ratio score of the hit words in the corresponding hit fields in each of the N structured information.
[0200] It should be noted that for the functions of the functional units in the search engine device 01-C described in the embodiments of the present application, reference can be made to the relevant descriptions of steps S300 - S302 in the method embodiments described above and Figure 3A the relevant descriptions of steps S400 - S403 in the method embodiments described in Figure 3B which will not be elaborated here.
[0201] Next, the terminal device provided in the embodiments of the present application is introduced.
[0202] Figure 6A The hardware structure diagram of the terminal device 03 provided in the embodiments of the present application is shown. The terminal device 03 is used to execute the image recommendation method provided in the foregoing method embodiments.
[0203] The terminal device 03 may include a processor 101, a memory 102, a wireless communication module 103, a mobile communication module 104, an antenna 103A, an antenna 104A, a power switch 105, a sensor module 106, a focusing motor 107, a camera 108, a display screen 109, etc. Among them, the sensor module 106 may include a gyroscope sensor 106A, an acceleration sensor 106B, an ambient light sensor 106C, an image sensor 106D, a distance sensor 106E, etc. Among them, the wireless communication module 103 may include a WLAN communication module, a Bluetooth communication module, etc. The above-mentioned multiple parts can transmit data through a bus.
[0204] The processor 101 may include various processing units to optimize the performance of the Lucene search engine when processing local data retrieval. For example, these processing units may include a main application processor (AP) responsible for complex data processing tasks, a graphics processing unit (GPU) to optimize the display of search results, and a specially designed digital signal processor (DSP) and neural network processor (NPU) to accelerate the execution of search algorithms. These units work together to improve the efficiency and accuracy of the Lucene search engine when processing queries, indexing, and sorting search results. In addition, this multi-processing unit design can support the Lucene search engine to process multiple tasks in parallel, thus providing fast response and accurate search results when processing large data sets.
[0205] The memory 102 is used to store the operating system, application code, and other data of the terminal device 03. For example, it may include local applications (such as various business applications, games, or tool applications, etc.) and the program code and data required by the Lucene search engine. The processor 101 realizes device functions and data processing tasks by executing the program code stored in the memory 102. The storage program area of the memory 102 is used to place the operating system and application code, while the storage data area is used to save the data generated when running the application. In addition to including a random access memory (RAM) for fast access, the memory 102 may also include one or more forms of non-volatile memory, such as a hard disk drive, a solid-state drive, or a flash device, to provide a persistent storage solution.
[0206] The wireless communication function of the terminal device 03 can be realized through the antenna 103A, the antenna 104A, the mobile communication module 104, the wireless communication module 103, a modulation and demodulation processor, and a baseband processor, etc.
[0207] The antennas 103A and 104A can be used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device 03 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.
[0208] The mobile communication module 104 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device 03. The mobile communication module 104 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 104 can receive electromagnetic waves through the antenna 104A, filter and amplify the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 104 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves through the antenna 104A and radiate it out.
[0209] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through the audio device, or displays images or videos through the display screen 109.
[0210] The wireless communication module 103 can provide solutions for wireless communications such as wireless local area networks (WLAN), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the terminal device 03. The wireless communication module 103 may be one or more devices integrating at least one communication processing module. The wireless communication module 103 receives electromagnetic waves through the antenna 103A, frequency-modulates and filters the electromagnetic wave signal, and transmits the processed signal to the processor 101. The wireless communication module 103 can also receive the signal to be transmitted from the processor 101, frequency-modulate and amplify it, and convert it into electromagnetic waves through the antenna 103A and radiate it out.
[0211] The power switch 105 can be used to control the power supply to the terminal device 03.
[0212] The gyroscope sensor 106A can be used to determine the motion posture of the terminal device 03. In some embodiments, the angular velocity of the terminal device 03 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 106A. The gyroscope sensor 106A can be used for anti-shake shooting. Exemplarily, when the shutter is pressed, the gyroscope sensor 106A detects the angle of the terminal device 03 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the terminal device 03 through reverse movement to achieve anti-shake. The gyroscope sensor 106A can also be used for navigation and somatosensory game scenes.
[0213] The acceleration sensor 106B can detect the magnitude of the acceleration of the terminal device 03 in various directions (generally three axes). When the terminal device 03 is stationary, the magnitude and direction of gravity can be detected. It can also be used to identify the posture of the terminal device. For example, the acceleration sensor 106B can be applied to applications such as horizontal and vertical screen switching, pedometers, etc.
[0214] The ambient light sensor 106C is used to sense the brightness of the ambient light. The terminal device 03 can adaptively adjust the brightness of the display screen 109 according to the sensed brightness of the ambient light. The ambient light sensor 106C can also be used to automatically adjust the white balance when taking pictures.
[0215] The image sensor 106D, also known as a photosensitive element, can utilize the photoelectric conversion function of the photoelectric device to convert the light image on the photosensitive surface into an electrical signal that is proportional to the light image. The image sensor can be a charge coupled device (CCD) sensor or a complementary metal-oxide-semiconductor (CMOS) sensor.
[0216] The distance sensor 106E can be used to measure the distance. The terminal device 03 can measure the distance by infrared or laser. In some shooting scenes, the terminal device 03 can use the distance sensor 106E to measure the distance to achieve fast focusing.
[0217] The focus motor 107 can be used for fast focusing. The terminal device 03 can control the movement of the lens through the focus motor 107 to achieve automatic focusing.
[0218] The terminal device 03 can realize the shooting function through ISP, camera 108, video codec, GPU, display screen 109 and application processor.
[0219] The ISP is used to process the data fed back by the camera 108. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and conversion into an image visible to the naked eye. The ISP can also optimize the noise and brightness of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 108.
[0220] The camera 108 can be used to capture still images or videos. An object generates an optical image through the lens and projects it onto the image sensor. The image sensor can convert the light signal into an electrical signal and then transmit the electrical signal to the ISP for conversion into a digital image signal. The ISP can output the digital image signal to the DSP for processing. The DSP converts the digital image signal into a standard image signal in formats such as RGB and YUV. In some embodiments, the terminal device 03 can include one or N cameras 108, where N is a positive integer greater than 1.
[0221] The video codec is used to compress or decompress digital images. The terminal device 03 can support one or more image codecs. In this way, the terminal device 03 can open or save pictures or videos in multiple coding formats.
[0222] The terminal device 03 can achieve the display function through the GPU, the display screen 109, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 109 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 101 can include one or more GPUs, which execute program instructions to generate or change the display information.
[0223] The display screen 109 is used to display images, videos, etc. The display screen 109 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the terminal device 03 can include one or N display screens 109, where N is a positive integer greater than 1.
[0224] It is to be understood that the structure illustrated in the embodiment of the present application does not constitute a specific limitation on the terminal device 03. In other embodiments of the present application, the terminal device 03 may include more or fewer components than shown in the figure, or combine some components, or split some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0225] The operations performed by each device in terminal device 03 can be found in Figure 3A - Figure 3B The relevant description of the embodiments in the embodiment will not be expanded in detail here.
[0226] The software system of the terminal device 03 can adopt one or more of a layered architecture, an event-driven architecture, a micro-core architecture, a micro-service architecture, or a cloud architecture, and can support a local search function based on a Lucene search engine. The embodiment of the present application takes a mobile operating system with a layered architecture as an example to illustrate the software structure of the terminal device 03.
[0227] Figure 6B It is a software structure block diagram of the terminal device 03 of the embodiment of the present application.
[0228] The layered architecture divides the software into several layers, each with clear roles and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the mobile operating system is divided into four layers, from top to bottom: application layer, application framework layer / core service layer, system library and runtime, and kernel layer.
[0229] The application layer can include a series of application packages.
[0230] like Figure 6B As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0231] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0232] like Figure 6B As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0233] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0234] The content provider can be used to store and retrieve data, making this data accessible to applications. It can be used to access and index local data such as the phone book, browsing history, and bookmarks, providing a data source for Lucene. The data can include videos, images, audio, incoming and outgoing calls, browsing history, bookmarks, phone books, etc.
[0235] The view system includes visual controls, such as controls for displaying text and controls for displaying pictures. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.
[0236] The phone manager is used to provide the communication functions of the terminal device. For example, the management of call states (including connection, disconnection, etc.).
[0237] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.
[0238] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of background-running applications, and can also be a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt sound, cause the terminal device to vibrate, and the indicator light to flash, etc.
[0239] Runtime can refer to all code libraries, frameworks, etc. required when the program is running. For example, it provides a Java virtual machine and core libraries, which can provide a necessary environment for the operation of the Lucene search engine.
[0240] The system library can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (such as: OpenGL ES), 2D graphics engine (such as: SGL), etc. Lucene can use these libraries to index and search image, audio, and video files on the device.
[0241] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0242] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0243] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0244] The 2D graphics engine is a drawing engine for 2D drawing.
[0245] The kernel layer is the layer between hardware and software. The kernel layer at least includes a display driver, a camera driver, an audio driver, and a sensor driver. For example, it can provide underlying hardware support for a search engine, enabling the Lucene search engine to access and index files in the file system.
[0246] It should be understood that each step in the above method embodiments can be completed by the integrated logic circuit of the hardware in the processor or the instructions in the form of software. The method steps disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor.
[0247] The present application also provides a terminal device, which may include: a memory and a processor. Among them, the memory can be used to store a computer program; the processor can be used to call the computer program in the memory, so that the terminal device executes the method executed on the terminal device side in any one of the above embodiments.
[0248] The present application also provides a terminal device, which may include: a memory and a processor. Among them, the memory can be used to store a computer program; the processor can be used to call the computer program in the memory, so that the terminal device executes the method executed on the terminal device side in any one of the above embodiments.
[0249] The present application also provides a chip system, which includes at least one processor for implementing the functions involved on the terminal device side in any one of the above embodiments.
[0250] In a possible design, the chip system further includes a memory, and the memory is used to store program instructions and data, and the memory is located inside or outside the processor.
[0251] The chip system can be composed of chips, or can include chips and other discrete devices.
[0252] Optionally, there may be one or more processors in the chip system. The processor may be implemented by hardware or by software. When implemented by hardware, the processor may be a logic circuit, an integrated circuit, etc. When implemented by software, the processor may be a general-purpose processor that implements its functions by reading software code stored in a memory.
[0253] Optionally, there may also be one or more memories in the chip system. The memory may be integrated with the processor or may be separately provided from the processor, which is not limited in the embodiments of the present application. Exemplarily, the memory may be a non-transitory processor, such as a read-only memory (ROM). It may be integrated with the processor on the same chip or may be separately provided on different chips. The embodiments of the present application do not specifically limit the type of the memory and the setting manner of the memory and the processor.
[0254] Exemplarily, the chip system may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), or other integrated chips.
[0255] The present application also provides a computer program product, which includes a computer program (which may also be referred to as code or instruction). When the computer program is run, the computer is caused to execute the method performed on the terminal device side in any one of the above embodiments.
[0256] The present application also provides a computer-readable storage medium that stores a computer program (which may also be referred to as code or instruction). When the computer program is run, the computer is caused to execute the method performed on the terminal device side in any one of the above embodiments.
[0257] The various embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0258] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state disk (SSD)), etc.
[0259] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware with a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.
[0260] In summary, the above description is only an embodiment of the technical solution of this application and is not used to limit the protection scope of this application. Any modifications, equivalent replacements, improvements, etc. made according to the disclosure of this application shall be included in the protection scope of this application.
Claims
1. A method for data retrieval, characterized in that, applied to a search engine, the method includes: Receiving a search condition and generating a first structured object, where the first structured object includes a query statement contained in the search condition; Creating a scoring calculator based on the first structured object, where the scoring calculator includes the first structured object, M document identifiers matching the query statement, and a scoring component, and M is a positive integer greater than 0; Obtaining M first documents matching the M document identifiers, and generating scoring description information corresponding to each of the M first documents through the scoring component; Generating M second structured objects based on the scoring description information corresponding to each of the first documents and the first structured object, where one second structured object corresponds to one first document.
2. The method according to claim 1, characterized in that, the search engine includes an inverted index, and the inverted index includes a mapping relationship between a document identifier and a corresponding document; the obtaining M first documents matching the M document identifiers includes: Searching for M first documents corresponding to the M document identifiers from the inverted index.
3. The method according to claim 1 or 2, characterized in that, the scoring description information corresponding to each of the first documents includes hit information and scoring information; the generating of the scoring description information corresponding to each of the M first documents through the scoring component includes: For the scoring calculator, screening out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q; Generating hit information and scoring information corresponding to each of the M first documents through the scoring component, and the scoring information is a score of the matching degree between the corresponding first document and the search condition.
4. The method according to claim 3, characterized in that, the hit information in the J first documents includes one or more of a hit field, a hit word, and a hit position that match the search condition in the corresponding first document.
5. The method according to claim 3, characterized in that, the first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document of each of the Q first documents, the set of document identifier sequences of the one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document.
6. The method according to any one of claims 1-5, characterized in that, the generating of the first structured object includes: Obtaining K sub-query statements included in the query statement, where K is a positive integer greater than 0; Generating K third structured objects based on the K sub-query statements, where each of the third structured objects includes a corresponding sub-query statement, and the first structured object includes the K third structured objects.
7. The method according to claim 6, wherein, the creating a score calculator based on the first structured object includes: creating K corresponding sub-score calculators based on the K third structured objects, each sub-score calculator including a corresponding third structured object, L document identifiers matching the corresponding third structured object, and the scoring component; creating the score calculator based on the K corresponding sub-score calculators, where L is a positive integer greater than 0.
8. The method according to any one of claims 1-7, wherein, the method further includes: for each score calculator, performing a first sorting on the corresponding first document based on the scoring description information; based on the result of the first sorting, screening out N first documents whose scoring description information meets a preset condition from the first documents corresponding to each score calculator, where N is a positive integer less than or equal to M.
9. The method according to claim 8, wherein, a preset storage component is included in the search engine; for each of the first documents in the score calculator, saving the corresponding second structured object to the storage component; the method further includes: based on the document identifiers of the N first documents, retrieving the corresponding N second structured objects from the storage component; generating N corresponding structured information based on the N second structured objects.
10. The method according to claim 9, wherein, the method further includes: performing a second sorting on the N first documents based on the N corresponding structured information, and outputting the sorting result to the search results.
11. The method according to claim 10, wherein, the performing a second sorting on the N first documents based on the N structured information includes: performing a second sorting on the N corresponding first documents based on the similarity score between each of the N structured information and the corresponding first document; alternatively, performing a second sorting on the N corresponding first documents based on the occupancy ratio score of the hit words in the corresponding hit fields in each of the N structured information.
12. A data retrieval device, wherein, applied to a search engine, the device includes: a first parsing unit, configured to receive a search condition and generate a first structured object, the first structured object including a query statement included in the search condition; a first creating unit, creating a score calculator based on the first structured object, the score calculator including the first structured object and M document identifiers matching the first structured object, where M is a positive integer greater than 0; an information obtaining unit, obtaining M first documents matching the M document identifiers, and generating scoring description information corresponding to each of the M first documents through the scoring component; a first generating unit, generating M second structured objects based on the scoring description information corresponding to each first document and the first structured object, where one second structured object corresponds to one first document.
13. The device according to claim 12, wherein, The search engine includes an inverted index, and the inverted index includes a mapping relationship between a document identifier and a corresponding document; The information acquisition unit is specifically configured to: Search the inverted index for M first documents corresponding to the M document identifiers.
14. The apparatus according to claim 12 or 13, wherein, The scoring description information corresponding to each of the first documents includes hit information and scoring information; The information acquisition unit is specifically configured to: For each score calculator, screen out Q first documents with a nested structure and J first documents with a non-nested structure from the M first documents, where Q is an integer not greater than M and not less than 0, and J is the difference between M and Q; Generate the hit information and scoring information corresponding to each of the M first documents through the scoring component, and the scoring information is the score of the matching degree between the corresponding first document and the search condition.
15. The apparatus according to claim 14, wherein, The hit information in the J first documents includes one or more of a hit field, a hit word, and a hit position that match the search condition in the corresponding first document.
16. The apparatus according to claim 14, wherein, The first document with a nested structure includes a main document and one or more sub-documents; the hit information in the Q first documents includes one or more of the document identifier of the main document of each of the Q first documents, the set of document identifier sequences of the one or more sub-documents, the document identifier of the sub-document hit in the main document, and the hit position of the hit sub-document.
17. The apparatus according to any one of claims 11-16, wherein, The first generation unit is specifically configured to: Obtain K sub-query statements included in the query statement, where K is a positive integer greater than 0; Generate K third structured objects based on the K sub-query statements, where each of the third structured objects includes a corresponding sub-query statement, and the first structured object includes the K third structured objects.
18. The apparatus according to claim 17, wherein, The first creation unit is specifically configured to: Create K corresponding sub-score calculators based on the K third structured objects, and each sub-score calculator includes the corresponding third structured object, L document identifiers matching the corresponding third structured object, and the scoring component; Create the score calculator based on the K corresponding sub-score calculators, where L is a positive integer greater than 0.
19. The apparatus according to any one of claims 11-18, wherein, The apparatus further includes: A first sorting unit that, for each score calculator, performs a first sorting on the corresponding first documents based on the scoring description information; A screening unit that, based on the result of the first sorting, screens out N first documents whose scoring description information meets a preset condition from the first documents corresponding to each score calculator, where N is a positive integer less than or equal to M.
20. The apparatus according to claim 19, wherein, The search engine includes a preset storage component; For each of the first documents in the scoring device, save the corresponding second structured object to the storage component; The device further includes: An extraction unit, based on the document identifiers of the N first documents, retrieve the corresponding N second structured objects from the storage component; An information generation unit, based on the N second structured objects, generate N corresponding structured information.
21. The device according to claim 20, wherein, The device further includes: A second sorting unit, based on the N corresponding structured information, perform a second sorting on the N first documents, and output the sorting result to the search results.
22. The device according to claim 21, wherein, The second sorting unit is specifically configured to: Based on the similarity score between each of the N structured information and the corresponding first document, perform a second sorting on the N corresponding first documents; Or, based on the ratio score of the hit words in each of the N structured information in the corresponding hit fields, perform a second sorting on the N corresponding first documents.
23. A computer storage medium, wherein, The computer storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method described in any one of claims 1-11 above.
24. A computer program, wherein, The computer program includes instructions, and when the computer program is executed by a computer, it causes the computer to execute the method described in any one of claims 1-11.
Citation Information
Patent Citations
Data processing method, server and computer storage medium
CN108520002A
Seat taking query method and device, and server
CN108664509A
Intelligent document retrieval method based on big data
CN115617957A
Structured document processing apparatus, structured document search apparatus, structured document system, method, and program
US20070027671A1
Document ranking with sub-query series
US20080010268A1
Cited By
Data networking search sorting method and system based on data pragmatics
CN120354018A
Literature semantic search method and system based on elastic search
CN120429311A