Information processing method and device, equipment and medium

By constructing and traversing query trees based on dynamically perceived node hot and cold attributes, the system solves the problem of poor efficiency and effectiveness in querying mixed hot and cold data sources in existing retrieval systems. This enables rapid response to hot data and reduces the cost of cold data, thereby improving the overall performance of the retrieval system.

CN121301437APending Publication Date: 2026-01-09BAIDU (CHINA) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511484821.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-16
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing retrieval systems are inefficient and ineffective when handling complex queries that combine hot and cold data sources.

Method used

A method that dynamically senses the hot and cold attributes of each node is adopted. A query tree is constructed based on the query expression, and the retrieval strategy is optimized by traversing and merging the query tree, including a hot data priority strategy and a cold data truncation strategy, to improve retrieval efficiency and effectiveness.

Benefits of technology

By optimizing retrieval strategies, we can ensure rapid response to frequently accessed hot data, reduce the storage cost of cold data, and improve the overall resource utilization and retrieval effectiveness of the retrieval system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121301437A_ABST
    Figure CN121301437A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method and device, equipment and a medium, relates to the technical field of data processing, in particular to the technical field of retrieval systems and the like, and can be used for application scenes such as generative retrieval, document intelligent editing, intelligent assistants, virtual assistants and intelligent e-commerce. The method comprises the steps that a query tree is constructed based on a query expression, the query expression comprises a plurality of recall conditions and at least one logic operator, the query tree comprises a plurality of leaf nodes corresponding to the recall conditions respectively, and each leaf node is based on a source of a document index list corresponding to the leaf node; the data list corresponds to a hot data list and / or a cold data list; and at least one internal node corresponding to the at least one logic operator; the merging query tree is traversed to obtain a retrieval result, and the merging strategy and the temperature attribute of each internal node are determined based on the logic operator corresponding to the internal node and the temperature attribute of the child node of the internal node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to the field of retrieval systems, which can be used in application scenarios such as generative retrieval, intelligent document editing, intelligent assistants, virtual assistants, and intelligent e-commerce. Specifically, it relates to an information processing method, an information processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Technology

[0002] Retrieval systems process, organize, and index unstructured or structured data to quickly locate and return relevant information based on user queries. Existing retrieval systems are still evolving and developing in terms of data processing efficiency and the relevance of search results.

[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention

[0004] This disclosure provides an information processing method, an information processing apparatus, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to one aspect of this disclosure, an information processing method for a retrieval system is provided. The retrieval system includes a first storage structure for storing hot data and a second storage structure for storing cold data. The method includes: constructing a query tree based on a query expression, the query expression including multiple recall conditions and at least one logical operator, wherein the query tree includes: multiple leaf nodes corresponding to the multiple recall conditions, each leaf node corresponding to a list of hot data obtained from the first storage structure and / or a list of cold data obtained from the second storage structure based on the source of the document index list corresponding to the leaf node; and at least one internal node corresponding to the at least one logical operator; and traversing and merging the query tree to obtain retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attributes of the child nodes of the internal node.

[0006] According to another aspect of this disclosure, an information processing apparatus for a retrieval system is provided. The retrieval system includes a first storage structure for storing hot data and a second storage structure for storing cold data. The apparatus includes: a construction unit configured to construct a query tree based on a query expression, the query expression including multiple recall conditions and at least one logical operator, wherein the query tree includes: multiple leaf nodes corresponding to the multiple recall conditions, each leaf node corresponding to a list of hot data obtained from the first storage structure and / or a list of cold data obtained from the second storage structure based on the source of the document index list corresponding to the leaf node; and at least one internal node corresponding to the at least one logical operator; and a traversal and merging unit configured to traverse and merge the query tree to obtain retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attributes of the child nodes of the internal node.

[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the methods described above.

[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described method.

[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described method when executed by a processor.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0012] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown; Figure 2 A flowchart of an information processing method according to an embodiment of the present disclosure is shown; Figure 3 A flowchart illustrating the traversal of a merge query tree to obtain retrieval results according to an embodiment of the present disclosure is shown; Figure 4 A structural block diagram of an information processing apparatus according to embodiments of the present disclosure is shown; and Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.

[0015] The terminology used in the description of the various examples in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.

[0016] In related technologies, retrieval systems that simultaneously include a first storage structure for storing hot data and a second storage structure for storing cold data perform poorly when completing query tasks.

[0017] To address the aforementioned issues, this disclosure proposes a method that dynamically senses the hot / cold attributes of each node during query tree merging and selects the optimal merging strategy accordingly. This significantly improves retrieval efficiency and effectiveness when handling complex queries with mixed hot / cold data sources.

[0018] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0019] Figure 1A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.

[0020] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the methods of this disclosure.

[0021] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.

[0022] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.

[0023] Users can use client devices 101, 102, 103, 104, 105, and / or 106 for human-computer interaction. The client devices provide interfaces that enable users to interact with them. The client devices can also output information to the user through these interfaces. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.

[0024] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.

[0025] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WiFi), and / or any combination of these and / or other networks.

[0026] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0027] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0028] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.

[0029] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.

[0030] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.

[0031] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.

[0032] Figure 1The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.

[0033] According to one aspect of this disclosure, an information processing method is provided. The retrieval system includes a first storage structure for storing hot data and a second storage structure for storing cold data. For example... Figure 2 As shown, method 200 includes: step S201, constructing a query tree based on a query expression, the query expression including multiple recall conditions and at least one logical operator, wherein the query tree includes: multiple leaf nodes corresponding to the multiple recall conditions, each leaf node corresponding to a hot data list obtained from a first storage structure and / or a cold data list obtained from a second storage structure based on the source of the document index list corresponding to the leaf node; and at least one internal node corresponding to at least one logical operator; and step S202, traversing and merging the query tree to obtain the retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attribute of the child nodes of the internal node.

[0034] Therefore, by adopting the above methods, we can ensure that frequently accessed hot data has a faster retrieval response, while reducing the storage cost of massive amounts of cold data, thereby improving the overall resource utilization of the retrieval system.

[0035] In some embodiments, the retrieval system disclosed herein can be a search engine that provides general retrieval capabilities for various services. The read speed of the first storage structure can be greater than that of the second storage structure. Exemplarily, the first storage structure can be memory, and the second storage structure can be a hard disk, such as a solid-state drive (SSD).

[0036] The format of a retrieval expression can be: cond1 op1 cond2 op2 ..., where cond is the recall condition and op is the logical operator. The retrieval system can employ a multi-path recall approach to obtain candidate documents, i.e., parallel retrieval using multiple recall conditions. These recall conditions can include: term recall based on inverted indexes, docvalue recall based on document attributes, and vector recall based on semantic vectors. To balance system performance and storage costs, data from any of the above recall paths may be divided into hot and cold parts. For example, the inverted list (i.e., the document index list) corresponding to a term can have its frequently accessed portion stored as hot data in memory, while its less frequently accessed portion is stored as cold data on disk. Therefore, a leaf node may correspond only to the hot data list, only to the cold data list, or simultaneously to both hot and cold data lists that need to be merged. These recall results (leaf nodes) will be combined using a query tree. The internal nodes of the query tree are logical operators such as "AND," "OR," and "NOT," used to merge the document index lists returned by the leaf nodes.

[0037] In an exemplary embodiment, the query expression "term1=weather&temperature>30&vector1=(...)" is received in step S201. The system constructs a query tree based on this expression. The leaf nodes of this query tree are, respectively, the term (term) recall node corresponding to the recall condition "term1=weather", the docvalue (attribute value) recall node corresponding to the recall condition "temperature>30", and the vector (semantic vector) recall node corresponding to the recall condition "vector1=(...)". The internal nodes of the query tree are two "AND" logical operation nodes, used to sequentially perform intersection operations on the results of the above three leaf nodes. During the construction process, the system determines the temperature attribute of each leaf node. For example, the "weather" recall condition may retrieve document indexes from both memory and hard disk simultaneously; in this case, the leaf node will be determined as a cold data node. The range query "temperature>30" primarily accesses the hard disk and is also determined as a cold data node. The query "vector1" may completely hit the vector index in memory, and therefore is determined as a hot data node.

[0038] According to some embodiments, step S201, constructing a query tree based on a query expression, may include: in response to a recall condition among multiple recall conditions that simultaneously corresponds to obtaining a hot data list from a first storage structure and a cold data list from a second storage structure, performing a logical OR operation on the hot data list and the cold data list to generate a leaf node corresponding to the recall condition, wherein the temperature attribute of the leaf node is determined to be cold data.

[0039] When a document index list for a recall condition (e.g., term="weather") is distributed across both the first and second storage structures, the system first merges the list of hot data retrieved from the first storage structure and the list of cold data retrieved from the second storage structure using an implicit "OR" logical operation to generate a complete leaf node. Since this merged list contains data that needs to be retrieved from the slower second storage structure, the temperature attribute of this leaf node is identified as cold data.

[0040] Therefore, this method ensures the integrity of the results of a single recall condition and provides its parent node in the query tree with accurate temperature attributes, providing a reliable basis for the parent node to correctly select the merging strategy.

[0041] According to some embodiments, step S202, traversing the merge query tree to obtain the retrieval result, may include: for a first internal node in at least one internal node corresponding to the logical OR operator, in response to determining that the child nodes of the first internal node include both cold data child nodes and hot data child nodes, determining the merging strategy of the first internal node as a hot data priority strategy, wherein the hot data priority strategy indicates that after the document index corresponding to the hot data child node has been returned, the document index corresponding to the cold data child node is returned.

[0042] The hot data priority strategy is applied to internal nodes (e.g., OrList) representing "OR" logic in the query tree. In some embodiments, when performing a merge, a regular "OR" operation (union) requires fetching data from all its child nodes simultaneously. If a child node contains a cold data node with high access latency (e.g., one that needs to be read from disk), the initial return latency of the entire merge process may be dragged down by this slow node. The hot data priority strategy changes the default scheduling method for internal nodes to fetch data from their child nodes. Specifically, the "OR" logic node first fully traverses and returns the complete document index of its hot data child nodes (from the first storage structure with faster read speed). Only after the list of hot data child nodes is completely exhausted will it begin to fetch and return document indexes from its cold data child nodes (from the second storage structure with slower read speed). This asymmetric merge method ensures that users or upper-layer applications can receive a portion of the most popular results as quickly as possible, optimizing the retrieval experience in scenarios such as first-screen loading.

[0043] According to some embodiments, the temperature attribute of the first internal node is determined to be cold data.

[0044] In some embodiments, the performance bottleneck of an "OR" logic operation node is often determined by its slowest-accessed child node. To obtain the complete result set of that node, its parent node needs to consider the potential time overhead of retrieving document indexes from cold data child nodes during traversal and merging. Therefore, identifying the temperature attribute of this internal node as cold data ensures that higher-level nodes can correctly select and execute subsequent merging strategies based on this cost information.

[0045] According to some embodiments, step S202, traversing the merge query tree to obtain the retrieval result, may include: for a second internal node corresponding to the logical AND operator in at least one internal node, in response to determining that the child nodes of the second internal node include both cold data child nodes and hot data child nodes, determining the merging strategy of the second internal node as a cold data truncation strategy, wherein the cold data truncation strategy indicates that the document index list corresponding to the cold data child node is truncated before merging the hot data child nodes and the cold data child nodes.

[0046] The cold data truncation strategy is applied to internal nodes representing AND logic. Performing a regular intersection operation directly between a memory list and a disk list results in poor performance due to frequent random disk lookups. The cold data truncation strategy introduces a preprocessing step: before performing the actual intersection calculation, it reads only a pre-defined subset of data from the document index list corresponding to the cold data child nodes and loads it into memory. This allows the subsequent intersection operation to be efficiently completed between two lists both residing in the first storage structure, avoiding slow accesses to the second storage structure during the merge process.

[0047] According to some embodiments, truncating the document index list corresponding to cold data child nodes may include: loading the document index list corresponding to cold data child nodes from the second storage structure to the first storage structure; and retaining a preset number of document indexes from the loaded document index list for merging with the document index list corresponding to hot data child nodes.

[0048] Therefore, by preloading a limited amount of cold data, it is ensured that subsequent merging operations can be performed entirely within the high-speed storage structure, thus guaranteeing the efficient execution of the cold data truncation strategy.

[0049] In some embodiments, the specific value of the preset quantity can be calculated or determined a priori or by other means, and is not limited here.

[0050] According to some embodiments, the temperature attribute of the second internal node is determined as thermal data.

[0051] In some embodiments, in response to determining that the child nodes of an internal node include only cold data child nodes, the temperature attribute of the internal node is determined to be cold data; in response to determining that the child nodes of an internal node include only hot data child nodes, the temperature attribute of the internal node is determined to be hot data.

[0052] The temperature attribute of an internal node is designed to reflect its actual performance when providing data to its parent node. For a typical AND operation node, if its child nodes contain cold data, the performance bottleneck lies in the slow access to the second storage structure. However, in this scheme, the second internal node has implemented a cold data truncation strategy for its cold data child nodes. This strategy, through a preprocessing step, pre-emptively restricts and limits access operations to slow storage, ensuring that the subsequent iterative process of providing merge results to the parent node is entirely completed within the first storage structure (memory). Therefore, from the parent node's perspective, the access cost and response speed of this AND node exhibit characteristics of hot data, thus its temperature attribute is determined as "hot data."

[0053] According to some embodiments, such as Figure 3 As shown, step S202, traversing and merging the query tree to obtain the retrieval results, may include: step S301, performing multiple rounds of iterative traversal and merging of the query tree; step S302, in each round of iteration, counting in real time the number of document indexes originating from each leaf node that have been merged into the retrieval results, and calculating the average value; and step S303, in the next round of iteration, pausing the retrieval of document indexes from leaf nodes where the number of merged document indexes exceeds the average value.

[0054] The steps described above provide a global balancing mechanism. This mechanism is a dynamic scheduling strategy during query merging. In conventional multi-path recall, fast-responding recall conditions (such as term recall) often occupy the head of the result set, leading to a lack of diversity in results. The global balancing mechanism, in each iteration of the merging process, continuously counts the number of document indexes originating from each leaf node that have been adopted into the final result and calculates the average number. In the next iteration, this mechanism pauses the retrieval of document indexes from leaf nodes that have contributed more than the average number, thus providing a chance for slower nodes with lagging contributions to catch up.

[0055] Thus, this global balancing mechanism ensures that the final retrieval results are a uniform mixture of document indexes from different recall channels, improving the diversity and richness of the results during the recall phase and providing a higher quality candidate set for the subsequent ranking process.

[0056] In some embodiments, to further enhance the shuffling effect of the results, the global balancing mechanism may also include a preprocessing step for the document index list of leaf nodes. In this preprocessing step, before performing multiple rounds of iterative traversal and merging, the system may reorder the document index lists within one or more leaf nodes according to a preset shuffling rule. This approach more thoroughly breaks the inherent order of the original lists, resulting in a more uniform mixture of final search results, thereby further improving the diversity of the results.

[0057] In some embodiments, in step S301, the system can perform traversal merging in a multi-round iterative manner, with each round of iteration aiming to retrieve and merge one or more document indexes from the query tree.

[0058] In some embodiments, in step S302, the system can, in each iteration, count in real time the specific number of document indexes from each leaf node that have been adopted and included in the retrieval results up to the current round, and dynamically calculate an average value based on these numbers.

[0059] In some embodiments, in step S303, the system may, when starting the next iteration, compare the number of contributions made by each leaf node with the calculated average value, and pause the acquisition of new document indexes from those leaf nodes whose contributions have exceeded the average value.

[0060] According to some embodiments, the information processing method may further include: counting the co-occurrence frequency of multiple cold data nodes in historical queries; and updating the temperature attributes of the multiple cold data nodes to hot data in response to determining that the co-occurrence frequency exceeds a preset threshold.

[0061] In some embodiments, the system can continuously monitor and count the frequency of multiple different cold data recall conditions co-occurring in historical queries, i.e., the co-occurrence frequency. When the system detects that the co-occurrence frequency of a certain set of cold data recall conditions exceeds a preset threshold within a preset statistical period, it will automatically update the temperature attributes of the leaf nodes corresponding to these recall conditions to hot data.

[0062] Therefore, the above method enables the retrieval system to learn common complex query patterns from user behavior and proactively warm up relevant cold data, avoiding inefficient access to cold data repeatedly for specific query combinations, realizing system self-optimization, and improving response performance to high-frequency complex queries.

[0063] In some embodiments, the preset statistical period and the preset threshold can be calculated or determined a priori or by other means, and are not limited herein.

[0064] According to some embodiments, step S202, traversing and merging the query tree to obtain the retrieval result, may include: dividing the document index list corresponding to each leaf node into a preset K fragments based on a hash function, where K is an integer greater than or equal to 1; constructing K replicas of the query tree, where the leaf node of the i-th replica consists of the i-th fragment of multiple document index lists, where i is an integer from 1 to K; and performing parallel traversal and merging on the K replicas, and merging the obtained K merging results to obtain the retrieval result.

[0065] In some embodiments, before performing the merge, the system can horizontally divide the long list of document indexes corresponding to each leaf node into K preset data fragments based on a hash function. Then, the system can construct K replica trees with the same logical structure as the original query tree and assign each set of data fragments (e.g., the i-th fragment of all long lists) to the i-th replica tree. Finally, the system traverses and merges these K replica trees in parallel and combines the results from each replica tree to obtain the final complete search result.

[0066] By using the above method, a large-scale monolithic query task can be decomposed into multiple small-scale subtasks that can be processed in parallel, thereby making full use of the system's multi-core computing resources, reducing query latency, and improving the system's throughput.

[0067] According to another aspect of this disclosure, an information processing apparatus is provided. The retrieval system includes a first storage structure for storing hot data and a second storage structure for storing cold data. For example... Figure 4 As shown, the apparatus 400 includes: a construction unit 410 configured to construct a query tree based on a query expression, the query expression including multiple recall conditions and at least one logical operator, wherein the query tree includes: multiple leaf nodes corresponding to the multiple recall conditions, each leaf node corresponding to a hot data list obtained from a first storage structure and / or a cold data list obtained from a second storage structure based on the source of the document index list corresponding to the leaf node; and at least one internal node corresponding to the at least one logical operator; and a traversal and merging unit 420 configured to traverse and merge the query tree to obtain retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attribute of the child nodes of the internal node.

[0068] It is understood that the operation and effects of units 410 to 420 in device 400 can be referred to the above description of steps S201 to S202.

[0069] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0070] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0071] refer to Figure 5 The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0072] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0073] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, hard disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, WiFi devices, WiMAX devices, cellular communication devices, and / or the like.

[0074] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods, processes, and / or processes described above. For example, in some embodiments, these methods, processes, and / or processes may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods, processes, and / or processes described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform these methods, processes, and / or processes by any other suitable means (e.g., by means of firmware).

[0075] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0076] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0077] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0078] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0079] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet, and blockchain networks.

[0080] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server, or simply "VPS") services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0081] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0082] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.

Claims

1. An information processing method for a retrieval system, the retrieval system comprising a first storage structure for storing hot data and a second storage structure for storing cold data, the method comprising: A query tree is constructed based on a query expression, wherein the query expression includes multiple recall conditions and at least one logical operator, and the query tree includes: Multiple leaf nodes corresponding to the plurality of recall conditions, wherein each leaf node corresponds to a hot data list obtained from the first storage structure and / or a cold data list obtained from the second storage structure, based on the source of the document index list corresponding to that leaf node; and At least one internal node corresponding to the at least one logical operator; and The query tree is traversed and merged to obtain the retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attribute of the child nodes of the internal node.

2. The method according to claim 1, wherein, Constructing a query tree based on a query expression includes: In response to one of the plurality of recall conditions simultaneously obtaining a hot data list from the first storage structure and a cold data list from the second storage structure, a logical OR operation is performed on the hot data list and the cold data list to generate a leaf node corresponding to the recall condition, wherein the temperature attribute of the leaf node is determined to be cold data.

3. The method according to claim 1 or 2, wherein, Traversing and merging the query tree to obtain the retrieval results includes: For the first internal node corresponding to the logical OR operator in the at least one internal node, in response to determining that the child nodes of the first internal node include both cold data child nodes and hot data child nodes, the merging strategy of the first internal node is determined to be a hot data priority strategy, wherein the hot data priority strategy indicates that the document index corresponding to the cold data child node is returned after the document index corresponding to the hot data child node has been returned.

4. The method according to claim 3, wherein, The temperature attribute of the first internal node is determined to be cold data.

5. The method according to claim 1 or 2, wherein, Traversing and merging the query tree to obtain the retrieval results includes: For the second internal node corresponding to the logical AND operator in the at least one internal node, in response to determining that the child nodes of the second internal node include both cold data child nodes and hot data child nodes, the merging strategy of the second internal node is determined to be a cold data truncation strategy, wherein the cold data truncation strategy indicates that the document index list corresponding to the cold data child node is truncated before merging the hot data child nodes and the cold data child nodes.

6. The method according to claim 5, wherein, Truncation of the document index list corresponding to the cold data child node includes: Load the document index list corresponding to the cold data child node from the second storage structure to the first storage structure; and A preset number of document indexes are retained from the loaded document index list for merging with the document index list corresponding to the hot data sub-node.

7. The method according to claim 5, wherein, The temperature attribute of the second internal node is determined to be thermal data.

8. The method according to claim 1 or 2, wherein, Traversing and merging the query tree to obtain the retrieval results includes: The query tree is iterated and merged multiple times. In each iteration, the number of document indexes originating from each leaf node that have been merged into the retrieval results is counted in real time, and an average value is calculated; and In the next iteration, retrieval of document indexes is paused from leaf nodes where the number of merged document indexes exceeds the average value.

9. The method according to claim 1 or 2, further comprising: Calculate the co-occurrence frequency of multiple cold data nodes in historical queries; as well as In response to determining that the co-occurrence frequency exceeds a preset threshold, the temperature attributes of the multiple cold data nodes are updated to hot data.

10. The method according to claim 1 or 2, wherein, Traversing and merging the query tree to obtain the retrieval results includes: Based on a hash function, the document index list corresponding to each leaf node is divided into a preset K segments, where K is an integer greater than or equal to 1; Construct K replicas of the query tree, wherein the leaf nodes of the i-th replica are composed of the i-th fragment of the plurality of document index lists, where i is an integer from 1 to K; and The K replicas are traversed and merged in parallel, and the resulting K merged results are combined to obtain the retrieval result.

11. An information processing apparatus for a retrieval system, the retrieval system comprising a first storage structure for storing hot data and a second storage structure for storing cold data, the apparatus comprising: A construction unit is configured to construct a query tree based on a query expression, the query expression including multiple recall conditions and at least one logical operator, wherein the query tree includes: Multiple leaf nodes corresponding to the multiple recall conditions, each leaf node corresponding to a hot data list obtained from the first storage structure and / or a cold data list obtained from the second storage structure, based on the source of the document index list corresponding to that leaf node; and At least one internal node corresponding to the at least one logical operator; and The traversal and merging unit is configured to traverse and merge the query tree to obtain the retrieval results, wherein the merging strategy and temperature attribute of each internal node are determined based on the logical operator corresponding to the internal node and the temperature attribute of the child nodes of the internal node.

12. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.

14. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-10.