Data processing method and device

By independently processing offline and real-time data indexing in the vector search system, the problem of insufficient real-time data processing in the existing technology is solved, and the efficient retrieval of real-time data and the optimization of the retrieval system is realized, which is suitable for large-scale data environments.

CN120353799APending Publication Date: 2025-07-22BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510450127.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing vector search technology assumes that the data entering the system is static and cannot effectively process the data flowing in in real time, which limits the real-time search function, and the internal data index construction is closely related to the search service, resulting in insufficient flexibility of the search architecture and difficulty in optimizing and scaling.

Method used

By collecting offline data within a preset time to a persistent database and generating offline data indexes, real-time data is acquired and stored in a temporary database in real time, real-time data index is generated only when the real-time data volume reaches the threshold, the construction of offline and real-time data indexes is independently processed, and the retrieval performance is optimized using distributed architecture and multi-threaded parallel computing.

Benefits of technology

It realizes the index construction of offline and real-time data, improves the performance of real-time data retrieval, optimizes the retrieval system architecture, supports real-time high-performance retrieval of large-scale data, and improves the flexibility of the system and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353799A_ABST
    Figure CN120353799A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps of collecting off-line data within a preset time, storing the off-line data into a persistent database, and generating an off-line data index corresponding to the off-line data based on the persistent database; acquiring real-time data after a preset time in real time, and storing the real-time data in a temporary database; and when the data volume of the real-time data reaches a threshold value, generating a real-time data index corresponding to the real-time data based on the temporary database. According to the embodiment, the index of the offline and real-time data is constructed, necessary support is provided for real-time data retrieval, and the retrieval performance of the real-time data can be greatly improved through the construction of the index of the real-time data; in addition, according to the embodiment of the invention, the construction of the data index is independently processed, an optimization direction is provided for the optimization of the retrieval system architecture, and the optimization of the retrieval system architecture is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular, to a data processing method and apparatus. Background Art

[0002] With the development of deep learning technology, vector retrieval, as a key technology therein, has been widely applied in text retrieval, image / video retrieval, recommendation systems, and anomaly detection. Existing vector retrieval usually assumes that the data entering the system is static, and cannot effectively process the data flowing in real time, which limits the real-time retrieval function. Moreover, the internal data index construction and retrieval service are tightly dependent, resulting in insufficient flexibility of the retrieval architecture and difficulty in optimizing and expanding resources. Summary of the Invention

[0003] In view of this, embodiments of the present invention provide a data processing method and apparatus, which realize the construction of indexes for offline and real-time data. It not only provides necessary support for real-time data retrieval, but also can greatly improve the retrieval performance of real-time data through the construction of real-time data indexes. Additionally, embodiments of the present invention separately process the construction of data indexes, which also provides an optimization direction for the optimization of the retrieval system architecture and helps to optimize the retrieval system architecture.

[0004] To achieve the above object, according to one aspect of the embodiments of the present invention, there is provided a data processing method, including:

[0005] Collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database;

[0006] Real-time acquiring real-time data after the preset time, storing the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, generating a real-time data index corresponding to the real-time data based on the temporary database.

[0007] Optionally, generating an offline data index corresponding to the offline data based on the persistent database includes: performing normalization processing on the offline data in the persistent database, storing the normalized offline standard data in a message middleware; consuming the offline standard data by subscribing to the message middleware, and generating the offline data index according to the offline standard data.

[0008] Optionally, real-time data after the preset time is obtained in real time, including: obtaining the real-time data in real time and performing normalization processing on the real-time data; the normalization processing and the generation of the offline data index are respectively implemented by a data processing component and an offline index construction component, and the data processing component and the offline index construction component are independent of each other, so as to allocate resources by respectively configuring the concurrency degrees of the data processing component and the offline index construction component.

[0009] Optionally, the generation of the real-time data index is implemented by an online service node; after the real-time data index corresponding to the real-time data is generated, the method further includes: storing the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index into the online service node, so that the online service node uses the stored data to execute corresponding services; the online function node is applied to a data retrieval service, and multiple online service nodes cooperate to execute the data retrieval service, and each online service node supports sharding and backing up the stored data in other online service nodes.

[0010] According to a second aspect of an embodiment of the present invention, a data retrieval method is provided, including:

[0011] Obtaining a retrieval statement corresponding to a retrieval task;

[0012] Retrieving target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index; the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index are determined according to any method provided in the first aspect of an embodiment of the present invention.

[0013] Optionally, retrieving target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index includes: dividing the offline data index, the real-time data index, and the real-time data that does not have a real-time data index into multiple data blocks; splitting and vector-converting the retrieval statement to obtain a corresponding retrieval vector set; determining the number of grouped vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task, and then grouping the retrieval vector set according to the grouped vector data to obtain retrieval vector groups; allocating a corresponding retrieval thread to each data block, retrieving each retrieval vector group in a multi-thread parallel manner, placing the retrieval results of each thread in the corresponding result storage space, and based on the result storage space, merging the retrieval results of each thread, and then obtaining the target data similar to the retrieval statement.

[0014] Optionally, determining the number of grouped vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task includes: calculating the memory of the retrieval vectors according to the dimensions of the retrieval vectors in the retrieval vector set; calculating the memory of the retrieval results corresponding to the retrieval vectors according to the preset number of similar retrievals and the number of the data blocks; calculating the sum of the memory of the retrieval vectors and the memory of the retrieval results corresponding to the retrieval vectors to obtain the memory requirement of the retrieval vectors; and calculating the number of grouped vectors of the retrieval vector set according to the memory requirement and the storage capacity of the cache space corresponding to the retrieval task.

[0015] According to a third aspect of an embodiment of the present invention, there is provided a data processing apparatus, including:

[0016] An off-line index generation module, configured to collect off-line data within a preset time, store the off-line data in a persistent database, and generate an off-line data index corresponding to the off-line data based on the persistent database;

[0017] A real-time index generation module, configured to obtain real-time data after the preset time in real time, store the real-time data in a temporary database; and when the data volume of the real-time data reaches a threshold, generate a real-time data index corresponding to the real-time data based on the temporary database.

[0018] According to a fourth aspect of an embodiment of the present invention, there is provided a data retrieval apparatus, including:

[0019] A statement acquisition module, configured to acquire a retrieval statement corresponding to a retrieval task;

[0020] A data retrieval module, configured to retrieve target data similar to the retrieval statement from an off-line data index, a real-time data index, and real-time data without a real-time data index in a temporary database; the off-line data index, the real-time data index, and the real-time data without a real-time data index in the temporary database are determined according to any method described in the first aspect of an embodiment of the present invention.

[0021] According to a fifth aspect of an embodiment of the present invention, there is provided a data processing electronic device, including:

[0022] One or more processors;

[0023] A storage device, configured to store one or more programs,

[0024] When the one or more programs are executed by the one or more processors, the one or more processors implement the methods provided in the first aspect and the second aspect of an embodiment of the present invention.

[0025] According to a sixth aspect of an embodiment of the present invention, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, the methods provided in the first aspect and the second aspect of the embodiments of the present invention are implemented.

[0026] According to a seventh aspect of an embodiment of the present invention, there is provided a computer program product including a computer program, and when the computer program is executed by a processor, the methods provided in the first aspect and the second aspect of the embodiments of the present invention are implemented.

[0027] One embodiment of the invention has the following advantages or beneficial effects: By collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database; obtaining real-time data after the preset time in real time and storing the real-time data in a temporary database; when the amount of real-time data reaches a threshold, generating a real-time data index corresponding to the real-time data based on the temporary database, the technical solution realizes the construction of indexes for offline and real-time data, not only provides necessary support for real-time data retrieval, but also can greatly improve the retrieval performance of real-time data through the construction of real-time data indexes; additionally, the embodiment of the present invention separately processes the construction of data indexes, which also provides an optimization direction for the optimization of the retrieval system architecture and helps with the optimization of the retrieval system architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings are used to better understand the present invention and do not constitute an improper limitation of the present invention. Among them:

[0029] Figure 1 is a schematic diagram of the main process of the data processing method according to an embodiment of the present invention;

[0030] Figure 2 is a flowchart of the data processing method according to a referenceable embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of the composition of the retrieval system according to an embodiment of the present invention;

[0032] Figure 4 is a flowchart of the data retrieval method according to an embodiment of the present invention;

[0033] Figure 5 is a schematic diagram of the principle of the data retrieval method according to an embodiment of the present invention;

[0034] Figure 6 is a flowchart of the data retrieval method according to a referenceable embodiment of the present invention;

[0035] Figure 7 is a schematic diagram of the structure of the data vector parallel retrieval calculation according to an embodiment of the present invention;

[0036] Figure 8 It is a schematic diagram of the main modules of a data processing device according to an embodiment of the present invention;

[0037] Figure 9 It is a schematic diagram of the main modules of a data retrieval device according to an embodiment of the present invention;

[0038] Figure 10 It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;

[0039] Figure 11 It is a schematic diagram of the structure of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners

[0040] It should be noted that in the technical solution of the present invention, in terms of the collection / collection, update, analysis, use, transmission, storage, etc. of the user's personal information involved, it complies with the provisions of relevant laws and regulations, is used for legal and reasonable purposes, is not shared, leaked or sold outside these legal uses, etc., and is subject to the supervision and management of the national regulatory authorities. Necessary measures should be taken for the user's personal information to selectively prevent the use or access to personal information data to prevent illegal access to such personal information data, ensure that the personnel authorized to access the personal information data comply with the provisions of relevant laws and regulations, and ensure the security of the user's personal information. In addition, once these user personal information data are no longer needed, the risk should be minimized by restricting or even prohibiting data collection and / or deleting data.

[0041] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.

[0042] Existing vector retrieval usually assumes that the data entering the system is static, and cannot effectively process the data flowing in real-time, which limits the real-time retrieval function. Moreover, the internal data index construction and retrieval service are closely dependent, resulting in insufficient flexibility of the retrieval architecture, making it difficult to optimize and expand resources, and unable to well meet the actual needs.

[0043] To solve the above problems existing in the prior art, the present invention proposes a data processing method. For the offline data in the retrieval system, the corresponding offline data index is generated, and the real-time data in the retrieval system is obtained in real time. A certain scale of real-time data is used to generate a real-time data index online, realizing the construction of indexes for offline and real-time data. This not only provides necessary support for real-time data retrieval, but also can greatly improve the retrieval performance of real-time data through the construction of real-time data indexes. Additionally, the embodiments of the present invention separately process the construction of data indexes, which also provides an optimization direction for the optimization of the retrieval system architecture and helps optimize the retrieval system architecture.

[0044] Figure 1 is a schematic diagram of the main process of the data processing method according to an embodiment of the present invention. As Figure 1 shown, the data processing method according to the embodiment of the present invention includes the following steps S101 to S102.

[0045] Step S101: Collect the offline data within a preset time, store the offline data in a persistent database, and generate an offline data index corresponding to the offline data based on the persistent database.

[0046] Specifically, in the field of data retrieval (such as using data indexes for data retrieval), the field of database optimization (such as uniquely identifying data through data indexes or establishing associations between various tables), the field of file system management (such as quickly locating file metadata and file data through indexes), etc., data indexes play a key role in data retrieval, positioning, relationship maintenance, and resource management. The embodiments of the present invention provide a processing method for real-time data to construct data indexes with real-time requirements, thereby meeting the real-time requirements of its application fields (such as the field of data retrieval, the field of database optimization, the field of file system management, etc.).

[0047] It can be understood that the data of various systems basically includes offline data. The offline data in the embodiments of the present invention is usually the full offline data, and the collection of offline data is generally timed. Offline data is calculated and generated at fixed time intervals. For example, the order volume of the previous day is statistically collected at 6:00 every morning. Therefore, according to the timing of the offline data, the offline data within a preset time is collected, and the obtained offline data is stored in a persistent database by the system. And according to the preset index construction rules, such as hash index, B-tree index, etc., an offline data index corresponding to the offline data in the persistent database is generated.

[0048] Step S102: Real-time obtain the real-time data after the preset time, store the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, generate a real-time data index corresponding to the real-time data based on the temporary database.

[0049] Specifically, the above only completes the construction of the data index for offline data. In a system that supports real-time performance, real-time data in the current period will also be received. The received real-time data is stored in a temporary database. Of course, the real-time data in the temporary database and the above offline data index can be directly used as the current full-scale data resource of the system. However, directly using real-time data will consume more resources and is not conducive to ensuring the system efficiency, especially for large amounts of data such as pictures, audio, and video. After receiving the real-time data, the embodiment of the present invention generates an online data index for the real-time data. In addition, considering the importance of online resources to the system, the embodiment of the present invention presets a threshold. Only when the data volume of the real-time data in the temporary database reaches this threshold, will the data index generation interface be called for this part of the real-time data to generate the corresponding real-time data index in real time.

[0050] It can be understood that the temporary database can store the real-time data that already has a real-time data index in a separate table, or store the real-time data whose data volume has not reached the threshold in a pre-allocated real-time storage area, so as to achieve the purpose of distinguishing the real-time data with a real-time data index and the real-time data that currently does not have a real-time data index.

[0051] Through the construction of the offline data index for offline data and the construction of the real-time data index for real-time data as described above, it can be seen that all data resources in the system are divided into offline data indexes, real-time data indexes, and a small amount of real-time data in the temporary database that does not have a real-time data index.

[0052] The embodiment of the present invention considers the real-time data in the system, and when the data volume of the real-time data reaches the threshold, it will generate a real-time data index corresponding to the real-time data, which is convenient for the system to directly use the real-time data index to efficiently execute tasks, and avoids the impact of frequently generating real-time data indexes on online tasks, achieving a balance between efficiency and resources and providing a reliable solution for the real-time performance of the system.

[0053] Figure 2 is a flowchart of a data processing method according to a reference embodiment of the present invention. As another embodiment of the present invention, as Figure 2 shown, the data processing method may include:

[0054] Step S201, collect offline data within a preset time, and store the offline data in a persistent database.

[0055] Step S202: Perform standardization processing on the offline data in the persistent database, store the standardized offline standard data in the message middleware; consume the offline standard data by subscribing to the message middleware, and generate the offline data index according to the offline standard data.

[0056] Specifically, in the embodiments of the present invention, the method for constructing offline and real-time data indexes is applied to a retrieval system, and the construction of specific data indexes is described by taking the retrieval system as an example. Considering that the amount of offline data collected by the retrieval system is large, although the processing of the offline data index has relatively low requirements for timeliness, in order to ensure the integrity of the data, in the embodiments of the present invention, a message middleware is created inside the retrieval system to ensure the integrity of the data and play a role in peak shaving and valley filling during large traffic. Additionally, considering that the data sources are generally relatively rich and the data formats from different sources are different, in order to facilitate the subsequent generation of the offline data index, first perform standardization processing on the offline data in the persistent database, uniformly convert the offline data into the structure required by the code program for generating the offline data index, and then store the standardized offline standard data in the message middleware of the retrieval system.

[0057] Furthermore, the offline data index generation module in the retrieval system will subscribe to the message middleware, obtain the standardized offline standard data stored in the message middleware, and based on the structure of the offline standard data, call the code interface of the offline data index to obtain the offline data index corresponding to the offline standard data.

[0058] Step S203: Obtain the real-time data in real time, perform standardization processing on the real-time data, and store it in the temporary database.

[0059] Under normal circumstances, the real-time data and the offline data are obtained through the same data interface. Then, by comparing whether the timestamp of the received data is within the acquisition time of the offline data, it is determined whether the data is real-time data. For real-time data, similar to the above, in order to facilitate the subsequent construction of the real-time data index, data standardization processing also needs to be performed, and the standardized real-time data is stored in the temporary database.

[0060] It can be understood that the real-time data is only the temporary data used to implement the real-time retrieval of the retrieval system at the current moment. After the acquisition time of the next offline data arrives, the real-time data will be persisted into a new round of offline data. At the same time, the offline data index of the new round of offline data will be compared with the real-time data index of the previous round and the real-time data stored in the temporary database without a real-time data index for duplicate data comparison and deduplication to ensure the uniqueness of the data.

[0061] According to an embodiment of the present invention, the standardization process and the generation of the offline data index are respectively implemented by a data processing component and an offline index construction component. The data processing component and the offline index construction component are independent of each other, so as to allocate resources by respectively configuring the concurrency of the data processing component and the offline index construction component.

[0062] Specifically, considering that the methods for standardizing offline data and real-time data are similar, the retrieval system integrates the standardization processing ability into a data processing component. The data processing component completes the standardization processing of offline data and real-time data, and after completing the data standardization processing, the data processing component stores the standardized data in a message middleware. Additionally, the data processing component also has the ability to identify offline data and real-time data. The data processing component obtains data from the data source of the file system and determines whether the data is offline data or real-time data based on the timestamp of the data.

[0063] Additionally, the retrieval system uses an offline index construction component to construct the offline data index for offline data. The offline index construction component subscribes to the message middleware, obtains the offline standard data, and then generates the offline data index corresponding to the offline standard data.

[0064] It should be noted that the data processing component and the offline index construction component in the embodiment of the present invention are independent and decoupled from each other. The message middleware serves as the data channel between the data processing component and the offline index construction component. The data processing component stores the processed data in the message middleware, and the offline index construction component subscribes to the message middleware and generates the corresponding offline data index based on the processed data. The independence between the data processing component and the offline index construction component facilitates the retrieval system to allocate appropriate concurrency according to the status of the data processing component and the offline index construction component to achieve the purpose of system resource allocation. For example, if the current processing speed of the data processing component is a bit slow, the concurrency of the data processing component can be increased at this time to allocate more resources to the data processing component to ensure the processing speed of the data processing component, realizing the flexible and independent allocation of the resources required by the data processing component and the offline index construction component according to needs.

[0065] It can be understood that in order to improve the performance of the retrieval system in a large-scale data environment in the embodiment of the present invention, a distributed retrieval system architecture is constructed. The data processing component and the offline index construction component are encapsulated in an offline node. The retrieval system internally includes multiple offline nodes, and each offline node is uniformly managed and allocated by a distributed coordination service.

[0066] Step S204, when the data volume of the real-time data reaches a threshold, generate a real-time data index corresponding to the real-time data based on the temporary database.

[0067] According to an embodiment of the present invention, the generation of the real-time data index is implemented by an online service node; after generating the real-time data index corresponding to the real-time data, the method further includes: storing the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index into the online service node, so that the online service node uses the stored data to execute the corresponding service; the online function node is applied to a data retrieval service, and multiple online service nodes cooperate to execute the data retrieval service, and each online service node supports sharding and backing up the stored data in other online service nodes.

[0068] Specifically, the retrieval system integrates the generation capability of the above real-time data index into an online service node, and after generating the real-time data index, stores the already generated offline data index, the real-time data in the temporary database that does not have a real-time data index, together with the real-time data index, into the online service node as the data source of the online service node, to execute its corresponding service for the online service node, such as the database optimization service or the data retrieval service mentioned above.

[0069] In addition, in order to support the retrieval of large-scale data, the data retrieval service of the embodiment of the present invention adopts distributed data retrieval, that is, multiple online service nodes are deployed, and multiple online service nodes cooperate to execute the data retrieval service. And in order to improve the reliability of the data source stored in each online service node, the embodiment of the present invention also horizontally scales each online service node. In the data retrieval service cluster, each online service node simultaneously undertakes two roles of sharding and backup. A node can be both the main storage and processing unit of the data source of this node and the backup node of the data source of other online service nodes. The system can flexibly adjust the number of each node according to factors such as data scale and request frequency to achieve horizontal expansion, so as to meet the real-time high-performance retrieval requirements of large-scale data.

[0070] Table 1 below lists an example of the storage sharding of the online service node:

[0071] Table 1 shows the sharding allocation results of 5 online service nodes. Each online service node can store multiple shards, and each shard is a backup of the shard data of other nodes. When node 0 fails and goes offline, other online service nodes can still provide the retrieval service for all the shards stored by node 0. Through sharding, large-scale data can be distributed among multiple online service nodes, improving the cluster's ability to process large-scale data.

[0072] Understandably, as the specific execution entity of the service, the online service node in the embodiments of the present invention is mainly responsible for generating real-time data indexes, and aggregating offline data indexes and real-time data without real-time data indexes to execute corresponding services, such as data retrieval services; and the generation of the above-mentioned offline data indexes is jointly completed by the data processing component and the offline data index construction component. The embodiments of the present invention decouple these two parts and define them as the online part and the offline part. The offline part is responsible for identifying real-time data and offline data, data standardization processing, and the construction of offline data indexes; the online part is responsible for the construction of real-time data search indexes and the execution of specific system functions, such as data index services.

[0073] That is, the embodiments of the present invention propose a system architecture with separated storage and computing, which decouples the index construction and the main services provided by the system. Although the online part is also involved in the construction of real-time data indexes, the weight of the construction of real-time data indexes is much lighter than that of the construction of offline data indexes. Generally, it is a system architecture with separated storage and computing. In this way, the system can flexibly and independently control the resource occupancy of index construction and online services. For example, for a retrieval system, when the index construction task volume is small and the QPS (number of query requests processed per second) of the retrieval is high, the resources for index construction can be reduced and the resources for retrieval services can be increased, realizing the optimization and elastic expansion of resources, so that the computing and storage resources can be independently expanded according to needs, improving the flexibility and efficiency of the system. This architecture with separated storage and computing helps to reduce costs and improve performance, and is particularly advantageous in scenarios where resources need to be dynamically adjusted.

[0074] Figure 3 It is a schematic diagram of the composition of the retrieval system in the embodiments of the present invention. The retrieval system is mainly divided into an offline part and an online part, and the real-time message stream and the file system serve as the data transmission channels between the offline part and the online part. The retrieval system in the embodiments of the present invention is a distributed real-time retrieval system, which supports real-time retrieval of large-scale data.

[0075] The offline part is mainly used to generate offline data indexes. The distributed coordination service Coordinator is used to dynamically allocate and manage each offline node, ensuring the efficient cooperation of each offline node (data processing component + offline index construction component) in the distribution and the stable operation of the system. The distributed coordination service provides a centralized service, responsible for functions such as the registration of each offline node, the issuance of scheduling tasks, status query, and alarm. When each offline node in the distribution starts, they first register with the distributed coordination service Coordinator, providing detailed information of the offline node (such as ID and available resources). Coordinator assigns appropriate tasks to each offline node to ensure load balancing. At the same time, Coordinator also regularly polls the status of all offline nodes to monitor their health and task progress. When an anomaly (such as high load or downtime) is detected, Coordinator immediately issues an alarm, notifying the administrator by means of logs, emails or text messages, etc., for a quick response. The green arrow in the figure represents the control link.

[0076] For each offline node, it includes a data processing component and an offline index construction component. The data processing component reads data from the data source where the original data is stored in the file system, and determines whether it is offline data or real-time data according to the timestamp of the data. For real-time data, after standardizing it according to the standard processing requirements of real-time data, it is stored in the real-time data message queue in the real-time message stream; for offline data, the offline data is first stored in the persistent database, and then standardized according to the standard processing requirements of offline data (of course, the standard processing requirements of offline data may be the same as those of real-time data), and the standardized offline standard data is stored in the message middleware. The offline data index construction component in the offline node subscribes to this message middleware, obtains the processed offline standard data, generates the corresponding offline data index according to the offline standard data, and stores the generated offline data index in the index file in the file system. The red arrow in the figure represents the data link.

[0077] The online part is mainly used for online retrieval of data and consists of a query-proxy and searchers. The query-proxy is responsible for request routing and forwarding, specifically including the processing of upstream queries, request load balancing, and the merging of results from multiple searchers. The query-proxy is also responsible for monitoring the status of searcher nodes. In case of abnormal status, it will give an alarm in time and stop sending retrieval requests to the abnormal searcher nodes. The query-proxy has no concept of sharding and improves the response processing ability by increasing the number of backups. The searcher nodes store sharded data, are responsible for processing real-time requests, and perform data retrieval. They play both the roles of sharding and backup. The searcher nodes are the computing nodes of the online retrieval system, capable of supporting real-time retrieval and ensuring the real-time nature and integrity of data.

[0078] The searcher, an online service node, mainly performs data retrieval tasks. First, it reads the real-time data from the real-time data message queue in the file system and stores it in a temporary database. When the amount of real-time data reaches a threshold, corresponding real-time data indexes are generated for the real-time data that reaches the threshold. At the same time, it also obtains the offline data indexes from the file system. At this time, the data in the online service node includes real-time data indexes, offline data indexes, and a small amount of real-time data whose data volume has not reached the threshold and does not have real-time data indexes. When constructing the real-time data indexes, low-concurrency computing tasks can be preferably used to avoid affecting the online retrieval service. Then, according to the received data retrieval task, data retrieval is performed using the data stored in the online service node (real-time data indexes, offline data indexes, and a small amount of real-time data whose data volume has not reached the threshold and does not have real-time data indexes). The blue arrows in the figure represent the retrieval link.

[0079] Based on the construction of offline data indexes and real-time data indexes, the embodiment of the present invention provides a composition architecture of a retrieval system, which not only supports distributed large-scale data real-time retrieval but also has a separated storage and computing architecture, facilitating flexible and independent regulation of the resource allocation for index construction and online retrieval, thereby enhancing the resource optimization ability and expansion ability of the retrieval system and better meeting the needs of large-scale data scenarios.

[0080] Based on the above architecture of the retrieval system, the embodiment of the present invention also provides a data retrieval method. Figure 4 It is a flowchart of the data retrieval method of the embodiment of the present invention. As Figure 4 shown, the data retrieval method may include:

[0081] Step S401, obtain the retrieval statement corresponding to the retrieval task.

[0082] Specifically, based on the architecture of the above-provided retrieval system, an embodiment of the present invention further provides a corresponding data retrieval method. The online service node receives a real-time data retrieval request and obtains a retrieval statement from the data retrieval request.

[0083] Step S402: Retrieve target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index; the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index are determined by the above data processing method.

[0084] Specifically, according to the obtained retrieval statement, using a retrieval statement executor, the retrieval statement is split into independent execution operations, and each execution operation will call a data query view to respectively retrieve similar target indexes from the offline data index and the real-time data index stored in the online service node using approximate nearest neighbor search ANN; at the same time, for the part of the real-time data stored in the online service node that does not have a real-time data index, a traditional data retrieval method is used for similarity retrieval. Finally, the retrieval results of these three parts are merged to obtain the target data finally retrieved.

[0085] Figure 5 It is a schematic diagram of the principle of the data retrieval method according to an embodiment of the present invention. Taking an online service node of the retrieval system as an example, it receives a real-time data retrieval request, obtains a retrieval statement from it, uses a retrieval statement executor to split the retrieval statement into independent execution operations, calls a data query view, and respectively queries and retrieves similar data from the real-time data that does not have a real-time data index, the real-time data index, and the offline data index stored in the current online service node, and merges the similar data of the three parts to obtain target data similar to the retrieval statement. It can be understood that the real-time data comes from the real-time data message queue of the retrieval system, and the offline data index comes from the index file of the retrieval system. The blue arrow in the figure represents the query link, the green arrow represents the construction of the real-time data index, the red arrow represents the real-time data, and the yellow arrow represents the offline data.

[0086] Based on the real-time retrieval system architecture with separation of storage and computing, an embodiment of the present invention uses the offline data index generated by the offline part, and the real-time data index generated by the online part with low concurrency, together with the current small amount of real-time data that has not generated a real-time data index as retrieval data sources, and respectively retrieves similar data similar to the retrieval statement from each retrieval data source, and then forms the target data of the retrieval result. An embodiment of the present invention uses a real-time data index instead of real-time data for data retrieval, which can greatly improve the retrieval efficiency and thus save precious online resources.

[0087] Figure 6is a flowchart of a data retrieval method according to a reference embodiment of the present invention. As another embodiment of the present invention, as Figure 6 shown, the data retrieval method may include:

[0088] Step S601, obtaining a retrieval statement corresponding to a retrieval task.

[0089] Step S602, dividing the offline data index, the real-time data index, and the real-time data without a real-time data index into multiple data blocks; splitting and vector-converting the retrieval statement to obtain a corresponding retrieval vector set.

[0090] Specifically, in the embodiment of the present invention, a vector retrieval method is used to find target data similar to the retrieval statement in the stored data of the online service node. Therefore, each offline data index, real-time data index, and real-time data without a real-time data index needs to be converted into a corresponding vector representation. For the sake of convenience of expression, the vectors corresponding to each offline data index, real-time data index, and real-time data without a real-time data index are uniformly represented by a data vector set. Similarly, the retrieval statement is also split to obtain multiple retrieval elements, and then vector-converted to obtain a retrieval vector set.

[0091] The retrieval based on vector index can be described as: given a set of retrieval vector sets with a length of M, and a set of data vector sets with a length of N, how to quickly find the k most similar (nearest in distance) data vectors for each retrieval vector. In the embodiment of the present invention, the multi-threaded parallel computing of the system is utilized to improve the efficiency of data vector retrieval, and the data vector set is divided into multiple data blocks, and each data block contains the same number of data vectors.

[0092] Step S603, determining the number of grouped vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task, and then grouping the retrieval vector set according to the grouped vector data to obtain a retrieval vector group.

[0093] Specifically, the embodiment of the present invention is based on optimizing the cache miss of the online service node. Considering that the three-level L3 cache space is much larger than the L1 / L2 cache, in order to store the accessed data vectors in the cache as much as possible, it is desired that the three-level L3 cache can store the data of a retrieval vector group and all its corresponding retrieval results for multiple parallel retrievals, minimizing the synchronization overhead. The embodiment of the present invention queries and determines the number of grouping vectors of the current retrieval vector set according to the number of data blocks and the storage capacity of the three-level L3 cache space of the online service node where the retrieval task is located, in combination with a reference table of the preset recommended number of grouping vectors. It should be noted that the online service node in the embodiment of the present invention is usually a machine device. Then, according to the determined grouping vector data, the retrieval vector set is grouped to obtain multiple retrieval vector groups. Assuming the number of grouping vectors is s, each retrieval vector group containing s retrieval vectors is obtained.

[0094] According to an embodiment of the present invention, determining the number of grouping vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task includes: calculating the memory of the retrieval vector according to the dimension of the retrieval vector in the retrieval vector set; calculating the memory of the retrieval result corresponding to the retrieval vector according to the preset number of similar retrievals and the number of data blocks; calculating the sum of the memory of the retrieval vector and the memory of the retrieval result corresponding to the retrieval vector to obtain the memory requirement of the retrieval vector; and calculating the number of grouping vectors of the retrieval vector set according to the memory requirement and the storage capacity of the cache space corresponding to the retrieval task.

[0095] Specifically, it is set that each retrieval vector in the retrieval vector set has the same dimension d, and the vector distances between the retrieval vectors in the online service node are the same. Since it is of floating-point type, is used. Correspondingly, the memory representation of a retrieval vector in the retrieval vector set can be obtained as: . It is set that the number of similar retrievals is k (the meaning of similar retrieval is to take the top k with the highest similarity), the number of data blocks is t, and the storage position offset of the retrieval result in the storage space is of int64 type, so is used. Then, the memory of the retrieval result corresponding to a retrieval vector is: . Furthermore, calculating the sum of the memory of a retrieval vector and the memory of the retrieval result corresponding to a retrieval vector is the memory requirement of a retrieval vector. Finally, calculating the ratio of the storage capacity of the three-level L3 cache space to the memory requirement is the number of grouping vectors s of the retrieval vector set, and there is a formula: .

[0096] Step S604: Allocate corresponding retrieval threads for each of the data blocks, retrieve each retrieval vector group in a multi-thread parallel manner, place the retrieval results of each thread in the corresponding result storage space, and based on the result storage space, merge the retrieval results of each thread, thereby obtaining target data similar to the retrieval statement.

[0097] Specifically, the embodiment of the present invention provides a more fine-grained multi-thread parallel method, which allocates threads to the data vector set (with a length of N), rather than the retrieval vector set (with a length of M) in the traditional implementation. Because in practice, the length N of the data vector set is usually much larger than the length M of the retrieval vector set. When the length M of the retrieval vector set is relatively small, it is difficult to fully utilize the multi-core parallelism of the CPU. Therefore, the embodiment of the present invention allocates corresponding retrieval threads for each data block.

[0098] Furthermore, retrieve each retrieval vector group in a multi-thread parallel manner, and place the retrieval results of each thread separately in its corresponding result storage space (result heap). This can minimize the synchronization overhead. For example, when obtaining the retrieval result corresponding to the second retrieval vector of the retrieval vector group, only the second retrieval results in the result heaps corresponding to each thread need to be merged. Finally, based on the combination of the retrieval results of each retrieval vector, the target data similar to the retrieval statement in the online service node is obtained.

[0099] Figure 7 is a schematic structural diagram of the parallel retrieval calculation of data vectors in the embodiment of the present invention. The data vector set is divided into t data blocks, and each data block contains b data vectors; the retrieval vector set is divided into w retrieval vector groups, and each retrieval vector group contains s retrieval vectors. Allocate corresponding retrieval threads for each data block, retrieve each retrieval vector group in a multi-thread parallel manner, use a heap to calculate the top k (topk) retrieval results of each retrieval vector, record the retrieval results in the result heap, and the H of the result heap i,j represents the topk retrieval results of the retrieval vector q calculated by the i-th thread j . After the parallel retrieval calculation is completed, merge the heaps H 1,j , H 2,j , H 3,j , ……, H t-1,j to obtain the result of the q j vector.

[0100] From the perspective of minimizing the CPU cache misses of the online service node and making full use of multi-core parallelism, the embodiment of the present invention optimizes the calculation process of data vector retrieval in a parallel manner by calculating the number s of grouped vectors of the retrieval vector set and allocating corresponding retrieval threads based on data blocks, thereby improving the speed and efficiency of data retrieval.

[0101] Figure 8 It is a schematic diagram of the main modules of the data processing device according to an embodiment of the present invention. As Figure 8 shown, the data processing device 800 mainly includes an offline index generation module 801 and a real-time index generation module 802.

[0102] The offline index generation module 801 is used to collect offline data within a preset time, store the offline data in a persistent database, and generate an offline data index corresponding to the offline data based on the persistent database;

[0103] The real-time index generation module 802 is used to obtain real-time data after the preset time in real time and store the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, a real-time data index corresponding to the real-time data is generated based on the temporary database.

[0104] According to an embodiment of the present invention, the offline index generation module 801 is further used to: perform standardization processing on the offline data in the persistent database, store the standardized offline standard data in a message middleware; consume the offline standard data by subscribing to the message middleware, and generate the offline data index according to the offline standard data.

[0105] According to another embodiment of the present invention, the real-time index generation module 802 is further used to: obtain the real-time data in real time and perform standardization processing on the real-time data; the standardization processing and the generation of the offline data index are respectively implemented by a data processing component and an offline index construction component, and the data processing component and the offline index construction component are independent of each other, so as to allocate resources by respectively configuring the concurrency degrees of the data processing component and the offline index construction component.

[0106] According to still another embodiment of the present invention, the generation of the real-time data index is implemented by an online service node; the data processing device 800 further includes an index data storage module (not shown in the figure), which is used to: after generating the real-time data index corresponding to the real-time data, store the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index in the online service node, so that the online service node uses the stored data to execute corresponding services; the online function node is applied to data retrieval services, and multiple online service nodes cooperate to execute the data retrieval service, and each online service node supports sharding and backing up the stored data in other online service nodes.

[0107] Figure 9 It is a schematic diagram of the main modules of the data retrieval device according to an embodiment of the present invention. As Figure 9As shown, the data retrieval device 900 mainly includes a statement acquisition module 901 and a data retrieval module 902.

[0108] The statement acquisition module 901 is used to acquire a retrieval statement corresponding to a retrieval task;

[0109] The data retrieval module 902 is used to retrieve target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index; the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index are determined according to any of the above data processing methods.

[0110] According to an embodiment of the present invention, the data retrieval module 902 is further used to: divide the offline data index, the real-time data index, and the real-time data that does not have a real-time data index into multiple data blocks; split and vector-convert the retrieval statement to obtain a corresponding retrieval vector set; determine the number of grouped vectors of the retrieval vector set according to the number of data blocks and the storage capacity of the cache space corresponding to the retrieval task, and then group the retrieval vector set according to the grouped vector data to obtain retrieval vector groups; allocate corresponding retrieval threads to each data block, retrieve each retrieval vector group in a multi-thread parallel manner, place the retrieval results of each thread in the corresponding result storage space, and based on the result storage space, merge the retrieval results of each thread to obtain the target data similar to the retrieval statement.

[0111] According to another embodiment of the present invention, the data retrieval module 902 is further used to: calculate the memory of the retrieval vector according to the dimension of the retrieval vector in the retrieval vector set; calculate the memory of the retrieval result corresponding to the retrieval vector according to the preset number of similar retrievals and the number of data blocks; calculate the sum of the memory of the retrieval vector and the memory of the retrieval result corresponding to the retrieval vector to obtain the memory requirement of the retrieval vector; calculate the number of grouped vectors of the retrieval vector set according to the memory requirement and the storage capacity of the cache space corresponding to the retrieval task.

[0112] Figure 10 It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied.

[0113] As Figure 10As shown, the system architecture 1000 may include terminal devices 1001, 1002, 1003, a network 1004, and a server 1005. The network 1004 is used to provide a medium for communication links between the terminal devices 1001, 1002, 1003 and the server 1005. The network 1004 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0114] Users can use the terminal devices 1001, 1002, 1003 to interact with the server 1005 through the network 1004 to receive or send messages, etc. Various communication client applications, such as data processing applications (only as an example), may be installed on the terminal devices 1001, 1002, 1003.

[0115] The terminal devices 1001, 1002, 1003 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.

[0116] The server 1005 may be a server that provides various services, such as a background management server (only as an example) that provides support for data processing performed by users using the terminal devices 1001, 1002, 1003. The background management server may collect offline data within a preset time, store the offline data in a persistent database, and generate an offline data index corresponding to the offline data based on the persistent database; obtain real-time data after the preset time in real time, store the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, generate a real-time data index corresponding to the real-time data based on the temporary database, etc., and feedback the processing result to the terminal device.

[0117] It should be noted that the data processing method provided by the embodiments of the present invention is generally executed by the server 1005. Correspondingly, the data processing device is generally set in the server 1005.

[0118] It should be understood that Figure 10 the numbers of terminal devices, networks, and servers in

[0119] are merely illustrative. According to actual needs, there may be any number of terminal devices, networks, and servers. Figure 11 Figure 11 Figure 11 is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. The terminal device or server shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0120] As Figure 11As shown, computer system 1100 includes a central processing unit (CPU) 1101, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1102 or a program loaded from a storage section 1108 into a random access memory (RAM) 1103. In the RAM 1103, various programs and data required for the operation of the system 1100 are also stored. The CPU 1101, ROM 1102, and RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.

[0121] The following components are connected to the I / O interface 1105: an input section 1106 including a keyboard, a mouse, etc.; an output section 1107 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1108 including a hard disk, etc.; and a communication section 1109 including a network interface card such as a LAN card, a modem, etc. The communication section 1109 performs communication processing via a network such as the Internet. A drive 1110 is also connected to the I / O interface 1105 as needed. A removable medium 1111, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 1110 as needed so that a computer program read therefrom can be installed into the storage section 1108 as needed.

[0122] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1109, and / or installed from the removable medium 1111. When the computer program is executed by a central processing unit (CPU) 1101, the above functions defined in the system of the present invention are executed.

[0123] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination thereof.

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The units involved in the embodiments of the present invention can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: A processor includes: an offline index generation module and a real-time index generation module.

[0126] Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the offline index generation module can also be described as "a module for collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database".

[0127] On the other hand, the present invention also provides a computer-readable medium, which can be included in the device described in the embodiments; or can exist alone without being assembled into the device. The computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database; obtaining real-time data after the preset time in real time, storing the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, generating a real-time data index corresponding to the real-time data based on the temporary database.

[0128] According to the technical solution of the embodiments of the present invention, the following advantages or beneficial effects are achieved: By collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database; obtaining real-time data after the preset time in real time, storing the real-time data in a temporary database; when the data volume of the real-time data reaches a threshold, generating a real-time data index corresponding to the real-time data based on the temporary database, the construction of indexes for offline and real-time data is realized, which not only provides necessary support for real-time data retrieval, but also can greatly improve the retrieval performance of real-time data through the construction of real-time data indexes; in addition, the embodiments of the present invention separately process the construction of data indexes, which also provides an optimization direction for the optimization of the retrieval system architecture and helps to optimize the retrieval system architecture.

[0129] The specific implementation manners do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A data processing method, characterized in that, Including: Collecting offline data within a preset time, storing the offline data in a persistent database, and generating an offline data index corresponding to the offline data based on the persistent database; Obtaining real-time data after the preset time in real time, and storing the real-time data in a temporary database; When the data volume of the real-time data reaches a threshold, generating a real-time data index corresponding to the real-time data based on the temporary database.

2. The method according to claim 1, wherein Generating an offline data index corresponding to the offline data based on the persistent database, including: Performing normalization processing on the offline data in the persistent database, and storing the normalized offline standard data in a message middleware; Consuming the offline standard data by subscribing to the message middleware, and generating the offline data index according to the offline standard data.

3. The method according to claim 2, characterized in that, Obtaining real-time data after the preset time in real time, including: Obtaining the real-time data in real time and performing normalization processing on the real-time data; the normalization processing and the generation of the offline data index are respectively implemented by a data processing component and an offline index construction component, and the data processing component and the offline index construction component are independent of each other, so as to allocate resources by respectively configuring the concurrency degrees of the data processing component and the offline index construction component.

4. The method according to claim 1, characterized in that The generation of the real-time data index is implemented by an online service node; After generating the real-time data index corresponding to the real-time data, the method further includes: Storing the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index into the online service node, so that the online service node uses the stored data to execute corresponding services; the online function node is applied to a data retrieval service, and multiple online service nodes cooperate to execute the data retrieval service, and each online service node supports sharding and backing up the stored data in other online service nodes.

5. A data retrieval method, characterized in that, Including: Obtaining a retrieval statement corresponding to a retrieval task; Retrieving target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index; the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index are determined according to any one of the methods described in claims 1 to 4.

6. The method according to claim 5, wherein Retrieving target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data in the temporary database that does not have a real-time data index, including: Dividing the offline data index, the real-time data index, and the real-time data that does not have a real-time data index into multiple data blocks; splitting and vector-converting the retrieval statement to obtain a corresponding retrieval vector set; Determining the number of grouped vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task, and then grouping the retrieval vector set according to the grouped vector data to obtain retrieval vector groups; Allocate a corresponding retrieval thread for each of the data blocks, retrieve each retrieval vector group in a multi-threaded parallel manner, place the retrieval results of each thread in the corresponding result storage space, and based on the result storage space, merge the retrieval results of each thread, so as to obtain the target data similar to the retrieval statement.

7. The method according to claim 6, wherein Determine the number of grouped vectors of the retrieval vector set according to the number of the data blocks and the storage capacity of the cache space corresponding to the retrieval task, including: Calculate the memory of the retrieval vectors according to the dimensions of the retrieval vectors in the retrieval vector set. Calculate the memory of the retrieval results corresponding to the retrieval vectors according to the preset number of similar retrievals and the number of the data blocks. Calculate the sum of the memory of the retrieval vectors and the memory of the retrieval results corresponding to the retrieval vectors to obtain the memory requirement of the retrieval vectors; calculate the number of grouped vectors of the retrieval vector set according to the memory requirement and the storage capacity of the cache space corresponding to the retrieval task.

8. A data processing device, characterized in that, Including: An offline index generation module, configured to collect offline data within a preset time, store the offline data in a persistent database, and generate an offline data index corresponding to the offline data based on the persistent database. A real-time index generation module, configured to obtain real-time data after the preset time in real time and store the real-time data in a temporary database. When the data volume of the real-time data reaches a threshold, generate a real-time data index corresponding to the real-time data based on the temporary database.

9. A data retrieval device, characterized in that, Including: A statement acquisition module, configured to acquire a retrieval statement corresponding to a retrieval task. A data retrieval module, configured to retrieve target data similar to the retrieval statement from the offline data index, the real-time data index, and the real-time data without a real-time data index in the temporary database; the offline data index, the real-time data index, and the real-time data without a real-time data index in the temporary database are determined according to any one of the methods described in claims 1 to 4.

10. A mobile electronic device terminal, characterized in that, Including: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-7.

11. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1-7.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1-7.

Citation Information

Cited By

  • Video retrieval method and device based on three-dimensional fragmentation and electronic equipment

    CN121524393A

  • Three-dimensional slice-based video retrieval method, device and electronic equipment

    CN121524393B