Data query method, database system, readable medium, and electronic device

By selecting the query data set according to the migration status of the target data set in the database system, data query and migration are executed in parallel, the problem of slow query speed in the database system during data type migration is solved, and the user experience is improved.

CN114691720BActive Publication Date: 2025-06-27SHANGHAI XUYU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210292093.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-23
Publication Date
2025-06-27
Estimated Expiration
2042-03-23

AI Technical Summary

Technical Problem

During the data type migration process, the database system needs to wait for the migration to complete before querying data, resulting in slow query speed and reducing the user's query experience.

Method used

According to the migration status of the target data set, select query data from the pre-migrated data set or the migrated data set, so that data query and data migration are executed in parallel.

Benefits of technology

It improves the speed of querying data from the database, improves the user's experience of querying data, and avoids the problem of incomplete query results during data migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691720B_ABST
    Figure CN114691720B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data processing, and discloses a data query method, a database system, a readable medium, and an electronic device. The method includes: determining that at least one first target data set in the target data set corresponding to the received first data query request is in the process of migrating from a first data type to a second data type; and determining, according to the migration status of the first target data set, to query and obtain a sub-query result associated with the first data query request in the first target data set of the first data type, or to query and obtain a sub-query result associated with the first data query request in the first target data set of the second data type. In this way, without interrupting the data migration operation, the sub-query result associated with the first data query request can be queried in parallel from the first target data set of the first data type or the first target data set of the first data type, improving the query speed and enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a data query method, a database system, a readable medium, and an electronic device. Background Art

[0002] With the development of artificial intelligence technology, deep learning models are increasingly widely used in daily life. As a result, there are more and more database systems for storing data related to deep learning models, such as vector database systems for storing vectors.

[0003] Database systems usually support multiple data types, such as streaming data, batch data, etc. Moreover, in order to improve the performance of the database system, such as data update speed, data query speed, database throughput, etc., one data type is usually migrated to another data type, such as migrating streaming data to batch data. When the target data set corresponding to the data query request received by the database system is in the process of being migrated from one data type to another data type, the database system needs to wait for the migration of the data set to complete and then query the query result of the data query request from the target data set of the migrated data type, resulting in slow query speed and reducing the query speed of the database system. Summary of the Invention

[0004] In view of this, embodiments of this application provide a data query method, a database system, a readable medium, and an electronic device. This method can query data from the data set before migration or the data set after migration according to the migration progress of the target data set corresponding to the data query request, enabling parallel execution of data query and data migration, which is beneficial to improving the speed of querying data from the database and enhancing the user experience of querying data.

[0005] In a first aspect, embodiments of this application provide a data query method applied to an electronic device. The method includes: determining that at least one first target data set in the target data set corresponding to the received first data query request is in the process of being migrated from a first data type to a second data type; and determining to query and obtain a sub-query result associated with the first data query request in the first target data set of the first data type, or query and obtain a sub-query result associated with the first data query request in the first target data set of the second data type according to the migration status of the first target data set.

[0006] That is to say, in the embodiments of the present application, the electronic device can select to query and obtain a sub-query result associated with the first data query request from the first target data set of the first data type or the target data set of the second data type according to the migration status of the first target data set (for example, the progress of the migration of the first target data set from the first data type to the second data type), rather than waiting for the migration of the first target data set to be completed and then querying and obtaining the sub-query result associated with the first data query request from the target data set of the second data type. This enables parallel execution of data migration and data query, improving the speed of obtaining the query result of the first data query request and enhancing the user experience of querying data.

[0007] In a possible implementation of the above first aspect, the migration status of the first target data set includes a first migration status and a second migration status: The first migration status indicates that some data in the first target data set of the first data type has been increased to the second data type, and the first target data set of the first data type has not been deleted yet; The second migration status indicates that all data in the first target data set has been increased to the second data type, and at least some data in the target data set of the first data type has not been deleted.

[0008] In a possible implementation of the above first aspect, determining to query and obtain a sub-query result associated with the first data query request in the first target data set of the first data type or in the first target data set of the second data type according to the migration status of the first target data set includes: When the migration status of the first target data set is the first migration status, determining to query and obtain the sub-query result associated with the first data query request from the first target data set of the first data type, and not interrupting the operation of adding the first target data set of the first data type to the first target data set of the second data type; When the migration status of the first target data set is the second migration status, determining to query and obtain the sub-query result associated with the first data query request from the first target data set of the second data type, and not interrupting the deletion operation of the first target data set of the first data type.

[0009] In an embodiment of the present application, the electronic device may perform a query operation for obtaining a query result of a first query request and a migration operation on a first target data set in parallel according to the migration state of the first target data set. For example, when the migration state of the first target data set is a first migration state, an addition operation in the query operation and the migration operation may be performed in parallel; when the migration state of the first target data set is a second migration state, a deletion operation in the query operation and the migration operation may be performed in parallel, without waiting for the migration operation to complete before performing the query operation, thereby improving the speed of obtaining the query result corresponding to the first query request and enhancing the user experience of querying data.

[0010] In a possible implementation of the first aspect above, during the process of querying and obtaining a sub-query result associated with a first data query request from a first target data set of a first data type, it further includes: detecting that all the data in the first target data set of the first data type has been increased to a second data type, and keeping the data in the first target data set of the first data type from being deleted.

[0011] In an embodiment of the present application, if it is detected that all the data in the first target data set of the first data type has been increased to a second data type and the sub-query result associated with the first data query request has not been obtained from the first target data set of the first data type (i.e., during the process of querying and obtaining the sub-query result associated with the first data query request from the first target data set of the first data type), then the data in the first target data set of the first data type is not operated on and deleted, so that it is possible to avoid the first target data set being deleted during the process of querying the query result of the first query request, ensure the integrity of the target data set, and improve the accuracy of the query result.

[0012] In a possible implementation of the first aspect above, the method further includes: when it is detected that there is a second data query request accessing the first target data set of the first data type, keeping the data in the first target data set of the first data type from being deleted before obtaining the query result of the second data query request.

[0013] In an embodiment of the present application, before deleting the first target data set of the first data type, it may be detected whether there are other data query requests (such as a second data query request) accessing the first target data set of the first data type, and when there are other data query requests accessing the first target data set of the first data type, the data in the first target data set of the first data type is kept from being deleted before obtaining the query result of the other data query request. In this way, the integrity of the target data set of the second data query request can be ensured, and the accuracy of the query result can be enhanced.

[0014] In a possible implementation of the first aspect above, during the process of querying and obtaining a sub-query result associated with the first data query request from the first target data set of the first data type, the following is further included: detecting that all the data in the target data set of the first data type has been increased to the second data type, and the operating state of the electronic device where the first target data set of the first data type is located satisfies a preset condition, pausing querying and obtaining the sub-query result associated with the first data query request from the first target data set of the first data type; deleting the target data set of the first data type, and continuing to query and obtain the sub-query result associated with the first data query request in the first target data set of the second data type.

[0015] In the embodiments of the present application, if it is determined that all the data in the target data set of the first data type has been increased to the second data type, and the operating state of the electronic device where the first target data set of the first data type is located satisfies a preset condition, it is possible to switch to querying and obtaining the sub-query result associated with the first data query request from the first target data set of the second data type. This can prevent the electronic device where the first target data set is located from crashing due to insufficient hardware resources, excessive access volume, etc., and improve the stability of the database system.

[0016] In a possible implementation of the first aspect above, the operating state of the electronic device where the first target data set of the first data type is located satisfying a preset condition includes at least one of the following conditions: the free memory of the electronic device where the first target data set is located is less than a first threshold or the memory usage is greater than a second threshold; the free amount of the processor of the electronic device where the first target data set is located is less than a third threshold or the processor usage is greater than a fourth threshold; the number of first data query requests processed by the electronic device where the first target data set is located exceeds a preset number.

[0017] In a possible implementation of the first aspect above, the first data type includes any one of the following data types: streaming data, batch data, bulk import data, and the second data type includes any one of the following data types: streaming data, batch data, bulk import data.

[0018] In a possible implementation of the first aspect above, the target data set corresponding to the first data query request is stored in the database system, the first target data set of the first data type is stored in the first storage area of the database system, and the first target data set of the second data type is stored in the second storage area of the database system; and during the process of the first target data set migrating from the first data type to the second data type, it includes: at least part of the data in the first target data of the first data type has been increased to the data in the first target data set of the second data type, and at least part of the data in the first target data of the first data type has not been deleted yet.

[0019] In a possible implementation of the above first aspect, the database system is a vector database system.

[0020] In a second aspect, an embodiment of the present application provides a database system, which includes: a coordination unit, configured to determine, when it is determined that at least one first target data set in the target data set corresponding to the received first data query request is in the process of migrating from a first data type to a second data type, according to the migration state of the first target data set, to query and obtain a sub-query result associated with the first data query request in the first target data set of the first data type, and send an instruction to query and obtain the sub-query result associated with the first data query request from the first target data set of the first data type to the query unit where the first target data set of the first data type is located, or to query and obtain the sub-query result associated with the first data query request in the first target data set of the second data type, and send an instruction to query and obtain the sub-query result associated with the first data query request from the first target data set of the second data type to the query unit where the first target data set of the second data type is located; at least one query unit, configured to query and obtain the sub-query result associated with the first data query request from the target data set of the first data type or the target data set of the second data type according to the instruction sent by the coordination unit.

[0021] In the embodiment of the present application, the database system can obtain the sub-query result of the first query request from the first target data set of the first data type or the first target data set of the second data type according to the migration state of the first target data set, without completely interrupting the process of migrating the first target data set from the first data type to the second data type, improving the speed of querying data from the database system and enhancing the user experience of querying data through the database system.

[0022] In a third aspect, an embodiment of the present application provides a readable medium, on which instructions are stored, and when the instructions are executed on an electronic device, the electronic device is enabled to implement any one of the data query methods provided by the above first aspect and various possible implementations of the above first aspect.

[0023] In a fourth aspect, an embodiment of the present application provides an electronic device, which includes: a memory, configured to store instructions executed by one or more processors of the electronic device; and a processor, which is one of the processors of the electronic device, configured to execute the instructions stored in the memory to implement any one of the data query methods provided by the above first aspect and various possible implementations of the above first aspect. Description of the Drawings

[0024] Figure 1According to some embodiments of the present application, a schematic structural diagram of a database system 10 is shown;

[0025] Figure 2A According to some embodiments of the present application, a schematic diagram of a migration operation;

[0026] Figure 2B According to some embodiments of the present application, a schematic diagram of an addition operation in a migration operation is shown;

[0027] Figure 2C According to some embodiments of the present application, a schematic diagram of a deletion operation in a migration operation is shown;

[0028] Figure 3A According to some embodiments of the present application, a schematic diagram of a data set S5 in an incomplete addition state is shown;

[0029] Figure 3B According to some embodiments of the present application, a schematic diagram of a data set S5 in an incomplete deletion state is shown;

[0030] Figure 3C According to some embodiments of the present application, another schematic diagram of a data set S5 in an incomplete deletion state is shown;

[0031] Figure 4 According to some embodiments of the present application, a schematic structural diagram of a data query method is shown;

[0032] Figure 5A According to some embodiments of the present application, a schematic diagram of setting a node replica for a query node is shown;

[0033] Figure 5B According to some embodiments of the present application, a schematic diagram of setting a data set replica for a data set is shown;

[0034] Figure 6 According to some embodiments of the present application, a schematic diagram of data in a data set S7” is shown;

[0035] Figure 7 According to some embodiments of the present application, a schematic diagram of the structure of a database system 200 is shown;

[0036] Figure 8 According to some embodiments of the present application, a schematic diagram of the structure of an electronic device 100 is shown. Detailed implementation manners

[0037] Illustrative embodiments of the present application include but are not limited to data query methods, database systems, readable media, and electronic devices.

[0038] The technical solutions of the embodiments of the present application will be introduced below with reference to the accompanying drawings.

[0039] Figure 1 According to some embodiments of the present application, a schematic structural diagram of a database system 10 is shown.

[0040] As Figure 1 shown, the database system 10 includes an access layer 11, a coordination service 12, execution nodes 13, and a storage service 14.

[0041] The access layer 11 includes multiple proxies, such as proxy 110, proxy 111, etc., and provides an interface for external user connections, such as the Application Program Interface (API) of the database system 10, for receiving data query requests from users, returning query results to users, etc.

[0042] The coordination service 12 is used to allocate tasks to the execution nodes 13, including but not limited to cluster topology node management, load balancing, timestamp generation, data declaration, and data management, etc.

[0043] In some embodiments, the coordination service 12 may include:

[0044] A root coordinator 120, which is used to process Data Definition Language (DDL) and Data Control Language (DCL) requests, such as creating, migrating, and deleting data sets, etc., and is also responsible for maintaining the central timing service and the advancement of time windows;

[0045] A query coordinator 121, which is used to manage the topology structure and load balancing of the query nodes 131 and migrate the data set from the source data type to the target data type;

[0046] A data coordinator 122, which is used to manage the topology structure of the data nodes 131, maintain the meta-information of the data, and trigger background data operations such as data flushing (transferring cached data to disk or storage system when the cached data reaches a preset size), data compaction, etc.;

[0047] An index coordinator 123 is used to manage the topology structure of the index nodes 133, build indexes, and maintain index meta-information.

[0048] The execution node 13 is used to execute the instructions issued by the coordination service 12 and the Data Manipulation Language (DML) instructions initiated by the agents 110 / 111.

[0049] In some embodiments, the execution node 13 may include:

[0050] A data node 130, which can be used to store the dataset of the target data type migrated from the dataset of the source data type in the object storage 141. It can be understood that in some embodiments, there may be multiple data nodes, such as data node 130A, data node 130B, etc.

[0051] A query node 131, which can be used to perform data operations according to the instructions of the coordination service 12 or the agents in the access layer 10. For example, according to the data query instructions of the agents 110 / 111, query the query results corresponding to the data query requests and send the query results to the agents 110 / 111. Another example is that according to the data migration instructions of the coordination service 12, convert the dataset in the query node from the source data type to the target data type.

[0052] It can be understood that in some embodiments, the query node is stateless, that is, each query node can only respond to the instructions sent by the query coordinator 121 without paying attention to the running status of other query nodes. In this way, users can set different numbers of query nodes in the database system 10 according to needs, such as the amount of data, the concurrent access quantity of data query requests, etc., improving the scalability of the database system 10.

[0053] An index node 132, which can be used to build indexes for the data / datasets in the database system to facilitate the storage and query of the data / datasets. It can be understood that in some embodiments, there may be multiple index nodes, such as index node 132A, index node 132B, etc.

[0054] The storage service 14 is used to store the data of the database and realize the persistence of the data in the database.

[0055] In some embodiments, the storage service 14 may include:

[0056] A meta store 140, which can be used to store the metadata in the database system, such as the status information of each node, datasets, etc. In some embodiments, the meta store 140 can be implemented through ETCD TM(A distributed key-value storage system can be used to store critical core data in a distributed system) to implement. In some other embodiments, the metadata storage 140 can also be implemented in other ways, which are not limited herein.

[0057] The object storage 141 can be used to store files of the database system's logs, scalar / vector index files, and intermediate processing results of queries. In some embodiments, the object storage 141 can be implemented using the Simple Storage Service (AWS S3), or it can also be implemented through Azure Blob (Microsoft TM object storage solution for the cloud). In some other embodiments, the object storage 141 can also be implemented based on a cache pool of memory or solid-state disks (SSD) in a form of separating hot and cold data. The implementation manner of the object storage in the embodiments of the present application is not limited.

[0058] The message storage (log broker) 142 can be a publish-subscribe system that supports playback. It can be used to ensure the integrity of incremental data by playing back the message storage during the recovery of an execution node going down, persist the data written in a streaming manner, asynchronously execute queries, event notifications, and result returns. In some embodiments, the message storage 142 can be implemented through apache pulsar TM , RocksDB TM , apache kafka TM , Pravega TM and other streaming storage systems, or it can also be implemented through other streaming storage systems. The implementation manner of the message storage 142 in the embodiments of the present application is not limited.

[0059] In the database system 10, after the proxy 110 / 111 receives a data query request sent by the client 00, it forwards the data query request to the query coordinator 121. The query coordinator 121 determines at least one query node (such as query nodes 131A and 132B) where the target data set corresponding to the data query request is located according to the query features (such as feature vectors, result quantities, feature types (i.e., the types of the aforementioned unstructured data), etc.) in the data query request, and forwards the data query request to each query node. Each query node queries the sub-query result of the query request in the target data set of each query node and returns the sub-query result to the proxy 110 / 111. Then, after the proxy 110 / 111 summarizes the sub-query results of each query node, it obtains the query result of the data query request and returns the query result to the user, such as sending it to the client 00.

[0060] It can be understood that Figure 1 The structure of the database system 10 shown is only an example. In some other embodiments, the database system 10 may include more or fewer modules, or some modules may be combined or split, which is not limited herein.

[0061] It can be understood that Figure 1 Each module in the database system 10 shown may be deployed on the same electronic device or on different electronic devices, which is not limited herein. Moreover, the database system 10 may be implemented as a single electronic device or in the form of a server cluster.

[0062] It can be understood that in some embodiments, the database system 10 may be a vector database. The feature vectors of various unstructured data (such as text, images, audio, video, DNA sequences, product information, material structures, etc.) stored in this database can obtain results related / similar to the search data or the feature vector of the search data according to the search data provided by the user or the feature vector of the search data, and can be used in scenarios such as instantaneously returning the most similar picture to the uploaded picture from a vast database, retrieving similar videos according to video key frames or performing real-time video recommendations, quickly retrieving vast amounts of audio data such as speeches / music / sound effects and returning similar audio, retrieving similar chemical molecular structures / superstructures / substructures, recommending relevant information or products according to user behavior and needs, an interactive intelligent Q&A robot can automatically answer questions for users, classifying genes by comparing similar DNA sequences, and helping users search for the required information through keywords in a text database.

[0063] It can be understood that the data query method provided by the embodiments of the present application is applicable to any database system with data migration operations. For the convenience of description, the following will introduce the technical solutions of the embodiments of the present application in combination with Figure 1 the structure of the database system 10 shown.

[0064] As described above, when the target data set corresponding to the data query request received by the database system is in the process of migrating from one data type to another data type, the database system needs to wait for the migration of this data set to be completed, and then query the query result of this data query request from the target data set of the migrated data type, resulting in slow query speed and reducing the query speed of the database system.

[0065] For example, after the proxy 110 receives a data query request, it forwards the data query request to the query coordinator 121. The query coordinator 121 determines that the target data set corresponding to the data query request is in the data migration process of migrating from the source data type to the target data type (for example, refer to Figure 2A, assuming that dataset S5 in the target dataset is in the process of migrating from the streaming data in query node 131A to the batch data in query node 131B), it is necessary to wait until the source data type is migrated to the target data type (for example, after the dataset S5 of streaming data is migrated to the dataset S5 of batch data), so that the query node where the target dataset of the target data type is located (for example, query node 131B where the dataset S5 of batch data is located) can query the query result of the data query request from the target dataset of the target data type (for example, the dataset S5 of batch data), and then return the query result to the proxy 110, and the proxy 110 provides the query result to the sender of the data query request (for example, client 00). Since the data migration operation process includes operations such as data writing and data deletion, which takes a long time, the database system 10 needs to wait for the target dataset migration to complete before returning the query result after receiving the data query request, which reduces the query speed of the database system 10 and affects the user experience of querying data from the database system.

[0066] To this end, the embodiments of the present application provide a data query method, which is applied to a database system. When the database system determines that the target dataset corresponding to the received data query request is in the process of migrating from the source data type to the target data type, according to the migration status of the target dataset, it is determined to query the query result corresponding to the data query request from the target dataset of the source data type or the target dataset of the target data type. In this way, the data query operation can be executed in parallel with the data migration operation, improving the query speed of the database system and enhancing the user experience of querying data from the database system.

[0067] It can be understood that the data migration operation (handoff) refers to converting a dataset of one data type (source data type) into a dataset of another data type (target data type). Usually, it includes two steps: converting and storing the dataset of the source data type as the dataset of the target data type (add operation), and deleting the dataset of the source data type (delete operation). Generally, to avoid data loss caused by data migration failure, usually after all the add operations of the dataset of the source data type are completed and stored as the dataset of the target data type, the delete operation is then performed on the source data type. For example, referring to Figure 2B and Figure 2C , for the aforementioned migration process of migrating the dataset S5 of streaming data in query node 131A to the dataset S5' of batch data in query node 131B, the add operation is to convert the dataset S5 of streaming data into the dataset S5' of batch data and store it in query node 131B (as shown in Figure 2B ), and the delete operation is to delete the dataset S5 of streaming data from query node 131A (as shown in Figure 2Cas shown).

[0068] It can be understood that in some embodiments, the data set of the source data type and the data set of the target type are stored in different storage areas in the database system 1. For example, the data set of the source data type and the data set of the target type can be stored in different storage areas of the same query node, or can be stored in different query nodes in the database system 1.

[0069] It can be understood that the number of data in the data set of the source data type and the data set of the target data type and the specific content of each data are the same.

[0070] It can be understood that the data migration operation can migrate the data set of the source data type to the query node where the data set of the source data type is located, or can also be migrated to other query nodes, which is not limited herein.

[0071] It can be understood that in some embodiments, when the size of the cached streaming data reaches the preset size, the message storage 142 will send a data migration request to the nodes in the query node 131. In response to this data migration request, the nodes in the query node 131 will migrate the cached streaming data into batch data.

[0072] It can be understood that the target data type and the source data type can be any data types supported by the database system, including but not limited to streaming data (also known as real-time data, incremental data, etc.), batch data (also known as historical data, etc.), batch import data, etc., which is not limited herein. For the convenience of description, in the following embodiments, the source data type is streaming data and the target data type is batch data for introduction.

[0073] It can be understood that the migration state of the data set is used to indicate the progress of the data migration of the data set, including the add incomplete state (the first migration state) and the delete incomplete state (the second migration state). Among them:

[0074] The add incomplete state represents the following state: some data sets in the target data set of the source data type have been added as data of the target data type. For example, referring to Figure 3A , the data set S5 of the streaming data in the query node 131A includes 4 pieces of streaming data: S5-1, S5-2, S5-3, and S5-4. Among them, the streaming data S5-1 and the streaming data S5-2 have been migrated to the batch data S5-1' and the batch data S5-2' in the data set S5' of the batch data in the query node 131B, but the streaming data S5-3 and the streaming data S5-4 have not been migrated yet. At this time, the migration state of the data set S5 is the add incomplete state.

[0075] As described above, the target data set corresponding to the data query request includes data sets S1, S2, S3, S4, S5, S6, S7, and S8. When the query coordinator 121 determines that the migration status of data set S5 in the target data set is the increasing incomplete status, it determines to query the sub-query result of the data query request in data set S5 from data set S5 of the source data type. That is to say, referring to Figure 3A , the target data set 30 corresponding to the data query request includes: data set S1 of batch data, data set S2 of batch data, data set S3 of batch data, data set S4 of batch data, data set S5 of streaming data, data set S6 of streaming data, data set S7 of streaming data, and data set S8 of streaming data. The query coordinator 121 forwards the data query request to the query node 131A and the data query node 131B, and then the query node 131A and the data query node 131B obtain the sub-query results corresponding to the data query request from the corresponding data sets, and then send each sub-query result to the proxy 110 / 111, and then the proxy 110 / 111 forwards it to the user.

[0076] That is to say, the migration operation of data set S5 does not interrupt the acquisition of the sub-query result of the data query request from data set S5. The migration operation of migrating data S5-3 and data S5-4 in data set S5 of streaming data to data set S5' of batch data in the query node 131B and the data query operation in data set S5 of streaming data can be executed in parallel, which improves the query speed of the database system and enhances the user experience of querying data from the database system.

[0077] It can be understood that in some embodiments, if the query operation in data set S5 of streaming data has not been completed, and all the data in data set S5 of streaming data has been added to data set S5' of batch data, the query coordinator 121 can block the deletion operation of data set S5 of streaming data until all the query operations in data set S5 of streaming data are completed to ensure the integrity of the target data set.

[0078] It can be understood that in some embodiments, if the query operation in data set S5 of streaming data has not been completed, and all the data in data set S5 of streaming data has been added to data set S5' of batch data, if the running status of the query node (such as query node 131A) where data set S5 of streaming data is located meets the preset conditions, the query coordinator 121 can interrupt the query operation in data set S5 of streaming data, perform a deletion operation on data set S5 of streaming data, and resume the interrupted query operation in data set S5' of batch data.

[0079] Among them, the preset conditions may include at least one of the following conditions: a condition for indicating insufficient hardware resources of the electronic device where the query node is located, such as the free memory amount of the electronic device where the query node is located being less than the memory free amount threshold / the memory usage amount being greater than the memory usage amount threshold, the free processor amount of the electronic device where the query node is located being less than the processor free amount threshold / the processor usage amount being greater than the processor usage amount threshold, etc. It can be understood that in some other embodiments, the preset conditions may further include other conditions, such as the number of data query requests processed by the electronic device where the query node is located being greater than a preset number, etc., which will not be limited here.

[0080] In this way, it is possible to avoid insufficient hardware resources of the query node, or the query node crashing due to an excessive amount of query request data responded by the query node, which affects the normal operation of the database system. In some embodiments, the user can set the above preset conditions according to the hardware resource configuration of each query node, the concurrent access requirements of the database system, etc., such as setting the memory free amount threshold / memory usage amount threshold, the processor free amount threshold / processor usage amount threshold, the preset number of data query requests, etc., so as to improve the speed of querying data in the database system while ensuring the normal operation of the query node.

[0081] The deletion incomplete state represents the following state: all the data in the target dataset of the source data type has been added as data of the target data type, and some or none of the data in the target dataset of the source data type has been deleted. For example, referring to Figure 3B , the dataset S5 of the streaming data in the query node 131A includes 4 pieces of streaming data: S5-1, S5-2, S5-3, and S5-4, and the streaming data S5-1, streaming data S5-2, streaming data S5-3, and streaming data S5-4 have been migrated to the batch data S5-1', batch data S5-2', batch data S5-3', and batch data S5-4' in the dataset S5' of the batch data in the query node 131B, but the streaming data S5-1, streaming data S5-2, streaming data S5-3, and streaming data S5-4 in the dataset S5 of the streaming data in the query node 131A have not been deleted. At this time, the migration state of the dataset S5 is the deletion incomplete state. Another example, referring to Figure 3C , the streaming data S5-1, streaming data S5-2, streaming data S5-3, and streaming data S5-4 in the dataset S5 of the streaming data in the query node 131A have been migrated to the batch data S5-1', batch data S5-2', batch data S5-3', and batch data S5-4' in the dataset S5' of the batch data in the query node 131B, and the streaming data S5-1 and streaming data S5-2 in the dataset S5 of the streaming data in the query node 131A have been deleted. At this time, the migration state of the dataset S5 is also the deletion incomplete state.

[0082] As described above, when the target data set corresponding to the data query request includes data sets S1, S2, S3, S4, S5, S6, S7, and S8, and the query coordinator 121 determines that the data set S5 in the target data set is in the state of incomplete deletion, it determines to query the sub-query result of the data query request in the data set S5 from the data set S5 of the target data type. That is to say, referring to Figure 3B and Figure 3C , the target data set 31 or the target data 32 corresponding to the data query request includes: the data set S1 of batch data, the data set S2 of batch data, the data set S3 of batch data, the data set S4 of batch data, the data set S5' of batch data, the data set S6 of streaming data, the data set S7 of streaming data, and the data set S8 of streaming data. The query coordinator 121 forwards the data query request to the query node 131A and the data query node 131B. Then, the query node 131A and the data query node 131B obtain the sub-query results corresponding to the data query request from the corresponding data sets, and then send each sub-query result to the proxy 110 / 111. Finally, the proxy 110 / 111 aggregates each sub-query result and forwards it to the user, for example, sending it to the client 00.

[0083] That is to say, the migration operation of the data set S5 does not interrupt the query operation in the data set S5' of batch data. The deletion operation of the data in the data set S5 of streaming data can be executed in parallel with the data query operation of obtaining the sub-query result of the data query request from the data set S5' of batch data. In this way, the database system 10 does not need to wait until the deletion operation of the data in the data set S5 of streaming data is completed, and then execute the query operation in the data set S5' of batch data, which improves the speed of querying data by the database system and enhances the user experience of using the database system.

[0084] It can be understood that the foregoing classification of the migration state of the data set into the state of incomplete addition and the state of incomplete deletion is only an example. In other embodiments, other classifications may also be used to indicate the progress of data migration of the data set, and some migration states may also be split or combined, which is not limited herein.

[0085] Next, in combination with Figure 1 the database system 10 shown in Figure 2A the target data set shown in Figures 3A to 3B and the migration state of the data set shown in

[0086] Specifically, Figure 4 According to some embodiments of the present application, an interaction flow diagram of a data query method is shown. As Figure 4As shown, the interaction process includes the following steps.

[0087] S401: The agent 110 receives a data query request and forwards the data query request to the query coordinator 121.

[0088] After the agent 110 receives the data query request, it immediately forwards the data query request to the query coordinator.

[0089] In some embodiments, the agent 110 receiving the data query request may be that the user can send a data query request to the agent 110 through the client 00 of the database system 10. In other embodiments, the agent 110 receiving the data query request may also be that the user sends a data query request to the agent 110 through the access interface of the database system 10. The embodiments of the present application do not limit the source and form of the data query request received by the agent 110.

[0090] In some embodiments, the data query request may include a feature type (such as an image, sound, video, text, DNA sequence, etc.), the features of the target data (such as the feature vector N, feature tensor, etc.), the number K of the data in the query result, etc.

[0091] S402: The query coordinator 121 determines the target data set corresponding to the data query request according to the data query request.

[0092] The query coordinator 121 determines the target data set corresponding to the data query request according to the data query request.

[0093] For example, in some embodiments, each data set in the database system has a central vector, and this central vector represents the average features of the data in this data set (for example, in some embodiments, the central vector may be the clustering center of the data set, the average value of the feature vectors in this data set, etc.). The query coordinator 121 can compare the feature vector N in the data query request with the central vectors of each data set whose feature type is the same as the feature type in the data query request, and use the data sets corresponding to the preset number of central vectors whose similarity with the feature vector in the data query request is greater than the preset value and / or whose similarity with the feature vector in the data query request is the largest as the target data set. For example, assuming that the feature type in the data query request is an image and the feature vector is N, the query coordinator 121 uses the data sets in the database system 10 whose feature type is an image and the similarity between the central vector and the feature vector N is greater than the preset value (such as the similarity is greater than 0.9, the distance is less than 0.25, etc.) as the target data set.

[0094] It can be understood that the similarity between the feature vector and the center vector in the data query request can be represented by the distance between the vectors, including but not limited to Euclidean distance, inner product, Jaccard distance, Tanimoto distance, Hamming distance, etc., which will not be limited here. Among them, for Euclidean distance, Jaccard distance, and Hamming distance, the smaller the distance between the feature vector and the center vector in the data query request, the higher the similarity; for inner product and Tanimoto distance, the larger the distance between the feature vector and the center vector in the data query request, the higher the similarity.

[0095] For example, assume that the feature vector N in the data query request is [1, 3, 5, 7, 9]. The center vectors of the aforementioned datasets S1, S2, S3, S4, S5, S6, S7, and S8 are S8-N = [1, 3, 5, 7.1, 9.2], S2-N = [1, 3.1, 5, 7, 9], S3-N = [1, 3, 5.2, 7, 9], S4-N = [1.2, 3, 5, 7, 9], S5-N = [1, 3, 5, 7.1, 9], S6-N = [1, 3.1, 5.2, 7, 9], S7-N = [1, 3, 5.2, 7, 9.1], S8-N = [1.2, 3, 5.1, 7, 9] respectively. According to the formula for the Euclidean distance between vectors shown in formula (1) below.

[0096]

[0097] In formula (1), D(a, b) is the Euclidean distance between vector a and vector b, n is the dimension of vector a and vector b, a i is the value corresponding to the i-th dimension of vector a, and b i is the value corresponding to the i-th dimension of vector b.

[0098] The distances between the center vectors of the aforementioned datasets S1, S2, S3, S4, S5, S6, S7, and S8 and the feature vector N in the data query request can be obtained as D(N, S1-N) = 0.224, D(N, S2-N) = 0.1, D(N, S3-N) = 0.2, D(N, S4-N) = 0.2, D(N, S5-N)

[0099] = 0.1, D(N, S6-N) = 0.224, D(N, S7-N) = 0.224, D(N, S8-N) = 0.224, all of which are less than the Euclidean distance threshold of 0.25. Therefore, the datasets corresponding to the data query request include datasets S1, S2, S3, S4, S5, S6, S7, and S8.

[0100] In some other embodiments, the query coordinator 121 can also determine the target data set corresponding to the data query request in other ways, which is not limited herein.

[0101] S403: The query coordinator 121 determines whether each target data set is in the process of migration.

[0102] In some embodiments, after the query coordinator 121 determines the target data set corresponding to the data query request, the query coordinator 121 obtains the status identifier of each target data set from the query node 131 or the message storage 142. The status identifier is used to indicate whether the target data set is in the process of migration. If it is determined that the target data set is in the process of migration, it is necessary to determine from which type of target data set to query the sub-query result of the data query request according to the migration status of the target data set, and proceed to step S405; otherwise, it means that the sub-query result of the data query request can be directly queried through the target data set, and proceed to step S404A.

[0103] For example, in the Figures 3A to 3C scenario shown, the data set S5 is in the process of migrating from the data set S5 of streaming data to the data set S5' of batch data, then proceed to step S405 to perform a query operation on the data set S5, while the data sets S1, S2, S3, S4, S6, S7, and S8 are not in the process of migration, and proceed to step S404A to perform a query operation on the data sets S1, S2, S3, S4, S6, S7, and S8.

[0104] In some embodiments, the query coordinator 121 can also determine whether each target data set is in the process of migration according to whether the database system is performing the foregoing add operation or delete operation on the data in the target data set. It can be understood that in some other embodiments, the query coordinator 121 can also determine whether the target data set is in the process of migration in other ways, which is not limited herein.

[0105] It can be understood that in some embodiments, when there are multiple target data sets, the query coordinator 121 can determine whether each target data set is in the process of migration one by one, or can determine whether each target data set is in the process of migration in batches, which is not limited herein.

[0106] S404A: The query coordinator 121 sends a data query sub-request to the query node corresponding to the target data set.

[0107] For a target data set that is not in the process of migration, the query coordinator 121 directly sends a data query sub-request to the query node corresponding to the target data set. For example, when the query coordinator 121 determines that the aforementioned data sets S1, S2, S3, S4, S6, S7, and S8 are not in the process of migration, it directly sends a sub-query request to query node 131A to query the query results corresponding to the aforementioned feature vector N from data sets S1, S2, S3, and S6, and sends a sub-query request to query node 131B to query the query results corresponding to the aforementioned feature vector N from data sets S4, S7, and S8.

[0108] It can be understood that in some embodiments, multiple copies can be set for query nodes and / or data sets in the database system 10, so that multiple query requests for the same data set can be executed in parallel to improve the query speed. For example, in some embodiments, the database system 10 can set copies for the entire query node. Refer to Figure 5A , a copy query node 131A' can be set for query node 131A. The data sets in query node 131A and the corresponding copy query node 131A' are exactly the same. Again, for example, in some other embodiments, at least one copy can also be set for a data set in the database system 10 according to the number of query requests for the data set, etc., to avoid setting too many copies for a data set with fewer query requests when setting copies for query nodes, thus wasting the resources of the database system 10. Refer to Figure 5B , corresponding copies, namely copy data set S1', copy data set S4', and copy data set S7', are set for data sets S1, S4, and S7 in query node 131C, and corresponding copies, namely copy data set S1” and copy data set S7”, are set for data sets S1 and S7 in query node 131D.

[0109] For a target data set with multiple copies, the query coordinator 121 can obtain the hardware resource usage, response speed, etc. of the query nodes where each copy of the target data is located, and send a data query sub-request to the query node that can respond to the data query sub-request faster, has more idle hardware resources, and has fewer data query nodes that respond, so as to further improve the data query speed. For example, for data set S7, the response time of query node 131B is 20 milliseconds, the response time of query node 131C is 15 milliseconds, and the response time of query node 131D is 10 milliseconds. Then the query coordinator 121 can send a data query sub-request to query node 131D to query the sub-query results corresponding to the above-mentioned feature vector N from copy data set S7”.

[0110] S404B: The query nodes in query node 131 obtain sub-query results from corresponding data sets according to the data query sub-requests, and send the sub-query results to proxy 110.

[0111] Each query node in query node 131 obtains sub-query results from corresponding data sets according to the received data query sub-requests, and sends the sub-query results to proxy 110. For example, after the aforementioned query node 131D receives the data query sub-request for querying the sub-query results corresponding to the foregoing feature vector N from data set S7”, it can query, from data set S7”, the preset number of data with the highest similarity to feature vector N.

[0112] Specifically, for example, referring to Figure 6 , assume that data set S7” includes five pieces of data: data S7-1, data S7-2, data S7-3, data S7-4, and data S7-5. Query node 131D can obtain, according to the foregoing formula 1, that the Euclidean distances between the feature vectors of data S7-1, data S7-2, data S7-3, data S7-4, and data S7-5 and the feature vector N in the data query request are respectively: 0.014, 0.071, 0.143, 0.3, 0.412. Assume that the number of data in the sub-query results returned by each data set is 2. Then query node 131D returns, as the sub-query results of the data query request in data set S7, the two pieces of data with the smallest Euclidean distance between the feature vector N in data set S7”, namely data S7-1 and data S7-2, to proxy 110.

[0113] It can be understood that in some embodiments, each query node may also first send the sub-query results to query coordinator 121, and then query coordinator 121 forwards the sub-query results to proxy 110, which is not limited herein.

[0114] It can be understood that in some embodiments, after sending the sub-query results, each query node may also send a notification message indicating the completion of the data query sub-request to query coordinator 121, so that query coordinator 121 can schedule each query node based on this notification message, such as allocating query nodes / data sets for responding to other data query requests.

[0115] S405: Query coordinator 121 obtains the migration status of the target data set during the migration process, and determines the type of the target data set for responding to the data query request according to the migration status.

[0116] For the target data set during the migration process, query coordinator 121 obtains the migration status of the target data set, and determines the type of the target data set for responding to the data query request according to the migration status. For example, referring to Figure 3A, when the query coordinator 121 determines that the data set S5 is in the aforementioned incomplete addition state, it determines that the data set S5 of the streaming data is used to respond to the data query request. For another example, refer to Figure 3B and Figure 3C , when the query coordinator 121 determines that the data set S5 is in the aforementioned incomplete deletion state, it determines that the data set S5' of the batch data is used to respond to the data query request.

[0117] It can be understood that in some other embodiments, the query coordinator 121 can also use other methods to determine the type of the target data set for responding to the data query request according to the migration state, which is not limited herein.

[0118] S406A: The query coordinator 121 sends a data query sub-request to the query node corresponding to the target data set of the determined type.

[0119] After the query coordinator 121 determines the type of the target data set for responding to the data query request, it sends a data query sub-request to the query node where the target data set of the determined type is located. For example, after the query coordinator 121 determines to respond to the data query request through the data set S5 of the streaming data, it sends a data query sub-request to the query node 131A to query the sub-query result corresponding to the aforementioned feature vector N from the data set S5 of the streaming data. For another example, after the query coordinator 121 determines to respond to the data query request through the data set S5' of the batch data, it sends a data query sub-request to the query node 131B to query the sub-query result corresponding to the aforementioned feature vector N from the data set S5' of the batch data.

[0120] It can be understood that in some embodiments, after the query coordinator 121 determines that the target data set of the source data type (such as the target data set of the streaming data) responds to the data query request, it can also monitor the execution progress of the corresponding query node for this data query request. If, after the target data set of the source data type has been added as the target data set of the target data type (such as the target data set of the batch data), the query node has not obtained the sub-query result of this data query request, the query coordinator 121 can interrupt the deletion operation of the target data set of the source data type until the query node obtains the sub-query result of this data query request, and then resume the deletion operation of the target data set of the source data type.

[0121] Thereby, it can avoid the incomplete query results caused by deleting the data in the target data set of the source data type during the data query process, and improve the reliability of the query results of the database system 10.

[0122] Moreover, in the case where the target data set of the source data type responds to multiple data query requests, the query coordinator 121 may also resume the deletion operation on the target data set of the source data type when the sub-query results corresponding to all of the multiple data query requests have been obtained. For example, the query coordinator 121 may record the number of data query requests responded by each data set, increment the number when a new data query request arrives, decrement the number when a data query request is completed, and only allow the deletion operation on the target data set of the source data type when the number is 0.

[0123] This can avoid incomplete query results caused by deleting data in the target data set of the source data type during the data query process, and improve the reliability of the query results of the database system 10.

[0124] It can be understood that in some embodiments, after the query coordinator 121 determines that the target data set of the source data type (such as the target data set of streaming data) responds to a data query request, it may also monitor the operating conditions of the electronic device where the corresponding query node is located. After the target data set of the source data type has been increased to the target data set of the target data type (such as the target data set of batch data), if the operating conditions of the electronic device where the query node is located meet the foregoing preset conditions (such as the free memory of the electronic device where the query node is located is less than the memory free threshold / the memory usage is greater than the memory usage threshold, the free processor of the electronic device where the query node is located is less than the processor free threshold / the processor usage is greater than the processor usage threshold, the number of data query requests processed by the electronic device where the query node is located is greater than the preset number, etc.), suspend the unfinished data query requests, execute the deletion operation on the target data set of the source data type, and resume the suspended unfinished data query requests from the target data set of the target data type.

[0125] This can avoid the query node crashing due to excessive use of hardware resources, which affects the stability of the database system 10.

[0126] For example, in some embodiments, all data in the target data set of the target data type may be re-queryed. For another example, in some embodiments, assuming that the target data set of the source data type includes n pieces of data, and the first m (m < n) pieces of data among the n pieces of data have been queried when the query from the target data set of the source data type is suspended, then continue to query the query results of the data query request in the subsequent n - m pieces of data from the target data set of the target data type.

[0127] In this way, it is not necessary to repeatedly query the data that has been queried, which can save query time and improve query speed.

[0128] S406B: The query node in query node 131 obtains a sub-query result from the corresponding data set according to the data query sub-request, and sends the sub-query result to proxy 110.

[0129] The query node in query node 131 obtains a sub-query result from the corresponding data set according to the data query sub-request, and sends the sub-query result to proxy 110. For example, after query node 131A receives a data query sub-request for querying the sub-query result corresponding to the foregoing feature vector N from data set S5 of streaming data, it queries the sub-query result corresponding to the foregoing feature vector N from data set S5 of streaming data, and sends the sub-query result to proxy 110. The specific query method can refer to the relevant description in step S404B, which will not be elaborated here.

[0130] For another example, after query node 131A receives a data query sub-request for querying the sub-query result corresponding to the foregoing feature vector N from data set S5' of batch data, it queries the sub-query result corresponding to the foregoing feature vector N from data set S5' of batch data, and sends the sub-query result to proxy 110. The specific query method can refer to the relevant description in step S404B, which will not be elaborated here.

[0131] S407: Proxy 110 aggregates the sub-query results to obtain the query result of the data query request, and sends the query result to the user.

[0132] After receiving the sub-query results sent by each query node, proxy 110 aggregates the sub-query results to obtain the query result of the data query request, and sends the query result to the user. It can be understood that in some embodiments, the number of data in each sub-query result may be greater than the number K of data in the query result required to be returned by the data query request. At this time, proxy 110 can perform another screening on the sub-query results to determine the K data most similar to the foregoing feature vector N in the data query request as the query result of the data query request.

[0133] It can be understood that the query result of the data query request includes the data identifier of the data, and this data identifier can uniquely determine a piece of data in database system 10. In some other databases, the corresponding relationship between the data identifier and unstructured data (such as text, image, audio, video, DNA sequence, commodity information, substance structure, etc.) can be stored. Then, after the client 00 receives the query result, it can obtain the corresponding unstructured data from the database storing the corresponding relationship between the data identifier and unstructured data and the database storing unstructured data through the data identifiers of the data in the query result, and display the unstructured data to the user.

[0134] It can be understood that the execution order of the foregoing steps S401 to S407 is only an example. In some other embodiments, the execution order of some steps can be adjusted, and some steps can be combined or split, which is not limited herein. For example, in some embodiments, the foregoing steps S404A and S406A can be executed in parallel, and the foregoing step S404B can be executed in parallel with step S406B.

[0135] Through the method provided by the embodiments of the present application, when the database system 10 is in the process of processing and migrating the target data set corresponding to the received data query request, it does not need to wait for the migration process to be completed before performing data query, but instead executes the data query and the migration of the target data set in parallel. In this way, the speed of querying data by the database system 10 can be increased, and the experience of users using the database system 10 to query data can be improved.

[0136] Furthermore, the embodiments of the present application provide a database system 200. As Figure 7 shown, the database system 200 at least includes a coordination unit 201, a query unit 202, and a storage unit 203. Among them,

[0137] The coordination unit 201 can be used to determine the target data set corresponding to the data query request according to the received data query request, and determine whether the target data set is in the migration process, determine the migration status of the target data set, and determine the type of the target data set for responding to the data query request according to the migration status of the target data set, and the sub-query unit corresponding to the target data set of this type. For the specific functions of the coordination unit 201 and the methods for implementing these specific functions, reference can be made to the relevant descriptions of the foregoing query coordinator 121 (such as the foregoing description of the database system 10, the relevant descriptions of the foregoing steps S402, S403, S404A, S405, S406A, etc.), which will not be elaborated herein.

[0138] The query unit 202 can include at least one sub-query unit, which is used to obtain the query result of the data query request from the data set of the corresponding data type according to the data query request sent by the coordination unit 201. Specifically, reference can be made to the relevant descriptions of the foregoing query nodes 131 / 131A / 131B (such as the foregoing description of the database system 10, the relevant descriptions of the foregoing steps S404B, S406B, etc.), which will not be elaborated herein.

[0139] The storage unit 203 can be used to specifically store the data in the database system 10, such as the feature vectors of unstructured data, unstructured data, etc. Specifically, reference can be made to the relevant descriptions in the foregoing storage service 14, which will not be elaborated herein.

[0140] It can be understood, Figure 7The structure of the database system 200 shown is only an example. In some other embodiments, the database system 200 may also include more or fewer modules, and some modules may be combined or split. This is not limited herein.

[0141] When the target data set corresponding to the received data query request in the database system 200 provided by the embodiments of the present application is in the migration process, the database system 200 can select the target data set of the corresponding data type according to the migration state of the target data set to respond to the data query request, enabling parallel execution of data migration and data query, improving the data query speed of the database system 200, and enhancing the user experience of using the database system 200 to query data.

[0142] Furthermore, Figure 8 According to some embodiments of the present application, a schematic structural diagram of an electronic device 100 is shown. It can be understood that the electronic device 100 may be an electronic device that runs the nodes / modules in the database system 10, or may be the aforementioned client 00. As Figure 7 shown, the electronic device 100 may include one or more processors 101, a system memory 102, a non-volatile memory (NVM) 103, a communication interface 104, an input / output (I / O) device 105, and a system control logic 106 for coupling the processor 101, the system memory 102, the non-volatile memory 103, the communication interface 104, and the input / output (I / O) device 105. Among them:

[0143] The processor 101 may include one or more single-core or multi-core processors. In some embodiments, the processor 101 may include any combination of a general-purpose processor and a dedicated processor (e.g., a graphics processor, an application processor, a baseband processor, etc.). In some embodiments, the processor 101 may execute the instructions corresponding to the data query methods provided in the foregoing embodiments. For example, when the electronic device 100 is used to run the query coordinator 121, the processor 101 may be used to run instructions such as determining the target data set of the data query request, obtaining the migration state of the target data set, and determining the type of the target data set for responding to the data query request according to the migration state of the target data set. Another example is that when the electronic device is used to run the query node 131A, the processor 101 may be used to run instructions such as determining multiple data similar to the feature vector N in the data sets S1, S2, S3, S5, and S6 according to the feature vector N in the data query request.

[0144] The system memory 102 is a volatile memory, such as a Random-Access Memory (RAM), a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), etc. The system memory is used to temporarily store data and / or instructions. For example, in some embodiments, the system memory 102 can be used to temporarily store the instructions of the data query method provided in the foregoing embodiments, can also be used to store temporary copies of each data set, and can also be used to temporarily store query results of data query requests, etc.

[0145] The non-volatile memory 103 can include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the non-volatile memory 103 can include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as a Hard Disk Drive (HDD), a Compact Disc (CD), a Digital Versatile Disc (DVD), a Solid-State Drive (SSD), etc. The non-volatile memory 103 can also be a removable storage medium, such as a Secure Digital (SD) memory card, etc. In some embodiments, the non-volatile memory 103 can be used to store the instructions of the data query method provided in the foregoing embodiments, and can also permanently store the foregoing data sets, indexes of each data set, etc.

[0146] In particular, the system memory 102 and the non-volatile memory 103 can respectively include: a temporary copy and a permanent copy of the instruction 107. The instruction 107 can include: when executed by at least one of the processors 101, enabling the electronic device 100 to implement the data query method provided in the embodiments of the present application.

[0147] The communication interface 104 may include a transceiver for providing a wired or wireless communication interface for the electronic device 100, and thus communicating with any other suitable device via one or more networks. In some embodiments, the communication interface 104 may be integrated with other components of the electronic device 100. For example, the communication interface 104 may be integrated in the processor 101. In some embodiments, the electronic device 100 may communicate with other devices via the communication interface 104. For example, the communication interface 104 is used for communication between nodes / modules deployed in different electronic devices in the database system 10 (such as transmitting data query requests, transmitting query results, transmitting the migration status of each data set, etc.), and communicating with the client 00 (such as receiving data query requests and returning query results corresponding to the data query requests, etc.).

[0148] The input / output (I / O) device 105 may include a user interface that enables a user to interact with the electronic device 100. For example, in some embodiments, the input / output (I / O) device 105 may include output devices such as a display. For example, when the electronic device 100 is the client 00, the user may send unstructured data, such as images, audio, text, etc., as part of a data query request to the database system 10 via the input / output (I / O) device 105.

[0149] The system control logic 106 may include any suitable interface controller to provide any suitable interface for other modules of the electronic device 100. For example, in some embodiments, the system control logic 106 may include one or more memory controllers to provide an interface connected to the system memory 102 and the non-volatile memory 103.

[0150] In some embodiments, at least one of the processors 101 may be logically packaged with one or more controllers for the system control logic 106 to form a System in Package (SiP). In other embodiments, at least one of the processors 101 may also be integrated with the logic of one or more controllers for the system control logic 106 on the same chip to form a System-on-Chip (SoC).

[0151] It can be understood that the electronic device 100 can be any electronic device, including but not limited to tablet computers, desktop computers, servers / server clusters, laptop computers, handheld computers, notebooks, desktop computers, ultra-mobile personal computers (UMPCs), netbooks, mobile phones, cellular phones, personal digital assistants (PDAs), augmented reality (AR) / virtual reality (VR) devices, media players, smart TVs, smart speakers, smart watches, etc. The embodiments of the present application do not make any limitations in this regard.

[0152] It can be understood that the structure of the electronic device 100 shown in the embodiments of the present application does not constitute a specific limitation on the electronic device 100. In some other embodiments of the present application, the electronic device 100 may include more or fewer components than those shown in the figures, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figures can be implemented in hardware, software, or a combination of software and hardware.

[0153] Embodiments of the mechanisms disclosed in the present application can be implemented in hardware, software, firmware, or a combination of these implementation methods. Embodiments of the present application can be implemented as a computer program or program code executed on a programmable system, which includes at least one processor, a storage system (including volatile and non-volatile memories and / or storage elements), at least one input device, and at least one output device.

[0154] The program code can be applied to the input instructions to perform the various functions described in the present application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of the present application, the processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0155] The program code can be implemented in a high-level procedural language or an object-oriented programming language in order to communicate with the processing system. When necessary, it can also be implemented in assembly language or machine language. In fact, the mechanisms described in the present application are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.

[0156] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more transitory or non-transitory machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, the instructions may be distributed via a network or via other computer-readable media. Thus, machine-readable media may include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including but not limited to, floppy disks, optical disks, optical discs, CD-ROMs, magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, magnetic or optical cards, flash memory, or tangible machine-readable memories for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) in electrical, optical, acoustic, or other forms via the Internet. Thus, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0157] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such a particular arrangement and / or ordering may not be required. Rather, in some embodiments, these features may be arranged in a manner and / or order different from that shown in the illustrative drawings. Additionally, the inclusion of a structural or method feature in a particular figure does not imply that such a feature is required in all embodiments, and in some embodiments, these features may not be included or may be combined with other features.

[0158] It should be noted that each unit / module mentioned in the device embodiments of this application is a logical unit / module. Physically, a logical unit / module may be a physical unit / module, may be a part of a physical unit / module, or may also be implemented as a combination of multiple physical units / module. The physical implementation manner of these logical units / module themselves is not the most important. The combination of the functions implemented by these logical units / module is the key to solving the technical problems proposed in this application. In addition, in order to highlight the innovative part of this application, the above device embodiments of this application do not introduce units / modules that are not closely related to solving the technical problems proposed in this application, which does not mean that there are no other units / modules in the above device embodiments.

[0159] It should be noted that in the examples and the description of this patent, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0160] Although this application has been illustrated and described by reference to certain preferred embodiments thereof, those of ordinary skill in the art should understand that various changes may be made therein in form and detail without departing from the spirit and scope of this application.

Claims

1. A data query method, applied to an electronic device, characterized in that, The method includes: Determining that at least one first target data set in the target data set corresponding to the received first data query request is in the process of migrating from a first data type to a second data type; When the migration state of the first target data set is a first migration state, determining to query and obtain a sub-query result associated with the first data query request from the first target data set of the first data type, and not interrupting the operation of adding the first target data set of the first data type as the first target data set of the second data type; When the migration state of the first target data set is a second migration state, determining to query and obtain a sub-query result associated with the first data query request from the first target data set of the second data type, and not interrupting the deletion operation of the first target data set of the first data type; The first migration state indicates that some data in the first target data set of the first data type has been added as the second data type, and the first target data set of the first data type has not been deleted yet; The second migration state indicates that all data in the first target data set has been added as the second data type, and at least some data in the first target data set of the first data type has not been deleted yet.

2. The method according to claim 1, characterized in that, In the process of querying and obtaining a sub-query result associated with the first data query request from the first target data set of the first data type, it further includes: Detecting that all data in the first target data set of the first data type has been added as the second data type, and keeping the data in the first target data set of the first data type from being deleted.

3. The method according to claim 2, wherein The method further includes: When detecting that there is a second data query request accessing the first target data set of the first data type, before obtaining the query result of the second data query request, keeping the data in the first target data set of the first data type from being deleted.

4. The method according to claim 1, wherein In the process of querying and obtaining a sub-query result associated with the first data query request from the first target data set of the first data type, it further includes: Detecting that all data in the target data set of the first data type has been added as the second data type, and the operating state of the electronic device where the first target data set of the first data type is located meets a preset condition, pausing the querying and obtaining of the sub-query result associated with the first data query request from the first target data set of the first data type; Deleting the target data set of the first data type, and continuing to query and obtain the sub-query result associated with the first data query request in the first target data set of the second data type.

5. The method according to claim 4, wherein The operating state of the electronic device where the first target data set of the first data type is located meets a preset condition, including at least one of the following conditions: The free memory amount of the electronic device where the first target data set is located is less than a first threshold or the memory usage amount is greater than a second threshold; The free capacity of the processor of the electronic device where the first target data set is located is less than a third threshold or the usage of the processor is greater than a fourth threshold; The number of first data query requests processed by the electronic device where the first target data set is located exceeds a preset number.

6. The method according to any one of claims 1 to 5, characterized in that, The first data type includes any one of the following data types: streaming data, batch data, bulk import data, and the second data type includes any one of the following data types: streaming data, batch data, bulk import data.

7. The method according to any one of claims 1 to 5, characterized in that The target data set corresponding to the first data query request is stored in the database system. The first target data set of the first data type is stored in the first storage area of the database system, and the first target data set of the second data type is stored in the second storage area of the database system; And The first target data set is in the process of migrating from the first data type to the second data type, including: At least part of the data in the first target data of the first data type has been added to the data in the first target data set of the second data type, and at least part of the data in the first target data of the first data type has not been deleted.

8. The method according to claim 7, wherein The database system is a vector database system.

9. A database system, characterized in that, The database system includes: A coordination unit, configured to determine that when at least one first target data set is included in the target data set corresponding to the received first data query request, and the first target data set is in the process of migrating from the first data type to the second data type, when the migration state of the first target data set is the first migration state, determine to query and obtain a sub-query result associated with the first data query request in the first target data set of the first data type, without interrupting the operation of adding the first target data set of the first data type to the first target data set of the second data type, and send an instruction to query and obtain a sub-query result associated with the first data query request from the first target data set of the first data type to the query unit where the first target data set of the first data type is located, or, when the migration state of the first target data set is the second migration state, query and obtain a sub-query result associated with the first data query request in the first target data set of the second data type, without interrupting the deletion operation of the first target data set of the first data type, and send an instruction to query and obtain a sub-query result associated with the first data query request from the first target data set of the second data type to the query unit where the first target data set of the second data type is located, where the first migration state indicates that part of the data in the first target data set of the first data type has been added to the second data type and the first target data set of the first data type has not been deleted; the second migration state indicates that all the data in the first target data set has been added to the second data type and at least part of the data in the first target data set of the first data type has not been deleted; At least one query unit, configured to query and obtain a sub-query result associated with the first data query request from the target data set of the first data type or the target data set of the second data type according to an instruction sent by the coordination unit.

10. A readable medium, characterized in that, Instructions are stored on the readable medium, and when the instructions are executed on an electronic device, the electronic device implements the data query method according to any one of claims 1 to 8.

11. An electronic device, characterized in that, Comprising: A memory, configured to store instructions executed by one or more processors of the electronic device; And a processor, which is one of the processors of the electronic device, configured to execute the instructions stored in the memory to implement the data query method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Data inquiry condition generation method, and device, storage medium and electronic device

    CN108984623A

  • Data query method and device, electronic device and storage medium

    CN110674177A