Data query method, device and equipment based on data lake, medium and product
By creating index files in the data lake, the problem that data lakes are difficult to meet sequential reading and real-time retrieval in a single system is solved, and efficient data reading and retrieval is achieved, reducing storage redundancy and cost.
Patent Information
- Application Number
- CN202510435590.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to meet the needs of batch training data sequential reading and real-time inverted indexing and vector retrieval in a single data lake system, resulting in complex data processing links and redundant storage, limiting the efficient application of data lakes in multi-mode data processing.
The index files corresponding to multiple storage files are created in the data lake in advance, and the data is queried through the index files to realize random reading and sequential reading in the data lake, reducing the redundancy of data storage.
The random reading functions of batch training sequential reading, inverted order and vector search of data in the data lake are realized, reducing data storage redundancy and reducing the storage cost of the search engine.
Smart Images

Figure CN120353977A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to a data query method, apparatus, device, medium, and product based on a data lake. Background Art
[0002] Currently, as the core infrastructure supporting the storage and analysis of massive multi-modal data (structured data, semi-structured data, unstructured data), data lakes are widely used in large language models for processing multi-modal tasks (such as text-image understanding, cross-modal generation, intelligent retrieval, etc.).
[0003] The training and inference scenarios of large language models pose dual requirements for the efficient utilization of data: on the one hand, it is necessary to support the sequential reading of batch training data to achieve efficient distributed computing; on the other hand, it is necessary to have real-time inverted index and vector retrieval capabilities to meet the low-latency interactive query requirements.
[0004] However, existing technical architectures are difficult to meet the above requirements in a single data lake system, resulting in complex data processing links and redundant data storage, which severely restricts the efficient application of data lakes in multi-modal data processing. Summary of the Invention
[0005] Embodiments of the present disclosure provide a data query method, apparatus, device, medium, and product based on a data lake, which can reduce the redundancy of data storage.
[0006] In a first aspect, embodiments of the present disclosure provide a data query method based on a data lake, including:
[0007] In response to receiving a query request, obtaining index files corresponding to a plurality of pre-created storage files in the data lake, where the index files include a plurality of index information, and one index information is associated with the data row number information of a row of data in the storage file;
[0008] For each index file, determining at least one candidate index information that matches the query request from the plurality of index information included in the index file, and determining the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request;
[0009] According to the data row number information associated with each candidate index information, determining the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake;
[0010] Obtain corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information.
[0011] In a second aspect, an embodiment of the present disclosure provides a data query device based on a data lake, including:
[0012] An obtaining unit, configured to, in response to receiving a query request, obtain index files corresponding to a plurality of pre-created storage files from the data lake, where the index files include a plurality of index information, and one index information is associated with the data row number information of a row of data in the storage file;
[0013] A first determining unit, configured to, for each index file, determine at least one candidate index information that matches the query request from the plurality of index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request;
[0014] A second determining unit, configured to determine the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information according to the data row number information associated with each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake;
[0015] A query unit, configured to obtain corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information.
[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0017] The memory stores computer execution instructions;
[0018] The processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the data query method based on the data lake described in the first aspect and various possible designs of the first aspect above.
[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored, and when the processor executes the computer execution instructions, the data query method based on the data lake described in the first aspect and various possible designs of the first aspect above is implemented.
[0020] Fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which when executed by a processor implements the data query method based on a data lake as described in the first aspect above and various possible designs of the first aspect.
[0021] The data query method, device, equipment, medium and product based on a data lake provided in this embodiment include: in response to receiving a query request, obtaining index files respectively corresponding to a plurality of pre-created storage files from the data lake, where the index file includes a plurality of index information, and one index information is associated with the data line number information of a line of data in the storage file; for each index file, determining at least one candidate index information that matches the query request from the plurality of index information included in the index file, and determining the data line number information associated with each candidate index information and the matching value between each candidate index information and the query request; according to the data line number information associated with each candidate index information, determining the line number of the row data corresponding to each candidate index information in the data lake as the global line number information corresponding to each candidate index information, and the global line number information is used to represent the line number of a line of data in the storage file in the data lake; according to the matching value between each candidate index information and the query request and the global line number information corresponding to each candidate index information, obtaining the corresponding row data as the query result of the query request. In this technical solution, since index files respectively corresponding to a plurality of storage files are pre-created in the data lake that supports sequential reading, thus, the row data corresponding to the query request can be queried through the index files, realizing random reading of the data in the data lake. It can be seen that a set of data in the data lake can implement sequential reading for batch training and random reading functions of inverted index and vector retrieval, so the data storage redundancy of the data lake is reduced; moreover, the data lake can implement the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required to be used in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 It is a schematic diagram showing meeting schedule reminder information in the prior art provided by an embodiment of the present disclosure;
[0024] Figure 2 It is the flow of the data query method based on a data lake provided by an embodiment of the present disclosure Figure 1 ;
[0025] Figure 3 Schematic of the data query method based on the data lake provided by the embodiments of the present disclosure Figure 1 ;
[0026] Figure 4 Schematic of the data query method based on the data lake provided by the embodiments of the present disclosure Figure 2 ;
[0027] Figure 5 Block diagram of the structure of the information display device provided by the embodiments of the present disclosure;
[0028] Figure 6 Schematic diagram of the structure of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners
[0029] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0030] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0031] Currently, as the core infrastructure for supporting the storage and analysis of massive multi-modal data (structured data, semi-structured data, unstructured data), the data lake is widely used in large language models for processing multi-modal tasks (such as text and image understanding, cross-modal generation, intelligent retrieval, etc.).
[0032] The training and inference scenarios of large language models pose dual requirements for the efficient utilization of data: on the one hand, it is necessary to support the sequential reading of batch training data to achieve efficient distributed computing; on the other hand, it is necessary to have real-time inverted index and vector retrieval capabilities to meet the low-latency interactive query requirements.
[0033] However, the existing technical architecture is difficult to meet the above requirements in a single data lake system. The data for data retrieval in the data lake needs to be copied to the retrieval engine. In this way, the same data often needs to be stored twice, increasing the data storage redundancy. Moreover, the cost of storing data through the retrieval engine is also relatively high. Therefore, it severely restricts the efficient application of the data lake in multi-modal data processing.
[0034] To address the technical problems in the prior art, the inventors' technical concept is as follows: Multiple index files corresponding to respective storage files are pre-created in the data lake. Data is queried through the index files without the need to export the data to the retrieval engine. Moreover, the data lake itself supports sequential reading, so it is possible to achieve simultaneous sequential reading and retrieval query of a single piece of data within the data lake. Further, due to the distributed storage in columnar format in the data lake, the index query performance of the distributed storage in columnar format is optimized in the embodiments of the present disclosure.
[0035] Specifically, the steps for the data lake to query data through the index file may include: First, in response to receiving a query request, obtain the index files corresponding to the multiple pre-created storage files in the data lake. The index file includes multiple index information, and one index information is associated with the data row number information of a row of data in the storage file. Then, for each index file, determine at least one candidate index information that matches the query request from the multiple index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request; according to the data row number information associated with each candidate index information, determine the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information. The global row number information is used to represent the row number of a row of data in the storage file in the data lake. Finally, according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information, obtain the corresponding row data as the query result of the query request.
[0036] In this technical solution, since multiple index files corresponding to respective storage files are pre-created in the data lake that supports sequential reading, in this way, the row data corresponding to the query request can be queried through the index files, realizing the random reading of the data in the data lake. Thus, a set of data in the data lake can achieve the sequential reading for batch training and the random reading functions of inverted index and vector retrieval. Therefore, the data storage redundancy of the data lake is reduced. Moreover, the data lake can achieve the random reading function of retrieval, so there is no need to export the data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0037] The application scenarios of the embodiments of the present disclosure are explained below:
[0038] The data query method based on a data lake provided by an embodiment of the present disclosure can be applied to scenarios for optimizing the performance of various data storage systems that support sequential reading. Figure 1 FIG. is a schematic diagram of an application scenario of a data query method based on a data lake provided by an embodiment of the present disclosure. As Figure 1 shown, a user can send a data query request to a server 102 through a terminal 101. After receiving the data query request, the server 102, through the data query method provided by an embodiment of the present disclosure, based on index files respectively corresponding to multiple storage files pre-created in the data lake, obtains row data corresponding to the query request from the data lake to obtain a query result. The server 102 returns the query result to the terminal 101.
[0039] It should be noted that the above data query request can obtain row data corresponding to the query request from the data lake through inverted index retrieval or vector retrieval to implement the random reading function of the data lake.
[0040] Among them, inverted index retrieval is suitable for keyword retrieval to locate the position of data through keywords. At this time, the above data query method can be applied to multiple scenarios. For example, in a big data processing scenario, documents (such as web pages, product titles) containing specific keywords can be obtained from the data lake through inverted index retrieval. Another example is that in a log analysis scenario, error logs can be obtained from the data lake through an inverted index to support time range filtering.
[0041] Among them, vector indexing is suitable for similarity retrieval to retrieve similar data through similarity (for example, vector distance). At this time, the above data query method can be applied to multiple scenarios. For example, it can be applied to a scenario of searching for similar pictures through text. By inputting "sunset beach", relevant pictures can be obtained from the data lake through vector indexing. Another example is that in a big data processing scenario, the meaning of a user's question can be understood through vector indexing, and relevant answers can be obtained from the data lake.
[0042] Moreover, as Figure 1 shown, the data lake itself has a sequential reading function and can be used for a training engine to match and sequentially read data from the data lake. For example, when training an image classification model, sequential reading of image data in the data lake can be performed to avoid delays caused by random jumps.
[0043] It can be seen that by pre-creating index files in the data lake, a set of data in the data lake supports sequential reading for model training and also supports random reading through inverted index retrieval or vector retrieval using a data query request, thereby reducing the data storage redundancy of the data lake.
[0044] The following is the specific implementation process of the data query method, device, equipment, medium and product based on the data lake involved in the embodiments of the present disclosure. Some examples are for illustration only and are not limited. The execution subject of the data query method based on the data lake involved in the embodiments of the present disclosure is an electronic device, which can be a terminal, a server, etc.
[0045] Figure 2 The flow of the data query method based on the data lake provided by the embodiments of the present disclosure Figure 1 , such as Figure 2 shown, the data query method based on the data lake may include:
[0046] S201. In response to receiving a query request, obtain the index files corresponding to a plurality of pre-created storage files in the data lake, where the index file includes a plurality of index information, and one index information is associated with the data row number information of a row of data in the storage file.
[0047] In the embodiments of the present disclosure, the data lake includes a plurality of storage files, and each storage file includes a plurality of rows of data. The query request is used to read the rows of data in one or more storage files.
[0048] Optionally, the query request includes one or more filtering condition information. Optionally, the filtering condition information may include a retrieval condition. Among them, the retrieval condition in the query request is used to screen out candidate index information that matches the query request from the index file. Optionally, the retrieval condition may include one or more types of information such as text, image, etc.
[0049] For example, the retrieval condition includes the text information "iceberg". The index file corresponding to storage file 1 includes the index information 1 of the first row of data, the index information 2 of the second row of data, and the index information 3 of the third row of data; among them, the index information 1 of the first row of data is: "iceberg--[1]", the index information 2 of the second row of data is "sunset--[2]", and the index information 3 of the third row of data is "glacier--[3]". Among them, the matching value between the retrieval condition "iceberg" and the index information 1 "iceberg--[1]" is 1, the matching value between the retrieval condition "iceberg" and the index information 2 "sunset--[2]" is 0, and the matching value between the retrieval condition "iceberg" and the index information 3 "glacier--[3]" is 0.5. Among them, a matching value greater than 0 indicates that the candidate index information matches the query request. In this case, according to the retrieval condition "iceberg", the candidate index information with a matching value greater than 0 can be screened out from the index file: the index information 1 "iceberg--[1]" of the first row of data and the index information 3 "glacier--[3]" of the third row of data.
[0050] Among them, when the matching value is 0, it indicates that the index information does not match the query request. At this time, the index information will not be filtered as candidate index information. When the matching value is greater than 0, it indicates that the index information matches the query request. At this time, the index information will be filtered as candidate index information. It can be understood that the closer the matching value is to 1, the higher the matching degree between the retrieval condition and the index information.
[0051] In the embodiments of the present disclosure, the matching value between the retrieval condition and the index information can be determined according to the vector similarity between the retrieval condition and the index information. Specifically, the retrieval condition can be converted into a first word vector, and the index information can be converted into a second word vector; the vector similarity between the first word vector and the second word vector is determined as the matching value between the retrieval condition and the index information. It should be noted that when converting the index information into the second word vector, only the text in the index information needs to be converted into a word vector, and the line number associated with the index information does not need to be converted. For example, the index information 1 of the first row of data is: "iceberg--[1]". Only "iceberg" needs to be converted into a word vector, and the line number "[1]" associated with the index information does not need to be converted.
[0052] Among them, the filtering condition information can also include a limiting condition. Optionally, the limiting condition is used to limit the query result. Optionally, the limiting condition can be the number of query results, the storage date of the query results, etc.
[0053] Exemplarily, the retrieval condition includes the text information "iceberg". The filtering condition information is that the number of query results is 1. At this time, according to the retrieval condition "iceberg", 1 candidate index information with the highest matching value with the retrieval condition can be filtered out from the index file: the index information 1 of the first row of data "iceberg--[1]". Among them, the limiting condition can be expressed as "limit 1".
[0054] In the embodiments of the present disclosure, the storage file can be structured data, semi-structured data or unstructured data. Optionally, the storage file includes one or more of text, image, video, animation. Optionally, the data lake stores data in a columnar storage format and in a distributed manner. Among them, the storage format of the storage file can be any type of storage format. For example, the storage format of the storage file is the parquet (columnar storage) format. Among them, the distributed storage can be HDFS (Hadoop Distributed File System, a highly fault-tolerant distributed file system).
[0055] It should be noted that storing data in a columnar storage format can compress the data storage space, and storing the index file through HDFS distributed storage can reduce the storage cost.
[0056] Optionally, the index file includes multiple index information, which corresponds one-to-one with the row data in the storage file and is associated with the data row number information of a row of data in the storage file. Among them, the index information may be inverted index information or vector index information. Among them, the data row number information represents the unique identifier of the row data in the storage file. Optionally, the data row number information may be a row number, a file path, or a UUID (Universally Unique Identifier).
[0057] In the embodiments of the present disclosure, an index file for each storage file can be pre-created by a search engine. Exemplarily, as Figure 3 shown, the data lake includes Storage File 1, Storage File 2, and Storage File 3, and the respective index files corresponding to the multiple storage files are: Index File 1, Index File 2, and Index File 3.
[0058] In some embodiments, as Figure 1 shown, a query engine is deployed inside the front-end query node. When a query request is received, the data in the above data lake can be sliced into multiple query tasks according to file granularity by the query engine, and the query tasks are scheduled to execution nodes for execution.
[0059] Exemplarily, the data lake includes Storage File 1, Storage File 2, and Storage File 3, and the multiple query tasks sliced according to file granularity include Query Task 1, Query Task 2, and Query Task 3.
[0060] S202. For each index file, determine at least one candidate index information that matches the query request from the multiple index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request.
[0061] In the embodiments of the present disclosure, multiple query tasks can be concurrently executed by a distributed query engine. Correspondingly, for each index file, determining at least one candidate index information that matches the query request from the multiple index information included in the index file includes: generating a query task corresponding to each index file by the distributed query engine; concurrently executing the generated multiple query tasks by the distributed query engine, and by executing the query task corresponding to each index file, determining at least one candidate index information that matches the query request from the multiple index information included in the index file.
[0062] For example, if there are 10,000 storage files in the data lake, then for the index file of each storage file, 1 corresponding query task is generated, and a total of 10,000 query tasks are generated. The 10,000 query tasks are concurrently executed by the distributed query engine, and multiple candidate index information is determined from the index files corresponding to the 10,000 storage files.
[0063] In the disclosed embodiments, the execution order of the steps of sorting multiple storage files and the step of generating query tasks is not specifically limited. Optionally, the step of sorting multiple storage files may be executed first, and then the step of generating query tasks corresponding to each storage file may be executed; alternatively, the step of generating query tasks corresponding to each storage file may be executed first, and then the step of sorting multiple storage files may be executed; or the step of generating query tasks corresponding to each storage file and the step of sorting multiple storage files may be executed simultaneously.
[0064] In the disclosed embodiments of the present disclosure, since query tasks corresponding to each index file are first generated by the distributed query engine, and then the multiple generated query tasks are concurrently executed by the distributed query engine, when there are many storage files, the efficiency of determining multiple candidate index information from the index files corresponding to each storage file can be improved.
[0065] In some embodiments, the query request includes a retrieval condition and a limit condition; correspondingly, this step may include: for each index file, determining the matching value between each index information included in the index file and the retrieval condition; sorting in descending order of the matching value, and selecting at least one candidate index information that meets the limit condition from the index file as at least one candidate index information that matches the query request. Wherein, for the same index information, the matching value between the index information and the retrieval condition is equal to the matching value between the index information and the query request.
[0066] In the disclosed embodiments of the present disclosure, the limit condition in the query request is not specifically limited. Optionally, the limit condition may be the number of query results, the storage date of the query results, etc.
[0067] Exemplarily, as Figure 3 shown, the query request is: where match_any(text, 'iceberg') limit 1. Wherein, the text information in the query request is "iceberg" (iceberg). The limit condition in the query request is "limit 1", indicating that the number of candidate index information is 1.
[0068] Exemplarily, the data lake includes storage file 1, storage file 2, and storage file 3, and the index files corresponding to the multiple storage files are respectively: index file 1, index file 2, and index file 3.
[0069] Query the index file 1 to obtain a candidate index information 1. Determine that the data row number information associated with the candidate index information 1 is 10, and the matching value between the candidate index information and the query request is 0.5.
[0070] Query the index file 2 to obtain a candidate index information 2. Determine that the data row number information associated with the candidate index information 2 is 20, and the matching value between the candidate index information and the query request is 0.6.
[0071] Query the index file 3 to obtain a candidate index information 3. Determine that the data row number information associated with the candidate index information 3 is 10, and the matching value between the candidate index information and the query request is 0.4.
[0072] In the embodiments of the present disclosure, in the index query stage, by using the retrieval conditions and restriction conditions in the query request, the filtering conditions of the query request are increased, thereby improving the accuracy and efficiency of retrieving candidate index information that meets the conditions from the index file.
[0073] S203. According to the data row number information associated with each candidate index information, determine the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information. The global row number information is used to represent the row number of a row of data in the storage file in the data lake.
[0074] In the embodiments of the present disclosure, the data row number information associated with multiple candidate index information selected from different index files may be the same. At this time, the global row number information can be used to distinguish the data row number information associated with different candidate index information.
[0075] In some embodiments, multiple storage files can be sorted first, and then the global row number information can be obtained by adding up the data row number information and the file row number information of each storage file in front of the storage file position corresponding to the data row number information in the sorting. The purpose of sorting the multiple storage files is to obtain the global row number information corresponding to each candidate index information.
[0076] Optionally, the data row number information is used to represent the row number of the row data in the storage file in the storage file; correspondingly, according to the data row number information associated with each candidate index information, determining the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information may include the following steps (1) to (3):
[0077] (1) Determine the file row number information of each storage file in the data lake. The file row number information is used to represent the row number of the storage file in the data lake.
[0078] Optionally, this step may include: for each stored file, obtaining the total number of lines of the data included in the stored file, and sorting the multiple stored files in the data lake according to the unique file identifier of the stored file; for each stored file, according to the sorting of the multiple stored files, obtaining the total number of lines of the data included in each of the stored files sorted before the position of this stored file; determining the sum of the total number of lines of the data included in each stored file as the file line number information corresponding to the stored file.
[0079] Among them, the multiple stored files in the data lake can be sorted according to various methods. In some embodiments, the unique file identifier is used to distinguish different stored files. The unique file identifier may include one or more of the file name, file number, and file storage time. Optionally, the multiple stored files in the data lake can be sorted by the file name of the stored file. Optionally, the multiple stored files in the data lake can be sorted by the storage time of the stored file.
[0080] Optionally, if the stored file is the first stored file in the sorting, then determine that the file line number information corresponding to the stored file is 0.
[0081] Exemplarily, as Figure 3 shown, the data lake includes stored file 1, stored file 2, and stored file 3. The total number of lines of the data included in stored file 1 is 1000, the total number of lines of the data included in stored file 2 is 1000, and the total number of lines of the data included in stored file 3 is 1000.
[0082] Among them, since stored file 1 is the first stored file in the sorting, it is determined that the file line number information corresponding to stored file 1 is 0. The file line number information corresponding to stored file 2 is: the total number of lines of the data included in stored file 1 (1000) sorted before stored file 2. The file line number information corresponding to stored file 3 is: the sum of the total number of lines of the data included in stored file 1 and the total number of lines of the data included in stored file 2 sorted before stored file 3 (1000 + 1000).
[0083] (2) For each candidate index information, determine the target index file to which the candidate index information belongs.
[0084] Exemplarily, as Figure 3 shown, it is determined that the target index file to which the candidate index information 1 obtained by querying index file 1 belongs is index file 1. It is determined that the target index file to which the candidate index information 2 obtained by querying index file 2 belongs is index file 2. It is determined that the target index file to which the candidate index information 3 obtained by querying index file 3 belongs is index file 3.
[0085] (3) Determine the line number of the row data corresponding to the candidate index information in the data lake according to the sum of the line number information of the target storage file corresponding to the target index file and the line number information of the data rows associated with the candidate index information, and use it as the global line number information corresponding to the candidate index information.
[0086] Exemplarily, as Figure 3 shown, the line number information of the data rows associated with candidate index information 1 is 10, and the line number information of index file 1 to which candidate index information 1 belongs is 0. At this time, determine the line number 10 (i.e., 10 + 0) of the row data corresponding to candidate index information 1 in the data lake as the global line number information corresponding to candidate index information 1.
[0087] The line number information of the data rows associated with candidate index information 2 is 20, and the line number information of index file 2 to which candidate index information 2 belongs is 1000. At this time, determine the line number 1020 (i.e., 1000 + 20) of the row data corresponding to candidate index information 2 in the data lake as the global line number information corresponding to candidate index information 2.
[0088] The line number information of the data rows associated with candidate index information 3 is 10, and the line number information of index file 3 to which candidate index information 3 belongs is 2000. At this time, determine the line number 2010 (i.e., 2000 + 10) of the row data corresponding to candidate index information 3 in the data lake as the global line number information corresponding to candidate index information 3.
[0089] In the embodiments of the present disclosure, in view of the characteristic that the data lake stores data in a columnar storage format, global line number information is introduced. By introducing the global line number information, the candidate index information filtered out by each index file can be distinguished, and further, the line number information of the data rows associated with each candidate index information can be distinguished. In this way, the query accuracy can be improved.
[0090] It should be noted that if the global line number information is not introduced, after determining the data line number information associated with each candidate index information, since it is not known which storage file the data line number information corresponds to, it is necessary to check the corresponding records in each storage file in the data lake. Exemplarily, if there are 10,000 storage files and the row data corresponding to 10 candidate index information is retrieved, at this time, 10,000 * 10 records need to be checked. Therefore, it will bring a serious read amplification problem. In the embodiments of the present disclosure, by introducing the global line number information, it is possible to first sort the global line number information, select the top N (such as top 10) according to the sorting, and then check the corresponding storage file. Among them, the target storage file corresponding to the data line number information can be determined through the global line number information. In this way, no matter how many storage files there are in the data lake, the number of records for checking the target storage file is N, so the read amplification problem of data can be avoided. It can be seen that by introducing the global line number information, the read amplification problem of data can also be avoided, and the query efficiency can be improved.
[0091] S204. Obtain the corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global line number information corresponding to each candidate index information.
[0092] In some embodiments, this step may include the following steps (1) to (3):
[0093] (1) Select a preset number of candidate index information as the target index information from the determined multiple candidate index information according to the matching value between each candidate index information and the query request.
[0094] Optionally, this step includes: sorting in descending order according to the matching value between each candidate index information and the query request, and selecting a preset number of candidate index information from the determined multiple candidate index information as the target index information.
[0095] Exemplarily, as Figure 3 shown, the preset number is 1. Sorting in descending order according to the matching value, 1 candidate index information (1020, 0.6) is selected from the 3 determined candidate index information as the target index information.
[0096] (2) Determine the target storage file corresponding to the global line number information and determine the data line number information corresponding to the global line number information according to the global line number information corresponding to each target index information.
[0097] Optionally, this step may include: determining the line number range interval of each storage file in the data lake, determining the target line number range interval to which the global line number information belongs according to the global line number information corresponding to each target index information, determining the storage file corresponding to the target line number range interval as the target storage file corresponding to the global line number information; and subtracting the minimum line number information in the target line number range interval from the global line number information to obtain the data line number information corresponding to the global line number information.
[0098] In some embodiments, the target storage file corresponding to the global line number information and the data line number information corresponding to the global line number information may be determined according to the file sorting information. Correspondingly, determining the line number range interval of each storage file in the data lake includes: obtaining the total number of lines of data included in each storage file, and sorting the multiple storage files in the data lake according to the file unique identifier of the storage file; for each storage file, obtaining the total number of lines of data included in each storage file sorted before the position of this storage file according to the sorting of the multiple storage files; determining the sum of the total number of lines of data included in each storage file as the minimum line number information corresponding to the line number range interval of this storage file; determining the sum of the minimum line number information and the total number of lines of data included in this storage file as the maximum line number information corresponding to the line number range interval of this storage file.
[0099] Exemplarily, the data lake includes storage file 1, storage file 2, and storage file 3. The total number of lines of data included in storage file 1 is 1000, the total number of lines of data included in storage file 2 is 1000, and the total number of lines of data included in storage file 3 is 1000. In this case, it is determined that the line number range interval corresponding to storage file 1 is (0, 1000). It is determined that the line number range interval corresponding to storage file 1 is (1000, 2000). It is determined that the line number range interval corresponding to storage file 1 is (2000, 3000).
[0100] As Figure 3 shown, if the global line number information is 1020, then it is determined that the interval to which 1020 belongs is (1000, 2000), and the target storage file corresponding to this interval is "storage file 2". Subtracting the minimum line number information 1000 from the global line number information 1020, the data line number information corresponding to the global line number information is obtained as "20".
[0101] (3) According to the data line number information corresponding to the global line number information, obtain the line data corresponding to the data line number information as the query result of the query request.
[0102] In an embodiment of the present disclosure, after determining the target storage file corresponding to the global line number information and determining the data line number information corresponding to the global line number information, the record corresponding to this line can be point-checked from the corresponding target storage file through the data line number information, completing the entire query process.
[0103] Optionally, the data line number information corresponds to a line of data in the storage file. Accordingly, this step is: according to the data line number information corresponding to the global line number information, query the line data corresponding to the data line number information from the target storage file to obtain the line data corresponding to the query request.
[0104] Among them, the target storage file is stored in the data lake. Before querying the line data corresponding to the data line number information from the target storage file, it is necessary to first load the target storage file from the data lake.
[0105] It can be understood that when the number of target storage files is large, it needs to be loaded from the data lake multiple times, thus affecting the query efficiency. In some embodiments, multiple storage files in the data lake are stored in a columnar format. At this time, the storage files can be divided into multiple page files, and the query performance of the columnar format of the data lake can be optimized through page indexing.
[0106] Optionally, the target storage file includes multiple page files. Accordingly, this step may include: according to the data line number information corresponding to the global line number information, determine the target page file corresponding to the data line number information from the multiple page files included in the target storage file; determine whether the target page file has been cached in the local cache; if so, directly obtain the line data corresponding to the data line number information from the target page file cached in the local cache as the query result of the query request, if not, first load the target page file from the data lake into the local cache, and then obtain the line data corresponding to the data line number information from the target page file cached in the local cache as the query result of the query request.
[0107] Optionally, the storage file can be logically sliced into several page files page according to a fixed size (such as 1MB).
[0108] Exemplarily, such as Figure 4As shown, the stored file is divided into page file 1, page file 2, and page file 3. Among them, the r1, r2, r3, r4, and r5 block areas respectively represent the row data corresponding to 5 query requests. In the random reading stage of index data, the row data can be aligned to the page files first. As shown in the figure, r1 and r4 can be aligned to page file 1. r2 and r5 can be aligned to page file 2. r3 can be aligned to page file 3. When reading r1, r2, and r3, the corresponding page file 1, page file 2, and page file 3 can be loaded from the data lake into the local cache. When reading r4, since page file 1 has been cached, the page file 1 in the local cache can be directly reused, so there is no need to reload page file 1 from the data lake. When reading r5, since page file 2 has been cached, the page file 2 in the local cache can be directly reused, so there is no need to reload page file 2 from the data lake.
[0109] In the embodiment of the present disclosure, since the target page file is loaded from the data lake into the local cache when the target page file is read for the first time, in this way, when the target page file is read again, it can be read from the local cache without reloading from the data lake. Among them, the speed of reading data from the local cache is relatively fast, so the processing efficiency of the query request can be improved.
[0110] In some embodiments, the metadata corresponding to the target storage file includes the offset and page file size of each page file; the target page file can be loaded from the data lake into the local cache through the offset and page file size of the page file. Correspondingly, loading the target page file from the data lake into the local cache may include the following steps (a) to (c):
[0111] (a) Determine the number of rows of data included in each page file, divide the data row number information corresponding to the global row number information by this number of rows to obtain a quotient value and a remainder.
[0112] (b) Determine the target page file corresponding to the data row number information according to the quotient value and the page row number information of the data row number information in the target page file according to the remainder.
[0113] Optionally, determining the target page file corresponding to the data row number information according to the quotient value includes: determining the quotient value obtained by dividing the data row number information corresponding to the global row number information as the page index of the target page file; determining the page file corresponding to the target page index as the target page file corresponding to the data row number information.
[0114] Exemplarily, as Figure 4 shown, the stored file is divided into page file 1, page file 2, and page file 3. The number of rows of data included in page file 1, page file 2, and page file 3 is 100. Among them, the page index corresponding to page file 1 is 0, the page index corresponding to page file 2 is 1, and the page index corresponding to page file 3 is 2.
[0115] At this time, if the data line number information corresponding to the global line number information is 120, divide the data line number information to obtain a quotient value of 1 and a remainder of 20. Determine the page file 2 corresponding to the quotient value 1 (i.e., page index 1) as the target page file corresponding to the data line number information. Determine that the remainder 20 is the page line number information of the data line number information in the page file 2.
[0116] (c) Obtain the offset and page file size of the target page file from the metadata corresponding to the target storage file, and load the target page file from the data lake into the local cache according to the offset and page file size of the target page file.
[0117] In the embodiments of the present disclosure, as Figure 1 shown, the metadata corresponding to the storage file may be created when the storage file is stored in the data lake. The index file corresponding to the storage file may be created when the storage file is stored in the data lake.
[0118] In the embodiments of the present disclosure, since the storage file is divided into multiple page files, the target page file corresponding to the data line number information and the page line number information of the data line number information in the target page file can be determined through the data line number information and the page index. When reading the target page file, unnecessary data can be skipped, so that the read amplification of the columnar storage format can be effectively reduced, thereby improving the query performance of the columnar storage format of the data lake.
[0119] It should be noted that before obtaining the offset and page file size of the target page file from the metadata corresponding to the target storage file and loading the target page file from the data lake into the local cache according to the offset and page file size of the target page file, the method further includes: determining whether the target page file has been cached in the local cache; if so, directly read the row data corresponding to the page line number information from the target page file in the local cache to obtain the row data corresponding to the query request, and if not, perform the steps of obtaining the offset and page file size of the target page file from the metadata corresponding to the target storage file and loading the target page file from the data lake into the local cache according to the offset and page file size of the target page file.
[0120] In the embodiments of the present disclosure, since when the target page file is read for the first time, the target page file can be loaded from the data lake into the local cache according to the offset and page file size of the target page file, so that when the target page file is read again, it can be read from the local cache instead of being reloaded from the data lake. Among them, the speed of reading data from the local cache is relatively fast, so the processing efficiency of the query request can be improved.
[0121] An embodiment of the present disclosure provides a data query method based on a data lake. Since index files corresponding to multiple storage files are pre-created in the data lake that supports sequential reading, thus, row data corresponding to a query request can be queried through the index files, realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can achieve sequential reading for batch training and random reading functions of inverted index and vector retrieval. Therefore, the data storage redundancy of the data lake is reduced; moreover, the data lake can achieve the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0122] Figure 5 The following is a structural block diagram of a data query device based on a data lake provided by an embodiment of the present disclosure. As Figure 5 shown, the data query device based on the data lake includes:
[0123] An acquisition unit 501, configured to, in response to receiving a query request, acquire index files corresponding to multiple pre-created storage files in the data lake, where the index files include multiple index information, and one index information is associated with the data row number information of a row of data in the storage file;
[0124] A first determination unit 502, configured to, for each index file, determine at least one candidate index information that matches the query request from the multiple index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request;
[0125] A second determination unit 503, configured to determine the row number of the row data corresponding to each candidate index information in the data lake according to the data row number information associated with each candidate index information, as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake;
[0126] A query unit 504, configured to obtain the corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information.
[0127] According to one or more embodiments of the present disclosure, the query request includes a retrieval condition and a restriction condition; correspondingly, the first determination unit 502, for each index file, determines at least one candidate index information that matches the query request from multiple index information included in the index file, specifically including: for each index file, determining a matching value between each index information included in the index file and the retrieval condition; sorting in descending order according to the matching value, and selecting at least one candidate index information that meets the restriction condition from the index file as at least one candidate index information that matches the query request.
[0128] In the embodiments of the present disclosure, since in the index query stage, the screening conditions of the query request are increased through the retrieval condition and the restriction condition in the query request, the accuracy and efficiency of retrieving candidate index information that meets the conditions from the index file are improved.
[0129] According to one or more embodiments of the present disclosure, the data line number information is used to represent the line number of the line data in the storage file in the storage file; correspondingly, the second determination unit 503 determines the global line number information corresponding to the data line number information associated with each candidate index information as the global line number information corresponding to the candidate index information, specifically including: determining the file line number information of each storage file in the data lake, where the file line number information is used to represent the line number of the storage file in the data lake; for each candidate index information, determining the target index file to which the candidate index information belongs; determining the sum of the file line number information of the target storage file corresponding to the target index file and the data line number information associated with the candidate index information as the global line number information corresponding to the candidate index information.
[0130] According to one or more embodiments of the present disclosure, the second determination unit 503 determines the file line number information of each storage file in the data lake, specifically including: for each storage file, obtaining the total number of lines of the data included in the storage file, and sorting the multiple storage files in the data lake according to the file unique identifier of the storage file; for each storage file, obtaining the total number of lines of the data included in each storage file sorted before the position of the storage file according to the sorting of the multiple storage files; determining the sum of the total number of lines of the data included in each storage file as the file line number information corresponding to the storage file.
[0131] According to one or more embodiments of the present disclosure, the query unit 504 obtains corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information, specifically including: selecting a preset number of candidate index information as target index information from the determined multiple candidate index information according to the matching value between each candidate index information and the query request; determining the target storage file corresponding to the global row number information and determining the data row number information corresponding to the global row number information according to the global row number information corresponding to each target index information; and obtaining the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information.
[0132] In the embodiments of the present disclosure, in view of the characteristic that data in the data lake is stored in a columnar storage format, global row number information is introduced. By introducing the global row number information, the candidate index information screened out by each index file can be distinguished, and further, the data row number information associated with each candidate index information can be distinguished, so that the query accuracy can be improved; moreover, by introducing the global row number information, the problem of data read amplification can be avoided, and the query efficiency can be improved.
[0133] According to one or more embodiments of the present disclosure, the target storage file includes multiple page files; correspondingly, the query unit 504 obtains the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information, including: determining the target page file corresponding to the data row number information from the multiple page files included in the target storage file according to the data row number information corresponding to the global row number information; determining whether the target page file has been cached in the local cache; if so, directly obtaining the row data corresponding to the data row number information from the target page file cached in the local cache as the query result of the query request, if not, first loading the target page file from the data lake into the local cache, and then obtaining the row data corresponding to the data row number information from the target page file cached in the local cache as the query result of the query request.
[0134] In the embodiments of the present disclosure, since the target page file is loaded from the data lake into the local cache when the target page file is read for the first time, when the target page file is read again, it can be read from the local cache instead of being reloaded from the data lake. Among them, the speed of reading data from the local cache is relatively fast, so the processing efficiency of the query request can be improved.
[0135] Embodiments of the present disclosure provide a data query device based on a data lake. Since index files corresponding to multiple storage files are pre-created in the data lake that supports sequential reading, data rows corresponding to a query request can be queried through the index files, thus realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can implement sequential reading for batch training and random reading functions of inverted index and vector retrieval, thereby reducing the data storage redundancy of the data lake; moreover, the data lake can implement the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0136] To implement the above embodiments, embodiments of the present disclosure also provide an electronic device.
[0137] Referring to Figure 6 , which shows a schematic structural diagram of an electronic device 600 suitable for implementing embodiments of the present disclosure. The electronic device 600 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers, portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of embodiments of the present disclosure.
[0138] As Figure 6 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0139] Typically, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and a communication device 609. The communication device 609 can allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 6 an electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included.
[0140] Specifically, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0141] It should be noted that the computer-readable storage medium described above in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0142] The above computer-readable storage medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0143] The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.
[0144] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0145] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0146] The units involved in the embodiments described in this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0147] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on a Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0148] The electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present disclosure pre-create index files corresponding to multiple storage files in advance in a data lake that supports sequential reading. In this way, row data corresponding to a query request can be queried through the index files, realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can implement sequential reading for batch training and random reading functions of inverted index and vector retrieval, thus reducing the data storage redundancy of the data lake; moreover, the data lake can implement the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0149] In a first aspect, according to one or more embodiments of the present disclosure, a data query method based on a data lake is provided, including:
[0150] In response to receiving a query request, obtain index files corresponding to multiple pre-created storage files in the data lake, where the index files include multiple index information, and one index information is associated with the data row number information of a row of data in the storage file;
[0151] For each index file, determine at least one candidate index information that matches the query request from the multiple index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request;
[0152] According to the data row number information associated with each candidate index information, determine the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake;
[0153] According to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information, obtain the corresponding row data as the query result of the query request.
[0154] According to one or more embodiments of the present disclosure, the query request includes a retrieval condition and a limit condition; correspondingly, for each index file, determining at least one candidate index information that matches the query request from the multiple index information included in the index file includes: for each index file, determining the matching value between each index information included in the index file and the retrieval condition; sorting in descending order according to the matching value, and selecting at least one candidate index information that meets the limit condition from the index file as at least one candidate index information that matches the query request.
[0155] According to one or more embodiments of the present disclosure, the data line number information is used to represent the line number of the line data in the storage file in the storage file; correspondingly, determining the global line number information corresponding to the data line number information associated with each candidate index information as the global line number information corresponding to the candidate index information includes: determining the file line number information of each storage file in the data lake, where the file line number information is used to represent the line number of the storage file in the data lake; for each candidate index information, determining the target index file to which the candidate index information belongs; determining the sum of the file line number information of the target storage file corresponding to the target index file and the data line number information associated with the candidate index information as the global line number information corresponding to the candidate index information.
[0156] According to one or more embodiments of the present disclosure, determining the file line number information of each storage file in the data lake includes: for each storage file, obtaining the total number of rows of data included in the storage file, and sorting the multiple storage files in the data lake according to the file unique identifier of the storage file; for each storage file, according to the sorting of the multiple storage files, obtaining the total number of rows of data included in each storage file before the position of the storage file in the sorting; determining the sum of the total number of rows of data included in each storage file as the file line number information corresponding to the storage file.
[0157] According to one or more embodiments of the present disclosure, obtaining the corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global line number information corresponding to each candidate index information includes: selecting a preset number of candidate index information as target index information from the determined multiple candidate index information according to the matching value between each candidate index information and the query request; determining the target storage file corresponding to the global line number information and determining the data line number information corresponding to the global line number information according to the global line number information corresponding to each target index information; obtaining the row data corresponding to the data line number information as the query result of the query request according to the data line number information corresponding to the global line number information.
[0158] According to one or more embodiments of the present disclosure, the target storage file includes a plurality of page files; accordingly, obtaining the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information includes: determining, according to the data row number information corresponding to the global row number information, a target page file corresponding to the data row number information from the plurality of page files included in the target storage file; determining whether the target page file has been cached in the local cache; if so, directly obtaining the row data corresponding to the data row number information from the target page file cached in the local cache as the query result of the query request, and if not, first loading the target page file from the data lake into the local cache, and then obtaining the row data corresponding to the data row number information from the target page file cached in the local cache as the query result of the query request.
[0159] Embodiments of the present disclosure provide a data query method based on a data lake. Since index files corresponding to a plurality of storage files are pre-created in the data lake that supports sequential reading, thus, row data corresponding to a query request can be queried through the index files, realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can realize sequential reading for batch training and random reading functions of inverted index and vector retrieval, so the data storage redundancy of the data lake is reduced; moreover, the data lake can realize the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0160] In a second aspect, according to one or more embodiments of the present disclosure, there is provided a data query device based on a data lake, including:
[0161] An obtaining unit, configured to, in response to receiving a query request, obtain index files corresponding to a plurality of pre-created storage files from the data lake, where the index files include a plurality of index information, and one index information is associated with the data row number information of a row of data in the storage file;
[0162] A first determining unit, configured to, for each index file, determine at least one candidate index information that matches the query request from the plurality of index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request;
[0163] A second determining unit, configured to determine, according to the data row number information associated with each candidate index information, the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake;
[0164] A query unit, configured to obtain corresponding row data as a query result of the query request according to a matching value between each candidate index information and the query request and global row number information corresponding to each candidate index information.
[0165] According to one or more embodiments of the present disclosure, the query request includes a retrieval condition and a restriction condition; correspondingly, the first determination unit, for each index file, determines at least one candidate index information that matches the query request from a plurality of index information included in the index file, specifically including: for each index file, determining a matching value between each index information included in the index file and the retrieval condition; sorting in descending order according to the matching value, and selecting at least one candidate index information that satisfies the restriction condition from the index file as at least one candidate index information that matches the query request.
[0166] According to one or more embodiments of the present disclosure, the data row number information is used to represent the row number of the row data in the storage file in the storage file; correspondingly, the second determination unit determines the global row number information corresponding to the data row number information associated with each candidate index information as the global row number information corresponding to the candidate index information, specifically including: determining the file row number information of each storage file in the data lake, where the file row number information is used to represent the row number of the storage file in the data lake; for each candidate index information, determining the target index file to which the candidate index information belongs; determining the sum of the file row number information of the target storage file corresponding to the target index file and the data row number information associated with the candidate index information as the global row number information corresponding to the candidate index information.
[0167] According to one or more embodiments of the present disclosure, the second determination unit determines the file row number information of each storage file in the data lake, specifically including: for each storage file, obtaining the total number of rows of the data included in the storage file, and sorting a plurality of storage files in the data lake according to the unique file identifier of the storage file; for each storage file, obtaining the total number of rows of the data included in each storage file sorted before the position of the storage file according to the sorting of the plurality of storage files; determining the sum of the total number of rows of the data included in each storage file as the file row number information corresponding to the storage file.
[0168] According to one or more embodiments of the present disclosure, the query unit obtains corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information. Specifically, it includes: selecting a preset number of candidate index information as target index information from the determined multiple candidate index information according to the matching value between each candidate index information and the query request; determining the target storage file corresponding to the global row number information and determining the data row number information corresponding to the global row number information according to the global row number information corresponding to each target index information; obtaining the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information.
[0169] According to one or more embodiments of the present disclosure, the target storage file includes multiple page files; correspondingly, the query unit obtains the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information, including: determining the target page file corresponding to the data row number information from the multiple page files included in the target storage file according to the data row number information corresponding to the global row number information; determining whether the target page file has been cached in the local cache; if so, directly obtaining the row data corresponding to the data row number information from the target page file in the local cache as the query result of the query request, if not, first loading the target page file from the data lake into the local cache, and then obtaining the row data corresponding to the data row number information from the target page file in the local cache as the query result of the query request.
[0170] The embodiments of the present disclosure provide a data query device based on a data lake. Since index files corresponding to multiple storage files are pre-created in the data lake that supports sequential reading, thus, row data corresponding to a query request can be queried through the index files, realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can realize sequential reading for batch training and random reading functions of inverted index and vector retrieval, thus reducing the data storage redundancy of the data lake; and, the data lake can realize the random reading function of retrieval, so there is no need to export data to the retrieval engine, thus reducing the storage cost of the retrieval engine.
[0171] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;
[0172] The memory stores computer execution instructions;
[0173] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the data query method based on the data lake as described in the first aspect above and various possible designs of the first aspect.
[0174] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the data query method based on the data lake as described in the first aspect above and various possible designs of the first aspect.
[0175] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product including a computer program, which, when executed by a processor, implements the data query method based on the data lake as described in the first aspect above and various possible designs of the first aspect.
[0176] In summary, since index files corresponding to multiple storage files are pre-created in a data lake that supports sequential reading, thus, row data corresponding to a query request can be queried through the index files, realizing random reading of data in the data lake. It can be seen that a set of data in the data lake can achieve sequential reading for batch training and random reading functions for inverted index and vector retrieval, thus reducing the data storage redundancy of the data lake; and, the data lake can achieve the random reading function of retrieval, so there is no need to export data to a retrieval engine, thus reducing the storage cost of the retrieval engine.
[0177] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solution formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the present disclosure.
[0178] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0179] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A data query method based on a data lake, characterized in that, including: In response to receiving a query request, obtaining index files corresponding to a plurality of pre-created storage files from a data lake, wherein each of the index files includes a plurality of index information, and one index information is associated with the data row number information of a row of data in the storage file; For each index file, determining at least one candidate index information that matches the query request from the plurality of index information included in the index file, and determining the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request; According to the data row number information associated with each candidate index information, determining the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake; According to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information, obtaining the corresponding row data as the query result of the query request.
2. The method according to claim 1, wherein The query request includes a retrieval condition and a limit condition; Accordingly, the determining, for each index file, at least one candidate index information that matches the query request from the plurality of index information included in the index file includes: For each index file, determining the matching value between each index information included in the index file and the retrieval condition; Sorting in descending order according to the matching value, and selecting at least one candidate index information that meets the limit condition from the index file as at least one candidate index information that matches the query request.
3. The method according to claim 1, characterized in that, The data row number information is used to represent the row number of the row data in the storage file in the storage file; Accordingly, the determining, according to the data row number information associated with each candidate index information, the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information includes: Determining the file row number information of each storage file in the data lake, where the file row number information is used to represent the row number of the storage file in the data lake; For each candidate index information, determining the target index file to which the candidate index information belongs; According to the sum of the file row number information of the target storage file corresponding to the target index file and the data row number information associated with the candidate index information, determining the row number of the row data corresponding to the candidate index information in the data lake as the global row number information corresponding to the candidate index information.
4. The method according to claim 3, wherein The determining the file row number information of each storage file in the data lake includes: For each storage file, obtaining the total number of rows of the data included in the storage file, and sorting the plurality of storage files in the data lake according to the unique file identifier of the storage file; For each storage file, obtaining the total number of rows of the data included in each storage file sorted before the position of the storage file according to the sorting of the plurality of storage files; Determining the sum of the total number of rows of the data included in each of the storage files as the file row number information corresponding to the storage file.
5. The method according to claim 1, wherein Obtaining corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information includes: Selecting a preset number of candidate index information as target index information from the determined multiple candidate index information according to the matching value between each candidate index information and the query request; Determining the target storage file corresponding to the global row number information and determining the data row number information corresponding to the global row number information according to the global row number information corresponding to each target index information; Obtaining the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information.
6. The method according to claim 5, characterized in that The target storage file includes multiple page files; Correspondingly, obtaining the row data corresponding to the data row number information as the query result of the query request according to the data row number information corresponding to the global row number information includes: Determining the target page file corresponding to the data row number information from the multiple page files included in the target storage file according to the data row number information corresponding to the global row number information; Determining whether the target page file has been cached in the local cache; If so, directly obtaining the row data corresponding to the data row number information from the target page file in the local cache as the query result of the query request, and if not, first loading the target page file from the data lake into the local cache, and then obtaining the row data corresponding to the data row number information from the target page file in the local cache as the query result of the query request.
7. A data query device based on a data lake, characterized in that Including: An obtaining unit, configured to, in response to receiving a query request, obtain index files respectively corresponding to multiple pre-created storage files from a data lake, where the index file includes multiple index information, and one index information is associated with the data row number information of a row of data in the storage file; A first determining unit, configured to, for each index file, determine at least one candidate index information that matches the query request from the multiple index information included in the index file, and determine the data row number information associated with each candidate index information and the matching value between each candidate index information and the query request; A second determining unit, configured to determine, according to the data row number information associated with each candidate index information, the row number of the row data corresponding to each candidate index information in the data lake as the global row number information corresponding to each candidate index information, where the global row number information is used to represent the row number of a row of data in the storage file in the data lake; A query unit, configured to obtain corresponding row data as the query result of the query request according to the matching value between each candidate index information and the query request and the global row number information corresponding to each candidate index information.
8. An electronic device, characterized in that, Including: A processor and a memory; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the data query method based on a data lake according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the data query method based on the data lake as described in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the data query method based on the data lake as described in any one of claims 1 to 6 is implemented.